system

The system addresses home training challenges by using video and image analysis to provide real-time posture correction and nutritional guidance, ensuring effective and motivated home workouts.

JP2026069165APending Publication Date: 2026-04-23SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-11
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Individuals face challenges in effective home training without professional guidance, including incorrect postures, risk of injury, motivation maintenance, and complex nutritional management, which existing systems fail to address comprehensively.

Method used

A system comprising video acquisition, image analysis, difference measurement, feedback generation, and meal analysis means to provide real-time personalized fitness guidance and dietary management, utilizing machine learning and natural language feedback.

Benefits of technology

Enables efficient, safe, and comprehensive home training and health management by correcting postures and providing nutritional advice, enhancing training efficiency and motivation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026069165000001_ABST
    Figure 2026069165000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] Means of acquiring video, Image analysis means for analyzing the posture of the human body acquired by the image acquisition means and identifying joint positions, A difference measurement means that measures the difference between pre-registered ideal posture data and analyzed posture data, A feedback generation means that generates natural language feedback based on the difference obtained by the difference measurement means, A presentation means for presenting the generated feedback, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In modern times, it is difficult to succeed in home training without guidance from a professional trainer. There is an increasing need for the time and cost of going to the gym, and the need to exercise effectively without worrying about the eyes of others. However, in self-styled training, not only is it difficult to achieve results due to incorrect postures, but there is also a risk of injury. In addition, it is difficult to maintain motivation and there is a problem that it is difficult to continue. Furthermore, although it is important to grasp the content of meals and manage nutritional values for dieting and health management, personal management is complicated. Means for solving these problems are required.

Means for Solving the Problems

[0005] This invention provides a system comprising video acquisition means, image analysis means, difference measurement means, feedback generation means, and presentation means. When a user films their own training using a smartphone with the video acquisition means, the video data is analyzed by the image analysis means to identify joint positions. The difference measurement means then measures the difference between the trainer's ideal posture data and the user's posture. Based on the obtained difference, the feedback generation means generates instruction content in natural language and presents it to the user in real time using the presentation means, providing support similar to that of a professional trainer. Furthermore, by including a meal analysis means, the nutritional value is evaluated from images of meals taken by the user, assisting in health management. This makes it possible to train and manage health effectively and safely even at home.

[0006] "Video acquisition means" refers to a device or program that has the function of capturing the user's training posture in real time and acquiring it as video data.

[0007] "Image analysis means" refers to software or hardware that analyzes the posture of the human body from acquired video footage, identifies joint positions, and processes them as digital data.

[0008] A "difference measurement means" is a device or program that compares the analyzed user's posture data with the ideal trainer's posture data and measures the difference between the two.

[0009] A "feedback generation means" is a device or program that has the function of generating advice and guidance for the user in natural language format based on difference data obtained by a difference measurement means.

[0010] "Presentation means" refers to a device or function that provides the generated feedback information to the user visually or audibly.

[0011] A "meal analysis tool" is software or hardware that analyzes images of meals uploaded by users and evaluates the nutritional components based on their content. [Brief explanation of the drawing]

[0012] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14]It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when a sentiment engine is combined.

Embodiments for Carrying Out the Invention

[0013] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0014] First, the terms used in the following description will be explained.

[0015] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0016] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0017] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0018] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0019] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0020] [First Embodiment]

[0021] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0022] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0023] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0024] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0025] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0026] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0027] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0028] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0029] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0030] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0031] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0032] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0033] The fitness support system of the present invention allows users to perform efficient training at home without going to a gym, using a smartphone or tablet device. This system is configured as follows.

[0034] First, the user launches the application using their device. This application incorporates a video acquisition mechanism that allows the user's training session to be recorded in real time using the device's camera. The recorded video is then transmitted to a server via the network.

[0035] The server is equipped with image analysis capabilities to analyze the received video data. Specifically, it uses machine learning-based image recognition technology to identify the position of each joint in the user's body and convert it into posture data. At this stage, the user's movements during training are recorded as digital data.

[0036] Next, the server uses a difference measurement device to compare pre-registered ideal trainer posture data with the analyzed user posture data and measures the difference between the two. This difference data forms the basis for the feedback provided to the user.

[0037] The feedback generation mechanism generates natural language feedback based on the measured difference. For example, if the user is not raising their arm high enough, it will generate specific advice such as, "Raise your arm a little higher." The generated feedback is immediately provided to the user through the presentation mechanism. The presentation not only displays the feedback as text on the device screen but also plays it as audio.

[0038] Furthermore, the system includes a meal analysis function, which analyzes images of meals taken by the user on a server. This analysis uses image recognition technology to identify ingredients and evaluate their nutritional value. Based on this, the system provides the user with suggestions and advice for healthy eating.

[0039] In this way, this system enhances training efficiency and supports a healthy lifestyle by providing users with comprehensive and personalized fitness guidance and dietary management.

[0040] The following describes the processing flow.

[0041] Step 1:

[0042] The user launches the application on their device and prepares to start recording training videos. The user positions the camera appropriately and presses the record button to begin recording.

[0043] Step 2:

[0044] The device uses its camera to capture the user's training in real time and saves the video data to its internal storage. After recording begins, the video data is sent to the server frame by frame.

[0045] Step 3:

[0046] The server analyzes the received video data and uses image analysis tools to identify the joint positions and angles of the user's body for each frame. This generates the user's posture data.

[0047] Step 4:

[0048] The server compares the ideal trainer's posture data with the user's posture data using a difference measurement mechanism. This comparison yields specific difference data.

[0049] Step 5:

[0050] The server uses feedback generation methods based on differential data to generate natural language feedback to be provided to the user. The generated feedback is then put into text form and created as personalized advice tailored to the user.

[0051] Step 6:

[0052] Feedback is sent to the device and provided to the user in real time via a presentation method. The device displays it on the screen as a text message and also plays it as audio feedback through the speaker.

[0053] Step 7:

[0054] Users take photos of their meals using their devices and upload them to the server. By recording each meal, users can manage their eating habits.

[0055] Step 8:

[0056] The server analyzes uploaded meal images using a meal analysis tool to evaluate their nutritional components. The component information identified by image recognition technology is cross-referenced with a database and provided to the user as specific nutritional values ​​and advice necessary for health management.

[0057] By following these steps, the system provides training support and dietary management, offering users a real-time and comprehensive fitness experience.

[0058] (Example 1)

[0059] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0060] In modern times, providing individuals with effective exercise guidance and healthy dietary management is difficult for many. Traditional methods require significant time and expert guidance to obtain personalized feedback and dietary suggestions, posing a challenge to sustainable health management, especially for those leading busy lives.

[0061] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0062] In this invention, the server includes a shooting means for acquiring video data, an image analysis means for analyzing human body movements acquired by the shooting means and identifying joint positions, a comparison means for measuring the difference between registered ideal movement data and the analyzed movement data, and a generation means for generating advice using a generation AI model. This enables the provision of immediately customized exercise and dietary feedback to individuals, making it possible to maintain health efficiently.

[0063] "Video data" refers to digital information that visually captures a user's movements and activities.

[0064] "Filming means" refers to a device or function that films the user's actions in real time and acquires them as video data.

[0065] "Image analysis means" refers to a technology that analyzes the user's body movements and posture from acquired video data and identifies joint positions.

[0066] A "comparison method" is a process for comparing ideal exercise data with the user's exercise data and measuring the difference between them.

[0067] The "generation method" refers to a function that generates advice and feedback for the user based on the measured differential data.

[0068] "Presentation means" refers to a device or function that presents the generated advice to the user visually or audibly.

[0069] A "meal image" is digital information that visually captures the contents of a meal as photographed by the user.

[0070] "Methods for evaluating nutritional value" refer to technologies that identify the ingredients contained in a meal based on an image of the meal and evaluate its nutritional value.

[0071] This fitness support system enables users to efficiently manage their training and diet using their devices. Users can easily access this system using smartphones or tablets. The server at the core of this system is responsible for processing and analyzing the received data.

[0072] When the application is launched by the user, the device first uses its built-in camera to capture real-time video data of the user's training. The captured video data is then transmitted to a server via the network. During this process, the video may be compressed for more efficient transfer.

[0073] The server performs image analysis on the received video data. Specifically, it uses a machine learning platform such as TENSORFLOW® to perform image recognition to identify the position of each joint in the user's body. This technology allows the user's movements to be digitally recorded and stored as posture data.

[0074] Next, the server compares ideal exercise data with the user's analyzed exercise data. The difference measured by the comparison device forms the basis for identifying areas for improvement in the user's training.

[0075] Furthermore, the server uses a generative AI model to generate feedback for the user based on the measured differential data. The generated feedback is sent to the user's device and displayed visually as text on the screen, as well as played as audio. For example, if the user's arm is held low, advice such as "Please raise your arm a little higher" is provided.

[0076] Users can also take photos of their meals with their smartphones and send the images to the server. The server analyzes these meal images and evaluates their nutritional value. Based on the evaluation results, the user is provided with suggestions and advice for healthy eating. For example, if it is determined that there is a vitamin deficiency, advice such as "Add some fruit" will be generated.

[0077] An example of a prompt for the generating AI model would be: "Analyze the user's training video to obtain posture data and tell me the difference from the ideal posture. Also, generate healthy meal suggestions based on the food images."

[0078] In this way, the system functions as a comprehensive tool to support the user's health, providing personalized fitness guidance and dietary management.

[0079] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0080] Step 1:

[0081] The user launches a fitness application using the device and activates the camera. The input for this step is the user's workout scene, and the output is the video data that visually captures it. The device temporarily records this video data and prepares to send it to the server when ready.

[0082] Step 2:

[0083] The terminal transmits recorded video data to the server via the network. The input is temporarily recorded compressed video data, and the output is the video data transmitted to the server. The terminal compresses the data as needed to improve the efficiency of data transfer.

[0084] Step 3:

[0085] The server analyzes the received video data. The input is the transmitted video data, and the output is posture data including the joint positions of the user's body. The server uses machine learning tools such as TensorFlow for this analysis to automatically identify the user's joints from the video.

[0086] Step 4:

[0087] The server compares the analyzed posture data with existing ideal movement data. The input is the analyzed user posture data and the ideal posture data, and the output is the difference between the two. The server quantifies this difference and identifies areas that need improvement.

[0088] Step 5:

[0089] The server generates feedback using a generative AI model based on differential data. The input is differential data, and the output is specific feedback sentences for the user. The generated feedback provides the user with suggestions for improvement and advice regarding their movement in natural language.

[0090] Step 6:

[0091] The terminal presents the user with feedback received from the server. The input is the generated feedback data, and the output is the feedback text and synthesized speech feedback displayed on the terminal screen. The terminal displays this quickly so that the user can easily review it.

[0092] Step 7:

[0093] Users photograph their daily meals and send these images from their device to the server. The input is image data of the meal, and the output is an image file that can be processed by the server. The device transmits images quickly through a stable connection.

[0094] Step 8:

[0095] The server analyzes received meal images and uses image recognition technology to identify ingredients. The input is the submitted meal image, and the output is a list of ingredients and their nutritional information. Based on this, the server generates health-conscious meal suggestions for the user.

[0096] (Application Example 1)

[0097] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0098] Improving work efficiency within a factory requires optimizing worker movements. However, conventional methods have made it difficult to obtain detailed, real-time motion analysis and accurate feedback. A technology is needed to solve this problem.

[0099] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0100] In this invention, the server includes a video acquisition device, a pixel analysis device that analyzes the movement of a workpiece and identifies joint positions, and a difference measurement device that measures differences. This makes it possible to analyze the worker's movements in real time and provide optimal feedback.

[0101] A "video acquisition device" is a device used to film the movements of a work object and acquire them as video data.

[0102] A "pixel analysis device" is a device that analyzes acquired video data to identify the movement and joint positions of a workpiece.

[0103] A "difference measurement device" is a device that measures the difference between ideal operating data and analyzed actual operating data.

[0104] A "feedback generation device" is a device that generates feedback in natural language to a work object based on the difference in measured motion data.

[0105] A "display device" is a device that presents generated feedback to a work object visually or audibly.

[0106] The system implementing this invention supports workers in efficiently performing their tasks within a factory. A specific embodiment is shown below.

[0107] First, the worker puts on smart glasses. These smart glasses are equipped with cameras to record their movements. The smart glasses, acting as a video acquisition device, record the worker's movements in real time and transmit the acquired video data to a server. The server processes the video data using high-performance image analysis software. Specifically, it uses the OpenCV library to identify joint positions from the video data.

[0108] Next, the server uses TensorFlow, a deep learning framework, to analyze joint positions. Based on the analyzed motion data, it functions as a difference measurement device to measure the difference between the current motion data and ideal motion data. Based on the measured difference, the server acts as a feedback generator, producing specific improvement suggestions for the worker in natural language. This can utilize the GPT natural language processing model.

[0109] The generated feedback is sent to smart glasses, which act as a display device. The smart glasses not only display the feedback as text but also output it as audio using speech synthesis technology. Google® Text-to-Speech API can be used for speech synthesis.

[0110] As a concrete example, workers in a packaging factory may use smart glasses, and the analysis system may detect inefficient movements during their work. For instance, if an arm movement is inefficient, real-time feedback such as, "Moving your hand 30cm to the left from that position will improve efficiency," can be provided, enabling improvements to the work process. By using this invention, both work efficiency and safety can be improved simultaneously.

[0111] Examples of prompt statements to input into the generative AI model are as follows:

[0112] "Generate feedback based on motion analysis of factory workers. If the current motion is not optimal, specify areas for improvement."

[0113] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0114] Step 1:

[0115] The user (worker) puts on smart glasses and begins work. The smart glasses' camera captures their actions in real time during the work.

[0116] The input is video footage of the worker's movements, and the output is video data. The video data is sent to a server for subsequent analysis.

[0117] Step 2:

[0118] The server processes the received video data using the OpenCV library to identify the joint positions of the worker in each frame.

[0119] The input is the video data acquired in step 1, and the output is the joint coordinate data. This data is used for motion analysis.

[0120] Step 3:

[0121] The server uses TensorFlow to analyze joint coordinate data and generate actual motion data.

[0122] The input is the joint coordinate data obtained in step 2, and the output is detailed motion data. This analysis allows us to understand the overall movements of the worker during the task.

[0123] Step 4:

[0124] The server compares the data to ideal operating data and calculates the difference, acting as a difference measurement device.

[0125] The input consists of the actual operation data obtained in step 3 and pre-registered ideal operation data. The output is the difference data of the operation. This difference is used for feedback generation.

[0126] Step 5:

[0127] The server generates feedback from the differential data using a natural language processing model.

[0128] The input is the difference data obtained in step 4, and the output is text data of specific feedback. The feedback includes specific advice for improving the work.

[0129] Step 6:

[0130] The server converts the generated feedback into audio data using speech synthesis technology and sends it to the smart glasses.

[0131] The input is the text data of the feedback generated in step 5, and the output is audio data. The smart glasses enable real-time instruction by presenting this audio data to the worker.

[0132] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0133] The fitness support system of the present invention employs groundbreaking technology that combines an emotional engine to maximize the user's training effectiveness. This system has the following components:

[0134] The user launches the application on their device and prepares for training. The device's camera is used to record the user's training in real time using a video acquisition system. The acquired video data is transmitted to a server via the internet.

[0135] The server is equipped with image analysis capabilities to identify the position of each joint in the user's body. This information is compared with pre-stored ideal trainer posture data by a difference measurement system to generate difference data for measuring the effectiveness of the training.

[0136] The feedback generation mechanism uses this differential data to create natural language feedback tailored to the user. A notable feature here is the use of an emotion engine. The emotion engine analyzes the user's emotional state from their facial expressions and voice, and adjusts the feedback accordingly. For example, if the emotion engine detects that the user is tired, it generates an encouraging message such as, "Let's keep going a little longer, or you can take a short break." This feedback is quickly sent to the device and presented both on the screen and audibly.

[0137] Furthermore, a meal analysis system analyzes images of meals taken by users on a server. When combined with an emotion engine, users in specific emotional states are provided with motivational advice regarding their meal choices. For example, if a stressed state is detected, advice such as "It would be good to increase the amount of foods that have a relaxing effect" may be given.

[0138] Thus, this system comprehensively supports users' physical training and emotional well-being, providing integrated fitness and nutritional management. This maximizes user motivation and results, enabling sustainable health maintenance.

[0139] The following describes the processing flow.

[0140] Step 1:

[0141] The user launches the fitness app installed on their device and starts a training session. The user positions the device's camera appropriately and starts recording by pressing the record button.

[0142] Step 2:

[0143] The device uses its camera to capture real-time video of the user during training and generates video data. This data is then transmitted to a server via the network.

[0144] Step 3:

[0145] The server processes the received video data using image analysis tools to identify the position of each joint in the user's posture. Based on these analysis results, posture data is obtained.

[0146] Step 4:

[0147] The server uses a difference measurement mechanism to compare the analyzed user's posture data with pre-registered ideal trainer posture data. This process generates specific difference data.

[0148] Step 5:

[0149] The emotion engine analyzes the user's facial expressions and voice to identify their current emotional state. This emotional state data is then used to generate subsequent feedback.

[0150] Step 6:

[0151] The server's feedback generation mechanism combines differential data and emotional states to generate natural language feedback tailored to each user's situation. The content and tone of the feedback are adjusted based on the emotional state.

[0152] Step 7:

[0153] The generated feedback information is sent to the terminal. The presentation device displays this information as text on the screen and also plays it as audio, providing it to the user in real time.

[0154] Step 8:

[0155] The user takes a picture of their meal with their device and uploads it to the server. The meal image data is used as data for subsequent nutritional analysis.

[0156] Step 9:

[0157] The server utilizes meal analysis tools to analyze uploaded meal images and identify their nutritional value. The analysis results are provided as nutritional advice based on the user's emotional state, supporting the maintenance of their health.

[0158] Through the steps outlined above, this system aims to integrate and manage the user's physical training and emotional support, ultimately striving for sustainable fitness improvement.

[0159] (Example 2)

[0160] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0161] In fitness activities, for users to achieve ideal training results, comprehensive support is needed that considers not only their physical posture but also their emotional state and dietary choices. However, current systems are unable to manage these elements in an integrated manner, limiting their ability to maintain and improve user motivation and health benefits.

[0162] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0163] In this invention, the server includes video acquisition means, image analysis means, difference measurement means, feedback generation means, emotion analysis means, and diet analysis means. This enables comprehensive analysis of the user's physical training data, emotional state, and information regarding dietary choices, and provides personalized feedback and advice.

[0164] A "video acquisition device" is a device that captures the user's physical movements in real time and acquires the video data.

[0165] "Image analysis means" refers to an analysis device used to identify the position of each joint in the user's body from acquired video data.

[0166] A "difference measurement device" is a device that measures the difference between the analyzed user's posture data and pre-registered ideal posture data, and generates difference data.

[0167] A "feedback generation means" is a device that generates natural language feedback for the user based on the difference data obtained by the difference measurement means.

[0168] An "emotion analysis device" is a device that analyzes the user's emotional state from their facial expressions and voice, and adjusts the generated feedback content according to that emotion.

[0169] A "meal analysis device" is a device that analyzes images of meals taken by users and evaluates their nutritional value.

[0170] A "presentation means" is a device that presents generated feedback or advice to the user, and performs functions such as display or audio output.

[0171] The fitness support system of the present invention is configured as a multi-functional platform including emotion analysis and dietary analysis to optimize the user's training effectiveness. The user launches a dedicated application using a terminal and acquires training video through the terminal's camera. This video data is transmitted to a server via the internet.

[0172] The server uses the received video data to identify the joint positions of the user's body through image analysis. Simultaneously with this analysis, a difference measurement means measures the difference between the user's posture data and pre-stored ideal posture data, generating difference data. Subsequently, a feedback generation means generates natural language feedback tailored to the user's training progress. During this process, an emotion analysis means analyzes the user's facial expressions and voice to determine their emotional state and adjusts the feedback content accordingly.

[0173] Furthermore, this system includes a meal analysis function that analyzes images of meals taken by the user on a server. It evaluates the nutritional value from the analysis results and generates advice on meal choices based on the user's emotional state.

[0174] For example, when a user performs squats at home, the device's camera records the movement. This video is sent to a server, which analyzes the user's knee and hip positioning and generates feedback such as, "It would be good to point your knees slightly outward." In another example, a picture of a salad taken by a user at their dinner table is analyzed by the server, and advice such as, "Adding avocado is good for stress reduction," is provided.

[0175] An example of a prompt using a generative AI model is: "Generate a suitable feedback message when the user is tired. Example: 'You're doing great. It's okay to take a break.'" This allows the user to receive comprehensive support from both a physical and emotional perspective.

[0176] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0177] Step 1:

[0178] The user launches a fitness application on their device and prepares to begin training using the camera. Inputs include the user's selected training mode and the press of a start button. Outputs include the activation of the camera device and the display of the training mode on the user's screen. This completes the process of preparing to begin training.

[0179] Step 2:

[0180] The device's camera captures the user's training movements in real time, acquiring video data. The input is the user's physical movements. This video data is compressed and sent to the server. The output is the captured video data, which is also sent to the server in real time. This allows the necessary data to be collected for analysis.

[0181] Step 3:

[0182] The server processes the received video data using image analysis to identify the position of each joint in the user's body. The input is the transmitted video data. The output is data on the joint positions. Through data analysis, the user's posture can be precisely understood.

[0183] Step 4:

[0184] The server measures the difference between identified joint position data and ideal posture data using a difference measurement device. The inputs are joint position data and ideal posture data. The output is the difference data between these two sets of data. This process allows for a quantitative evaluation of the user's training effectiveness.

[0185] Step 5:

[0186] The server uses a feedback generation mechanism based on differential data to create natural language feedback. The input is differential data. An emotion analysis mechanism identifies the user's emotional state and adjusts the feedback content accordingly. The output is an emotion-sensitive feedback message. This feedback enhances the user's training motivation.

[0187] Step 6:

[0188] The server sends the generated feedback to the terminal, which then presents it to the user. The input is the generated feedback message. Output from the terminal includes screen displays and audio feedback. This allows the user to receive feedback in real time and modify their training.

[0189] Step 7:

[0190] The user takes a picture of their meal with their device, and the device sends this data to the server. The input is the image of the meal taken by the user. The output is the captured image data sent to the server. Through this process, data for meal analysis is collected.

[0191] Step 8:

[0192] The server analyzes the received images using a meal analysis system and evaluates their nutritional value. The input is the transmitted meal image data. The output is evaluation data regarding the nutritional value of the meal. This allows the user to obtain detailed information about the meal they have consumed.

[0193] Step 9:

[0194] The server generates advice on food choices based on the results of a meal analysis and the user's emotional state. The inputs are nutritional value evaluation data for meals and the user's emotional state. The output is specific meal selection advice. This allows users to strive for a diet that considers both their health and emotional state.

[0195] Step 10:

[0196] The generated dietary advice is sent to the device and presented to the user. The input is the advice data. The output is the advice displayed to the user from the device. This process allows the user to adjust their diet and make healthier choices.

[0197] (Application Example 2)

[0198] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0199] Traditional fitness support systems have struggled to provide appropriate feedback tailored to each user's individual emotions and fitness level. Furthermore, the lack of features that integrate training and nutrition information and support online marketplace purchases hindered users' ability to maintain sustainable health. There is a need to address these issues and support users in engaging in fitness activities more effectively.

[0200] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0201] In this invention, the server includes a video acquisition means, an image analysis means, a difference measurement means, an emotion analysis means, a means for analyzing the user's emotional state using the emotion analysis means and adjusting the feedback content in the feedback generation means, and an online connection means for providing purchase support in the virtual market. This enables the provision of appropriate feedback according to the user's emotional state and support for purchasing fitness-related products through the online market.

[0202] A "video acquisition means" is a device that acquires video data in order to capture the user's body movements in real time.

[0203] "Image analysis means" refers to a technology that includes an algorithm for analyzing acquired video data and identifying the positions of human joints.

[0204] The "difference measurement method" is a function for evaluating training effectiveness by measuring the difference between the analyzed posture data and the ideal posture data that has been registered in advance.

[0205] A "feedback generation method" is a means of analyzing difference data using natural language processing technology in order to generate feedback to be provided to the user.

[0206] A "presentation means" is a device for providing the generated feedback to the user visually or audibly.

[0207] "Emotional analysis methods" refer to technologies that analyze a user's emotional state from their facial expressions and voice.

[0208] "Online connectivity" refers to internet-based communication functions that allow users to connect to a virtual marketplace and assist in purchasing fitness-related products.

[0209] In this invention, the system operates by integrating various means to support the user's fitness activities. The server captures the user's body movements in real time via "video acquisition means" for acquiring video. This data is analyzed by "image analysis means" to determine the position of the user's joints. The analyzed data is compared with ideal trainer posture data by "difference measurement means" to generate difference data.

[0210] Based on this differential data, the "feedback generation means" utilizes a generation AI model to create feedback tailored to the user. This feedback is then adjusted based on the user's emotional state, which is evaluated through the "emotion analysis means." The generated feedback is then quickly provided to the user as audio or text via the "presentation means" on the user's terminal.

[0211] Furthermore, the system incorporates "online connectivity," allowing users to easily purchase necessary training products through a virtual marketplace. This creates an environment where users can seamlessly integrate fitness activities with product purchases.

[0212] For example, when a user wears smart glasses and does push-ups, the camera captures the position of their hands and analyzes that information. Then, appropriate feedback is provided, such as "Try spreading your hands a little wider." When the user is tired, emotionally appropriate feedback is provided, such as "Let's slow down."

[0213] An example of a prompt message would be: "Based on the user's training data, generate feedback for posture improvement. Next, analyze his emotional data and adjust the feedback to match his emotions."

[0214] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0215] Step 1:

[0216] The user begins training using the camera built into the device. The device captures the user's movements in real time using video acquisition equipment and sends the video data to the server. The input is the user's video captured by the camera, and the output is the video data sent to the server.

[0217] Step 2:

[0218] The server analyzes the received video data using image analysis tools to identify the user's joint positions. The input here is the video data obtained in step 1, and the output is the identified joint position data. Specifically, the server analyzes each frame using a machine learning model and maps each part of the body.

[0219] Step 3:

[0220] The server uses a difference measurement mechanism to measure the difference between the analyzed joint position data and pre-registered ideal posture data. The input is the joint position data and ideal posture data obtained in step 2, and the output is the difference data. Specifically, a particular index is set, and the difference is quantified using linear algebra.

[0221] Step 4:

[0222] The server uses a generative AI model to generate natural language feedback based on the differential data. The input here is the differential data obtained in step 3, and the output is feedback in natural language. Specifically, prompt sentences are input to the generative AI model, which then generates the optimal feedback sentence.

[0223] Step 5:

[0224] The server further uses emotion analysis tools to determine the user's emotional state from their facial expressions and voice data. The input is the user's facial expressions and voice data, and the output is the user's emotional state data. Specifically, it uses speech recognition technology and facial expression recognition technology such as DeepFace to detect and analyze the state.

[0225] Step 6:

[0226] Based on the user's emotional state, the generated feedback is adjusted and presented to the user through the device's display methods. The input is the feedback obtained in step 4 and the emotional state data obtained in step 5, and the output is the adjusted feedback. Specifically, the tone and content of the feedback are fine-tuned to match the user's emotions and presented as text display or audio output.

[0227] Step 7:

[0228] Using the terminal's online connectivity, users can access a virtual marketplace and purchase fitness-related products based on their feedback. Inputs are user feedback and virtual marketplace access rights, while output is a list of purchasable products. Specifically, the terminal opens an internet browser and provides an interface for the user to select suggested products and proceed with the purchase.

[0229] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0230] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0231] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0232] [Second Embodiment]

[0233] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0234] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0235] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0236] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0237] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0238] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0239] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0240] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0241] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0242] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0243] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0244] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0245] The fitness support system of the present invention allows users to perform efficient training at home without going to a gym, using a smartphone or tablet device. This system is configured as follows.

[0246] First, the user launches the application using their device. This application incorporates a video acquisition mechanism that allows the user's training session to be recorded in real time using the device's camera. The recorded video is then transmitted to a server via the network.

[0247] The server is equipped with image analysis capabilities to analyze the received video data. Specifically, it uses machine learning-based image recognition technology to identify the position of each joint in the user's body and convert it into posture data. At this stage, the user's movements during training are recorded as digital data.

[0248] Next, the server uses a difference measurement device to compare pre-registered ideal trainer posture data with the analyzed user posture data and measures the difference between the two. This difference data forms the basis for the feedback provided to the user.

[0249] The feedback generation mechanism generates natural language feedback based on the measured difference. For example, if the user is not raising their arm high enough, it will generate specific advice such as, "Raise your arm a little higher." The generated feedback is immediately provided to the user through the presentation mechanism. The presentation not only displays the feedback as text on the device screen but also plays it as audio.

[0250] Furthermore, the system includes a meal analysis function, which analyzes images of meals taken by the user on a server. This analysis uses image recognition technology to identify ingredients and evaluate their nutritional value. Based on this, the system provides the user with suggestions and advice for healthy eating.

[0251] In this way, this system enhances training efficiency and supports a healthy lifestyle by providing users with comprehensive and personalized fitness guidance and dietary management.

[0252] The following describes the processing flow.

[0253] Step 1:

[0254] The user launches the application on their device and prepares to start recording training videos. The user positions the camera appropriately and presses the record button to begin recording.

[0255] Step 2:

[0256] The device uses its camera to capture the user's training in real time and saves the video data to its internal storage. After recording begins, the video data is sent to the server frame by frame.

[0257] Step 3:

[0258] The server analyzes the received video data and uses image analysis tools to identify the joint positions and angles of the user's body for each frame. This generates the user's posture data.

[0259] Step 4:

[0260] The server compares the ideal trainer's posture data with the user's posture data using a difference measurement mechanism. This comparison yields specific difference data.

[0261] Step 5:

[0262] The server uses feedback generation methods based on differential data to generate natural language feedback to be provided to the user. The generated feedback is then put into text form and created as personalized advice tailored to the user.

[0263] Step 6:

[0264] Feedback is sent to the device and provided to the user in real time via a presentation method. The device displays it on the screen as a text message and also plays it as audio feedback through the speaker.

[0265] Step 7:

[0266] Users take photos of their meals using their devices and upload them to the server. By recording each meal, users can manage their eating habits.

[0267] Step 8:

[0268] The server analyzes uploaded meal images using a meal analysis tool to evaluate their nutritional components. The component information identified by image recognition technology is cross-referenced with a database and provided to the user as specific nutritional values ​​and advice necessary for health management.

[0269] By following these steps, the system provides training support and dietary management, offering users a real-time and comprehensive fitness experience.

[0270] (Example 1)

[0271] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0272] In modern times, providing individuals with effective exercise guidance and healthy dietary management is difficult for many. Traditional methods require significant time and expert guidance to obtain personalized feedback and dietary suggestions, posing a challenge to sustainable health management, especially for those leading busy lives.

[0273] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0274] In this invention, the server includes a shooting means for acquiring video data, an image analysis means for analyzing human body movements acquired by the shooting means and identifying joint positions, a comparison means for measuring the difference between registered ideal movement data and the analyzed movement data, and a generation means for generating advice using a generation AI model. This enables the provision of immediately customized exercise and dietary feedback to individuals, making it possible to maintain health efficiently.

[0275] "Video data" refers to digital information that visually captures a user's movements and activities.

[0276] "Filming means" refers to a device or function that films the user's actions in real time and acquires them as video data.

[0277] "Image analysis means" refers to a technology that analyzes the user's body movements and posture from acquired video data and identifies joint positions.

[0278] A "comparison method" is a process for comparing ideal exercise data with the user's exercise data and measuring the difference between them.

[0279] The "generation method" refers to a function that generates advice and feedback for the user based on the measured differential data.

[0280] "Presentation means" refers to a device or function that presents the generated advice to the user visually or audibly.

[0281] A "meal image" is digital information that visually captures the contents of a meal as photographed by the user.

[0282] "Methods for evaluating nutritional value" refer to technologies that identify the ingredients contained in a meal based on an image of the meal and evaluate its nutritional value.

[0283] This fitness support system enables users to efficiently manage their training and diet using their devices. Users can easily access this system using smartphones or tablets. The server at the core of this system is responsible for processing and analyzing the received data.

[0284] When the application is launched by the user, the terminal first uses the built-in camera to obtain the user's training scenery in real time as video data. The captured video data is sent to the server through the network. In this process, the video may be compressed for efficient transfer.

[0285] The server performs image analysis on the received video data. Specifically, using a machine learning platform such as TensorFlow, image recognition is carried out to identify the positions of each joint of the user's body. By this technology, the user's movement is digitally recorded and saved as pose data.

[0286] Next, the server compares the ideal motion data with the analyzed motion data of the user. The difference measured by the comparison means serves as a basis for identifying areas for improvement in the user's training.

[0287] Furthermore, the server uses a generative AI model to generate feedback for the user based on the measured difference data. The generated feedback is sent to the user's terminal and is visually displayed as text on the screen and also played as audio. As a specific example, when the position of the user's arm is low, advice such as "Please raise your arm a little higher" is provided.

[0288] Also, the user can take a photo of a meal with a smartphone and send the image to the server. The server analyzes this meal image and evaluates the nutritional value. Based on the evaluation results, suggestions and advice on a healthy diet are provided to the user. For example, when it is determined that there is a vitamin deficiency, advice such as "Let's add some fruits" is generated.

[0289] An example of a prompt sentence for the generative AI model is "Analyze the user's training video to obtain pose data and tell me the differences from the ideal pose. Also, generate suggestions for a healthy diet based on the meal image."

[0290] In this way, the system functions as a comprehensive tool to support the user's health, providing personalized fitness guidance and dietary management.

[0291] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0292] Step 1:

[0293] The user launches a fitness application using the device and activates the camera. The input for this step is the user's workout scene, and the output is the video data that visually captures it. The device temporarily records this video data and prepares to send it to the server when ready.

[0294] Step 2:

[0295] The terminal transmits recorded video data to the server via the network. The input is temporarily recorded compressed video data, and the output is the video data transmitted to the server. The terminal compresses the data as needed to improve the efficiency of data transfer.

[0296] Step 3:

[0297] The server analyzes the received video data. The input is the transmitted video data, and the output is posture data including the joint positions of the user's body. The server uses machine learning tools such as TensorFlow for this analysis to automatically identify the user's joints from the video.

[0298] Step 4:

[0299] The server compares the analyzed posture data with existing ideal movement data. The input is the analyzed user posture data and the ideal posture data, and the output is the difference between the two. The server quantifies this difference and identifies areas that need improvement.

[0300] Step 5:

[0301] Based on the differential data, the server uses the generative AI model to generate feedback. The input is the differential data, and the output is the specific feedback text for the user. The generated feedback provides improvements and advice on the user's movement in natural language.

[0302] Step 6:

[0303] The terminal presents the feedback received from the server to the user. The input is the generated feedback data, and the output is the feedback text displayed on the terminal screen and the voice feedback by voice synthesis. The terminal quickly displays this so that the user can easily confirm it.

[0304] Step 7:

[0305] The user takes a photo of their daily meal and sends the meal image from the terminal to the server. The input is the image data of the taken meal state, and the output is the image file that can be processed by the server. The terminal quickly sends the image through a stable connection.

[0306] Step 8:

[0307] The server analyzes the received meal image and uses image recognition technology to identify the ingredients. The input is the sent meal image, and the output is a list of ingredients and their nutritional value information. Based on this, the server generates a meal proposal considering the user's health for the user.

[0308] (Application Example 1)

[0309] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".

[0310] Improving work efficiency within a factory requires optimizing worker movements. However, conventional methods have made it difficult to obtain detailed, real-time motion analysis and accurate feedback. A technology is needed to solve this problem.

[0311] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0312] In this invention, the server includes a video acquisition device, a pixel analysis device that analyzes the movement of a workpiece and identifies joint positions, and a difference measurement device that measures differences. This makes it possible to analyze the worker's movements in real time and provide optimal feedback.

[0313] A "video acquisition device" is a device used to film the movements of a work object and acquire them as video data.

[0314] A "pixel analysis device" is a device that analyzes acquired video data to identify the movement and joint positions of a workpiece.

[0315] A "difference measurement device" is a device that measures the difference between ideal operating data and analyzed actual operating data.

[0316] A "feedback generation device" is a device that generates feedback in natural language to a work object based on the difference in measured motion data.

[0317] A "display device" is a device that presents generated feedback to a work object visually or audibly.

[0318] The system implementing this invention supports workers in efficiently performing their tasks within a factory. A specific embodiment is shown below.

[0319] First, the worker puts on smart glasses. These smart glasses are equipped with cameras to record their movements. The smart glasses, acting as a video acquisition device, record the worker's movements in real time and transmit the acquired video data to a server. The server processes the video data using high-performance image analysis software. Specifically, it uses the OpenCV library to identify joint positions from the video data.

[0320] Next, the server uses TensorFlow, a deep learning framework, to analyze joint positions. Based on the analyzed motion data, it functions as a difference measurement device to measure the difference between the current motion data and ideal motion data. Based on the measured difference, the server acts as a feedback generator, producing specific improvement suggestions for the worker in natural language. This can utilize the GPT natural language processing model.

[0321] The generated feedback is sent to smart glasses, which act as a display device. The smart glasses not only display the feedback as text but also output it as audio using speech synthesis technology. The Google Text-to-Speech API can be used for speech synthesis.

[0322] As a concrete example, workers in a packaging factory may use smart glasses, and the analysis system may detect inefficient movements during their work. For instance, if an arm movement is inefficient, real-time feedback such as, "Moving your hand 30cm to the left from that position will improve efficiency," can be provided, enabling improvements to the work process. By using this invention, both work efficiency and safety can be improved simultaneously.

[0323] Examples of prompt statements to input into the generative AI model are as follows:

[0324] "Generate feedback based on motion analysis of factory workers. If the current motion is not optimal, specify areas for improvement."

[0325] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0326] Step 1:

[0327] The user (worker) puts on smart glasses and begins work. The smart glasses' camera captures their actions in real time during the work.

[0328] The input is video footage of the worker's movements, and the output is video data. The video data is sent to a server for subsequent analysis.

[0329] Step 2:

[0330] The server processes the received video data using the OpenCV library to identify the joint positions of the worker in each frame.

[0331] The input is the video data acquired in step 1, and the output is the joint coordinate data. This data is used for motion analysis.

[0332] Step 3:

[0333] The server uses TensorFlow to analyze joint coordinate data and generate actual motion data.

[0334] The input is the joint coordinate data obtained in step 2, and the output is detailed motion data. This analysis allows us to understand the overall movements of the worker during the task.

[0335] Step 4:

[0336] The server compares the data to ideal operating data and calculates the difference, acting as a difference measurement device.

[0337] The input consists of the actual operation data obtained in step 3 and pre-registered ideal operation data. The output is the difference data of the operation. This difference is used for feedback generation.

[0338] Step 5:

[0339] The server generates feedback from the differential data using a natural language processing model.

[0340] The input is the difference data obtained in step 4, and the output is text data of specific feedback. The feedback includes specific advice for improving the work.

[0341] Step 6:

[0342] The server converts the generated feedback into audio data using speech synthesis technology and sends it to the smart glasses.

[0343] The input is the text data of the feedback generated in step 5, and the output is audio data. The smart glasses enable real-time instruction by presenting this audio data to the worker.

[0344] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0345] The fitness support system of the present invention employs groundbreaking technology that combines an emotional engine to maximize the user's training effectiveness. This system has the following components:

[0346] The user launches the application on their device and prepares for training. The device's camera is used to record the user's training in real time using a video acquisition system. The acquired video data is transmitted to a server via the internet.

[0347] The server is equipped with image analysis capabilities to identify the position of each joint in the user's body. This information is compared with pre-stored ideal trainer posture data by a difference measurement system to generate difference data for measuring the effectiveness of the training.

[0348] The feedback generation mechanism uses this differential data to create natural language feedback tailored to the user. A notable feature here is the use of an emotion engine. The emotion engine analyzes the user's emotional state from their facial expressions and voice, and adjusts the feedback accordingly. For example, if the emotion engine detects that the user is tired, it generates an encouraging message such as, "Let's keep going a little longer, or you can take a short break." This feedback is quickly sent to the device and presented both on the screen and audibly.

[0349] Furthermore, a meal analysis system analyzes images of meals taken by users on a server. When combined with an emotion engine, users in specific emotional states are provided with motivational advice regarding their meal choices. For example, if a stressed state is detected, advice such as "It would be good to increase the amount of foods that have a relaxing effect" may be given.

[0350] Thus, this system comprehensively supports users' physical training and emotional well-being, providing integrated fitness and nutritional management. This maximizes user motivation and results, enabling sustainable health maintenance.

[0351] The following describes the processing flow.

[0352] Step 1:

[0353] The user launches the fitness app installed on their device and starts a training session. The user positions the device's camera appropriately and starts recording by pressing the record button.

[0354] Step 2:

[0355] The device uses its camera to capture real-time video of the user during training and generates video data. This data is then transmitted to a server via the network.

[0356] Step 3:

[0357] The server processes the received video data using image analysis tools to identify the position of each joint in the user's posture. Based on these analysis results, posture data is obtained.

[0358] Step 4:

[0359] The server uses a difference measurement mechanism to compare the analyzed user's posture data with pre-registered ideal trainer posture data. This process generates specific difference data.

[0360] Step 5:

[0361] The emotion engine analyzes the user's facial expressions and voice to identify their current emotional state. This emotional state data is then used to generate subsequent feedback.

[0362] Step 6:

[0363] The server's feedback generation mechanism combines differential data and emotional states to generate natural language feedback tailored to each user's situation. The content and tone of the feedback are adjusted based on the emotional state.

[0364] Step 7:

[0365] The generated feedback information is sent to the terminal. The presentation device displays this information as text on the screen and also plays it as audio, providing it to the user in real time.

[0366] Step 8:

[0367] The user takes a picture of their meal with their device and uploads it to the server. The meal image data is used as data for subsequent nutritional analysis.

[0368] Step 9:

[0369] The server utilizes meal analysis tools to analyze uploaded meal images and identify their nutritional value. The analysis results are provided as nutritional advice based on the user's emotional state, supporting the maintenance of their health.

[0370] Through the steps outlined above, this system aims to integrate and manage the user's physical training and emotional support, ultimately striving for sustainable fitness improvement.

[0371] (Example 2)

[0372] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0373] In fitness activities, for users to achieve ideal training results, comprehensive support is needed that considers not only their physical posture but also their emotional state and dietary choices. However, current systems are unable to manage these elements in an integrated manner, limiting their ability to maintain and improve user motivation and health benefits.

[0374] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0375] In this invention, the server includes video acquisition means, image analysis means, difference measurement means, feedback generation means, emotion analysis means, and diet analysis means. This enables comprehensive analysis of the user's physical training data, emotional state, and information regarding dietary choices, and provides personalized feedback and advice.

[0376] A "video acquisition device" is a device that captures the user's physical movements in real time and acquires the video data.

[0377] "Image analysis means" refers to an analysis device used to identify the position of each joint in the user's body from acquired video data.

[0378] A "difference measurement device" is a device that measures the difference between the analyzed user's posture data and pre-registered ideal posture data, and generates difference data.

[0379] A "feedback generation means" is a device that generates natural language feedback for the user based on the difference data obtained by the difference measurement means.

[0380] An "emotion analysis device" is a device that analyzes the user's emotional state from their facial expressions and voice, and adjusts the generated feedback content according to that emotion.

[0381] A "meal analysis device" is a device that analyzes images of meals taken by users and evaluates their nutritional value.

[0382] A "presentation means" is a device that presents generated feedback or advice to the user, and performs functions such as display or audio output.

[0383] The fitness support system of the present invention is configured as a multi-functional platform including emotion analysis and dietary analysis to optimize the user's training effectiveness. The user launches a dedicated application using a terminal and acquires training video through the terminal's camera. This video data is transmitted to a server via the internet.

[0384] The server uses the received video data to identify the joint positions of the user's body through image analysis. Simultaneously with this analysis, a difference measurement means measures the difference between the user's posture data and pre-stored ideal posture data, generating difference data. Subsequently, a feedback generation means generates natural language feedback tailored to the user's training progress. During this process, an emotion analysis means analyzes the user's facial expressions and voice to determine their emotional state and adjusts the feedback content accordingly.

[0385] Furthermore, this system includes a meal analysis function that analyzes images of meals taken by the user on a server. It evaluates the nutritional value from the analysis results and generates advice on meal choices based on the user's emotional state.

[0386] For example, when a user performs squats at home, the device's camera records the movement. This video is sent to a server, which analyzes the user's knee and hip positioning and generates feedback such as, "It would be good to point your knees slightly outward." In another example, a picture of a salad taken by a user at their dinner table is analyzed by the server, and advice such as, "Adding avocado is good for stress reduction," is provided.

[0387] An example of a prompt using a generative AI model is: "Generate a suitable feedback message when the user is tired. Example: 'You're doing great. It's okay to take a break.'" This allows the user to receive comprehensive support from both a physical and emotional perspective.

[0388] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0389] Step 1:

[0390] The user launches a fitness application on their device and prepares to begin training using the camera. Inputs include the user's selected training mode and the press of a start button. Outputs include the activation of the camera device and the display of the training mode on the user's screen. This completes the process of preparing to begin training.

[0391] Step 2:

[0392] The device's camera captures the user's training movements in real time, acquiring video data. The input is the user's physical movements. This video data is compressed and sent to the server. The output is the captured video data, which is also sent to the server in real time. This allows the necessary data to be collected for analysis.

[0393] Step 3:

[0394] The server processes the received video data using image analysis to identify the position of each joint in the user's body. The input is the transmitted video data. The output is data on the joint positions. Through data analysis, the user's posture can be precisely understood.

[0395] Step 4:

[0396] The server measures the difference between identified joint position data and ideal posture data using a difference measurement device. The inputs are joint position data and ideal posture data. The output is the difference data between these two sets of data. This process allows for a quantitative evaluation of the user's training effectiveness.

[0397] Step 5:

[0398] The server uses a feedback generation mechanism based on differential data to create natural language feedback. The input is differential data. An emotion analysis mechanism identifies the user's emotional state and adjusts the feedback content accordingly. The output is an emotion-sensitive feedback message. This feedback enhances the user's training motivation.

[0399] Step 6:

[0400] The server sends the generated feedback to the terminal, which then presents it to the user. The input is the generated feedback message. Output from the terminal includes screen displays and audio feedback. This allows the user to receive feedback in real time and modify their training.

[0401] Step 7:

[0402] The user takes a picture of their meal with their device, and the device sends this data to the server. The input is the image of the meal taken by the user. The output is the captured image data sent to the server. Through this process, data for meal analysis is collected.

[0403] Step 8:

[0404] The server analyzes the received images using a meal analysis system and evaluates their nutritional value. The input is the transmitted meal image data. The output is evaluation data regarding the nutritional value of the meal. This allows the user to obtain detailed information about the meal they have consumed.

[0405] Step 9:

[0406] The server generates advice on food choices based on the results of a meal analysis and the user's emotional state. The inputs are nutritional value evaluation data for meals and the user's emotional state. The output is specific meal selection advice. This allows users to strive for a diet that considers both their health and emotional state.

[0407] Step 10:

[0408] The generated dietary advice is sent to the device and presented to the user. The input is the advice data. The output is the advice displayed to the user from the device. This process allows the user to adjust their diet and make healthier choices.

[0409] (Application Example 2)

[0410] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0411] Traditional fitness support systems have struggled to provide appropriate feedback tailored to each user's individual emotions and fitness level. Furthermore, the lack of features that integrate training and nutrition information and support online marketplace purchases hindered users' ability to maintain sustainable health. There is a need to address these issues and support users in engaging in fitness activities more effectively.

[0412] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0413] In this invention, the server includes a video acquisition means, an image analysis means, a difference measurement means, an emotion analysis means, a means for analyzing the user's emotional state using the emotion analysis means and adjusting the feedback content in the feedback generation means, and an online connection means for providing purchase support in the virtual market. This enables the provision of appropriate feedback according to the user's emotional state and support for purchasing fitness-related products through the online market.

[0414] A "video acquisition means" is a device that acquires video data in order to capture the user's body movements in real time.

[0415] "Image analysis means" refers to a technology that includes an algorithm for analyzing acquired video data and identifying the positions of human joints.

[0416] The "difference measurement method" is a function for evaluating training effectiveness by measuring the difference between the analyzed posture data and the ideal posture data that has been registered in advance.

[0417] A "feedback generation method" is a means of analyzing difference data using natural language processing technology in order to generate feedback to be provided to the user.

[0418] A "presentation means" is a device for providing the generated feedback to the user visually or audibly.

[0419] "Emotional analysis methods" refer to technologies that analyze a user's emotional state from their facial expressions and voice.

[0420] "Online connectivity" refers to internet-based communication functions that allow users to connect to a virtual marketplace and assist in purchasing fitness-related products.

[0421] In this invention, the system operates by integrating various means to support the user's fitness activities. The server captures the user's body movements in real time via "video acquisition means" for acquiring video. This data is analyzed by "image analysis means" to determine the position of the user's joints. The analyzed data is compared with ideal trainer posture data by "difference measurement means" to generate difference data.

[0422] Based on this differential data, the "feedback generation means" utilizes a generation AI model to create feedback tailored to the user. This feedback is then adjusted based on the user's emotional state, which is evaluated through the "emotion analysis means." The generated feedback is then quickly provided to the user as audio or text via the "presentation means" on the user's terminal.

[0423] Furthermore, the system incorporates "online connectivity," allowing users to easily purchase necessary training products through a virtual marketplace. This creates an environment where users can seamlessly integrate fitness activities with product purchases.

[0424] For example, when a user wears smart glasses and does push-ups, the camera captures the position of their hands and analyzes that information. Then, appropriate feedback is provided, such as "Try spreading your hands a little wider." When the user is tired, emotionally appropriate feedback is provided, such as "Let's slow down."

[0425] An example of a prompt message would be: "Based on the user's training data, generate feedback for posture improvement. Next, analyze his emotional data and adjust the feedback to match his emotions."

[0426] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0427] Step 1:

[0428] The user begins training using the camera built into the device. The device captures the user's movements in real time using video acquisition equipment and sends the video data to the server. The input is the user's video captured by the camera, and the output is the video data sent to the server.

[0429] Step 2:

[0430] The server analyzes the received video data using image analysis tools to identify the user's joint positions. The input here is the video data obtained in step 1, and the output is the identified joint position data. Specifically, the server analyzes each frame using a machine learning model and maps each part of the body.

[0431] Step 3:

[0432] The server uses a difference measurement mechanism to measure the difference between the analyzed joint position data and pre-registered ideal posture data. The input is the joint position data and ideal posture data obtained in step 2, and the output is the difference data. Specifically, a particular index is set, and the difference is quantified using linear algebra.

[0433] Step 4:

[0434] The server uses a generative AI model to generate natural language feedback based on the differential data. The input here is the differential data obtained in step 3, and the output is feedback in natural language. Specifically, prompt sentences are input to the generative AI model, which then generates the optimal feedback sentence.

[0435] Step 5:

[0436] The server further uses emotion analysis tools to determine the user's emotional state from their facial expressions and voice data. The input is the user's facial expressions and voice data, and the output is the user's emotional state data. Specifically, it uses speech recognition technology and facial expression recognition technology such as DeepFace to detect and analyze the state.

[0437] Step 6:

[0438] Based on the user's emotional state, the generated feedback is adjusted and presented to the user through the device's display methods. The input is the feedback obtained in step 4 and the emotional state data obtained in step 5, and the output is the adjusted feedback. Specifically, the tone and content of the feedback are fine-tuned to match the user's emotions and presented as text display or audio output.

[0439] Step 7:

[0440] Using the terminal's online connectivity, users can access a virtual marketplace and purchase fitness-related products based on their feedback. Inputs are user feedback and virtual marketplace access rights, while output is a list of purchasable products. Specifically, the terminal opens an internet browser and provides an interface for the user to select suggested products and proceed with the purchase.

[0441] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0442] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0443] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0444] [Third Embodiment]

[0445] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0446] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0447] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0448] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0449] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0450] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0451] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0452] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0453] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0454] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0455] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0456] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0457] The fitness support system of the present invention allows users to perform efficient training at home without going to a gym, using a smartphone or tablet device. This system is configured as follows.

[0458] First, the user launches the application using their device. This application incorporates a video acquisition mechanism that allows the user's training session to be recorded in real time using the device's camera. The recorded video is then transmitted to a server via the network.

[0459] The server is equipped with image analysis capabilities to analyze the received video data. Specifically, it uses machine learning-based image recognition technology to identify the position of each joint in the user's body and convert it into posture data. At this stage, the user's movements during training are recorded as digital data.

[0460] Next, the server uses a difference measurement device to compare pre-registered ideal trainer posture data with the analyzed user posture data and measures the difference between the two. This difference data forms the basis for the feedback provided to the user.

[0461] The feedback generation mechanism generates natural language feedback based on the measured difference. For example, if the user is not raising their arm high enough, it will generate specific advice such as, "Raise your arm a little higher." The generated feedback is immediately provided to the user through the presentation mechanism. The presentation not only displays the feedback as text on the device screen but also plays it as audio.

[0462] Furthermore, the system includes a meal analysis function, which analyzes images of meals taken by the user on a server. This analysis uses image recognition technology to identify ingredients and evaluate their nutritional value. Based on this, the system provides the user with suggestions and advice for healthy eating.

[0463] In this way, this system enhances training efficiency and supports a healthy lifestyle by providing users with comprehensive and personalized fitness guidance and dietary management.

[0464] The following describes the processing flow.

[0465] Step 1:

[0466] The user launches the application on their device and prepares to start recording training videos. The user positions the camera appropriately and presses the record button to begin recording.

[0467] Step 2:

[0468] The device uses its camera to capture the user's training in real time and saves the video data to its internal storage. After recording begins, the video data is sent to the server frame by frame.

[0469] Step 3:

[0470] The server analyzes the received video data and uses image analysis tools to identify the joint positions and angles of the user's body for each frame. This generates the user's posture data.

[0471] Step 4:

[0472] The server compares the ideal trainer's posture data with the user's posture data using a difference measurement mechanism. This comparison yields specific difference data.

[0473] Step 5:

[0474] The server uses feedback generation methods based on differential data to generate natural language feedback to be provided to the user. The generated feedback is then put into text form and created as personalized advice tailored to the user.

[0475] Step 6:

[0476] Feedback is sent to the device and provided to the user in real time via a presentation method. The device displays it on the screen as a text message and also plays it as audio feedback through the speaker.

[0477] Step 7:

[0478] Users take photos of their meals using their devices and upload them to the server. By recording each meal, users can manage their eating habits.

[0479] Step 8:

[0480] The server analyzes uploaded meal images using a meal analysis tool to evaluate their nutritional components. The component information identified by image recognition technology is cross-referenced with a database and provided to the user as specific nutritional values ​​and advice necessary for health management.

[0481] By following these steps, the system provides training support and dietary management, offering users a real-time and comprehensive fitness experience.

[0482] (Example 1)

[0483] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0484] In modern times, providing individuals with effective exercise guidance and healthy dietary management is difficult for many. Traditional methods require significant time and expert guidance to obtain personalized feedback and dietary suggestions, posing a challenge to sustainable health management, especially for those leading busy lives.

[0485] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0486] In this invention, the server includes a shooting means for acquiring video data, an image analysis means for analyzing human body movements acquired by the shooting means and identifying joint positions, a comparison means for measuring the difference between registered ideal movement data and the analyzed movement data, and a generation means for generating advice using a generation AI model. This enables the provision of immediately customized exercise and dietary feedback to individuals, making it possible to maintain health efficiently.

[0487] "Video data" refers to digital information that visually captures a user's movements and activities.

[0488] "Filming means" refers to a device or function that films the user's actions in real time and acquires them as video data.

[0489] "Image analysis means" refers to a technology that analyzes the user's body movements and posture from acquired video data and identifies joint positions.

[0490] A "comparison method" is a process for comparing ideal exercise data with the user's exercise data and measuring the difference between them.

[0491] The "generation method" refers to a function that generates advice and feedback for the user based on the measured differential data.

[0492] "Presentation means" refers to a device or function that presents the generated advice to the user visually or audibly.

[0493] A "meal image" is digital information that visually captures the contents of a meal as photographed by the user.

[0494] "Methods for evaluating nutritional value" refer to technologies that identify the ingredients contained in a meal based on an image of the meal and evaluate its nutritional value.

[0495] This fitness support system enables users to efficiently manage their training and diet using their devices. Users can easily access this system using smartphones or tablets. The server at the core of this system is responsible for processing and analyzing the received data.

[0496] When the application is launched by the user, the device first uses its built-in camera to capture real-time video data of the user's training. The captured video data is then transmitted to a server via the network. During this process, the video may be compressed for more efficient transfer.

[0497] The server performs image analysis on the received video data. Specifically, it uses a machine learning platform such as TensorFlow to perform image recognition, identifying the position of each joint in the user's body. This technology allows the user's movements to be digitally recorded and stored as posture data.

[0498] Next, the server compares ideal exercise data with the user's analyzed exercise data. The difference measured by the comparison device forms the basis for identifying areas for improvement in the user's training.

[0499] Furthermore, the server uses a generative AI model to generate feedback for the user based on the measured differential data. The generated feedback is sent to the user's device and displayed visually as text on the screen, as well as played as audio. For example, if the user's arm is held low, advice such as "Please raise your arm a little higher" is provided.

[0500] Users can also take photos of their meals with their smartphones and send the images to the server. The server analyzes these meal images and evaluates their nutritional value. Based on the evaluation results, the user is provided with suggestions and advice for healthy eating. For example, if it is determined that there is a vitamin deficiency, advice such as "Add some fruit" will be generated.

[0501] An example of a prompt for the generating AI model would be: "Analyze the user's training video to obtain posture data and tell me the difference from the ideal posture. Also, generate healthy meal suggestions based on the food images."

[0502] In this way, the system functions as a comprehensive tool to support the user's health, providing personalized fitness guidance and dietary management.

[0503] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0504] Step 1:

[0505] The user launches a fitness application using the device and activates the camera. The input for this step is the user's workout scene, and the output is the video data that visually captures it. The device temporarily records this video data and prepares to send it to the server when ready.

[0506] Step 2:

[0507] The terminal transmits recorded video data to the server via the network. The input is temporarily recorded compressed video data, and the output is the video data transmitted to the server. The terminal compresses the data as needed to improve the efficiency of data transfer.

[0508] Step 3:

[0509] The server analyzes the received video data. The input is the transmitted video data, and the output is posture data including the joint positions of the user's body. The server uses machine learning tools such as TensorFlow for this analysis to automatically identify the user's joints from the video.

[0510] Step 4:

[0511] The server compares the analyzed posture data with existing ideal movement data. The input is the analyzed user posture data and the ideal posture data, and the output is the difference between the two. The server quantifies this difference and identifies areas that need improvement.

[0512] Step 5:

[0513] The server generates feedback using a generative AI model based on differential data. The input is differential data, and the output is specific feedback sentences for the user. The generated feedback provides the user with suggestions and advice on how to improve their movement in natural language.

[0514] Step 6:

[0515] The terminal presents the user with feedback received from the server. The input is the generated feedback data, and the output is the feedback text and synthesized speech feedback displayed on the terminal screen. The terminal displays this quickly so that the user can easily review it.

[0516] Step 7:

[0517] Users photograph their daily meals and send these images from their device to the server. The input is image data of the meal, and the output is an image file that can be processed by the server. The device transmits images quickly through a stable connection.

[0518] Step 8:

[0519] The server analyzes received meal images and uses image recognition technology to identify ingredients. The input is the submitted meal image, and the output is a list of ingredients and their nutritional information. Based on this, the server generates health-conscious meal suggestions for the user.

[0520] (Application Example 1)

[0521] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0522] Improving work efficiency within a factory requires optimizing worker movements. However, conventional methods have made it difficult to obtain detailed, real-time motion analysis and accurate feedback. A technology is needed to solve this problem.

[0523] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0524] In this invention, the server includes a video acquisition device, a pixel analysis device that analyzes the movement of a workpiece and identifies joint positions, and a difference measurement device that measures differences. This makes it possible to analyze the worker's movements in real time and provide optimal feedback.

[0525] A "video acquisition device" is a device used to film the movements of a work object and acquire them as video data.

[0526] A "pixel analysis device" is a device that analyzes acquired video data to identify the movement and joint positions of a workpiece.

[0527] A "difference measurement device" is a device that measures the difference between ideal operating data and analyzed actual operating data.

[0528] A "feedback generation device" is a device that generates feedback in natural language to a work object based on the difference in measured motion data.

[0529] A "display device" is a device that presents generated feedback to a work object visually or audibly.

[0530] The system implementing this invention supports workers in efficiently performing their tasks within a factory. A specific embodiment is shown below.

[0531] First, the worker puts on smart glasses. These smart glasses are equipped with cameras to record their movements. The smart glasses, acting as a video acquisition device, record the worker's movements in real time and transmit the acquired video data to a server. The server processes the video data using high-performance image analysis software. Specifically, it uses the OpenCV library to identify joint positions from the video data.

[0532] Next, the server uses TensorFlow, a deep learning framework, to analyze joint positions. Based on the analyzed motion data, it functions as a difference measurement device to measure the difference between the current motion data and ideal motion data. Based on the measured difference, the server acts as a feedback generator, producing specific improvement suggestions for the worker in natural language. This can utilize the GPT natural language processing model.

[0533] The generated feedback is sent to smart glasses, which act as a display device. The smart glasses not only display the feedback as text but also output it as audio using speech synthesis technology. The Google Text-to-Speech API can be used for speech synthesis.

[0534] As a concrete example, workers in a packaging factory may use smart glasses, and the analysis system may detect inefficient movements during their work. For instance, if an arm movement is inefficient, real-time feedback such as, "Moving your hand 30cm to the left from that position will improve efficiency," can be provided, enabling improvements to the work process. By using this invention, both work efficiency and safety can be improved simultaneously.

[0535] Examples of prompt statements to input into the generative AI model are as follows:

[0536] "Generate feedback based on motion analysis of factory workers. If the current motion is not optimal, specify areas for improvement."

[0537] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0538] Step 1:

[0539] The user (worker) puts on smart glasses and begins work. The smart glasses' camera captures their actions in real time during the work.

[0540] The input is video footage of the worker's movements, and the output is video data. The video data is sent to a server for subsequent analysis.

[0541] Step 2:

[0542] The server processes the received video data using the OpenCV library to identify the joint positions of the worker in each frame.

[0543] The input is the video data acquired in step 1, and the output is the joint coordinate data. This data is used for motion analysis.

[0544] Step 3:

[0545] The server uses TensorFlow to analyze joint coordinate data and generate actual motion data.

[0546] The input is the joint coordinate data obtained in step 2, and the output is detailed motion data. This analysis allows us to understand the overall movements of the worker during the task.

[0547] Step 4:

[0548] The server compares the data to ideal operating data and calculates the difference, acting as a difference measurement device.

[0549] The input consists of the actual operation data obtained in step 3 and pre-registered ideal operation data. The output is the difference data of the operation. This difference is used for feedback generation.

[0550] Step 5:

[0551] The server generates feedback from the differential data using a natural language processing model.

[0552] The input is the difference data obtained in step 4, and the output is text data of specific feedback. The feedback includes specific advice for improving the work.

[0553] Step 6:

[0554] The server converts the generated feedback into audio data using speech synthesis technology and sends it to the smart glasses.

[0555] The input is the text data of the feedback generated in step 5, and the output is audio data. The smart glasses enable real-time instruction by presenting this audio data to the worker.

[0556] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0557] The fitness support system of the present invention employs groundbreaking technology that combines an emotional engine to maximize the user's training effectiveness. This system has the following components:

[0558] The user launches the application on their device and prepares for training. The device's camera is used to record the user's training in real time using a video acquisition system. The acquired video data is transmitted to a server via the internet.

[0559] The server is equipped with image analysis capabilities to identify the position of each joint in the user's body. This information is compared with pre-stored ideal trainer posture data by a difference measurement system to generate difference data for measuring the effectiveness of the training.

[0560] The feedback generation mechanism uses this differential data to create natural language feedback tailored to the user. A notable feature here is the use of an emotion engine. The emotion engine analyzes the user's emotional state from their facial expressions and voice, and adjusts the feedback accordingly. For example, if the emotion engine detects that the user is tired, it generates an encouraging message such as, "Let's keep going a little longer, or you can take a short break." This feedback is quickly sent to the device and presented both on the screen and audibly.

[0561] Furthermore, a meal analysis system analyzes images of meals taken by users on a server. When combined with an emotion engine, users in specific emotional states are provided with motivational advice regarding their meal choices. For example, if a stressed state is detected, advice such as "It would be good to increase the amount of foods that have a relaxing effect" may be given.

[0562] Thus, this system comprehensively supports users' physical training and emotional well-being, providing integrated fitness and nutritional management. This maximizes user motivation and results, enabling sustainable health maintenance.

[0563] The following describes the processing flow.

[0564] Step 1:

[0565] The user launches the fitness app installed on their device and starts a training session. The user positions the device's camera appropriately and starts recording by pressing the record button.

[0566] Step 2:

[0567] The device uses its camera to capture real-time video of the user during training and generates video data. This data is then transmitted to a server via the network.

[0568] Step 3:

[0569] The server processes the received video data using image analysis tools to identify the position of each joint in the user's posture. Based on these analysis results, posture data is obtained.

[0570] Step 4:

[0571] The server uses a difference measurement mechanism to compare the analyzed user's posture data with pre-registered ideal trainer posture data. This process generates specific difference data.

[0572] Step 5:

[0573] The emotion engine analyzes the user's facial expressions and voice to identify their current emotional state. This emotional state data is then used to generate subsequent feedback.

[0574] Step 6:

[0575] The server's feedback generation mechanism combines differential data and emotional states to generate natural language feedback tailored to each user's situation. The content and tone of the feedback are adjusted based on the emotional state.

[0576] Step 7:

[0577] The generated feedback information is sent to the terminal. The presentation device displays this information as text on the screen and also plays it as audio, providing it to the user in real time.

[0578] Step 8:

[0579] The user takes a picture of their meal with their device and uploads it to the server. The meal image data is used as data for subsequent nutritional analysis.

[0580] Step 9:

[0581] The server utilizes meal analysis tools to analyze uploaded meal images and identify their nutritional value. The analysis results are provided as nutritional advice based on the user's emotional state, supporting the maintenance of their health.

[0582] Through the steps outlined above, this system aims to integrate and manage the user's physical training and emotional support, ultimately striving for sustainable fitness improvement.

[0583] (Example 2)

[0584] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0585] In fitness activities, for users to achieve ideal training results, comprehensive support is needed that considers not only their physical posture but also their emotional state and dietary choices. However, current systems are unable to manage these elements in an integrated manner, limiting their ability to maintain and improve user motivation and health benefits.

[0586] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0587] In this invention, the server includes video acquisition means, image analysis means, difference measurement means, feedback generation means, emotion analysis means, and diet analysis means. This enables comprehensive analysis of the user's physical training data, emotional state, and information regarding dietary choices, and provides personalized feedback and advice.

[0588] A "video acquisition device" is a device that captures the user's physical movements in real time and acquires the video data.

[0589] "Image analysis means" refers to an analysis device used to identify the position of each joint in the user's body from acquired video data.

[0590] A "difference measurement device" is a device that measures the difference between the analyzed user's posture data and pre-registered ideal posture data, and generates difference data.

[0591] A "feedback generation means" is a device that generates natural language feedback for the user based on the difference data obtained by the difference measurement means.

[0592] An "emotion analysis device" is a device that analyzes the user's emotional state from their facial expressions and voice, and adjusts the generated feedback content according to that emotion.

[0593] A "meal analysis device" is a device that analyzes images of meals taken by users and evaluates their nutritional value.

[0594] A "presentation means" is a device that presents generated feedback or advice to the user, and performs functions such as display or audio output.

[0595] The fitness support system of the present invention is configured as a multi-functional platform including emotion analysis and dietary analysis to optimize the user's training effectiveness. The user launches a dedicated application using a terminal and acquires training video through the terminal's camera. This video data is transmitted to a server via the internet.

[0596] The server uses the received video data to identify the joint positions of the user's body through image analysis. Simultaneously with this analysis, a difference measurement means measures the difference between the user's posture data and pre-stored ideal posture data, generating difference data. Subsequently, a feedback generation means generates natural language feedback tailored to the user's training progress. During this process, an emotion analysis means analyzes the user's facial expressions and voice to determine their emotional state and adjusts the feedback content accordingly.

[0597] Furthermore, this system includes a meal analysis function that analyzes images of meals taken by the user on a server. It evaluates the nutritional value from the analysis results and generates advice on meal choices based on the user's emotional state.

[0598] For example, when a user performs squats at home, the device's camera records the movement. This video is sent to a server, which analyzes the user's knee and hip positioning and generates feedback such as, "It would be good to point your knees slightly outward." In another example, a picture of a salad taken by a user at their dinner table is analyzed by the server, and advice such as, "Adding avocado is good for stress reduction," is provided.

[0599] An example of a prompt using a generative AI model is: "Generate a suitable feedback message when the user is tired. Example: 'You're doing great. It's okay to take a break.'" This allows the user to receive comprehensive support from both a physical and emotional perspective.

[0600] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0601] Step 1:

[0602] The user launches a fitness application on their device and prepares to begin training using the camera. Inputs include the user's selected training mode and the press of a start button. Outputs include the activation of the camera device and the display of the training mode on the user's screen. This completes the process of preparing to begin training.

[0603] Step 2:

[0604] The device's camera captures the user's training movements in real time, acquiring video data. The input is the user's physical movements. This video data is compressed and sent to the server. The output is the captured video data, which is also sent to the server in real time. This allows the necessary data to be collected for analysis.

[0605] Step 3:

[0606] The server processes the received video data using image analysis to identify the position of each joint in the user's body. The input is the transmitted video data. The output is data on the joint positions. Through data analysis, the user's posture can be precisely understood.

[0607] Step 4:

[0608] The server measures the difference between identified joint position data and ideal posture data using a difference measurement device. The inputs are joint position data and ideal posture data. The output is the difference data between these two sets of data. This process allows for a quantitative evaluation of the user's training effectiveness.

[0609] Step 5:

[0610] The server uses a feedback generation mechanism based on differential data to create natural language feedback. The input is differential data. An emotion analysis mechanism identifies the user's emotional state and adjusts the feedback content accordingly. The output is an emotion-sensitive feedback message. This feedback enhances the user's training motivation.

[0611] Step 6:

[0612] The server sends the generated feedback to the terminal, which then presents it to the user. The input is the generated feedback message. Output from the terminal includes screen displays and audio feedback. This allows the user to receive feedback in real time and modify their training.

[0613] Step 7:

[0614] The user takes a picture of their meal with their device, and the device sends this data to the server. The input is the image of the meal taken by the user. The output is the captured image data sent to the server. Through this process, data for meal analysis is collected.

[0615] Step 8:

[0616] The server analyzes the received images using a meal analysis system and evaluates their nutritional value. The input is the transmitted meal image data. The output is evaluation data regarding the nutritional value of the meal. This allows the user to obtain detailed information about the meal they have consumed.

[0617] Step 9:

[0618] The server generates advice on food choices based on the results of a meal analysis and the user's emotional state. The inputs are nutritional value evaluation data for meals and the user's emotional state. The output is specific meal selection advice. This allows users to strive for a diet that considers both their health and emotional state.

[0619] Step 10:

[0620] The generated dietary advice is sent to the device and presented to the user. The input is the advice data. The output is the advice displayed to the user from the device. This process allows the user to adjust their diet and make healthier choices.

[0621] (Application Example 2)

[0622] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0623] Traditional fitness support systems have struggled to provide appropriate feedback tailored to each user's individual emotions and fitness level. Furthermore, the lack of features that integrate training and nutrition information and support online marketplace purchases hindered users' ability to maintain sustainable health. There is a need to address these issues and support users in engaging in fitness activities more effectively.

[0624] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0625] In this invention, the server includes a video acquisition means, an image analysis means, a difference measurement means, an emotion analysis means, a means for analyzing the user's emotional state using the emotion analysis means and adjusting the feedback content in the feedback generation means, and an online connection means for providing purchase support in the virtual market. This enables the provision of appropriate feedback according to the user's emotional state and support for purchasing fitness-related products through the online market.

[0626] A "video acquisition means" is a device that acquires video data in order to capture the user's body movements in real time.

[0627] "Image analysis means" refers to a technology that includes an algorithm for analyzing acquired video data and identifying the positions of human joints.

[0628] The "difference measurement method" is a function for evaluating training effectiveness by measuring the difference between the analyzed posture data and the ideal posture data that has been registered in advance.

[0629] A "feedback generation method" is a means of analyzing difference data using natural language processing technology in order to generate feedback to be provided to the user.

[0630] A "presentation means" is a device for providing the generated feedback to the user visually or audibly.

[0631] "Emotional analysis methods" refer to technologies that analyze a user's emotional state from their facial expressions and voice.

[0632] "Online connectivity" refers to internet-based communication functions that allow users to connect to a virtual marketplace and assist in purchasing fitness-related products.

[0633] In this invention, the system operates by integrating various means to support the user's fitness activities. The server captures the user's body movements in real time via "video acquisition means" for acquiring video. This data is analyzed by "image analysis means" to determine the position of the user's joints. The analyzed data is compared with ideal trainer posture data by "difference measurement means" to generate difference data.

[0634] Based on this differential data, the "feedback generation means" utilizes a generation AI model to create feedback tailored to the user. This feedback is then adjusted based on the user's emotional state, which is evaluated through the "emotion analysis means." The generated feedback is then quickly provided to the user as audio or text via the "presentation means" on the user's terminal.

[0635] Furthermore, the system incorporates "online connectivity," allowing users to easily purchase necessary training products through a virtual marketplace. This creates an environment where users can seamlessly integrate fitness activities with product purchases.

[0636] For example, when a user wears smart glasses and does push-ups, the camera captures the position of their hands and analyzes that information. Then, appropriate feedback is provided, such as "Try spreading your hands a little wider." When the user is tired, emotionally appropriate feedback is provided, such as "Let's slow down."

[0637] An example of a prompt message would be: "Based on the user's training data, generate feedback for posture improvement. Next, analyze his emotional data and adjust the feedback to match his emotions."

[0638] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0639] Step 1:

[0640] The user begins training using the camera built into the device. The device captures the user's movements in real time using video acquisition equipment and sends the video data to the server. The input is the user's video captured by the camera, and the output is the video data sent to the server.

[0641] Step 2:

[0642] The server analyzes the received video data using image analysis tools to identify the user's joint positions. The input here is the video data obtained in step 1, and the output is the identified joint position data. Specifically, the server analyzes each frame using a machine learning model and maps each part of the body.

[0643] Step 3:

[0644] The server uses a difference measurement mechanism to measure the difference between the analyzed joint position data and pre-registered ideal posture data. The input is the joint position data and ideal posture data obtained in step 2, and the output is the difference data. Specifically, a particular index is set, and the difference is quantified using linear algebra.

[0645] Step 4:

[0646] The server uses a generative AI model to generate natural language feedback based on the differential data. The input here is the differential data obtained in step 3, and the output is feedback in natural language. Specifically, prompt sentences are input to the generative AI model, which then generates the optimal feedback sentence.

[0647] Step 5:

[0648] The server further uses emotion analysis tools to determine the user's emotional state from their facial expressions and voice data. The input is the user's facial expressions and voice data, and the output is the user's emotional state data. Specifically, it uses speech recognition technology and facial expression recognition technology such as DeepFace to detect and analyze the state.

[0649] Step 6:

[0650] Based on the user's emotional state, the generated feedback is adjusted and presented to the user through the device's display methods. The input is the feedback obtained in step 4 and the emotional state data obtained in step 5, and the output is the adjusted feedback. Specifically, the tone and content of the feedback are fine-tuned to match the user's emotions and presented as text display or audio output.

[0651] Step 7:

[0652] Using the terminal's online connectivity, users can access a virtual marketplace and purchase fitness-related products based on their feedback. Inputs are user feedback and virtual marketplace access rights, while output is a list of purchasable products. Specifically, the terminal opens an internet browser and provides an interface for the user to select suggested products and proceed with the purchase.

[0653] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0654] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0655] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0656] [Fourth Embodiment]

[0657] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0658] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0659] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0660] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0661] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0662] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0663] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0664] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0665] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0666] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0667] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0668] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0669] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0670] The fitness support system of the present invention allows users to perform efficient training at home without going to a gym, using a smartphone or tablet device. This system is configured as follows.

[0671] First, the user launches the application using their device. This application incorporates a video acquisition mechanism that allows the user's training session to be recorded in real time using the device's camera. The recorded video is then transmitted to a server via the network.

[0672] The server is equipped with image analysis capabilities to analyze the received video data. Specifically, it uses machine learning-based image recognition technology to identify the position of each joint in the user's body and convert it into posture data. At this stage, the user's movements during training are recorded as digital data.

[0673] Next, the server uses a difference measurement device to compare pre-registered ideal trainer posture data with the analyzed user posture data and measures the difference between the two. This difference data forms the basis for the feedback provided to the user.

[0674] The feedback generation mechanism generates natural language feedback based on the measured difference. For example, if the user is not raising their arm high enough, it will generate specific advice such as, "Raise your arm a little higher." The generated feedback is immediately provided to the user through the presentation mechanism. The presentation not only displays the feedback as text on the device screen but also plays it as audio.

[0675] Furthermore, the system includes a meal analysis function, which analyzes images of meals taken by the user on a server. This analysis uses image recognition technology to identify ingredients and evaluate their nutritional value. Based on this, the system provides the user with suggestions and advice for healthy eating.

[0676] In this way, this system enhances training efficiency and supports a healthy lifestyle by providing users with comprehensive and personalized fitness guidance and dietary management.

[0677] The following describes the processing flow.

[0678] Step 1:

[0679] The user launches the application on their device and prepares to start recording training videos. The user positions the camera appropriately and presses the record button to begin recording.

[0680] Step 2:

[0681] The device uses its camera to capture the user's training in real time and saves the video data to its internal storage. After recording begins, the video data is sent to the server frame by frame.

[0682] Step 3:

[0683] The server analyzes the received video data and uses image analysis tools to identify the joint positions and angles of the user's body for each frame. This generates the user's posture data.

[0684] Step 4:

[0685] The server compares the ideal trainer's posture data with the user's posture data using a difference measurement mechanism. This comparison yields specific difference data.

[0686] Step 5:

[0687] The server uses feedback generation methods based on differential data to generate natural language feedback to be provided to the user. The generated feedback is then put into text form and created as personalized advice tailored to the user.

[0688] Step 6:

[0689] Feedback is sent to the device and provided to the user in real time via a presentation method. The device displays it on the screen as a text message and also plays it as audio feedback through the speaker.

[0690] Step 7:

[0691] Users take photos of their meals using their devices and upload them to the server. By recording each meal, users can manage their eating habits.

[0692] Step 8:

[0693] The server analyzes uploaded meal images using a meal analysis tool to evaluate their nutritional components. The component information identified by image recognition technology is cross-referenced with a database and provided to the user as specific nutritional values ​​and advice necessary for health management.

[0694] By following these steps, the system provides training support and dietary management, offering users a real-time and comprehensive fitness experience.

[0695] (Example 1)

[0696] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0697] In modern times, providing individuals with effective exercise guidance and healthy dietary management is difficult for many. Traditional methods require significant time and expert guidance to obtain personalized feedback and dietary suggestions, posing a challenge to sustainable health management, especially for those leading busy lives.

[0698] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0699] In this invention, the server includes a shooting means for acquiring video data, an image analysis means for analyzing human body movements acquired by the shooting means and identifying joint positions, a comparison means for measuring the difference between registered ideal movement data and the analyzed movement data, and a generation means for generating advice using a generation AI model. This enables the provision of immediately customized exercise and dietary feedback to individuals, making it possible to maintain health efficiently.

[0700] "Video data" refers to digital information that visually captures a user's movements and activities.

[0701] "Filming means" refers to a device or function that films the user's actions in real time and acquires them as video data.

[0702] "Image analysis means" refers to a technology that analyzes the user's body movements and posture from acquired video data and identifies joint positions.

[0703] A "comparison method" is a process for comparing ideal exercise data with the user's exercise data and measuring the difference between them.

[0704] The "generation method" refers to a function that generates advice and feedback for the user based on the measured differential data.

[0705] "Presentation means" refers to a device or function that presents the generated advice to the user visually or audibly.

[0706] A "meal image" is digital information that visually captures the contents of a meal as photographed by the user.

[0707] "Methods for evaluating nutritional value" refer to technologies that identify the ingredients contained in a meal based on an image of the meal and evaluate its nutritional value.

[0708] This fitness support system enables users to efficiently manage their training and diet using their devices. Users can easily access this system using smartphones or tablets. The server at the core of this system is responsible for processing and analyzing the received data.

[0709] When the application is launched by the user, the device first uses its built-in camera to capture real-time video data of the user's training. The captured video data is then transmitted to a server via the network. During this process, the video may be compressed for more efficient transfer.

[0710] The server performs image analysis on the received video data. Specifically, it uses a machine learning platform such as TensorFlow to perform image recognition, identifying the position of each joint in the user's body. This technology allows the user's movements to be digitally recorded and stored as posture data.

[0711] Next, the server compares ideal exercise data with the user's analyzed exercise data. The difference measured by the comparison device forms the basis for identifying areas for improvement in the user's training.

[0712] Furthermore, the server uses a generative AI model to generate feedback for the user based on the measured differential data. The generated feedback is sent to the user's device and displayed visually as text on the screen, as well as played as audio. For example, if the user's arm is held low, advice such as "Please raise your arm a little higher" is provided.

[0713] Users can also take photos of their meals with their smartphones and send the images to the server. The server analyzes these meal images and evaluates their nutritional value. Based on the evaluation results, the user is provided with suggestions and advice for healthy eating. For example, if it is determined that there is a vitamin deficiency, advice such as "Add some fruit" will be generated.

[0714] An example of a prompt for the generating AI model would be: "Analyze the user's training video to obtain posture data and tell me the difference from the ideal posture. Also, generate healthy meal suggestions based on the food images."

[0715] In this way, the system functions as a comprehensive tool to support the user's health, providing personalized fitness guidance and dietary management.

[0716] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0717] Step 1:

[0718] The user launches a fitness application using the device and activates the camera. The input for this step is the user's workout scene, and the output is the video data that visually captures it. The device temporarily records this video data and prepares to send it to the server when ready.

[0719] Step 2:

[0720] The terminal transmits recorded video data to the server via the network. The input is temporarily recorded compressed video data, and the output is the video data transmitted to the server. The terminal compresses the data as needed to improve the efficiency of data transfer.

[0721] Step 3:

[0722] The server analyzes the received video data. The input is the transmitted video data, and the output is posture data including the joint positions of the user's body. The server uses machine learning tools such as TensorFlow for this analysis to automatically identify the user's joints from the video.

[0723] Step 4:

[0724] The server compares the analyzed posture data with existing ideal movement data. The input is the analyzed user posture data and the ideal posture data, and the output is the difference between the two. The server quantifies this difference and identifies areas that need improvement.

[0725] Step 5:

[0726] The server generates feedback using a generative AI model based on differential data. The input is differential data, and the output is specific feedback sentences for the user. The generated feedback provides the user with suggestions and advice on how to improve their movement in natural language.

[0727] Step 6:

[0728] The terminal presents the user with feedback received from the server. The input is the generated feedback data, and the output is the feedback text and synthesized speech feedback displayed on the terminal screen. The terminal displays this quickly so that the user can easily review it.

[0729] Step 7:

[0730] Users photograph their daily meals and send these images from their device to the server. The input is image data of the meal, and the output is an image file that can be processed by the server. The device transmits images quickly through a stable connection.

[0731] Step 8:

[0732] The server analyzes received meal images and uses image recognition technology to identify ingredients. The input is the submitted meal image, and the output is a list of ingredients and their nutritional information. Based on this, the server generates health-conscious meal suggestions for the user.

[0733] (Application Example 1)

[0734] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0735] Improving work efficiency within a factory requires optimizing worker movements. However, conventional methods have made it difficult to obtain detailed, real-time motion analysis and accurate feedback. A technology is needed to solve this problem.

[0736] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0737] In this invention, the server includes a video acquisition device, a pixel analysis device that analyzes the movement of a workpiece and identifies joint positions, and a difference measurement device that measures differences. This makes it possible to analyze the worker's movements in real time and provide optimal feedback.

[0738] A "video acquisition device" is a device used to film the movements of a work object and acquire them as video data.

[0739] A "pixel analysis device" is a device that analyzes acquired video data to identify the movement and joint positions of a workpiece.

[0740] A "difference measurement device" is a device that measures the difference between ideal operating data and analyzed actual operating data.

[0741] A "feedback generation device" is a device that generates feedback in natural language to a work object based on the difference in measured motion data.

[0742] A "display device" is a device that presents generated feedback to a work object visually or audibly.

[0743] The system implementing this invention supports workers in efficiently performing their tasks within a factory. A specific embodiment is shown below.

[0744] First, the worker puts on smart glasses. These smart glasses are equipped with cameras to record their movements. The smart glasses, acting as a video acquisition device, record the worker's movements in real time and transmit the acquired video data to a server. The server processes the video data using high-performance image analysis software. Specifically, it uses the OpenCV library to identify joint positions from the video data.

[0745] Next, the server uses TensorFlow, a deep learning framework, to analyze joint positions. Based on the analyzed motion data, it functions as a difference measurement device to measure the difference between the current motion data and ideal motion data. Based on the measured difference, the server acts as a feedback generator, producing specific improvement suggestions for the worker in natural language. This can utilize the GPT natural language processing model.

[0746] The generated feedback is sent to smart glasses, which act as a display device. The smart glasses not only display the feedback as text but also output it as audio using speech synthesis technology. The Google Text-to-Speech API can be used for speech synthesis.

[0747] As a concrete example, workers in a packaging factory may use smart glasses, and the analysis system may detect inefficient movements during their work. For instance, if an arm movement is inefficient, real-time feedback such as, "Moving your hand 30cm to the left from that position will improve efficiency," can be provided, enabling improvements to the work process. By using this invention, both work efficiency and safety can be improved simultaneously.

[0748] Examples of prompt statements to input into the generative AI model are as follows:

[0749] "Generate feedback based on motion analysis of factory workers. If the current motion is not optimal, specify areas for improvement."

[0750] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0751] Step 1:

[0752] The user (worker) puts on smart glasses and begins work. The smart glasses' camera captures their actions in real time during the work.

[0753] The input is video footage of the worker's movements, and the output is video data. The video data is sent to a server for subsequent analysis.

[0754] Step 2:

[0755] The server processes the received video data using the OpenCV library to identify the joint positions of the worker in each frame.

[0756] The input is the video data acquired in step 1, and the output is the joint coordinate data. This data is used for motion analysis.

[0757] Step 3:

[0758] The server uses TensorFlow to analyze joint coordinate data and generate actual motion data.

[0759] The input is the joint coordinate data obtained in step 2, and the output is detailed motion data. This analysis allows us to understand the overall movements of the worker during the task.

[0760] Step 4:

[0761] The server compares the data to ideal operating data and calculates the difference, acting as a difference measurement device.

[0762] The input consists of the actual operation data obtained in step 3 and pre-registered ideal operation data. The output is the difference data of the operation. This difference is used for feedback generation.

[0763] Step 5:

[0764] The server generates feedback from the differential data using a natural language processing model.

[0765] The input is the difference data obtained in step 4, and the output is text data of specific feedback. The feedback includes specific advice for improving the work.

[0766] Step 6:

[0767] The server converts the generated feedback into audio data using speech synthesis technology and sends it to the smart glasses.

[0768] The input is the text data of the feedback generated in step 5, and the output is audio data. The smart glasses enable real-time instruction by presenting this audio data to the worker.

[0769] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0770] The fitness support system of the present invention employs groundbreaking technology that combines an emotional engine to maximize the user's training effectiveness. This system has the following components:

[0771] The user launches the application on their device and prepares for training. The device's camera is used to record the user's training in real time using a video acquisition system. The acquired video data is transmitted to a server via the internet.

[0772] The server is equipped with image analysis capabilities to identify the position of each joint in the user's body. This information is compared with pre-stored ideal trainer posture data by a difference measurement system to generate difference data for measuring the effectiveness of the training.

[0773] The feedback generation mechanism uses this differential data to create natural language feedback tailored to the user. A notable feature here is the use of an emotion engine. The emotion engine analyzes the user's emotional state from their facial expressions and voice, and adjusts the feedback accordingly. For example, if the emotion engine detects that the user is tired, it generates an encouraging message such as, "Let's keep going a little longer, or you can take a short break." This feedback is quickly sent to the device and presented both on the screen and audibly.

[0774] Furthermore, a meal analysis system analyzes images of meals taken by users on a server. When combined with an emotion engine, users in specific emotional states are provided with motivational advice regarding their meal choices. For example, if a stressed state is detected, advice such as "It would be good to increase the amount of foods that have a relaxing effect" may be given.

[0775] Thus, this system comprehensively supports users' physical training and emotional well-being, providing integrated fitness and nutritional management. This maximizes user motivation and results, enabling sustainable health maintenance.

[0776] The following describes the processing flow.

[0777] Step 1:

[0778] The user launches the fitness app installed on their device and starts a training session. The user positions the device's camera appropriately and starts recording by pressing the record button.

[0779] Step 2:

[0780] The device uses its camera to capture real-time video of the user during training and generates video data. This data is then transmitted to a server via the network.

[0781] Step 3:

[0782] The server processes the received video data using image analysis tools to identify the position of each joint in the user's posture. Based on these analysis results, posture data is obtained.

[0783] Step 4:

[0784] The server uses a difference measurement mechanism to compare the analyzed user's posture data with pre-registered ideal trainer posture data. This process generates specific difference data.

[0785] Step 5:

[0786] The emotion engine analyzes the user's facial expressions and voice to identify their current emotional state. This emotional state data is then used to generate subsequent feedback.

[0787] Step 6:

[0788] The server's feedback generation mechanism combines differential data and emotional states to generate natural language feedback tailored to each user's situation. The content and tone of the feedback are adjusted based on the emotional state.

[0789] Step 7:

[0790] The generated feedback information is sent to the terminal. The presentation device displays this information as text on the screen and also plays it as audio, providing it to the user in real time.

[0791] Step 8:

[0792] The user takes a picture of their meal with their device and uploads it to the server. The meal image data is used as data for subsequent nutritional analysis.

[0793] Step 9:

[0794] The server utilizes meal analysis tools to analyze uploaded meal images and identify their nutritional value. The analysis results are provided as nutritional advice based on the user's emotional state, supporting the maintenance of their health.

[0795] Through the steps outlined above, this system aims to integrate and manage the user's physical training and emotional support, ultimately striving for sustainable fitness improvement.

[0796] (Example 2)

[0797] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0798] In fitness activities, for users to achieve ideal training results, comprehensive support is needed that considers not only their physical posture but also their emotional state and dietary choices. However, current systems are unable to manage these elements in an integrated manner, limiting their ability to maintain and improve user motivation and health benefits.

[0799] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0800] In this invention, the server includes video acquisition means, image analysis means, difference measurement means, feedback generation means, emotion analysis means, and diet analysis means. This enables comprehensive analysis of the user's physical training data, emotional state, and information regarding dietary choices, and provides personalized feedback and advice.

[0801] A "video acquisition device" is a device that captures the user's physical movements in real time and acquires the video data.

[0802] "Image analysis means" refers to an analysis device used to identify the position of each joint in the user's body from acquired video data.

[0803] A "difference measurement device" is a device that measures the difference between the analyzed user's posture data and pre-registered ideal posture data, and generates difference data.

[0804] A "feedback generation means" is a device that generates natural language feedback for the user based on the difference data obtained by the difference measurement means.

[0805] An "emotion analysis device" is a device that analyzes the user's emotional state from their facial expressions and voice, and adjusts the generated feedback content according to that emotion.

[0806] A "meal analysis device" is a device that analyzes images of meals taken by users and evaluates their nutritional value.

[0807] A "presentation means" is a device that presents generated feedback or advice to the user, and performs functions such as display or audio output.

[0808] The fitness support system of the present invention is configured as a multi-functional platform including emotion analysis and dietary analysis to optimize the user's training effectiveness. The user launches a dedicated application using a terminal and acquires training video through the terminal's camera. This video data is transmitted to a server via the internet.

[0809] The server uses the received video data to identify the joint positions of the user's body through image analysis. Simultaneously with this analysis, a difference measurement means measures the difference between the user's posture data and pre-stored ideal posture data, generating difference data. Subsequently, a feedback generation means generates natural language feedback tailored to the user's training progress. During this process, an emotion analysis means analyzes the user's facial expressions and voice to determine their emotional state and adjusts the feedback content accordingly.

[0810] Furthermore, this system includes a meal analysis function that analyzes images of meals taken by the user on a server. It evaluates the nutritional value from the analysis results and generates advice on meal choices based on the user's emotional state.

[0811] For example, when a user performs squats at home, the device's camera records the movement. This video is sent to a server, which analyzes the user's knee and hip positioning and generates feedback such as, "It would be good to point your knees slightly outward." In another example, a picture of a salad taken by a user at their dinner table is analyzed by the server, and advice such as, "Adding avocado is good for stress reduction," is provided.

[0812] An example of a prompt using a generative AI model is: "Generate a suitable feedback message when the user is tired. Example: 'You're doing great. It's okay to take a break.'" This allows the user to receive comprehensive support from both a physical and emotional perspective.

[0813] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0814] Step 1:

[0815] The user launches a fitness application on their device and prepares to begin training using the camera. Inputs include the user's selected training mode and the press of a start button. Outputs include the activation of the camera device and the display of the training mode on the user's screen. This completes the process of preparing to begin training.

[0816] Step 2:

[0817] The device's camera captures the user's training movements in real time, acquiring video data. The input is the user's physical movements. This video data is compressed and sent to the server. The output is the captured video data, which is also sent to the server in real time. This allows the necessary data to be collected for analysis.

[0818] Step 3:

[0819] The server processes the received video data using image analysis to identify the position of each joint in the user's body. The input is the transmitted video data. The output is data on the joint positions. Through data analysis, the user's posture can be precisely understood.

[0820] Step 4:

[0821] The server measures the difference between identified joint position data and ideal posture data using a difference measurement device. The inputs are joint position data and ideal posture data. The output is the difference data between these two sets of data. This process allows for a quantitative evaluation of the user's training effectiveness.

[0822] Step 5:

[0823] The server uses a feedback generation mechanism based on differential data to create natural language feedback. The input is differential data. An emotion analysis mechanism identifies the user's emotional state and adjusts the feedback content accordingly. The output is an emotion-sensitive feedback message. This feedback enhances the user's training motivation.

[0824] Step 6:

[0825] The server sends the generated feedback to the terminal, which then presents it to the user. The input is the generated feedback message. Output from the terminal includes screen displays and audio feedback. This allows the user to receive feedback in real time and modify their training.

[0826] Step 7:

[0827] The user takes a picture of their meal with their device, and the device sends this data to the server. The input is the image of the meal taken by the user. The output is the captured image data sent to the server. Through this process, data for meal analysis is collected.

[0828] Step 8:

[0829] The server analyzes the received images using a meal analysis system and evaluates their nutritional value. The input is the transmitted meal image data. The output is evaluation data regarding the nutritional value of the meal. This allows the user to obtain detailed information about the meal they have consumed.

[0830] Step 9:

[0831] The server generates advice on food choices based on the results of a meal analysis and the user's emotional state. The inputs are nutritional value evaluation data for meals and the user's emotional state. The output is specific meal selection advice. This allows users to strive for a diet that considers both their health and emotional state.

[0832] Step 10:

[0833] The generated dietary advice is sent to the device and presented to the user. The input is the advice data. The output is the advice displayed to the user from the device. This process allows the user to adjust their diet and make healthier choices.

[0834] (Application Example 2)

[0835] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0836] Traditional fitness support systems have struggled to provide appropriate feedback tailored to each user's individual emotions and fitness level. Furthermore, the lack of features that integrate training and nutrition information and support online marketplace purchases hindered users' ability to maintain sustainable health. There is a need to address these issues and support users in engaging in fitness activities more effectively.

[0837] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0838] In this invention, the server includes a video acquisition means, an image analysis means, a difference measurement means, an emotion analysis means, a means for analyzing the user's emotional state using the emotion analysis means and adjusting the feedback content in the feedback generation means, and an online connection means for providing purchase support in the virtual market. This enables the provision of appropriate feedback according to the user's emotional state and support for purchasing fitness-related products through the online market.

[0839] A "video acquisition means" is a device that acquires video data in order to capture the user's body movements in real time.

[0840] "Image analysis means" refers to a technology that includes an algorithm for analyzing acquired video data and identifying the positions of human joints.

[0841] The "difference measurement method" is a function for evaluating training effectiveness by measuring the difference between the analyzed posture data and the ideal posture data that has been registered in advance.

[0842] A "feedback generation method" is a means of analyzing difference data using natural language processing technology in order to generate feedback to be provided to the user.

[0843] A "presentation means" is a device for providing the generated feedback to the user visually or audibly.

[0844] "Emotional analysis methods" refer to technologies that analyze a user's emotional state from their facial expressions and voice.

[0845] "Online connectivity" refers to internet-based communication functions that allow users to connect to a virtual marketplace and assist in purchasing fitness-related products.

[0846] In this invention, the system operates by integrating various means to support the user's fitness activities. The server captures the user's body movements in real time via "video acquisition means" for acquiring video. This data is analyzed by "image analysis means" to determine the position of the user's joints. The analyzed data is compared with ideal trainer posture data by "difference measurement means" to generate difference data.

[0847] Based on this differential data, the "feedback generation means" utilizes a generation AI model to create feedback tailored to the user. This feedback is then adjusted based on the user's emotional state, which is evaluated through the "emotion analysis means." The generated feedback is then quickly provided to the user as audio or text via the "presentation means" on the user's terminal.

[0848] Furthermore, the system incorporates "online connectivity," allowing users to easily purchase necessary training products through a virtual marketplace. This creates an environment where users can seamlessly integrate fitness activities with product purchases.

[0849] For example, when a user wears smart glasses and does push-ups, the camera captures the position of their hands and analyzes that information. Then, appropriate feedback is provided, such as "Try spreading your hands a little wider." When the user is tired, emotionally appropriate feedback is provided, such as "Let's slow down."

[0850] An example of a prompt message would be: "Based on the user's training data, generate feedback for posture improvement. Next, analyze his emotional data and adjust the feedback to match his emotions."

[0851] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0852] Step 1:

[0853] The user begins training using the camera built into the device. The device captures the user's movements in real time using video acquisition equipment and sends the video data to the server. The input is the user's video captured by the camera, and the output is the video data sent to the server.

[0854] Step 2:

[0855] The server analyzes the received video data using image analysis tools to identify the user's joint positions. The input here is the video data obtained in step 1, and the output is the identified joint position data. Specifically, the server analyzes each frame using a machine learning model and maps each part of the body.

[0856] Step 3:

[0857] The server uses a difference measurement mechanism to measure the difference between the analyzed joint position data and pre-registered ideal posture data. The input is the joint position data and ideal posture data obtained in step 2, and the output is the difference data. Specifically, a particular index is set, and the difference is quantified using linear algebra.

[0858] Step 4:

[0859] The server uses a generative AI model to generate natural language feedback based on the differential data. The input here is the differential data obtained in step 3, and the output is feedback in natural language. Specifically, prompt sentences are input to the generative AI model, which then generates the optimal feedback sentence.

[0860] Step 5:

[0861] The server further uses emotion analysis tools to determine the user's emotional state from their facial expressions and voice data. The input is the user's facial expressions and voice data, and the output is the user's emotional state data. Specifically, it uses speech recognition technology and facial expression recognition technology such as DeepFace to detect and analyze the state.

[0862] Step 6:

[0863] Based on the user's emotional state, the generated feedback is adjusted and presented to the user through the device's display methods. The input is the feedback obtained in step 4 and the emotional state data obtained in step 5, and the output is the adjusted feedback. Specifically, the tone and content of the feedback are fine-tuned to match the user's emotions and presented as text display or audio output.

[0864] Step 7:

[0865] Using the terminal's online connectivity, users can access a virtual marketplace and purchase fitness-related products based on their feedback. Inputs are user feedback and virtual marketplace access rights, while output is a list of purchasable products. Specifically, the terminal opens an internet browser and provides an interface for the user to select suggested products and proceed with the purchase.

[0866] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0867] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0868] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0869] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0870] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0871] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0872] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0873] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0874] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0875] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0876] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0877] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0878] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0879] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0880] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0881] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0882] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0883] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0884] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0885] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0886] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0887] The following is further disclosed regarding the embodiments described above.

[0888] (Claim 1)

[0889] Means of acquiring video,

[0890] Image analysis means for analyzing the posture of the human body acquired by the image acquisition means and identifying joint positions,

[0891] A difference measurement means that measures the difference between pre-registered ideal posture data and analyzed posture data,

[0892] A feedback generation means that generates natural language feedback based on the difference obtained by the difference measurement means,

[0893] A presentation means for presenting the generated feedback,

[0894] A system that includes this.

[0895] (Claim 2)

[0896] The system according to claim 1, wherein the feedback generation means generates audio data and the presentation means outputs the audio data as audio.

[0897] (Claim 3)

[0898] Furthermore, the system according to claim 1 is further equipped with a meal analysis means for analyzing images of meals and evaluating their nutritional value.

[0899] "Example 1"

[0900] (Claim 1)

[0901] A means of capturing video data,

[0902] Image analysis means for analyzing human body movement acquired by the imaging means and identifying joint positions,

[0903] A comparison method for measuring the difference between registered ideal exercise data and analyzed exercise data,

[0904] A generation means that generates advice in natural language based on the difference obtained by the comparison means,

[0905] A presentation means for displaying or playing back the generated advice,

[0906] A method for analyzing food images and evaluating the nutritional value of the food content,

[0907] A system that includes this.

[0908] (Claim 2)

[0909] The system according to claim 1, wherein the generation means generates audio information and the presentation means outputs the audio information as audio.

[0910] (Claim 3)

[0911] Furthermore, the system according to claim 1, which generates advice using a generative AI model.

[0912] "Application Example 1"

[0913] (Claim 1)

[0914] Video acquisition device,

[0915] A pixel analysis device analyzes the movement of the workpiece acquired by the video acquisition device and identifies the joint positions,

[0916] A difference measurement device that measures the difference between pre-registered ideal operation data and analyzed operation data,

[0917] A feedback generation device that generates natural language feedback based on the difference obtained by the difference measurement device,

[0918] A display device that presents the generated feedback,

[0919] A system that includes this.

[0920] (Claim 2)

[0921] The system according to claim 1, wherein the feedback generation device generates audio data and outputs the audio data as audio using the display device.

[0922] (Claim 3)

[0923] Furthermore, the system according to claim 1 is equipped with an analysis device that generates improvement suggestions regarding the movement of the workpiece.

[0924] "Example 2 of combining an emotion engine"

[0925] (Claim 1)

[0926] Means of acquiring video,

[0927] Image analysis means for analyzing the posture of the human body acquired by the image acquisition means and identifying joint positions,

[0928] A difference measurement means that measures the difference between pre-registered ideal posture data and analyzed posture data,

[0929] A feedback generation means that generates natural language feedback based on the difference obtained by the difference measurement means,

[0930] A sentiment analysis tool that analyzes the user's emotional state and adjusts the feedback accordingly,

[0931] A meal analysis method that analyzes images of meals taken and evaluates their nutritional value,

[0932] A presentation means for presenting the generated feedback,

[0933] A system that includes this.

[0934] (Claim 2)

[0935] The system according to claim 1, wherein the feedback generation means generates audio data and the presentation means outputs the audio data as audio.

[0936] (Claim 3)

[0937] The system according to claim 1, wherein the emotion analysis means generates advice as motivation regarding food selection.

[0938] "Application example 2 of combining emotional engines"

[0939] (Claim 1)

[0940] Means of acquiring video,

[0941] Image analysis means for analyzing the posture of the human body acquired by the image acquisition means and identifying joint positions,

[0942] A difference measurement means that measures the difference between pre-registered ideal posture data and analyzed posture data,

[0943] A feedback generation means that generates natural language feedback based on the difference obtained by the difference measurement means,

[0944] A presentation means for presenting the generated feedback,

[0945] Emotion analysis methods,

[0946] The emotion analysis means analyzes the user's emotional state and the means adjusts the content of the feedback in the feedback generation means,

[0947] Online connectivity methods to support purchases in the virtual market,

[0948] A system that includes this.

[0949] (Claim 2)

[0950] The system according to claim 1, wherein the feedback generation means generates audio data and the presentation means outputs the audio data as audio.

[0951] (Claim 3)

[0952] Furthermore, the system according to claim 1, further comprising a meal analysis means for analyzing images of meals and evaluating their nutritional value. [Explanation of Symbols]

[0953] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. Means of acquiring video, Image analysis means for analyzing the posture of the human body acquired by the image acquisition means and identifying joint positions, A difference measurement means that measures the difference between pre-registered ideal posture data and analyzed posture data, A feedback generation means that generates natural language feedback based on the difference obtained by the difference measurement means, A presentation means for presenting the generated feedback, A system that includes this.

2. The system according to claim 1, wherein the feedback generation means generates audio data and the presentation means outputs the audio data as audio.

3. Furthermore, the system according to claim 1 is further equipped with a meal analysis means for analyzing images of meals and evaluating their nutritional value.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A