Method and system for providing exercise posture feedback by using vision-language model
A vision language model system analyzes user exercise movements to provide real-time, personalized feedback, improving exercise performance and program adaptation for musculoskeletal disorders without frequent hospital visits.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- EVEREX
- Filing Date
- 2025-10-27
- Publication Date
- 2026-05-07
AI Technical Summary
There is a need for personalized feedback on exercise movements to improve the accuracy and efficiency of exercise performance, particularly for individuals with musculoskeletal disorders, without requiring frequent hospital visits.
A method and system using a vision language model to analyze user exercise movements, generate feedback, and update exercise programs based on user information, including natural language processing and image analysis, to provide customized feedback in real-time.
Enhances exercise performance accuracy and efficiency by providing personalized feedback on user movements, allowing for flexible exercise program updates based on current user state and exercise history.
Smart Images

Figure KR2025017208_07052026_PF_FP_ABST
Abstract
Description
Method and System for Providing Motor Motion Feedback Using a Vision Language Model
[0001] The present invention relates to a method and system for providing feedback on a user's movement using a vision language model.
[0002] With the recent rapid advancement of artificial intelligence (AI) technology, generative AI based on Vision-Language Models (VLMs), capable of simultaneously understanding and processing visual information such as images and videos as well as human language, is garnering attention. In particular, by integratively analyzing natural language and visual information, Vision-Language Models are demonstrating advanced technological capabilities that go beyond existing text-based question-and-answer large-scale language models. They are capable of understanding users' actual behaviors or actions and generating customized responses.
[0003] Beyond simple interactive applications, these technologies are being utilized in healthcare fields such as sports, fitness, daily exercise, and rehabilitation management. In particular, there is a growing demand for technologies that provide personalized healthcare to users from exercise videos captured through various electronic devices, such as smart devices, mobile cameras, and wearable sensors.
[0004] For example, vision language models can be effectively utilized for the management and rehabilitation of musculoskeletal disorders. Musculoskeletal disorders refer to pain or injury occurring in the musculoskeletal system, including muscles, nerves, tendons, ligaments, bones, and surrounding tissues. As a principle, the treatment of musculoskeletal disorders should begin with less invasive procedures; non-pharmacological conservative treatments (e.g., exercise therapy and education, cognitive therapy, or relaxation therapy) should be implemented first, followed by pharmacological treatment and surgical treatment in sequence. Treatment guidelines strongly recommend non-pharmacological conservative treatment for musculoskeletal disorders, and active research on methods for implementing such treatments is being conducted, primarily in the United States and Europe. However, since continuous treatment and rehabilitation are crucial for non-pharmacological conservative treatment, the need for patients to visit the hospital frequently poses a significant burden.
[0005] To address these issues, there is a need to provide a personalized feedback service on users' exercise movements remotely using a vision language model.
[0006] The present invention is intended to provide a method and system for providing exercise motion feedback using a vision language model capable of providing customized feedback on a user's exercise motion.
[0007] Specifically, the present invention aims to provide a method and system for providing exercise motion feedback using a vision language model capable of natural language processing and image analysis, which analyzes a user's exercise motion based on user information and provides customized feedback information according to the analysis results.
[0008] Furthermore, the present invention aims to provide a method and system for providing exercise motion feedback using a vision language model, which generates a prompt based on a user's response and processes the prompt as input to the vision language model to update the user's exercise program.
[0009] To solve the problem described above, the present invention proposes a method that utilizes a vision language model to analyze a user's exercise movements and generate and provide feedback on the exercise movements. The method for providing exercise movement feedback using a vision language model according to the present invention may include the steps of: receiving an exercise video of a user from an electronic device; collecting user information related to the user; generating a prompt that enables analysis of the exercise video based on the user information using the user information and the exercise video; inputting the prompt and the exercise video into a pre-trained vision language model to perform an analysis of the user's exercise movements through the vision language model; obtaining feedback information regarding the user's exercise movements generated from the vision language model based on the analysis results of the user's exercise movements; and providing the feedback information to the electronic device.
[0010] Furthermore, the step of generating the above prompt may include the step of evaluating the user's exercise expectation level based on the user information, and the step of generating a first prompt that enables an analysis of the user's exercise movements to be performed according to the exercise expectation level.
[0011] Furthermore, in the step of evaluating the user's exercise expectation level, the user's exercise expectation level for at least one of a plurality of body parts can be evaluated based on at least one of the medical information and exercise history information included in the user information.
[0012] Furthermore, the step of generating the first prompt may include, for at least one of the plurality of body parts, specifying at least one body part satisfying a preset condition among the plurality of body parts based on a determined exercise expectation level of the user, setting a weight for the specified at least one body part, and generating the first prompt to analyze the user's exercise movement related to any one of the plurality of body parts based on the set weight.
[0013] Furthermore, the step of performing an analysis of the user's movement may include the step of extracting key points corresponding to each of a plurality of pre-set joint points in a specific object corresponding to the user included in the movement video using the previously trained vision language model, and the step of analyzing the relative positional relationship between the key points and, based on the analysis of the positional relationship, performing an analysis of the user's movement related to any one of the plurality of body parts.
[0014] Furthermore, the step of performing an analysis of the user's exercise movements may include: a step of calculating a performance score for the user's exercise movements based on the analysis results of the positional relationships; a step of comparing the performance score for the past exercise movements and the performance score for the user's exercise movements using the exercise history analysis results for past exercise movements performed in the past in relation to the user's exercise movements included in the exercise history information; and a step of generating a trend analysis result related to the user's exercise movements based on the performance score.
[0015] Furthermore, the step of generating the above prompt further includes the step of generating a second prompt that requests that feedback information regarding the user's exercise motion be generated according to the analysis result regarding the exercise motion, and in the step of generating the second prompt, the second prompt may be generated to generate the feedback information based on the analysis result including the trend analysis result.
[0016] Furthermore, the feedback information includes text feedback information containing a feedback message regarding the analysis result and voice feedback information corresponding to the text feedback information, and the feedback message may include at least one of a correction instruction for the exercise movement, the performance score for the exercise movement, comparison information with the exercise history analysis result, and alternative movement information related to the exercise movement, based on the analysis result.
[0017] Furthermore, the step of calculating the performance score may include a step of calculating the similarity between the relative positional relationship between the key points according to the user's exercise movements and the positional relationship corresponding to a preset correct posture, and a step of calculating the performance score according to the cumulative time in which the similarity satisfies a preset condition.
[0018] Furthermore, the method may further include a step of updating an exercise program performed by the user using the above-described vision language model, wherein the updating step may include: receiving voice data corresponding to voice received through a microphone provided in the electronic device; receiving user survey response data for at least one user survey provided in the electronic device; generating user response information using at least one of the voice data and the survey response data; generating a third prompt to update an exercise program assigned to a user account using the user response information; and inputting the third prompt into the vision language model to update the exercise program through the vision language model.
[0019] Meanwhile, the exercise motion feedback provision system using a vision language model according to the present invention includes a communication unit that receives an exercise video of a user from an electronic device and a control unit that collects user information related to the user. The control unit generates a prompt that enables analysis of the exercise video based on the user information using the user information and the exercise video, inputs the prompt and the exercise video into a pre-trained vision language model to perform an analysis of the user's exercise motion through the vision language model, obtains feedback information regarding the user's exercise motion generated based on the analysis result of the user's exercise motion from the vision language model, and provides the feedback information to the electronic device.
[0020] Meanwhile, the program is executed by one or more processes in an electronic device and is stored on a computer-readable recording medium, and the program may include instructions for performing the steps of: receiving a user’s exercise video from the electronic device; collecting user information related to the user; generating a prompt that enables analysis of the exercise video based on the user information using the user information and the exercise video; inputting the prompt and the exercise video into a pre-trained vision language model and performing an analysis of the user’s exercise movements through the vision language model; obtaining feedback information regarding the user’s exercise movements generated from the vision language model based on the analysis results of the user’s exercise movements; and providing the feedback information to the electronic device.
[0021] The method and system for providing exercise motion feedback using a vision language model according to the present invention can improve the accuracy and efficiency of exercise performance by using a video of a user's exercise performance as input, automatically performing an analysis of the exercise motion, and providing customized feedback information to the user in real time.
[0022] Furthermore, the method and system for providing exercise motion feedback using a vision language model according to the present invention can provide personalized feedback based on the user-customized motion analysis results by analyzing the user's exercise motions while reflecting the level of exercise expectation.
[0023] Furthermore, the method and system for providing exercise motion feedback using a vision language model according to the present invention generates a prompt reflecting a user response received from the user after performing an exercise, and processes it as input to the vision language model, thereby enabling the exercise program to be flexibly updated according to the current user state through the vision language model.
[0024] FIG. 1 is a conceptual diagram illustrating a motion motion feedback provision system using a vision language model according to the present invention.
[0025] FIG. 2 is a flowchart illustrating a method for providing motion feedback using a vision language model according to the present invention.
[0026] FIG. 3a is a flowchart illustrating the process of generating a first prompt that enables analysis of an exercise video using user information and an exercise video according to the present invention.
[0027] FIG. 3b is a flowchart illustrating the process of generating different first prompts according to user types according to the present invention.
[0028] FIGS. 4a to 4c are conceptual diagrams for explaining the process of analyzing a user's exercise movements according to the present invention.
[0029] FIGS. 5 and 6 are conceptual diagrams for explaining the process of generating feedback information according to the present invention and providing it to an electronic device.
[0030] FIGS. 7a and FIGS. 7b are conceptual diagrams illustrating the process of updating an exercise program based on a user's response according to the present invention.
[0031] FIG. 8 is a block diagram illustrating a computing system in which the present invention can be implemented.
[0032] FIGS. 9 and FIGS. 10 are block diagrams illustrating an embodiment of a computing device according to the present invention.
[0033] The present invention relates to a method and system for providing exercise motion feedback using a vision language model. More specifically, the present invention relates to a method and system for providing customized feedback information regarding a user's exercise motion using a vision language model capable of natural language understanding and image analysis.
[0034] The vision language model according to the present invention may refer to an intelligent system capable of autonomously performing specific tasks by understanding and processing visual information and linguistic information in combination without human intervention, based on a generative artificial intelligence model. Specifically, the vision language model is a multimodal generative model capable of processing visual inputs such as images and videos together with linguistic inputs such as text and voice, and can analyze video information related to a user's movement and generate natural language-based feedback information based on the results to provide to an electronic device.
[0035] The “feedback information” according to the present invention may refer to various information provided to an electronic device to assist and improve a user’s exercise performance, such as a feedback message based on user information and a video of the user’s exercise, a performance score, and a correct exercise video corresponding to the exercise movement performed by the user. In this case, the “feedback message” according to the present invention may include information related to at least one of an accuracy evaluation of an exercise movement, a request for modification of an exercise movement, and guidance on alternative movements and exercise programs based on the user's state.
[0036] A user (U, or patient) according to the present invention can perform an exercise program according to prescription information prescribed by a medical institution for the indications of the user (U) through an application or webpage provided by the feedback providing system (100) according to the present invention.
[0037] A user (or patient, U) may possess a user account registered in the feedback provision system (100) according to the present invention. For convenience of explanation, the account of a user who is a patient is referred to as a "user account (or patient account)." The "account" described above may be created through a page linked to the feedback provision system (100). Alternatively, the "account" may be created on at least one other server (e.g., a medical staff server) linked to the feedback provision system (100) according to the present invention. Accordingly, in this specification, without distinguishing the server where the account was issued, all accounts based on the feedback provision system (100) according to the present invention are referred to as "accounts already registered in the feedback provision system (100) according to the present invention."
[0038] Meanwhile, a doctor can issue a prescription related to rehabilitation treatment to a user (U) through a doctor terminal. At this time, the doctor (D) may possess a doctor account already registered in the feedback provision system (100) according to the present invention. In this specification, an electronic device logged in with a doctor account is referred to as a doctor terminal. As an example, the feedback provision system (100) according to the present invention may receive medical information prescribed by the doctor (D) to the user (U) by linking with a medical staff server.
[0039] Furthermore, according to the present invention, a user can execute an application of an electronic device (10) to perform exercise according to an exercise program assigned to a user account and receive feedback information regarding the user's exercise movements. At this time, the present invention can collect various information using a camera, a microphone, and a plurality of different sensors equipped in the electronic device, and process the collected information using a vision language model. Specifically, the present invention can collect exercise video of the exercise movements performed by the user using a camera equipped in the electronic device, and collect the user's voice data through a microphone equipped in the electronic device.
[0040] Furthermore, the present invention can analyze the user's exercise movements included in collected exercise videos based on a vision language model and user information. Here, "user information" may include various information related to the user, such as the user's medical information (or prescription information, disease information), exercise history information, user response information, exercise program information, and user type information.
[0041] The present invention focuses on feedback regarding exercise movements related to the treatment of “indications related to musculoskeletal disorders,” but is not necessarily limited thereto. As an example, the feedback regarding exercise movements described in the present invention may be for exercise intended for the treatment of a user with various diseases (e.g., cancer, diabetes, hypertension, etc.).
[0042] Furthermore, the present invention may provide feedback information regarding the user's exercise movements necessary for health promotion in daily life, rather than rehabilitation exercises for therapeutic purposes related to the user's indications. For example, the exercise according to the present invention is not limited to any specific purpose and may be an exercise performed for various purposes, such as rehabilitation exercises, fitness exercises, ball sports, or dance exercises, for therapeutic purposes, health promotion purposes, or beauty purposes. Moreover, the feedback on exercise movements according to the present invention may refer to feedback for various exercises, such as rehabilitation exercises, fitness exercises, ball sports, or dance exercises, for therapeutic purposes, health promotion purposes, or beauty purposes, and can be understood as not being limited to a specific category of exercise. Additionally, there is no limitation on the type of exercise according to the present invention, nor is there a limitation on the location of the exercise, such as indoor or outdoor exercise.
[0043] In the foregoing, motor motion feedback using a vision language model according to the present invention has been generally described, and this can be implemented by a feedback providing system described below. Below, with reference to FIG. 1, a motor motion feedback providing system using a vision language model according to the present invention will be described in detail. FIG. 1 is a conceptual diagram for explaining a motor motion feedback providing system using a vision language model according to the present invention.
[0044] As illustrated in FIG. 1, a motion motion feedback providing system using a vision language model according to the present invention (hereinafter referred to as the “feedback providing system,” 100) may include at least one of a communication unit (110), a storage unit (120), and a control unit (130). At this time, the feedback providing system (100) according to the present invention is not limited to the components described above and may further include components that perform the same or similar roles as the functions described in the specification. Meanwhile, the feedback providing system (100) according to the present invention may be implemented as an application or software. The feedback providing system (100) implemented as software in this manner may be downloaded via a program (e.g., Play Store) that allows the application to be downloaded on an electronic device (10), or implemented via an initial installation program on an electronic device (10). In this case, the communication unit (110), the storage unit (120), and the control unit (130) according to the present invention may be utilized as components of the electronic device (10). In the present invention, the electronic device (10) can be understood to mean an application installed on the electronic device (10). Such an application (or software) can be understood as a component of the feedback providing system (100) according to the present invention.
[0045] In the present invention, the electronic device (10) may also be named a "mobile terminal" or a "user terminal," and the electronic device (10) described in this specification may include a mobile phone, a smartphone, a smart TV, a laptop computer, a digital broadcasting terminal, a PDA (personal digital assistants), a PMP (portable multimedia player), a navigation device, a slate PC, a tablet PC, an ultrabook, a wearable device (e.g., a smartwatch, a smart glass, a head-mounted display), etc.
[0046] More specifically, the electronic device (10) according to the present invention is not limited to an electronic device with an application activated, but may also mean an electronic device connected to an electronic device with an application activated. As an example, based on the fact that the electronic device (10) according to the present invention is a smartphone, the electronic device (10) may also mean a smart TV connected to said smartphone.
[0047] Meanwhile, the feedback providing system (100) may exist inside a server (hereinafter referred to as the server) built to perform a specific purpose (e.g., providing feedback information regarding exercise movements), or it may exist as a separate device from the server. When the feedback providing system (100) exists inside the server, the feedback providing system (100) according to the present invention may provide feedback information regarding exercise movements through at least one component among a communication unit (110), a storage unit (120), and a control unit (130) located inside the server, or through a module that performs a function similar to each of the above components. In this case, the application may provide feedback information regarding exercise movements on an electronic device (10) on which the application is installed through communication with the server. Furthermore, the feedback providing system (100) according to the present invention may provide feedback information regarding exercise movements according to the present invention to the electronic device (10) by linking with a plurality of different external servers.
[0048] Meanwhile, the communication unit (110) of the feedback providing system (100) according to the present invention may be connected to an electronic device (10), a VLM server (140), a central server, a device, and at least one network via a wireless or wired network, and configured to receive or transmit overall data and information necessary for the operation of the feedback providing system (100) according to the present invention.
[0049] The communication unit (110) can receive user information corresponding to a user account logged into the electronic device (10). Additionally, the communication unit (110) can receive at least one of collected user voice data and exercise video using a microphone (11), a camera (12), and a sensor unit (13) provided in the electronic device (10). Here, the sensor unit (13) may include at least one sensor among an infrared sensor, a LiDAR sensor, an accelerometer, an illuminance sensor, a proximity sensor, a position sensor, a face recognition sensor, an iris scanner, a heart rate sensor, a touch sensor, and a pressure sensor.
[0050] At this time, the feedback providing system (100) may further include a module for converting voice data into text. For example, in order to convert voice data received from an electronic device (10) into text, the present invention may include a conversion module comprising at least one of a Hidden Markov Model (HMM), a Gaussian / Deep Neural Net (GMM / DNN), a Weighted Finite State Transducer (WFST), a Connectionist Temporal Classification (CTC), a Recurrent Neural Network Transducer (RNN-T), an Attention-based Seq2Seq, a Self-Supervised Learning-based model, and a Transformer-based model.
[0051] The communication unit (110) may include at least one communication module capable of wireless communication and wired communication between the feedback providing system (100) and the communication target. Additionally, the communication unit (110) may include a communication module that connects the feedback providing system (100) to at least one network.
[0052] Meanwhile, the communication unit (110) can support various communication methods depending on the communication standard of the communicating device. For example, the communication unit (110) may be configured to perform communication using at least one of the following technologies: WLAN (Wireless LAN), Wi-Fi (Wireless-Fidelity), Wi-Fi (Wireless Fidelity) Direct, DLNA (Digital Living Network Alliance), WiBro (Wireless Broadband), WiMAX (World Interoperability for Microwave Access), HSDPA (High Speed Downlink Packet Access), HSUPA (High Speed Uplink Packet Access), LTE (Long Term Evolution), LTE-A (Long Term Evolution-Advanced), 5G (5th Generation Mobile Telecommunication), Bluetooth (Bluetooth™ Frequency Identification), Infrared Communication (Infrared Data Association; IrDA), UWB (Ultra-Wideband), ZigBee, NFC (Near Field Communication), Wi-Fi Direct, and Wireless USB (Wireless Universal Serial Bus).
[0053] Next, the storage unit (120) may be configured to store various information related to the present invention. In the present invention, the storage unit (120) may be provided in the feedback providing system (100) itself, or alternatively, at least a part of the storage unit (120) may mean a database (Database: DB, 200).
[0054] The storage unit (120) may include one or more non-transient computer-readable storage media that can be read and / or accessed by at least one processor. One or more computer-readable storage media may include volatile and / or non-volatile storage components such as optical, magnetic, organic, or other memory or disk storage devices. In some examples, the storage unit (120) may be implemented using a single physical device (e.g., one optical, magnetic, organic, or other memory or disk storage device), whereas in other examples, the storage unit (120) may be implemented using multiple physical devices.
[0055] The storage unit (120) may include computer-readable instructions and additional data. The storage unit (120) may include a storage necessary to perform at least some of the methods and techniques described herein and / or at least some of the functions of the device and network.
[0056] Furthermore, at least a portion of the storage unit (120) may be a cloud storage or a cloud server. That is, the storage unit (120) is sufficient as a space where information necessary for the operation of the feedback providing system (100) according to the present invention is stored, and it can be understood that there are no restrictions on the physical space. Accordingly, the storage unit (120) and the database (200) may be used interchangeably below.
[0057] Specifically, the storage unit (120) may store user information corresponding to a user account logged into the electronic device (10). For example, the storage unit (120) may store user information including at least one of medical information, exercise history information, user response information, exercise program information, and user type information. As an example, the medical information according to the present invention may include at least one of the user's age, the user's gender, the user's medical history, prescription information, and treatment plan. Additionally, the exercise history information according to the present invention may include at least one of the date of exercise performance, the composition of exercise movements included in the exercise program performed, the results of motion analysis for each exercise movement, and the results of exercise history analysis.
[0058] Furthermore, the user response information according to the present invention may include at least one of the user's voice data received from a microphone (13) provided in the electronic device (10) and at least one response data to a survey provided to the electronic device (10) in relation to the exercise action performed by the user.
[0059] In addition, the exercise program information according to the present invention may include at least one of the name of each of a plurality of exercise items constituting an exercise program assigned to a user account, the number of exercises, the timing of the exercises, and information on the difficulty of the exercises, as well as an exercise video and an exercise description corresponding to each of the plurality of exercise items.
[0060] Furthermore, the user type information according to the present invention may include information related to at least one of the user's exercise purpose, exercise level, and restricted exercise body parts.
[0061] Meanwhile, user authentication information may be stored in the storage unit (120). Here, “user authentication information” may refer to information used in a user authentication process performed to log in to a user account on an electronic device (10). As an example, user authentication information may be various, such as i) ID, ii) password, iii) password pattern, iv) user’s fingerprint authentication information, v) face authentication information, vi) voice authentication information, vii) iris authentication information, viii) vein authentication information, etc., set by the user.
[0062] Commands for the operation of the prompt generation unit (131) may be stored in the storage unit (120) according to the present invention. The prompt generation unit (131) may refer to a module that generates a prompt to be input into the vision language model (132) based on at least one of user information, exercise performance results, and user response information. The method for generating a prompt to be input into the vision language model (132) according to the present invention may be very diverse, and in this specification, any module capable of generating a prompt to be input into the vision language model (132) is not limited to its type or method. The control unit (130) according to the present invention may further include at least one of a module and an algorithm that perform the same function as the prompt generation unit (131). At this time, the prompt generation unit (131) may generate a prompt to be input into the vision language model (132) through at least one artificial intelligence model or module.
[0063] For example, the prompt generation unit (131) may include at least one of a large language model based on T5 (Text-to-Text Transfer Transformer), BART (Bidirectional and Auto-Regressive Transformer), GPT (Generative Pre-trained Transformer), or LLaMA (Language Model for Many Applications), a rule-based template matching algorithm, a conditional prompting technique, a contextual embedding selection module, or a few-shot prompt generator. At this time, the prompt generation unit (131) according to the present invention is not limited to the models, algorithms, or modules described above, and the feedback providing system (100) according to the present invention may further include models, algorithms, or modules that perform the same function as the prompt generation unit (131).
[0064] Commands for the operation of a vision language model (132) may be stored in the storage unit (120) according to the present invention. The vision language model (132) may refer to an artificial intelligence model capable of performing analysis of a user's exercise movements included in the exercise video, generation of feedback information, and exercise program updates based on an exercise video received from an electronic device (10) and user information collected from the storage unit (120, or database, 200). The method for performing analysis of a user's exercise movements, generation of feedback information, and exercise program updates according to the present invention may be very diverse, and the present specification does not limit the method for performing analysis of a user's exercise movements, generation of feedback information, and exercise program updates.
[0065] That is, in the present invention, the vision language model (132) is not limited to any type or method as long as it is an artificial intelligence model capable of analyzing the user's exercise movements included in the exercise video, generating feedback information, and updating the exercise program based on the exercise video and user information received from the electronic device (10). The control unit (130) according to the present invention may further include at least one of a module and an algorithm that perform the same function as the vision language model (132). For example, the vision language model (132) may include at least one artificial intelligence model among Gemini-2.5, GPT-4o, DeepSeekVL2, Gemma 3, LLaMA3.2Vision, Flamingo, BLIP (Bootstrapping Language-Image Pretraining), CLIP (Contrastive Language-Image Pretraining), Kosmos-1, and Open Flamingo. At this time, the vision language model (132) according to the present invention is not limited to the model described above, and the feedback providing system (100) according to the present invention may further include a model having the same function as the vision language model (132).
[0066] Next, the control unit (130) may be configured to control the overall operation of the feedback providing system (100) related to the present invention. The control unit (130) may process signals, data, information, etc. that are input or output through the components described above, or provide or process appropriate information and functions to the user.
[0067] The control unit (130) can control the output of a service page for a feedback service regarding exercise movements through a display unit (or touchscreen) provided in the electronic device (10). Such a service page may be output on the electronic device (10) through an application or web page installed on the electronic device (10). The service page is a page linked to the feedback providing system (100) according to the present invention and is configured to be controlled by the feedback providing system (100) according to the present invention.
[0068] Furthermore, if the service page is provided in the form of an application, the service page can be controlled by the CPU (Central processing unit) of the electronic device (10) on which the application is installed. In this case, the CPU of the electronic device (10) can provide customized feedback information regarding the user's exercise movements based on information provided by the feedback providing system (100) according to the present invention.
[0069] Meanwhile, the control unit (130) can collect user information from the storage unit (120). Furthermore, the control unit (130) can collect data related to the user's exercise movements by using at least one of the sensor unit (11), camera (12), and microphone (13) provided in the electronic device (10). For example, the control unit (130) can collect exercise video of the user's exercise movements by using the camera (12) provided in the electronic device (10). In addition, the control unit (130) can collect voice data of the user during exercise movements by using the microphone (13) provided in the electronic device (10).
[0070] Furthermore, the control unit (130) can generate a first prompt that enables analysis of the exercise video based on the user information using the user information and the exercise video. Specifically, the control unit (130) can evaluate the user's exercise expectation level for at least one of a plurality of body parts based on at least one of the medical information and exercise history information included in the user information.
[0071] The control unit (130) can generate a first prompt to perform an analysis of the user's exercise movements according to the exercise expectation level. Specifically, the control unit (130) can specify at least one body part among a plurality of body parts that satisfies a preset condition based on the determined user's exercise expectation level for at least one of a plurality of body parts. The control unit (130) can set a weight for the specified at least one body part and generate a first prompt to perform an analysis of the user's exercise movements related to any one of the plurality of body parts based on the set weight.
[0072] Furthermore, the control unit (130) can process the generated first prompt as an input to the vision language model (132) to enable user-customized exercise motion analysis through the vision language model (132). At this time, although the present invention describes the vision language model (132) being included and operated in the control unit (130), it is not limited thereto and can also be understood as using a vision language model (132) included in a VLM server (140) that exists separately from the feedback providing system (100). The VLM server (140) may include at least one vision language model (132) among Gemini-2.5, GPT-4o, DeepSeekVL2, Gemma 3, LLaMA3.2Vision, Flamingo, BLIP (Bootstrapping Language-Image Pretraining), CLIP (Contrastive Language-Image Pretraining), Kosmos-1, and Open Flamingo. At this time, the vision language model (132) included in the VLM server (140) according to the present invention is not limited to the model described above, and the VLM server (140) according to the present invention may further include a model having the same function as the vision language model (132).
[0073] For convenience of explanation, the VLM server (140) and the vision language model (132) are used interchangeably below, and it can be understood that the control unit (130) uses the vision language model (132) as using at least one vision language model included in the VLM server (140).
[0074] Specifically, the vision language model (132) can analyze the user's exercise performance status from an exercise video containing exercise movements based on user information, and perform joint position extraction and motion tracking for analysis. To this end, the vision language model (132) may include a preprocessing algorithm that divides the input exercise video into frames and extracts position information of human joints from each frame, and can evaluate the user's movements based on the joint position information extracted from the exercise video.
[0075] The method for evaluating a user's movement based on joint position information extracted from a movement video according to the present invention may be very diverse, and the present specification does not limit the method for evaluating a user's movement based on joint position information extracted from a movement video. In the present invention, the vision language model (132) is not limited in type or method as long as it is a model capable of evaluating a user's movement based on joint position information extracted from a movement video. The control unit (130) according to the present invention may further include at least one of a module and an algorithm that perform the same function as the vision language model (132). For example, the vision language model (132) may include a pre-trained motion analysis model, and through the pre-trained motion analysis model, the user's movement can be evaluated based on joint position information extracted from a movement video.
[0076] The motion analysis model according to the present invention is an artificial intelligence model trained using a learning data set that includes position information for joint points, and can analyze the exercise posture of a user (U) from the exercise video data to be analyzed. In the present invention, the vision language model (132) is described as generating an analysis result for the user's exercise motion by including a pre-trained motion analysis model, but is not limited thereto. For example, the control unit (130) can train the vision language model (132) to analyze the user's exercise motion by linking with a motion analysis server (not shown) that includes a motion analysis model. That is, the control unit (130) can analyze the user's exercise motion using the vision language model (132) that includes a pre-trained motion analysis model, and the vision language model (132) can also analyze the exercise motion using a motion analysis model included in a separate motion analysis server.
[0077] Furthermore, the control unit (130) can generate a second prompt that generates feedback information related to the user's exercise motion by using the motion analysis results for the user's exercise motion. Specifically, the control unit (130) can generate a second prompt that generates feedback information related to the user's exercise motion by using the prompt generation unit (131), and input the generated second prompt into the vision language model (132).
[0078] The vision language model (132) can generate feedback information regarding the user's exercise movements based on the second prompt. Here, the feedback information may include at least one of an accuracy evaluation of the exercise movements, guidance information for the exercise movements, alternative movement information and exercise program information based on the user's state, a performance score, and a feedback message.
[0079] Furthermore, the control unit (130) may provide the generated feedback information to the electronic device (10). For example, the control unit (130) may output the generated text feedback information to a display unit provided in the electronic device (10). As another example, the control unit (130) may output the generated voice feedback information as sound (or voice) through a speaker provided in the electronic device (10). There may be a wide variety of methods for outputting the generated feedback information as sound (or voice) through a speaker provided in the electronic device (10) according to the present invention, and the present specification is not limited to any type or method as long as it is a module capable of outputting the generated feedback information as sound (or voice) through a speaker provided in the electronic device (10).
[0080] The control unit (130) according to the present invention may further include at least one of a module and an algorithm for outputting feedback information generated through a speaker provided in the electronic device (10) as sound (or voice). Specifically, the control unit (130) may further include a module for converting text into voice. For example, the control unit (130) may use a Text-to-Speech (TTS) module to output text included in the feedback information as sound (or voice) through a speaker provided in the electronic device (10).
[0081] Furthermore, the control unit (130) can generate a third prompt to update the exercise program assigned to the user account based on at least one of the user response information and user survey received from the electronic device (10). Specifically, the control unit (130) can receive voice data of the user regarding exercise movements using a microphone (13) provided in the electronic device (10).
[0082] Additionally, the control unit (130) may provide at least one user survey related to the user's exercise movements on a service page provided to the electronic device (10). Furthermore, the control unit (130) may receive survey response data for at least one user survey received through the electronic device (10). Furthermore, the control unit (130) may use the prompt generation unit (131) to generate a third prompt that updates the exercise program assigned to the user account based on at least one of voice data and survey response data. Here, “updating the exercise program” may mean changing at least one of the exercise movements, the number of repetitions of the movements, and the total exercise time that constitute the exercise program.
[0083] Meanwhile, the feedback providing system (100) may include one or more processors, and such processors may include one or more general-purpose processors and / or one or more special-purpose processors (e.g., digital signal processors, tensor processing units (TPUs), graphics processing units (GPUs), neural network processing units (NPUs), application integrated circuits, application semiconductors (ASICs), etc.). One or more processors may be configured to execute instructions, computer-readable instructions, and / or other instructions described herein that are stored (or included) in the storage unit (120). The feedback providing system (100) may perform data processing described below in cooperation with memory and at least one processor. The processor may perform a series of operations and data processing using data and information stored in memory. Here, “memory” may be a component of the storage unit (120), and “processor” may be used interchangeably with the control unit (130).
[0084] In the foregoing, the feedback providing system (100) of the present invention has been described, and it can be implemented based on the method of providing motion feedback using a vision language model described below.
[0085] Hereinafter, with reference to FIG. 2 together with FIG. 3a, FIG. 3b, FIG. 4a to FIG. 4c, FIG. 5, and FIG. 6, a method for providing exercise motion feedback using a vision language model according to the present invention will be described in more detail. FIG. 2 is a flowchart for explaining a method for providing exercise motion feedback using a vision language model according to the present invention, FIG. 3a is a flowchart for explaining a process of generating a first prompt that enables analysis of an exercise video using user information and an exercise video according to the present invention. FIG. 3b is a flowchart for explaining a process of generating different first prompts according to user types according to the present invention, FIG. 4a to FIG. 4c are conceptual diagrams for explaining a process of analyzing a user's exercise motion according to the present invention, and FIG. 5 and FIG. 6 are conceptual diagrams for explaining a process of generating feedback information and providing it to an electronic device according to the present invention.
[0086] In the present invention, a process of receiving a user's exercise video from an electronic device may be performed (S210, see FIG. 2).
[0087] As illustrated in FIG. 3a, the control unit (130) may activate a sensor unit (11) comprising a microphone (13), a camera (12), and multiple different sensors provided in the electronic device (10) to receive an exercise video of the user's exercise movements. Furthermore, the control unit (130) may receive sensing information including at least one of video data and voice data based on the activation of at least one of the sensor unit (11) comprising a microphone (13), a camera (12), and multiple different sensors provided in the electronic device (10). For example, the control unit (130) may activate a camera provided in the electronic device (10) to receive an exercise video (310) including the user's exercise movements from the camera. Here, the “exercise video (310)” may refer to video data captured during the process of the user performing a specific exercise movement according to an exercise item of an exercise program pre-set in the user account. At this time, the explanation assumes that the user account is pre-logged in to the application.
[0088] At this time, the present invention describes receiving an exercise video (310) including the user's exercise movements through a camera (13) provided in an electronic device (10), but is not limited thereto, and may also receive an exercise video from a separate video equipment connected to the electronic device in at least one of wireless communication and wired communication. The method for receiving an exercise video (310) including the user's exercise movements according to the present invention may be very diverse, and in this specification, any module capable of receiving an exercise video (310) including the user's exercise movements is not limited to any type or method.
[0089] Specifically, the control unit (130) can provide an exercise guide video corresponding to an exercise item matched to an exercise item included in an exercise program to the electronic device (10). More specifically, the control unit (130) can control the electronic device (10) so that an exercise guide video corresponding to at least one exercise item is played sequentially according to an exercise program assigned to a user account.
[0090] Furthermore, the control unit (130) can output an exercise video (310) received through a camera to the display unit of the electronic device (10) in real time. Specifically, based on camera activation, the control unit (130) can output an exercise video to the display unit in which a user performs an exercise movement according to an exercise guide video output to the display unit of the electronic device (10). Furthermore, the control unit (130) can store the exercise video received from the electronic device (10) in a database (200).
[0091] Next, in the present invention, a process of collecting user information related to the user may be carried out (S220, see FIG. 2).
[0092] The control unit (130) may collect user information (320) corresponding to a user account logged into the electronic device (10) from a database (200, or storage unit (120)). Here, “user information (320)” may include at least one of medical information (321), exercise history information (322), user response information (323), exercise program information (324), user type information (325), and user authentication information (not shown). As an example, the medical information (321) according to the present invention may include at least one of the user’s age, the user’s gender, past medical history related to the user’s indication, prescription information, and treatment plan. As previously described, “indication” refers to a symptom or clinical situation requiring specific treatment or examination, which can be understood as the user’s disease or symptom. The control unit (130) may collect medical information (321) including at least one of information related to the user’s indication and information regarding the treatment plan from at least one of a doctor terminal and a medical staff server.
[0093] Additionally, the control unit (130) may collect exercise history information from the database (200). Here, the exercise history information (322) may include at least one of the date of exercise performance, the composition of exercise items included in the exercise program performed, the motion analysis result (or analysis result) regarding the user's exercise movements, and the exercise history analysis result. Furthermore, the control unit (130) may collect user response information (323) from at least one of the electronic device (10) and the database (200). Here, the “user response information (323)” may include at least one of the user’s voice data received from the microphone (13) provided in the electronic device (10) and response data to at least one user survey provided to the electronic device (10) in relation to the exercise movements performed by the user.
[0094] Additionally, the control unit (130) may collect exercise program information (324) from a database. The exercise program information (324) may include at least one of the names of each of a plurality of exercise items constituting an exercise program assigned to a user account, the number of exercises, the timing of the exercises, and information regarding the difficulty level of the exercises, as well as an exercise guide video and an exercise description corresponding to each of the plurality of exercise items. Furthermore, the user type information according to the present invention may include information related to at least one of an exercise purpose type, an exercise level type, and restricted exercise body parts.
[0095] Additionally, user authentication information may refer to information used in a user authentication process performed to log in to a user account on an electronic device (10). As an example, user authentication information may be various, such as i) ID, ii) password, iii) password pattern, iv) user's fingerprint authentication information, v) face authentication information, vi) voice authentication information, vii) iris authentication information, viii) vein authentication information, etc., set by the user.
[0096] As previously explained, “indication” refers to a symptom or clinical situation requiring specific treatment or examination, which can be understood as the user’s disease or symptoms. The control unit (130) may collect medical information including at least one of information regarding the user’s indication and information regarding a treatment plan from at least one of a doctor terminal and a medical staff server.
[0097] Next, in the present invention, a process of generating a prompt that enables analysis of the exercise video based on user information using user information and exercise video may be performed (S230, see FIG. 2).
[0098] As illustrated in FIG. 3a, the control unit (130) can generate a prompt that enables analysis of exercise movements based on user information and exercise video using the prompt generation unit (131). Specifically, the control unit (130) can control the vision language model (132) to generate a first prompt (330) that enables analysis of the user's exercise movements included in the exercise video based on user information. The prompt generation unit (131) according to the present invention can utilize natural language processing (NLP) technology to receive multiple different data inputs, extract prompt information, and generate a prompt to be input to the vision language model (132). Here, natural language processing (NLP) technology may refer to technology capable of understanding, interpreting, and generating human language using artificial intelligence technology (e.g., deep learning). Specifically, the prompt generation unit (131) can generate a prompt based on prompt engineering using at least one of user information and exercise video. Here, prompt engineering can refer to the technique of designing and optimizing input text (or prompts) to effectively utilize natural language processing (NLP) models.
[0099] For example, the control unit (130) can generate a first prompt based on the collected user information, based on the prompt generation unit (131), to enable analysis of the exercise video. There may be many different methods for generating a first prompt based on the collected user information according to the present invention to enable analysis of the exercise video, and the present specification is not limited to any type or method as long as it is a module capable of generating a first prompt based on the collected user information to enable analysis of the exercise video.
[0100] For example, the control unit (130) may include a medical information prompt (331) generated by extracting at least one of the user's past medical history related to age, gender, and indications from the medical information (321) to enable analysis of the exercise video based on the collected medical information (321), in the first prompt. Additionally, the control unit (130) may include an exercise history prompt (332) generated by extracting at least one of the user's past exercise records, exercise history analysis results, and user's response data regarding past exercise from the exercise history information (322) to enable analysis of the exercise video based on the collected exercise history information (322), in the first prompt (330).
[0101] Furthermore, the control unit (130) may include a user response prompt (333) generated by extracting user voice data and survey response data from the user response information (323) based on the collected user response information (323) so that an analysis of the exercise video may be performed, in the first prompt (330). Additionally, the control unit (130) may include an exercise program prompt (334) generated by extracting information related to at least one of a plurality of exercise items, number of exercises, exercise time, and exercise equipment constituting the exercise program assigned to the user account from the exercise program information (324) based on the collected exercise program information (324) so that an analysis of the exercise video may be performed, in the first prompt (330).
[0102] The control unit (130) may include a user type prompt (335) generated by extracting information related to at least one of the type of exercise purpose, the type of exercise level, and the restricted exercise area from the user type information (325) so that analysis of the exercise video is performed based on the collected user type information (325), in the first prompt (330). Here, information related to at least one of the type of exercise purpose, the type of exercise level, and the restricted exercise area may be generated based on at least one of medical information (321) and user response information (323) and stored in the database (200).
[0103] The method for generating user type information (325) according to the present invention may be very diverse, and in this specification, any module capable of generating user type information (325) is not limited to any specific type or method. For example, the control unit (130) may use medical information (321) and user response information (323) to set various types for exercise purposes, such as rehabilitation type, muscle strengthening type, and diet type, and may set at least one type for exercise level among beginner, intermediate, and advanced. In addition, the control unit (130) may use medical information (321) and user response information (323) to set restricted exercise parts among a plurality of body parts where exercise movements are restricted. The user type information (325) according to the present invention is not limited to the examples described above and may be modified or expanded according to user information.
[0104] Meanwhile, the control unit (130) can evaluate the user's exercise expectation level based on user information. Specifically, the user's exercise expectation level for at least one of a plurality of body parts can be evaluated based on at least one of the medical information and exercise history information included in the user information. Furthermore, the control unit (130) can generate a first prompt that enables an analysis of the user's exercise movements to be performed according to the exercise expectation level. Here, "exercise expectation level" may refer to a numerical value of the user's expected exercise performance suitability based on at least one of the user's medical information and exercise history information. The method for evaluating the user's exercise expectation level according to the present invention may be very diverse, and the present specification is not limited to the type and method as long as it is a module capable of evaluating the user's exercise expectation level.
[0105] For example, the control unit (130) can calculate the user's exercise expectation level by assigning a pre-set weight to each of the user's medical information and exercise history information. As an example, the control unit (130) can calculate the user's exercise expectation level by assigning a pre-set weight to the user's age information among the medical information. Specifically, the control unit (130) can calculate the user's exercise expectation level according to the user's age information by assigning different weights to each of the first interval (under 18 years old, 18 to 39 years old), the second interval (40 to 64 years old), and the third interval (65 years old or older). The control unit (130) can assign a first age weight (ex, +1) to the user's exercise expectation level corresponding to the first interval, assign a second age weight (ex, 0) to the user's exercise expectation level corresponding to the second interval, and assign a third age weight (ex, -1) to the user's exercise expectation level corresponding to the third interval. Furthermore, the control unit (130) can calculate the user's exercise expectation level differently according to age information based on weights pre-set for each section. At this time, the age weights according to the present invention can be set in various ways by the user, medical staff, and the system.
[0106] Additionally, the control unit (130) can calculate the user's exercise expectation level by assigning a pre-set weight to the gender information. Specifically, the control unit (130) can calculate the exercise expectation level differently according to the gender information by assigning a first gender weight (e.g., +1) to the first gender (e.g., male) and a second gender weight (e.g., 0) to the user of the second gender. At this time, the gender weight according to the present invention can be set in various ways by the user, medical staff, and the system.
[0107] Additionally, the control unit (130) can calculate the user's exercise expectation level by assigning a pre-set weight to the exercise history using at least one of the exercise frequency, cumulative exercise time, and past performance score included in the exercise history information. For example, the control unit (130) can calculate the exercise expectation level by assigning a first exercise history weight (e.g., +1) if the user's exercise frequency or cumulative performance score over the last 30 days is above a specific threshold, a second exercise history weight (e.g., 0) if it is at an average level, and a third exercise history weight (e.g., -1) if it is below a certain threshold. Furthermore, the control unit (130) can calculate the user's exercise expectation level differently according to the exercise history information based on the pre-set weight for each interval. At this time, the exercise history weight and the specific threshold can be set in various ways by the user, medical staff, and the system.
[0108] Additionally, the control unit (130) can calculate the user's expected exercise level by assigning a pre-set weight to information such as disease history or injured areas included in the medical information. For example, the control unit (130) can specify at least one body part among a plurality of body parts that satisfies a pre-set condition based on the determined user's expected exercise level for at least one of the plurality of body parts. Here, the “pre-set condition” may refer to a condition satisfied when, based on the user's medical information, there is a history of disease, surgery, or injury in a specific body part, or when exercise restriction is required for that part.
[0109] Based on the user's exercise expectation level according to the present invention, there may be a wide variety of methods for specifying at least one body part among a plurality of body parts that satisfies a preset condition, and in this specification, any module capable of specifying at least one body part among a plurality of body parts that satisfies a preset condition based on the user's exercise expectation level is not limited to its type or method. As an example, the control unit (130) can process unstructured text data included in medical information to specify a body part that requires exercise restriction.
[0110] To this end, the control unit (130) may include a text analysis module based on natural language processing, and the text analysis module may analyze text included in medical information to determine whether to restrict movement for a specific body part. The control unit (130) may extract key keywords, such as disease names, symptom names, and anatomical part names, from sentences included in medical information through morphological analysis, stop word removal, and named entity recognition. Furthermore, the control unit (130) may normalize the extracted keywords based on a pre-established medical knowledge base (e.g., SNOMED-CT, ICD-10, or UMLS) and map them to at least one of a standard disease code and a body part code. For example, the control unit (130) may extract keywords such as “back pain,” “lumbar pain,” and “herniated disc” to normalize all body parts requiring movement restriction to the ‘lumbar’ part. The control unit (130) can determine that if the normalized expression corresponds to a specific body part (e.g., lumbar spine), movement restriction is required for that part.
[0111] Furthermore, the control unit (130) can set a weight for the specified at least one body part. Specifically, the control unit (130) can set a pre-set weight for the specified body part based on the fact that the specific body part where exercise is restricted is identified from the user's medical information. For example, the control unit (130) may assign a first restriction weight (e.g., -3) to the expected level of exercise for the specified body part, and assign a second restriction weight (e.g., 0) when there is no exercise restriction. Furthermore, the control unit (130) can calculate the user's expected level of exercise differently according to the medical information based on the pre-set weight for each section. At this time, the restriction weight can be set in various ways by the user, medical staff, and the system.
[0112] That is, the control unit (130) can analyze at least one of the medical information and exercise history information included in the user information, evaluate the user's exercise purpose and current performance ability according to a preset weight, and calculate the user's exercise expectation level based thereon. For example, the control unit (130) can calculate the final exercise expectation level by reflecting the preset weights to the medical information and exercise history information, respectively, based on a standard expectation level score (e.g., 5 points).
[0113] Furthermore, the control unit (130) can generate a first prompt to analyze the user's exercise movements related to any one of a plurality of body parts based on a set weight. As illustrated in (a) of FIG. 3b, the control unit (130) can evaluate the exercise expectation level of the first user to analyze the user's exercise movements based on the first user information (320a). Specifically, the control unit (130) can calculate the user's exercise expectation level as a first expectation level (341) by using at least one of the first medical information and the first exercise history information related to the first user. Furthermore, the control unit (130) can generate a first prompt (330a) to analyze the user's exercise movements related to any one of a plurality of body parts based on the first user's exercise expectation level by using the prompt input (131).
[0114] Alternatively, as illustrated in (b) of FIG. 3b, the control unit (130) may evaluate the exercise expectation level of the second user in order to analyze the user's exercise movements based on the second user information (320b). Specifically, the control unit (130) may calculate the user's exercise expectation level as the second expectation level (342) by using at least one of the second medical information and the second exercise history information related to the second user. Furthermore, the control unit (130) may generate a first prompt (330b) that analyzes the user's exercise movements related to any one of a plurality of body parts based on the second user's exercise expectation level using the prompt input (131).
[0115] Next, in the present invention, a process may be carried out in which a prompt and a motion video are input into a pre-trained vision language model to perform an analysis of the user's motion through the vision language model (S240, see FIG. 2).
[0116] As illustrated in FIG. 4a, the control unit (130) inputs the generated first prompt (330) and the motion video (310) into a pre-trained vision language model (132) to generate a motion analysis result (410) from the pre-trained vision language model (132). As previously described, the vision language model (132) may refer to an artificial intelligence model capable of performing analysis of the user's motion included in the motion video, generating feedback information, and updating the exercise program based on the motion video and user information received from the electronic device (10). For example, the vision language model (132) may refer to an artificial intelligence model capable of analyzing the user's motion from the motion video to be analyzed, which is trained based on a training data set containing location information for joint points.
[0117] In the present invention, the vision language model (132) is described as learning the operation of the motion analysis model to analyze the user's exercise motion, but it is not limited thereto. It may also be configured to link with a separately provided motion analysis server to call the function of the motion analysis model included in the motion analysis server or to analyze the exercise motion based thereon. In other words, the present invention may include various implementation forms in which the vision language model (132) directly performs the function of the motion analysis model by incorporating it, as well as various implementation forms in which it performs the analysis function in cooperation with an external server or an external analysis module.
[0118] Meanwhile, the vision language model (132) according to the present invention may include a motion generation model. Here, the “motion generation model” may refer to a model that learns the user’s movement and generates a vector corresponding to the user’s movement (or motion). The motion generation model may analyze the user’s movement in a video to extract time-series data, and using the extracted time-series data, generate vector data including at least one of a joint position, joint angle, velocity, and acceleration corresponding to the user’s skeletal structure. Furthermore, based on the generated vector data, the motion generation model may generate motion data in which the user’s movement over a specific period of time is vectorized, and may visually output the motion data in at least one of a 2D or 3D space. For example, the motion generation model may include at least one of an artificial intelligence model based on an RNN (Recurrent Neural Network), LSTM (Long Short-Term Memory), GRU (Gated Recurrent Unit), CNN (Convolutional Neural Network), Transformer, or GNN (Graph Neural Network).
[0119] In this way, the vision language model (132) according to the present invention can analyze the user's movements included in the image data using a motion generation model. Furthermore, the vision language model (132) can generate a motion analysis result (or analysis result) for the user's movement.
[0120] In the following, the process of a vision language model (132) performing an analysis of a user's movement by linking with a motion analysis server is described in detail. As previously explained, the vision language model (132) is trained to implement the motion analysis model included in the motion analysis server, and can perform an analysis of a user's movement independently. At this time, the control unit (130) can train the vision language model (132) according to the learning process of the motion analysis model described below. Here, the “motion analysis model (52)” is a motion analysis model trained using a learning data set containing position information for joint points, and can estimate the movement posture of the user (U) from the movement video to be analyzed.
[0121] Referring to FIG. 4b, the motion analysis server (50) according to the present invention may refer to a cloud server that performs motion analysis of a user from motion video data capturing the user's motion. The motion analysis server (50) can analyze the relative positional relationship between key points (P1, P2) corresponding to a plurality of joint points of the user (U) extracted from the motion video (400) through a motion analysis model learned using learning data related to joint points. Here, "joint point" may refer to a plurality of joints of the user (U) (or a part of the user (U)'s body including joints). And, "key point" may refer to an area corresponding to each of the plurality of joint points of the user (U) in the motion video (400). Accordingly, in the present invention, "joint point" and "key point" may be used interchangeably, and the same reference numeral "P2" may be assigned to each joint point and key point to explain them.
[0122] The control unit (130) can use a motion analysis model (52) to extract key points (P1, P2) corresponding to joint points from a user’s exercise video (400), and analyze the user’s (U) exercise motion based on an analysis of the positional relationship between the extracted key points (P1, P2). In the present invention, a series of processes for analyzing the user’s exercise motion from an exercise video (400) using key points extracted through the motion analysis model (52) can be named the “exercise motion analysis process.”
[0123] In the present invention, the physical space and subject where the motion analysis process takes place are not separately distinguished, and it can be described as taking place in the feedback providing system (100). The motion analysis process can be performed using key points extracted from the motion analysis model (52). As previously described, the control unit (130) can use a pre-trained vision language model to extract key points corresponding to each of a pre-set plurality of joint points from a specific object corresponding to the user included in the motion video.
[0124] More specifically, the vision language model (132) can analyze video data received from the camera on a frame-by-frame basis to extract key points corresponding to joint points corresponding to the user's movement. For example, the vision language model (132) can use various object detection algorithms. For example, the vision language model (132) can use an algorithm that ensembles multiple bounding boxes (Weighted Box Fusion, WBF). However, it is obvious that the vision language model (132) is not limited to the object detection algorithm described above, but can use various object detection algorithms capable of detecting objects corresponding to the user (U) from video data. As another example, the motion analysis model (52) included in the motion analysis server (50) can identify or estimate the user's joint points from the movement video (400) through learning on training data specialized for joint points, and extract key points corresponding thereto.
[0125] In the present invention, the training data for which the motion analysis model (52) performs training may be stored in a motion analysis database (60), and such a motion analysis database (60) may also be named a “training data DB.” Further details regarding the training data will be described later.
[0126] According to the present invention, the motion analysis server (50) may include at least one of a learning unit (51) and a motion analysis model (52). The motion analysis server (50) may be provided inside the feedback providing system (100) according to the present invention or may be an external server. That is, the motion analysis server (50) according to the present invention performs the function of learning about motion analysis in conjunction with the vision language model (132), and it can be understood that there are no physical space constraints. Detailed information regarding the motion analysis server (50) will be described later along with the learning data.
[0127] The motion analysis database (60) is a storage facility where a learning data set is stored, and may be provided within the feedback providing system (100) itself according to the present invention or may be an external storage facility (or external DB). It can be understood that the motion analysis database (60) according to the present invention is sufficient as long as it is a space where the learning data set is stored, and there are no restrictions on the physical space.
[0128] Meanwhile, the “exercise video (400)” described in the present invention may include at least one of “exercise video data to be analyzed” and “exercise video data to be learned.” The “exercise video data to be analyzed” is exercise video data that is the subject of posture estimation analysis of the user (U), and the “exercise video data to be learned” can be understood as an exercise video (400) that is the subject of machine learning for a motion analysis model. Here, “posture estimation analysis” may mean extracting key points from the exercise video data.
[0129] The learning unit (51) may be configured to perform learning for a motion analysis model (52) based on the exercise video data to be learned. The learning unit (51) may train the motion analysis model (52) using the learning data. The learning unit (51) may detect a user (U) in the exercise video data to be learned and extract various learning data used for estimating exercise posture from the detected user (U). Such learning data may be used interchangeably with “information,” “data,” “data value,” or “data value.”
[0130] Meanwhile, the extraction of training data may be performed by means other than the training unit (51). The training unit (51) may use various object detection algorithms to detect the user (U) from the training target motion video data. For example, the training unit (51) may use an algorithm that ensembles multiple bounding boxes (Weighted Box Fusion, WBF). However, it is obvious that the training unit (51) is not limited to the object detection algorithm described above and may use various object detection algorithms capable of detecting an object corresponding to the user (U) from the training target motion video data.
[0131] Furthermore, the learning unit (51) can perform learning for the motion analysis model (52) based on the learning data set existing in the motion analysis database (60). As previously explained, the learning data set may include location information of joint points. The motion analysis model (52) is a motion analysis model learned using the learning data set (Data set) containing location information for joint points, and can estimate the exercise posture of the user (U) from the exercise video data to be analyzed.
[0132] Meanwhile, the motion analysis model (52) can extract key points corresponding to the user's joint points from the motion video (400) using the learning data set generated by the learning unit (51). The motion analysis model (52) can analyze the user's motion in the motion video (400) using the extracted key points. Specifically, the motion analysis model (52) can analyze the relative positional relationship between key points and, based on the analysis of the positional relationship, perform an analysis on the user's motion related to any one of the plurality of body parts. For example, it can estimate and analyze information regarding at least one of i) the position of the joint point, ii) the range of motion of the joint point, iii) the movement path of the joint point, iv) the connection relationship between the joint points, and v) the symmetry relationship of the joint point for the user (U).
[0133] Furthermore, the motion analysis model (52) can perform an analysis of at least one of the following: the range of motion of a joint, the speed of movement (or acceleration) of a joint, the user's body balance, body equilibrium, and body alignment state (e.g., leg axis alignment state, spine alignment state, etc.) included in the motion video to be analyzed. In the present invention, the motion analysis model (52) may also be configured to include a learning unit (51). Furthermore, conversely, the learning unit (51) may include the motion analysis model (52), and in this case, the learning unit (51) can train the motion analysis model (52) to perform a posture estimation function. Accordingly, in the present invention, the function performed by the motion analysis model (52) may be described interchangeably as being performed by the learning unit (51).
[0134] That is, the control unit (130) according to the present invention can train the vision language model (132) to perform the same function as the previously learned motion analysis model (52), and can perform an analysis of the user's motion using the vision language model (132).
[0135] As illustrated in FIG. 4c, the control unit (130) can output a guidance message (e.g., “Please stand inside the screen”) on the electronic device (10) so that the user (U) is fully contained within a specific area of the exercise video (or the display of the electronic device) in order to detect the user (U) from the exercise video (400).
[0136] As previously explained, the vision language model (132) can detect the user (U) from image data using an object detection algorithm based on the user's entire body within a specific area. Furthermore, the vision language model (132) can receive an exercise video (400) of the user performing an exercise movement according to an exercise program through a camera based on the detection of the user's entire body within a specific area.
[0137] In this case, the control unit (130) can control the electronic device (10) to film a user performing an exercise while an exercise video corresponding to an exercise item assigned to a user account is being played. Then, the control unit (130) can match the exercise video captured by the camera of the electronic device (10) with each of the plurality of exercise items included in the exercise program and store it in the storage unit (120).
[0138] Furthermore, the control unit (130) can calculate a performance score for an exercise movement using a vision language model (132). Specifically, the vision language model (132) can calculate a performance score for a user's exercise movement based on the analysis results of the relative positional relationship between key points. Specifically, the control unit (130) can calculate the similarity between the relative positional relationship between key points according to the user's exercise movement and the positional relationship corresponding to a pre-set correct posture using the vision language model (132). For example, the vision language model (132) can quantitatively calculate the similarity between the user's exercise motion and the correct posture by calculating the relative position vector between the key points based on the key point coordinate information of each frame extracted from the exercise video, and by comparing this with the relative position vector of the reference posture to calculate a similarity value (e.g., Cosine Similarity and Euclidean Distance). At this time, the pre-set correct posture and the key point coordinate information of each frame in the correct posture may be matched with the exercise program assigned to the user account and stored in the database (200).
[0139] The control unit (130) can calculate a performance score for a user’s exercise movement included in an exercise video (400) based on a vision language model (132). There may be many different methods for calculating a performance score for an exercise movement according to the present invention, and the present specification is not limited to any type or method as long as it is a module capable of calculating a performance score for an exercise movement.
[0140] For example, the control unit (130) can calculate the performance score based on the vision language model (132) according to the cumulative time during which similarity satisfies a preset condition. Here, “preset condition” may refer to a condition in which the similarity between the relative positional relationship between key points according to the user’s exercise movement and the correct posture is satisfied at or above a preset threshold. Specifically, the vision language model (132) can calculate the user’s performance score based on the ratio by comparing the cumulative time during which the preset condition is satisfied with the total exercise performance time. As an example, if the total exercise performance time is 100 seconds and the cumulative time during which the preset condition is satisfied is 70 seconds, the vision language model (132) can calculate the performance score for the corresponding exercise movement as 70 points.
[0141] Meanwhile, the control unit (130) can compare the performance score for past exercise movements and the performance score for the user's exercise movements by using the exercise history analysis results for past exercise movements that were performed in the past in relation to the user's exercise movements included in the exercise history information, based on the vision language model (132). Specifically, the vision language model (132) can collect the performance score for past exercise movements by using the past analysis results for specific exercise movements from the exercise history information included in the user information. Furthermore, the vision language model (132) can compare the performance score for past exercise movements with the performance score of the exercise movement currently being performed. Based on the vision language model (132), the control unit (130) can generate a trend analysis result related to the user's exercise movements based on the performance score.
[0142] As illustrated in FIG. 4a, the control unit (130) can generate a motion analysis result (410) including at least one of an analysis result (411) and a trend analysis result (414) for an exercise motion being performed by a user by using a vision language model (132). Specifically, the vision language model (132) can generate an analysis result (411) for an exercise motion including analysis data (412) for the user's exercise motion and a performance score (413) for the user's exercise motion. Additionally, the vision language model (132) can generate a trend analysis result (414) including trend analysis data (415) for the user's exercise motion by comparing the performance score for past exercise motions with the performance score of the exercise motion being performed.
[0143] Meanwhile, in the present invention, a process of obtaining feedback information regarding the user's movement, generated based on the analysis result of the user's movement from a vision language model, may be performed (S250, see FIG. 2).
[0144] The control unit (130) can generate a second prompt (420) requesting that feedback information regarding the user's exercise movement be generated according to the analysis result of the exercise movement. Specifically, the control unit (130) can generate a second prompt (420) requesting that feedback information regarding the user's exercise movement be generated based on the motion analysis result (410) using a prompt generation unit (131).
[0145] As illustrated in FIG. 5, the prompt generation unit (131) can generate a second prompt (420) that generates feedback information (510) based on a motion analysis result (410, or analysis result) including a trend analysis result. Furthermore, the control unit (130) can generate a second prompt that generates feedback information (510) based on a motion analysis result including a trend analysis result.
[0146] Specifically, the prompt generation unit (131) may generate a second prompt (420) requesting that feedback information regarding the user's exercise movement be generated based on a motion analysis result (410) that includes at least one of an analysis result (411) and a trend analysis result (414) regarding the exercise movement being performed by the user. Here, the “feedback information” may include information related to at least one of a feedback message based on user information and the user’s exercise video, a performance score, and a correct exercise video corresponding to the exercise movement performed by the user. Specifically, the “feedback message” according to the present invention may include information related to at least one of an accuracy evaluation of the exercise movement, a request for modification of the exercise movement, and guidance on alternative movements and exercise programs according to the user's state.
[0147] At this time, the feedback information (510) may include text feedback information containing a feedback message regarding the analysis result and voice feedback information corresponding to the text feedback information. Additionally, the feedback message according to the present invention may include at least one of a correction instruction for an exercise movement, a performance score for an exercise movement, comparison information with the exercise history analysis result, and alternative movement information related to the exercise movement, based on the analysis result. For example, the control unit (130) may generate a feedback message (e.g., “Please straighten your upper body a little more. You are doing better than yesterday!! Keep up the good work!”) that includes content comparing the user’s exercise movement with past exercise movements using the exercise history analysis result.
[0148] Next, in the present invention, a process of providing feedback information to an electronic device may be carried out (S260, see FIG. 2).
[0149] As illustrated in FIG. 6, the control unit (130) can output feedback information regarding the user's exercise movements generated using a vision language model (132) to the electronic device (10). Specifically, the control unit (130) can generate feedback information regarding the user's exercise movements in real time using the vision language model (132) and output it to the electronic device (10) while the user's exercise movements are being performed. For example, the control unit (130) can output the generated text feedback information to a display unit provided in the electronic device (10). Specifically, the control unit (130) can visually display text feedback information, including at least one of a feedback message (530) related to an accuracy evaluation of the exercise movements, a request for correction of the exercise movements, and guidance on alternative movements and exercise programs according to the user's state, a performance score (413), and a correct exercise video (610) corresponding to the exercise movements performed by the user, along with an exercise video (310), in a part of the service page output to the display unit of the electronic device (10).
[0150] Additionally, the control unit (130) can output voice feedback information (520) generated through a speaker provided in the electronic device (10) as sound (or voice). There may be a wide variety of methods for outputting voice feedback information (520) generated through a speaker provided in the electronic device (10) according to the present invention as sound (or voice), and in this specification, any module capable of outputting voice feedback information (520) generated through a speaker provided in the electronic device (10) as sound (or voice) is not limited to any specific type or method.
[0151] The control unit (130) according to the present invention may further include at least one of a module and an algorithm for outputting feedback information generated through a speaker provided in the electronic device (10) as sound (or voice). Specifically, the control unit (130) may further include a module for converting text into voice. For example, the control unit (130) may use a Text-to-Speech (TTS) module to output text included in the feedback information as sound (or voice) through a speaker provided in the electronic device (10).
[0152] In the foregoing, a method for providing exercise motion feedback using a vision language model according to the present invention has been described in detail. Below, with reference to FIGS. 7a and 7b, the process of updating an exercise program using the method for providing exercise motion feedback using a vision language model will be described in detail. FIGS. 7a and 7b are conceptual diagrams for explaining the process of updating an exercise program based on a user's response according to the present invention.
[0153] Meanwhile, in the present invention, a process of updating an exercise program performed by a user can be carried out using a vision language model (132).
[0154] As illustrated in FIG. 7a, the control unit (130) can receive voice data (710) corresponding to voice received through a microphone (13) provided in the electronic device (10). Furthermore, the control unit (130) can analyze the voice data (710) to generate voice analysis data. More specifically, the control unit (130) according to the present invention may include a voice recognition model capable of analyzing voice data (710) corresponding to voice received through a microphone provided in the electronic device. There may be a wide variety of methods for analyzing voice data (710) corresponding to voice received through a microphone provided in the electronic device according to the present invention, and the present specification is not limited to the type and method of any module or model capable of analyzing voice data (710) corresponding to voice received through a microphone provided in the electronic device. The control unit (130) according to the present invention may further include at least one of a module and an algorithm that perform the same function as the voice recognition model. For example, a speech recognition model can generate speech analysis data by analyzing speech data (710) based on at least one of a STT (Speech-to-Text) algorithm, HMM (Hidden Markov Model), HMM-GMM (Hidden Markov Model-Gaussian Mixture Model), CTC (Connectionist Temporal Classification) based model, Beam Search based model, DNN (Deep Neural Network) based model, RNN (Recurrent Neural Network) based model, Seq2Seq (Sequence-to-Sequence) model, and Transformer based model.
[0155] Furthermore, the control unit (130) may provide at least one user survey (720) in relation to the user's exercise movements on a service page output to the electronic device (10). Here, the at least one user survey (720) may include at least one multiple-choice question item related to the user's exercise movements and a multiple-choice survey (721) composed of a plurality of selection items corresponding to the multiple-choice question item.
[0156] The control unit (130) may provide a survey page (or service page) containing a plurality of surveys on the electronic device (10). In this case, the page containing the plurality of surveys may be output via the touch screen (or display) of the electronic device (10). Here, the survey (or question, or problem, or item, or test) provided on the page may include a survey related to the user's exercise movements. In the present invention, the survey related to exercise movements may be diverse. For example, the survey related to exercise movements may include various elements for evaluating the user's condition related to exercise movements, such as an evaluation of the difficulty of the exercise movements, the presence or absence of pain, the duration of pain, mental health, and physical health. Such a survey related to exercise movements may include a credible survey actually used in psychiatry to diagnose the user's condition. Additionally, the control unit (130) may periodically update the survey by additionally collecting surveys for analyzing the user's condition through a central server, an external server, or a website.
[0157] Furthermore, the control unit (130) may receive objective survey response data corresponding to an objective survey included in at least one user survey provided to the electronic device (10). Specifically, the control unit (130) may receive response data for an objective survey (or objective survey response data) from the electronic device (10) that includes a response matched to an item selected by user input among a plurality of selection items. At this time, the objective survey response data may include natural language response data corresponding to a specific selection item selected by user input among a plurality of selection items.
[0158] Meanwhile, the control unit (130) may provide at least one open-ended survey (722) related to exercise movements on a service page output to the electronic device (10). Specifically, the at least one user survey (720) provided on the service page may further include an open-ended survey (722) capable of receiving natural language input for at least one open-ended question item from the user terminal in relation to exercise movements.
[0159] Furthermore, the control unit (130) can receive natural language input for a subjective question item as subjective survey response data. At this time, the control unit (130) can receive the natural language input entered into the electronic device (10) as survey response data for the user's subjective survey (722). At this time, if the user's voice is received by the electronic device (10) as survey response data for the subjective survey, the control unit (130) can receive voice data corresponding to the voice received from the electronic device (10) as subjective survey response data for the subjective survey (722).
[0160] The control unit (130) can generate user response information (730) using at least one of voice data (710) and survey response data. Furthermore, the control unit (130) can generate a prompt to be input to the vision language model (132) using the user response information (730). The control unit (130) can process the generated prompt as input to the vision language model (132) and update the exercise program assigned to the user account through the vision language model (132). Here, “updating the exercise program” may mean changing at least one of the exercise movements, the number of repetitions of the movements, and the total exercise time that constitute the exercise program.
[0161] In the following, an exercise program update that changes exercise movements is described as an example, but is not limited thereto, and at least one of the number of repetitions and total exercise time may also be changed.
[0162] For example, the control unit (130) can generate a third prompt (740) that updates the exercise program assigned to the user account using user response information (730). Furthermore, the control unit (130) can input the third prompt (740) into the vision language model (132) and update the exercise program through the vision language model (132).
[0163] As illustrated in FIG. 7b, the control unit (130) may extract text (731, 732) of a pre-set topic related to an exercise program update from user response information (730) and include the text of the pre-set topic in a third prompt. Here, “pre-set topic” may mean at least one of adjusting the difficulty of the exercise, selecting the type of exercise, and changing the type of exercise.
[0164] For example, the control unit (130) can generate a third prompt (740) to change the “lunge motion” to another exercise item based on receiving voice data such as “I felt pain during the lunge motion, please change to another exercise!” through a microphone. Furthermore, the control unit (130) can update the exercise program (750) assigned to the user account by inputting the third prompt (740) into the vision language model (132). At this time, the exercise program (750) assigned to the user account may be stored in the database (200), and the exercise program (750) may include multiple different exercise motions (751 to 753).
[0165] Furthermore, the control unit (130) can change a specific exercise movement (741) among the plurality of exercise movements to another exercise (761) according to a third prompt using a vision language model (132). Specifically, the vision language model (132) can change the exercise movement according to a third prompt (740). Similar exercise movements related to a specific exercise movement (751) may be matched and stored in the storage unit (120). The vision language model (132) can change the exercise items constituting the exercise program assigned to the user account by using similar exercise movements matched to the specific movement. For example, the control unit (130) can receive a third prompt and update the exercise program to provide a specific exercise movement (e.g., lunge) by changing it to a similar exercise movement (e.g., “hip extension exercise”) related to the specific exercise movement. The control unit (130) can obtain an updated exercise program (760) from a vision language model (140) and provide the updated exercise program (760) to the electronic device (10).
[0166] Furthermore, the vision language model (132) can update the exercise program by changing the exercise movements assigned after the next day of the specific exercise period (e.g., “Day 1 of Week 2”) based on user response information (730) regarding the exercise movements assigned during the specific exercise period (e.g., “Day 6 of Week 1”). The control unit (130) can provide the updated exercise program (760) on the electronic device (10) from the day after the specific day.
[0167] As described above, the method and system for providing exercise motion feedback using a vision language model according to the present invention can improve the accuracy and efficiency of exercise performance by using a video of a user's exercise performance as input, automatically performing an analysis of the exercise motion, and providing customized feedback information to the user in real time.
[0168] Furthermore, the method and system for providing exercise motion feedback using a vision language model according to the present invention can provide personalized feedback based on the user-customized motion analysis results by analyzing the user's exercise motions while reflecting the level of exercise expectation.
[0169] Furthermore, the method and system for providing exercise motion feedback using a vision language model according to the present invention generates a prompt reflecting a user response received from the user after performing an exercise, and processes it as input to the vision language model, thereby enabling the exercise program to be flexibly updated according to the current user state through the vision language model.
[0170] Furthermore, the motion motion feedback providing system (100) using a vision language model according to the present invention can be implemented through a computing system (or device) described below and can perform data processing related to the motion motion feedback providing method using the vision language model described above.
[0171] Meanwhile, FIG. 8 illustrates an example of a block diagram of a computing system in which the present invention can be implemented.
[0172] Referring to FIG. 8, a computing system (10000) that performs a method for providing motion feedback using a vision language model according to one embodiment of the present invention may include at least one computing device. At this time, the at least one computing device may be a single processor or a multiprocessor computing device.
[0173] The components of at least one computing device of the present invention may include various hardware components such as one or more processors, memory, other hardware, and a system bus (not shown) that connects various system components so that they can transmit and receive data to and from each other (e.g., telecommutatively connected, physically connected, electrically connected), and the components of at least one computing device are not limited thereto and may be very diverse.
[0174] Meanwhile, at least one computing device included in a computing system (10000) that performs a method of providing motion feedback using a vision language model may be connected to communicate via a network (1070). For example, at least one computing device included in the computing system (10000) may be clustered or may be part of a local area network (LAN). Additionally, at least one computing device may be part of a wide area network (WAN) or connected to at least one of a client-server network and a peer-to-peer network within the cloud.
[0175] Meanwhile, when at least one computing device is used in at least one of a network environment and a cloud computing environment, the at least one computing device may be connected to at least one of a public and private network through a network interface or adapter. In one embodiment, other communication connection devices, such as a modem, may be used to establish communication through the network. The modem may be at least one of an internal modem and an external modem, and may be connected to a system bus through a network interface or a specific mechanism, etc. A wireless network component consisting of an interface and an antenna may be coupled to the network through a device such as an access point, a peer computer, etc. In the present invention, the method of connecting at least one computing device to communicate through the network (1070) is not limited, and it may be connected to communicate in a manner different from the described example.
[0176] Furthermore, other computer-type devices and / or systems not shown in FIG. 8 may also interact technically with at least one computing device or other system through one or more connections to the network (1070) via a network interface. Here, the network interface may include network interface equipment such as a physical network interface controller (NIC) or a virtual network interface (VIF).
[0177] The network (1070) of the present invention may include various forms such as the Internet, WLAN (Wireless LAN), Wi-Fi (Wireless-Fidelity), Wi-Fi (Wireless Fidelity) Direct, DLNA (Digital Living Network Alliance), WiBro (Wireless Broadband), WiMAX (World Interoperability for Microwave Access), HSDPA (High Speed Downlink Packet Access), HSUPA (High Speed Uplink Packet Access), LTE (Long Term Evolution), LTE-A (Long Term Evolution-Advanced), 5G (5th Generation Mobile Telecommunication), Bluetooth (Bluetooth™ Frequency Identification), Infrared Communication (Infrared Data Association; IrDA), UWB (Ultra-Wideband), ZigBee, NFC (Near Field Communication), Wi-Fi Direct, Wireless USB (Wireless Universal Serial Bus), etc., and in the present invention, data transmission may be performed based on standard communication protocols such as TCP / IP, HTTP, SSL, etc.
[0178] A computing system (10000) that performs a method for providing exercise motion feedback using a vision language model according to the present invention may include at least one of a user computing device (1010, or system), a training computing system (1050, or device), and a server computing system (1030, or device).
[0179] A user computing device (1010) according to the present invention may be understood as a computing device comprising at least one processor (1011) and a memory (1012) for performing a method of providing motion feedback using a vision language model. For example, the user computing device (1010) may include at least one computing device among a smartphone, a smart TV, a laptop computer, a desktop computer, a digital broadcasting terminal, a PDA (personal digital assistant), a PMP (portable multimedia player), a navigation device, a slate PC, a tablet PC, an ultrabook, a wearable device (e.g., a smartwatch, a smart glass), and a head-mounted display (HMD).
[0180] At least one processor (1011) constituting the user computing device (1010) may include one or more general-purpose processors and / or one or more special-purpose processors. For example, at least one processor (1011) constituting the user computing device (1010) may be composed of at least one or a plurality of electrically connected processors among a central processing unit (CPU), a graphics processing unit (GPU), a tensor processing unit (TPU), a neural network processing unit (NPU), an arithmetic logic unit (ALU), a floating-point arithmetic unit (FPU), an application integrated circuit, an application semiconductor (ASIC), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, and / or other electrical units for performing functions.
[0181] Furthermore, at least one processor (1011) may be configured to execute computer-readable instructions contained in memory (1012) and / or other instructions described herein. Memory (1012) constituting a user computing system (1010) according to the present invention may include volatile memory, non-volatile memory, fixed media, removable media, magnetic media, optical media, semiconductor media and / or other types of physically durable storage media. For example, memory (1012) may include one or more non-transient / transient computer-readable storage media such as RAM, ROM, HDD (Hard Disk Drive), SSD (Solid State Disk), SSD (Silicon Disk Drive), EEPROM, EPROM, flash memory device, magnetic disk, and combinations thereof, and may include web storage of a server performing memory storage functions over the internet. This memory (1012) can store data and instructions necessary for the at least one processor (1011) to perform the operation of an application for providing motion feedback using a vision language model.
[0182] A user computing device (1010) may include one or more user input components (1021) that detect user input. For example, the user input component (1021) may also be referred to as a user interface module. The user input component (1021) may include a touch screen, a computer mouse, a keyboard, a keypad, a touchpad, a trackball, a joystick, a voice recognition module, or other similar devices. However, the present invention does not limit the type of user input component (1021). In this case, the user input component (1021) in the present invention does not necessarily mean a hardware means, but can be understood as a channel for receiving input from a user. Meanwhile, the user of the present invention may refer to an automated agent, script, playback software, etc., that operates on behalf of one or more people.
[0183] A user can interact with a computing system (10000) including at least one computing device through input text, touch, voice, movement, computer vision, gestures and / or other forms of input / output using a user input component (1021). For example, the user input component (1021) may include one or more of a command line interface (CLI), a graphical user interface (GUI), a natural user interface (NUI), a voice command interface and / or other user interface (UI) representations.
[0184] Between the user input component (1021) and the user computing device (1010), one or more application programming interface (API) calls may be made based on user input received from a user interface and / or a network. Here, the expression “based on” may be interpreted to include cases where it is based on the use of a specific configuration, modified from, derived from, influenced by, dependent on, or otherwise derived from a specific configuration. In some embodiments, an API call may be configured for a specific API, which may be interpreted or converted into an API call configured for another API. Here, an API may mean a defined interface or connection between computers or between computer programs.
[0185] In one embodiment, the user computing device (1010) may store at least one machine learning model (1020). For example, the user computing device (1010) may be various machine learning models, such as a plurality of neural networks (e.g., deep neural networks) that analyze the user's exercise movements based on exercise video and user information and provide feedback information on the exercise movements, or other types of machine learning models including non-linear models and / or linear models, and may be composed of a combination thereof.
[0186] According to an embodiment of the present invention, a user computing device (1010) may perform a method of providing exercise motion feedback using a vision language model by using a local or / and external machine learning model (1020). Alternatively, the user computing device (1010) may perform a method of providing exercise motion feedback using a vision language model by using a machine learning model (1040) provided by a server.
[0187] In addition, according to another embodiment of the present invention, a server computing system (1030) communicating with a user computing device (1010) may provide feedback information regarding the user's exercise movements to the user computing device (1010) via an application or / and the web in accordance with a request from the user received through the user computing device (1010).
[0188] In addition, according to another embodiment of the present invention, by linking at least a part of a user computing device (1010) and a server computing system (1030) with each other to perform a method of providing exercise motion feedback using a vision language model, feedback information regarding the user's exercise motion can be provided to the user.
[0189] Additionally, according to various embodiments of the present invention, a user computing device (1010) and / or a server computing system (1030) can learn machine learning models (1020, 1040) performed in a method for providing exercise motion feedback using a vision language model through interaction with a training computing system (1050) that is communicatedly connected via a network (1070). In this case, the training computing system (1050) may be a computing system separate from the server computing system (1030). Alternatively, in some embodiments, the training computing system (1050) may be part of the server computing system (1030) or part of the user computing device (1010).
[0190] Meanwhile, the server computing system (1030) may include at least one processor (1031) and memory (1032). Here, the processor (1031) may be composed of at least one or a plurality of electrically connected processors among a central processing unit (CPU), a graphics processing unit (GPU), a tensor processing unit (TPU), a neural network processing unit (NPU), an application integrated circuit, an application semiconductor (ASIC), an arithmetic logic unit (ALU), a floating-point arithmetic unit (FPU), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, and / or other electrical units for performing functions. For example, at least one processor (1031) may include a circuit and a transistor configured to execute instructions from memory (1032).
[0191] The memory (1032) constituting the server computing system (1030) according to the present invention may include volatile memory, non-volatile memory, fixed media, removable media, magnetic media, optical media, semiconductor media, and / or other types of physically durable storage media. For example, the memory (1032) may include one or more non-transient / transient computer-readable storage media such as RAM, ROM, HDD (Hard Disk Drive), SSD (Solid State Disk), SSD (Silicon Disk Drive), EEPROM, EPROM, flash memory device, magnetic disk, etc., and combinations thereof, and may include web storage of a server that performs the storage function of memory over the internet. Additionally, the server computing system (1030) may further include a data storage (data store). For example, the data storage may be composed of at least one of a relational database, a NoSQL database, a data warehouse, and a local file system.
[0192] In the memory (1032) constituting the server computing system (1030) according to the present invention, data and instructions necessary for the at least one processor (1031) to perform the operation of an application for providing motion feedback using a vision language model may be stored.
[0193] In one embodiment, the server computing system (1030) may be composed of a single device or a plurality of computing devices, and may be configured to operate according to a sequential or parallel computing architecture. Additionally, a distributed processing system may be configured with a plurality of networked devices.
[0194] Meanwhile, the training computing system (1050) may include at least one processor (1051) and memory (1052). The model trainer (1060) is a logical component that executes the training of at least one machine learning model (1020, 1040) and may be implemented in the form of hardware, firmware, or software. For example, the model trainer (1060) may be executed by the processor (1051) after loading training data (1061) stored in a storage device into memory (1052). For example, the model trainer (1060) may be configured to execute one or more operations (e.g., model training, model reconstruction, model validation, model testing) on at least one machine learning model.
[0195] The machine learning model of the present invention may include at least one of a statistical model, an algorithm, a neural network (NN), a convolutional neural network (CNN), a generative neural network (GNN), a Word2Vec model, a Bag of Words model, a TF-IDF (document frequency-inverse document frequency) model, a GPT (Generative Pre-trained Transformer) model (or other autoregressive models), a PPO (Proximal Policy Optimization) model, a nearest neighbor model (e.g., a k-nearest neighbor model), a linear regression model, a K-means clustering model, a Q-learning model, a TD (Temporal Difference) model, a Deep Adversarial Network model, and all other types of models further described herein.
[0196] Specifically, the model trainer (1060) may execute operations to train a machine learning model, and said operations may include at least one of adding, removing, and modifying model parameters. At this time, the training of the machine learning model may be at least one of supervised learning, semi-supervised learning, and unsupervised learning. In one embodiment, the training of the machine learning model may include the step of repeatedly inputting training data (1061) based on epochs and repeatedly performing the machine learning model training process configured in this way. Here, an epoch may refer to a unit in which the entire set of training data (1061) undergoes forward and backpropagation processing once. In some implementations, different levels of training methods (e.g., supervised learning, semi-supervised learning, unsupervised learning) may be used for different epochs.
[0197] The training data (1061) of the present invention may include input data and / or data previously output from at least one machine learning model (e.g., recursive learning feedback). The parameters of at least one machine learning model may include at least one of a seed value, a model node, a model layer, an algorithm, a function, connections between different machine learning models, connections between parameters, machine learning model constraints, and other digital components that influence the output of the machine learning model. In this case, the model connections between different machine learning models may include or represent relationships between model parameters and / or models, which may be dependent or interdependent, hierarchical, and / or static or dynamic. The combinations and configurations of model parameters described herein may be too complex to be maintained or used by human cognitive abilities.
[0198] In the present invention, the machine learning parameters described according to the embodiments are not limited, and a single machine learning model may further include a plurality of model parameters.
[0199] Meanwhile, FIG. 9 illustrates an example of a block diagram of a computing device (1100) that may be included in a user computing device (1010), a server computing system (1030), and a training computing system (1050), as an embodiment of a computing system (10000) in which the present invention can be implemented.
[0200] As illustrated in FIG. 9, the computing device (1100) may include at least one application (e.g., Application 1 to Application N), and each of the at least one application may include a machine learning library and a model execution environment for performing a method of providing motion feedback using a machine learning-based vision language model. The at least one application included in the computing device (1100) may communicate with the sensor, context manager, device state manager, or additional component(s) within the computing device (1100) via an Application Programming Interface (API). In one embodiment, the at least one application may interface with device components, such as receiving sensor data or state data or transmitting prediction results to an output device via a public or private API.
[0201] Meanwhile, FIG. 10 illustrates an example of a block diagram in another aspect of a computing device (1200), which is one of the components of a computing system (10000) that performs a method of providing motion feedback using a vision language model according to an embodiment of the present invention.
[0202] A computing device (1200) according to the present invention may include at least one application (e.g., Application 1 to Application N), and at least one application may communicate with a central intelligence layer (1210). Each application may interact with a shared model within the central intelligence layer (1210) through an API (e.g., a common API).
[0203] The central intelligence layer (1210) includes one or more machine learning models and may share them among multiple applications or provide them independently to each. In one embodiment, the central intelligence layer (1210) may be integrated as part of an operating system or implemented as a separate logical layer.
[0204] Additionally, the central intelligence layer (1210) can communicate with the central device data layer (1220). The central device data layer (1220) can integrate and store exercise videos, user information, and motion analysis results stored within the computing device (1200), and provide them as input data necessary for providing exercise motion feedback using a vision language model. Each device component (e.g., sensor, state manager, etc.) can communicate with the central device data layer (1220) through a private API, etc.
[0205] The technology described in this specification may be composed of a single or multiple computing devices, and a machine learning model that performs a method for providing motion feedback using a vision language model may be executed sequentially or in parallel on one component or multiple distributed components. The data storage, machine learning model, and application may be distributed and operated locally or over a network, and these configurations can be flexibly applied to various system architectures.
[0206] Meanwhile, computer-readable media include all types of recording devices in which data that can be read by a computer system is stored. Examples of computer-readable media include HDD (Hard Disk Drive), SSD (Solid State Disk), SSD (Silicon Disk Drive), ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, etc.
[0207] Furthermore, the computer-readable medium may be a server or cloud storage that includes a storage and is accessible to an electronic device via communication. In this case, the computer may download the program according to the present invention from the server or cloud storage via wired or wireless communication.
[0208] Furthermore, in the present invention, the computer described above is an electronic device equipped with a processor, namely a CPU (Central Processing Unit), and no special limitations are placed on its type.
[0209] Meanwhile, the above detailed description should not be interpreted restrictively in all respects but should be considered exemplary. The scope of the invention should be determined by a reasonable interpretation of the appended claims, and all modifications within the equivalent scope of the invention are included within the scope of the invention.
Claims
1. A step of receiving a user's exercise video from an electronic device; A step of collecting user information related to the above user; A step of generating a prompt that enables analysis of the exercise video based on the user information, using the user information and the exercise video; A step of inputting the above prompt and the above exercise video into a pre-trained vision language model, and performing an analysis of the user's exercise movements through the vision language model; A step of obtaining feedback information regarding the user's movement, generated from the above vision language model based on the analysis result of the user's movement; and A method for providing motor motion feedback using a vision language model, characterized by including the step of providing the above feedback information to the electronic device.
2. In Paragraph 1, The step of generating the above prompt is, A step of evaluating the user's exercise expectation level based on the above user information; A method for providing exercise motion feedback using a vision language model, characterized by including the step of generating a first prompt that enables analysis of the user's exercise motion to be performed according to the exercise expectation level.
3. In Paragraph 2, In the step of evaluating the exercise expectation level of the above user, A method for providing exercise motion feedback using a vision language model, characterized by evaluating the user's exercise expectation level for at least one of a plurality of body parts based on at least one of medical information and exercise history information included in the user information.
4. In Paragraph 3, The step of generating the first prompt above is, A step of specifying at least one body part among the plurality of body parts that satisfies a preset condition based on the determined exercise expectation level of the user, for at least one of the plurality of body parts; Step of setting a weight for the specified at least one body part; and A method for providing exercise motion feedback using a vision language model, characterized by including the step of generating a first prompt that analyzes the user's exercise motion related to any one of the plurality of body parts based on set weights.
5. In Paragraph 4, The step of performing an analysis of the exercise movements of the above user is: A step of extracting key points corresponding to each of a plurality of pre-set joint points in a specific object corresponding to the user included in the motion video using the above-mentioned pre-trained vision language model; and A method for providing exercise motion feedback using a vision language model, characterized by including the step of analyzing the relative positional relationship between the key points and, based on the analysis of the positional relationship, performing an analysis on the user's exercise motion related to any one of the plurality of body parts.
6. In Paragraph 5, The step of performing an analysis of the exercise movements of the above user is: A step of calculating a performance score for the user's exercise movements based on the analysis results of the above positional relationship; A step of comparing a performance score for a past exercise movement and a performance score for the user's exercise movement using an exercise history analysis result for a past exercise movement performed in the past in relation to the user's exercise movement included in the exercise history information; and A method for providing exercise motion feedback using a vision language model, characterized by including the step of generating a trend analysis result related to the user's exercise motion based on the above performance score.
7. In Paragraph 6, The step of generating the above prompt is, The method further includes the step of generating a second prompt requesting that feedback information regarding the user's exercise movement be generated according to the analysis result regarding the exercise movement. In the step of generating the second prompt mentioned above, A method for providing motion feedback using a vision language model, characterized by generating the second prompt that generates the feedback information based on the analysis results including the trend analysis results.
8. In Paragraph 7, The above feedback information is, It includes text feedback information containing a feedback message regarding the above analysis result and voice feedback information corresponding to the text feedback information, and The above feedback message is, A method for providing exercise motion feedback using a vision language model, characterized by including at least one of a modification instruction for the exercise motion, a performance score for the exercise motion, comparison information with the exercise history analysis result, and alternative motion information related to the exercise motion, based on the analysis result.
9. In Paragraph 7, The step of calculating the above performance score is, A step of calculating the similarity between the relative positional relationship between the key points according to the exercise movements of the user and the positional relationship corresponding to the pre-set correct posture; A method for providing motor motion feedback using a vision language model, characterized by including a step of calculating a performance score based on the cumulative time in which the similarity satisfies a preset condition.
10. In Paragraph 1, The method further includes the step of updating the exercise program performed by the user using the above vision language model, and The above update step is, A step of receiving voice data corresponding to voice received through a microphone provided in the electronic device; A step of receiving user survey response data for at least one user survey provided to the electronic device; A step of generating user response information using at least one of the voice data and the survey response data; A step of generating a third prompt to update the exercise program assigned to the user account using the above user response information; and A method for providing exercise motion feedback using a vision language model, characterized by including the step of inputting the third prompt into the vision language model and updating the exercise program through the vision language model.
11. A communication unit that receives a user's exercise video from an electronic device; and It includes a control unit that collects user information related to the above user, and The above control unit is, Using the above user information and the above exercise video, a prompt is generated to enable analysis of the exercise video based on the above user information, and The above prompt and the above movement video are input into a pre-trained vision language model, and through the vision language model, an analysis of the user's movement is performed. From the above vision language model, feedback information regarding the user's movement, generated based on the analysis result of the user's movement, is obtained, and A motion feedback provision system using a vision language model characterized by providing the above feedback information to the electronic device.
12. A program that is executed by one or more processes in an electronic device and stored on a computer-readable recording medium, The above program is, A step of receiving a user's exercise video from an electronic device; A step of collecting user information related to the above user; A step of generating a prompt that enables analysis of the exercise video based on the user information, using the user information and the exercise video; A step of inputting the above prompt and the above exercise video into a pre-trained vision language model, and performing an analysis of the user's exercise movements through the vision language model; A step of obtaining feedback information regarding the user's movement, generated from the above vision language model based on the analysis result of the user's movement; and A program stored on a computer-readable recording medium characterized by including instructions for performing the step of providing the above feedback information to the electronic device.