Method and system for evaluating exercise performance rate by using vision langauge model
The vision language model system accurately assesses exercise performance by analyzing user motion, ensuring genuine movement execution and providing personalized programs, addressing the limitations of perfunctory viewing in existing methods.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- EVEREX
- Filing Date
- 2025-10-28
- Publication Date
- 2026-05-07
AI Technical Summary
Existing video-based verification methods cannot accurately determine if a user has performed exercise movements, as they can be evaded through perfunctory viewing, making it difficult to assess exercise performance and treatment compliance.
A method and system using a vision language model for evaluating exercise performance by analyzing user motion information and reference motion information, identifying target body parts, and setting performance conditions, with feedback loops for re-performance and program updates.
Accurately evaluates exercise performance, ensures genuine movement execution, and provides personalized exercise programs based on performance rates, enhancing treatment compliance and effectiveness.
Smart Images

Figure KR2025017337_07052026_PF_FP_ABST
Abstract
Description
Method and System for Evaluating Exercise Performance Using Vision Language Models
[0001] The present invention relates to a method and system for evaluating the exercise performance rate of a user's exercise movements using a vision language model.
[0002] With the recent rapid advancement of artificial intelligence (AI) technology, Vision-Language Models (VLMs), capable of simultaneously understanding and processing visual information such as images and videos as well as human language, are garnering attention. In particular, by integratively analyzing natural language and visual information, VLMs are demonstrating advanced technological capabilities that go beyond existing text-based question-and-answer large-scale language models. They enable the understanding of actual user behavior or actions and the generation of customized responses.
[0003] Such artificial intelligence technology is being utilized in medical industries, including healthcare, sports, fitness, exercise for daily living, and rehabilitation management, going beyond simple conversational question-and-answer. In particular, there is a growing demand for AI technology that provides personalized healthcare to users based on exercise videos captured through various electronic devices, such as smart devices, mobile cameras, and wearable sensors.
[0004] Recently, the use of digital therapeutics, along with artificial intelligence technology, has been gaining attention in the medical industry. For instance, in the case of CPAP machines used to treat sleep apnea, insurance coverage applies only if the user consistently uses the device for a certain period of time; if usage falls below this standard, the patient is responsible for the cost. This system is designed to guarantee treatment effectiveness and compliance, with the patient's actual usage serving as a crucial criterion for determining eligibility for treatment support. Similarly, for rehabilitation exercises or healthcare services utilizing digital therapeutic devices, whether the user has actually performed the exercises can serve as a key indicator for evaluating treatment effectiveness.
[0005] However, existing video-based verification methods have a drawback in that they can only determine whether a user has watched a video, making it difficult to verify whether the user has actually performed the exercise. For example, users may evade authentication through perfunctory viewing behavior, such as playing an exercise video without actually performing the workout. To address this issue, there is a need to utilize a vision language model to evaluate whether the user performs the exercise movements according to execution conditions based on the user's exercise video, and to calculate the user's exercise performance rate accordingly.
[0006] The present invention is intended to provide a method and system capable of evaluating the exercise performance rate of a user's exercise movements using a vision language model.
[0007] Specifically, the present invention aims to provide a method and system for evaluating exercise performance using a vision language model capable of natural language processing and image analysis, which can evaluate a user's exercise performance rate by utilizing user motion information and reference motion information.
[0008] Furthermore, the present invention aims to provide a method and system for evaluating exercise performance using a vision language model, which can identify a target body part according to the main movements of a specific exercise motion from an exercise video and evaluate the user's exercise performance rate according to performance conditions set on the target body part.
[0009] Furthermore, the purpose is to provide a method and system for evaluating exercise performance using a vision language model that can request the re-performance of specific exercise movements that failed to meet performance conditions, based on the user's exercise performance rate.
[0010] Furthermore, the present invention aims to provide a method and system for evaluating exercise performance using a vision language model capable of receiving user feedback on exercise performance and updating an exercise program based on the user feedback.
[0011] To solve the problem described above, the present invention proposes a method for evaluating a motion performance rate using a vision language model, utilizing a user's motion and a reference motion for which performance is requested. The method for evaluating a motion performance rate using a vision language model according to the present invention may include the steps of: requesting an electronic device to perform a specific motion; receiving a user's motion video related to the specific motion from the electronic device; extracting user motion information regarding the user's motion corresponding to the specific motion from the motion video; generating a prompt to perform an analysis on whether the user's motion satisfies a preset performance condition using the user motion information and the reference motion information corresponding to the specific motion; inputting the generated prompt and the motion video into a pre-trained vision language model to obtain performance evaluation information regarding the user's motion from the pre-trained vision language model; and evaluating the user's motion performance rate based on the performance evaluation information.
[0012] Furthermore, in the above specific exercise motion, a target body part is defined and exists according to the main motion of the above specific exercise motion, and the target body part may include at least one of a main target body part and an auxiliary target body part depending on the degree of correlation with the main motion.
[0013] Furthermore, the step of extracting the user motion information may include the step of identifying the target body part, which is predefined for the specific exercise motion, from the exercise video, and the step of analyzing the movement of the target body part in the exercise video to extract the user motion information for the specific exercise motion.
[0014] Furthermore, the step of extracting the user motion information may include a step of determining whether the main target body part is identifiable in the exercise video, and a step of determining whether to extract the user motion information for at least one of the main target body part and the auxiliary target body part according to the result of the determination.
[0015] Furthermore, the step of determining whether the main target body part is identified further includes, if the main target body part is not identified in the exercise video, a step of requesting the electronic device to include the main target body part in the exercise video, and in the step of requesting the main target body part to be included in the exercise video, a guidance message may be output to the electronic device instructing to adjust at least one of the video shooting position and posture so that the main target body part is included in the exercise video.
[0016] Furthermore, the step of obtaining the performance evaluation information regarding the user's exercise movement may include the step of analyzing whether the performance condition set for the target body part is satisfied for a predetermined number of times for the specific exercise movement using the previously trained vision language model, and the step of obtaining the performance evaluation information regarding the user's exercise movement from the previously trained vision language model according to the analysis result regarding the satisfaction.
[0017] Furthermore, the step of evaluating the exercise performance rate may include a step of calculating the ratio of the number of times the exercise satisfies the performance conditions set for the target body part among the preset number of times the exercise is performed, and a step of evaluating the user's exercise performance rate as at least one of a performance score and a percentile according to the calculated ratio of the number of times the exercise is performed.
[0018] Furthermore, the method may further include a step of requesting the user to perform a specific exercise movement corresponding to the user's exercise movement again based on the exercise performance rate calculated for the user's exercise movement, and the step of requesting the user to perform the specific exercise movement again may include: providing a request message to the electronic device requesting the user to perform the specific exercise movement again based on the fact that the exercise performance rate did not satisfy a preset standard condition; receiving user input for a request icon displayed in a part area of the request message from the electronic device; and re-evaluating the exercise performance rate for the user's exercise movement corresponding to the specific exercise movement based on the occurrence of an activation event for the request icon according to the user input.
[0019] Furthermore, the method further includes a step of updating an exercise program set in the user account of the user based on the exercise performance rate calculated for the exercise movements of the user, and the step of updating the exercise program may include a step of generating feedback information related to the exercise movements of the user, a step of generating a feedback prompt for updating the exercise program using the feedback information, and a step of processing the feedback prompt as input to the pre-trained vision language model to update the exercise program from the pre-trained vision language model.
[0020] Furthermore, the step of generating the feedback information may include receiving voice data corresponding to voice received through a microphone provided in the electronic device, receiving user survey response data for at least one user survey provided in the electronic device, and generating feedback information using at least one of the voice data and the survey response data.
[0021] Meanwhile, the exercise performance rate evaluation system using a vision language model according to the present invention includes a communication unit that receives a user’s exercise video related to a specific exercise movement from an electronic device and a control unit that requests the electronic device to perform the exercise for the specific exercise movement. The control unit extracts user action information regarding the user’s exercise movement corresponding to the specific exercise movement from the exercise video, generates a prompt that performs an analysis of whether the user’s exercise movement satisfies a preset performance condition using the user action information and reference action information corresponding to the specific exercise movement, inputs the generated prompt and the exercise video into a pre-trained vision language model to obtain performance evaluation information regarding the user’s exercise movement from the pre-trained vision language model, and evaluates the user’s exercise performance rate based on the performance evaluation information.
[0022] Meanwhile, the program is executed by one or more processes in an electronic device and is stored on a computer-readable recording medium, and the program may include instructions for performing the following steps: requesting the electronic device to perform an exercise for a specific exercise movement; receiving a user’s exercise video related to the specific exercise movement from the electronic device; extracting user action information regarding the user’s exercise movement corresponding to the specific exercise movement from the exercise video; generating a prompt to perform an analysis of whether the user’s exercise movement satisfies pre-set performance conditions using the user action information and reference action information corresponding to the specific exercise movement; inputting the generated prompt and the exercise video into a pre-trained vision language model to obtain performance evaluation information regarding the user’s exercise movement from the pre-trained vision language model; and evaluating the user’s exercise performance rate based on the performance evaluation information.
[0023] The method and system for evaluating exercise performance rate using a vision language model according to the present invention can determine whether the user actually performs the exercise and prevent perfunctory exercise performance by performing an evaluation of the user's exercise movements included in an exercise video and calculating the exercise performance rate.
[0024] Furthermore, the method and system for evaluating exercise performance using a vision language model according to the present invention can contribute to determining whether a key movement is accurately performed and maximizing exercise effects by identifying a target body part corresponding to a key movement of a specific exercise motion from an exercise video and quantitatively evaluating whether the performance conditions set for the target body part are satisfied.
[0025] Furthermore, the method and system for evaluating exercise performance rate using a vision language model according to the present invention can contribute to preventing the user from abandoning exercise performance and inducing continuous exercise performance by inducing the user to re-perform the same exercise according to the user's exercise performance rate.
[0026] Furthermore, the method and system for evaluating exercise performance using a vision language model according to the present invention can generate a customized exercise program with adjusted difficulty, intensity, and number of repetitions based on the exercise performance rate. Through this, the exercise effect can be maximized by providing an individualized exercise program suitable for the user's physical condition, exercise performance ability, and recovery progress.
[0027] FIG. 1 is a conceptual diagram illustrating a system for evaluating the performance rate of motor movements using a vision language model according to the present invention.
[0028] FIG. 2 is a flowchart illustrating a method for evaluating exercise performance using a vision language model according to the present invention.
[0029] FIG. 3 is a flowchart illustrating the process of collecting exercise images related to a user's exercise movements from an electronic device according to the present invention.
[0030] FIGS. 4a and FIGS. 4b are conceptual diagrams for explaining the process of extracting user motion information from a motion video according to the present invention.
[0031] FIGS. 5A and FIGS. 5B are conceptual diagrams for explaining the process of generating a prompt according to the present invention.
[0032] FIG. 6 is a conceptual diagram illustrating the process of evaluating the exercise performance rate of a user's exercise movements using a vision language model according to the present invention.
[0033] FIGS. 7a to 7c are conceptual diagrams for explaining an embodiment of requesting the re-execution of an exercise movement according to the exercise performance rate according to the present invention, and updating an exercise program set in a user account using user feedback information.
[0034] FIG. 8 is a block diagram illustrating a computing system in which the present invention can be implemented.
[0035] FIGS. 9 and FIGS. 10 are block diagrams illustrating an embodiment of a computing device according to the present invention.
[0036] Hereinafter, embodiments disclosed in this specification will be described in detail with reference to the attached drawings. Identical or similar components are assigned the same reference number regardless of the drawing symbols, and redundant descriptions thereof will be omitted. The suffixes "module" and "part" used for components in the following description are assigned or used interchangeably solely for the ease of drafting the specification and do not have distinct meanings or roles in themselves. Furthermore, in describing the embodiments disclosed in this specification, if it is determined that a detailed description of related prior art could obscure the essence of the embodiments disclosed in this specification, such detailed description will be omitted. Additionally, the attached drawings are intended only to facilitate understanding of the embodiments disclosed in this specification; the technical concept disclosed in this specification is not limited by the attached drawings, and it should be understood that they include all modifications, equivalents, and substitutions that fall within the spirit and technical scope of the present invention.
[0037] Terms including ordinal numbers, such as first, second, etc., may be used to describe various components, but said components are not limited by said terms. These terms are used solely for the purpose of distinguishing one component from another.
[0038] When it is stated that one component is "connected" or "connected" to another component, it should be understood that while it may be directly connected or connected to that other component, there may also be other components in between. On the other hand, when it is stated that one component is "directly connected" or "directly connected" to another component, it should be understood that there are no other components in between.
[0039] A singular expression includes a plural expression unless the context clearly indicates otherwise.
[0040] In this application, terms such as “comprising” or “having” are intended to specify the existence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.
[0041] The present invention relates to a method and system for evaluating exercise performance using a vision language model. More specifically, the present invention relates to a method and system for evaluating the exercise performance of a user's exercise movements using a vision language model capable of natural language understanding and image analysis.
[0042] The vision language model according to the present invention may refer to an intelligent system capable of autonomously performing specific tasks by understanding and processing visual information and linguistic information in combination without human intervention, based on a generative artificial intelligence model. Specifically, the vision language model according to the present invention is a multimodal generative model capable of processing visual inputs such as images and videos together with linguistic inputs such as text and voice, and can analyze video information related to a user's exercise movements and evaluate the exercise performance rate of the user's exercise movements based on the results.
[0043] In the present invention, "exercise performance rate" may refer to an indicator calculated by evaluating a user's exercise movement based on user motion information extracted from a user's exercise video received through an electronic device, according to pre-set performance conditions for a reference exercise movement requested to be performed.
[0044] A user (U, or patient) according to the present invention can perform an exercise program according to prescription information prescribed by a medical institution for the indication of the user (U) through an application or webpage provided by the performance rate evaluation system (100) according to the present invention. Specifically, in the present invention, an exercise program including at least one exercise movement can be set on an electronic device logged into a user account.
[0045] A user (or patient, U) of the present invention may possess a user account registered in the performance rate evaluation system (100) according to the present invention. For convenience of explanation, the account of a user who is a patient in this specification is referred to as a "user account (or patient account)." The "account" described above may be created through a page linked to the performance rate evaluation system (100). Alternatively, the "account" may be created on at least one other server (e.g., a medical staff server) linked to the performance rate evaluation system (100) according to the present invention. Accordingly, in this specification, without distinguishing the server where the account was issued, all accounts based on the performance rate evaluation system (100) according to the present invention are referred to as "accounts already registered in the performance rate evaluation system (100) according to the present invention."
[0046] Meanwhile, a medical professional (Doctor) can issue a prescription related to rehabilitation treatment to a user (U) through a medical professional terminal. At this time, the medical professional (D) may possess a medical professional account already registered in the performance rate evaluation system (100) according to the present invention. In this specification, an electronic device logged in with a medical professional account is referred to as a medical professional terminal. As an example, the performance rate evaluation system (100) according to the present invention may receive medical information prescribed by the medical professional (D) to the user (U) by linking with a medical professional server.
[0047] In the present invention, a user may be requested to perform an exercise for a specific exercise movement included in an exercise program. Furthermore, the user according to the present invention may execute an application of an electronic device (10) and perform an exercise according to an exercise program assigned to a user account. At this time, the present invention may collect various information using a camera, a microphone, and a plurality of different sensors equipped in the electronic device, and process the collected information using a vision language model. Specifically, the present invention may collect an exercise video of a specific exercise movement performed by the user using a camera equipped in the electronic device, and collect the user's voice data through a microphone equipped in the electronic device.
[0048] Furthermore, the present invention can extract user motion information regarding an exercise motion from an exercise video received from an electronic device. Here, “user motion information” refers to information extracted based on a user exercise video received from an electronic device, and may mean motion information obtained by analyzing the movement of body parts related to the main movements of the exercise motion performed by the user. More specifically, the user motion information may include motion information corresponding to at least one of the following: a position corresponding to a user’s body part included in the user’s exercise video, the direction and speed of the movement of the body part, a change in relative position between different body parts, and a motion pattern between a plurality of frames constituting the user’s exercise video. Such user motion information can be utilized to evaluate whether the user’s exercise motion satisfies pre-set performance conditions through comparison with reference motion information.
[0049] In this case, the “reference motion information” may include motion information corresponding to at least one of predefined target body part information according to the main motion of a specific exercise motion, a reference position corresponding to the target body part, and performance condition information corresponding to the target body part. As an example, the reference motion information may be configured based on a reference video capturing the reference exercise motion and predefined expert motion data.
[0050] Reference motion information according to the present invention can be compared with user motion information and utilized as an evaluation criterion for the user's exercise performance. Specifically, the present invention can generate a prompt that performs an analysis of whether the user's exercise motion satisfies pre-set performance conditions by utilizing user motion information and reference motion information. Furthermore, the present invention can input the generated prompt into a vision language model and evaluate the exercise performance rate for an exercise motion related to the user's indication through the vision language model.
[0051] The present invention focuses on feedback regarding exercise movements related to the treatment of “indications related to musculoskeletal disorders,” but is not necessarily limited thereto. As an example, the feedback regarding exercise movements described in the present invention may be for exercise intended for the treatment of a user with various diseases (e.g., cancer, diabetes, hypertension, etc.).
[0052] In addition, the present invention may evaluate the exercise performance rate for exercise movements necessary for health promotion in daily life, rather than rehabilitation exercises for therapeutic purposes related to the user's indications. For example, the exercise according to the present invention is not limited to any specific purpose and may be an exercise performed for various purposes, such as rehabilitation exercises, fitness exercises, ball sports, or dance exercises, for therapeutic, health promotion, or cosmetic purposes.
[0053] Furthermore, the evaluation of the performance rate of exercise movements according to the present invention may be an evaluation for various types of exercise, such as rehabilitation exercises, fitness exercises, ball sports, and dance exercises, for various purposes including therapeutic, health promotion, and cosmetic purposes, and it can be understood that it is not limited to a specific category of exercise. In addition, there are no limitations on the type of exercise according to the present invention, nor are there any limitations on the location of the exercise, such as indoor or outdoor exercise.
[0054] In the foregoing, the exercise performance rate evaluation using a vision language model according to the present invention has been generally described, and this can be implemented by the exercise performance rate evaluation system using a vision language model described below. Below, with reference to FIG. 1, the exercise movement performance rate evaluation system using a vision language model according to the present invention will be described in detail. FIG. 1 is a conceptual diagram for explaining the exercise movement performance rate evaluation system using a vision language model according to the present invention.
[0055] As illustrated in FIG. 1, a motion performance evaluation system using a vision language model according to the present invention (hereinafter referred to as the “performance evaluation system,” 100) may include at least one of a communication unit (110), a storage unit (120), and a control unit (130). At this time, the performance evaluation system (100) according to the present invention is not limited to the components described above and may further include components that perform the same or similar roles as the functions described in the specification.
[0056] Meanwhile, the performance rate evaluation system (100) according to the present invention may be implemented as an application or software. The performance rate evaluation system (100) implemented as software in this manner may be downloaded via a program (e.g., Play Store) that allows the application to be downloaded on an electronic device (10), or implemented via an initial installation program on the electronic device (10). In this case, the communication unit (110), storage unit (120), and control unit (130) according to the present invention may be utilized as components of the electronic device (10). In the present invention, the electronic device (10) can be understood as referring to an application installed on the electronic device (10). Such an application (or software) can be understood as a component of the performance rate evaluation system (100) according to the present invention.
[0057] In the present invention, the electronic device (10) may also be named a "mobile terminal" or a "user terminal," and the electronic device (10) described in this specification may include a mobile phone, a smartphone, a smart TV, a laptop computer, a digital broadcasting terminal, a PDA (personal digital assistants), a PMP (portable multimedia player), a navigation device, a slate PC, a tablet PC, an ultrabook, a wearable device (e.g., a smartwatch, a smart glass, a head-mounted display), etc.
[0058] More specifically, the electronic device (10) according to the present invention is not limited to an electronic device with an application activated, but may also mean an electronic device connected to an electronic device with an application activated. As an example, based on the fact that the electronic device (10) according to the present invention is a smartphone, the electronic device (10) may also mean a smart TV connected to said smartphone.
[0059] Meanwhile, the performance rate evaluation system (100) may exist inside a server (hereinafter referred to as the server) built to perform a specific purpose (e.g., evaluation of the exercise performance rate for an exercise movement), or it may exist as a separate device from the server. When the performance rate evaluation system (100) exists inside the server, the performance rate evaluation system (100) according to the present invention may evaluate the exercise performance rate for an exercise movement through at least one component among the communication unit (110), storage unit (120), and control unit (130) located inside the server, or through a module that performs a function similar to each of the above components. In this case, the application may provide performance evaluation information for an exercise movement on an electronic device (10) on which the application is installed through communication with the server. Furthermore, the performance rate evaluation system (100) according to the present invention may provide performance evaluation information for an exercise movement according to the present invention to the electronic device (10) by linking with a plurality of different external servers.
[0060] Meanwhile, the communication unit (110) of the performance rate evaluation system (100) according to the present invention may be connected to an electronic device (10), a VLM server (140), a central server, a device, and at least one network via a wireless or wired network, and configured to receive or transmit overall data and information necessary for the operation of the performance rate evaluation system (100) according to the present invention.
[0061] The communication unit (110) can receive user information corresponding to a user account logged into the electronic device (10). Additionally, the communication unit (110) can receive at least one of collected user voice data and exercise video using a microphone (11), a camera (12), and a sensor unit (13) provided in the electronic device (10). Here, the sensor unit (13) may include at least one sensor among an infrared sensor, a LiDAR sensor, an accelerometer, an illuminance sensor, a proximity sensor, a position sensor, a face recognition sensor, an iris scanner, a heart rate sensor, a touch sensor, and a pressure sensor.
[0062] At this time, the performance evaluation system (100) may further include a module for converting voice data into text. For example, in order to convert voice data received from an electronic device (10) into text, the present invention may include a conversion module comprising at least one of a Hidden Markov Model (HMM), a Gaussian / Deep Neural Net (GMM / DNN), a Weighted Finite State Transducer (WFST), a Connectionist Temporal Classification (CTC), a Recurrent Neural Network Transducer (RNN-T), an Attention-based Seq2Seq, a Self-Supervised Learning-based model, a Transformer-based model, and a Mamba-based model.
[0063] The communication unit (110) may include at least one communication module capable of wireless communication and wired communication between the performance rate evaluation system (100) and the communication target. Additionally, the communication unit (110) may include a communication module that connects the performance rate evaluation system (100) to at least one network.
[0064] Meanwhile, the communication unit (110) can support various communication methods depending on the communication standard of the communicating device. For example, the communication unit (110) may be configured to perform communication using at least one of the following technologies: WLAN (Wireless LAN), Wi-Fi (Wireless-Fidelity), Wi-Fi (Wireless Fidelity) Direct, DLNA (Digital Living Network Alliance), WiBro (Wireless Broadband), WiMAX (World Interoperability for Microwave Access), HSDPA (High Speed Downlink Packet Access), HSUPA (High Speed Uplink Packet Access), LTE (Long Term Evolution), LTE-A (Long Term Evolution-Advanced), 5G (5th Generation Mobile Telecommunication), Bluetooth (Bluetooth™ Frequency Identification), Infrared Communication (Infrared Data Association; IrDA), UWB (Ultra-Wideband), ZigBee, NFC (Near Field Communication), Wi-Fi Direct, and Wireless USB (Wireless Universal Serial Bus).
[0065] Next, the storage unit (120) may be configured to store various information related to the present invention. In the present invention, the storage unit (120) may be provided in the performance rate evaluation system (100) itself, or alternatively, at least a part of the storage unit (120) may mean a database (Database: DB, 200).
[0066] The storage unit (120) may include one or more non-transient computer-readable storage media that can be read and / or accessed by at least one processor. One or more computer-readable storage media may include volatile and / or non-volatile storage components such as optical, magnetic, organic, or other memory or disk storage devices. In some examples, the storage unit (120) may be implemented using a single physical device (e.g., one optical, magnetic, organic, or other memory or disk storage device), whereas in other examples, the storage unit (120) may be implemented using multiple physical devices.
[0067] The storage unit (120) may include computer-readable instructions and additional data. The storage unit (120) may include a storage necessary to perform at least some of the methods and techniques described herein and / or at least some of the functions of the device and network.
[0068] Furthermore, at least a portion of the storage unit (120) may be a cloud storage or a cloud server. That is, the storage unit (120) is sufficient as long as it is a space where information necessary for the operation of the image generation system (100) according to the present invention is stored, and it can be understood that there are no restrictions on the physical space. Accordingly, the database (200) and the storage unit (120) may be used interchangeably below.
[0069] Data and commands necessary for the operation of the counseling service providing system (100) according to the present invention may be stored in the storage unit (120). Specifically, commands for the operation of the prompt generation unit (131) may be stored in the storage unit (120). The prompt generation unit (131) according to the present invention may refer to a module that generates a prompt to be input into a vision language model based on user motion information extracted from a motion video and motion information collected from a database. There may be a wide variety of methods for generating a prompt to be input into a vision language model according to the present invention, and the present specification is not limited to any type or method as long as it is a module capable of generating a prompt to be input into a vision language model.
[0070] The performance rate evaluation system (100) according to the present invention may further include at least one of a module and an algorithm that perform the same function as the prompt generation unit (131). At this time, the prompt generation unit (131) may generate a prompt to be input to a vision language model (140) through at least one artificial intelligence model or module. For example, the prompt generation unit (131) may include at least one of a large language model based on T5 (Text-to-Text Transfer Transformer), BART (Bidirectional and Auto-Regressive Transformer), GPT (Generative Pre-trained Transformer), or LLaMA (Language Model for Many Applications), a rule-based template matching algorithm, a conditional prompting technique, a structured output prompting technique, a contextual embedding selection module, or a few-shot prompt generator. At this time, the prompt generation unit (131) according to the present invention is not limited to the model, algorithm, or module described above, and the performance rate evaluation system (100) according to the present invention may further include a model, algorithm, or module that performs the same function as the prompt generation unit (131).
[0071] Furthermore, commands for the operation of the motion information extraction module may be stored in the storage unit (120) according to the present invention. Here, the motion information extraction module may refer to a module that extracts user motion information including at least one of target body part information and motion performance information from a motion video. The method for extracting user motion information from a motion video according to the present invention may be very diverse, and the present specification is not limited to the type and method as long as it is a module capable of extracting user motion information from a motion video.
[0072] The control unit (130) according to the present invention may further include at least one of a module and an algorithm that perform the same function as the motion information extraction module. At this time, the motion information module can extract user motion information from a motion video through at least one algorithm. For example, the motion information extraction module may include at least one of an object detection model, a motion analysis model, a body part segmentation algorithm, a pose estimation model, an inter-frame motion analysis algorithm, and a vision language model, including at least one of a Convolutional Neural Network (CNN) based neural network, a Recurrent Neural Network (RNN) based neural network, and a Transformer based neural network.
[0073] For example, an object detection model is a model for identifying body parts within an image, and CNN-based object detection neural networks such as YOLO (You Only Look Once), Faster R-CNN, and SSD (Single Shot MultiBox Detector) can be used. Additionally, a body part segmentation algorithm is an algorithm for recognizing human body parts by separating them at the pixel level, and CNN-based segmentation models such as DeepLab, HRNet, and Mask R-CNN can be applied.
[0074] In addition, the pose estimation model of the present invention is a model for estimating human joint positions (keypoints) in 2D or 3D coordinates, and at least one neural network-based model, such as a CNN-RNN-based neural network or a Transformer-based neural network, such as OpenPose, PoseNet, BlazePose, HRNet, etc., may be utilized. The inter-frame motion analysis algorithm of the present invention is an algorithm for analyzing motion patterns between consecutive frames in a time-series manner, and may include a CNN or RNN-based model for time-series image analysis, such as an Optical Flow-based algorithm, 3D CNN, ConvLSTM, and Temporal Shift Module (TSM). At this time, the motion information extraction module according to the present invention is not limited to the models, algorithms, or modules described above, and the performance rate evaluation system (100) according to the present invention may further include a model, algorithm, or module that performs the same function as the motion information extraction module.
[0075] Furthermore, commands for the operation of a vision language model (132) may be stored in the storage unit (120) according to the present invention. The vision language model (132) may refer to an artificial intelligence model capable of analyzing a user's exercise movements based on an exercise video received from an electronic device (10) and exercise program information collected from the storage unit (120, or database, 200), and generating performance evaluation information for the exercise movements according to the analysis results. There may be a wide variety of methods for evaluating a user's exercise performance rate according to the present invention, and the present specification does not limit the method for evaluating a user's exercise performance rate regarding a user's exercise movements.
[0076] That is, the vision language model (132) according to the present invention is not limited in type or method as long as it is an artificial intelligence model capable of evaluating the user's exercise performance rate for the user's exercise movements based on exercise video received from an electronic device (10) and exercise program information collected from a storage unit (120, or database, 200). The control unit (130) according to the present invention may further include at least one of a module and an algorithm that perform the same function as the vision language model (132). For example, the vision language model (132) may include at least one artificial intelligence model among Flamingo, BLIP (Bootstrapping Language-Image Pretraining), BLIP-2, Gemini, GPT-4 with Vision, Kosmos-1, and Open Flamingo. At this time, the vision language model (132) according to the present invention is not limited to the models described above, and the performance rate evaluation system (100) according to the present invention may further include a model that performs the same function as the vision language model (132).
[0077] Specifically, the storage unit (120) may store user information corresponding to a user account logged into the electronic device (10). For example, the storage unit (120) may store user information including at least one of medical information, exercise history information, feedback information, and exercise program information. As an example, the medical information according to the present invention may include at least one of the user's age, the user's gender, the user's medical history, prescription information, and treatment plan. Additionally, the exercise history information according to the present invention may include at least one of the date of exercise performance, the composition of exercise movements included in the exercise program performed, the results of motion analysis for each exercise movement, and the results of exercise history analysis.
[0078] Furthermore, the feedback information according to the present invention may include at least one of the user's voice data received from a microphone provided in the electronic device (10) and response data to at least one survey provided to the electronic device (10) in relation to the exercise movement performed by the user.
[0079] In addition, the exercise program information according to the present invention may include reference motion information for each of a plurality of exercise movements constituting an exercise program assigned to a user account. Here, the reference motion information may include information regarding a predefined target body part according to the main movement of a specific exercise movement. Specifically, the reference motion information may include target body part information for at least one of a main target body part and an auxiliary target body part, depending on the degree of relevance to the main movement of the specific exercise movement. Here, the main target body part may refer to a body part where a core movement takes place in relation to the main movement of the specific exercise movement, and which should be considered preferentially in the evaluation of whether performance conditions are satisfied; and the auxiliary target body part may refer to a target body part that assists the movement of the main target body part or contributes to securing balance, stability, and accuracy of the exercise.
[0080] Furthermore, the exercise program information according to the present invention may further include at least one of the name of each of a plurality of exercise movements constituting an exercise program assigned to a user account, the number of times the exercise is performed, the timing of the exercise, and the difficulty level of the exercise, as well as an exercise guide video and an exercise description corresponding to each of the plurality of exercise movements.
[0081] Meanwhile, the database (200) according to the present invention may store performance condition information for a target body part according to the main movement of a specific exercise movement. At this time, the “performance condition information” according to the present invention is condition information referenced to evaluate the exercise performance rate for a specific exercise movement, and may include performance condition information set for a target body part in relation to a specific exercise movement.
[0082] Meanwhile, user authentication information may be stored in the storage unit (120). Here, “user authentication information” may refer to information used in a user authentication process performed to log in to a user account on an electronic device (10). As an example, user authentication information may be various, such as i) ID, ii) password, iii) password pattern, iv) user’s fingerprint authentication information, v) face authentication information, vi) voice authentication information, vii) iris authentication information, viii) vein authentication information, etc., set by the user.
[0083] Next, the control unit (130) may be configured to control the overall operation of the performance rate evaluation system (100) related to the present invention. The control unit (130) may process signals, data, information, etc. that are input or output through the components described above, or provide or process appropriate information and functions to the user.
[0084] The control unit (130) can control the output of a service page for evaluating the exercise performance rate of an exercise movement through a display unit (or touchscreen) provided in the electronic device (10). This service page may be output on the electronic device (10) through an application or web page installed on the electronic device (10). The service page may be a page linked to the performance rate evaluation system (100) according to the present invention and may be configured to be controlled by the performance rate evaluation system (100) according to the present invention.
[0085] Furthermore, if the service page is provided in the form of an application, the service page may be controlled by the CPU (Central processing unit) of the electronic device (10) on which the application is installed. In this case, the CPU of the electronic device (10) may provide at least one of performance evaluation information regarding the user's exercise movement and the exercise performance rate based on information provided by the performance rate evaluation system (100) according to the present invention.
[0086] Meanwhile, the control unit (130) can collect exercise program information set in a user account from a database (200, or, the control unit (120)). Furthermore, the control unit (130) can receive a video of the user’s exercise related to a specific exercise movement requested by the user among a plurality of exercise movements included in the set exercise program. Specifically, the control unit (130) can collect data related to the user’s exercise movement by using at least one of a sensor unit (11), a camera (12), and a microphone (13) provided in the electronic device (10). For example, the control unit (130) can collect a video of the user’s exercise regarding a specific exercise movement requested by the user among a plurality of exercise movements included in the set exercise program by using the camera (12) provided in the electronic device (10). In addition, the control unit (130) can collect voice data related to the specific exercise movement through the microphone (13) provided in the electronic device (10).
[0087] The control unit (130) can extract target body part information from a movement video using a motion information extraction module. Specifically, the motion information extraction module can identify the target body part that is predefined for a specific movement from the movement video.
[0088] Furthermore, the control unit (130) can extract user motion information including at least one of target body part information and motion performance information from a motion video using a motion information extraction module. Specifically, the control unit (130) can determine whether to identify a main target body part in the motion video, and based on the determination result, determine whether to extract user motion information for at least one of the main target body part and the auxiliary target body part.
[0089] At this time, if the main target body part is not identified in the exercise video, the control unit (130) may request the electronic device to include the main target body part in the exercise video. For example, the control unit (130) may output a guidance message to the electronic device (10) instructing it to adjust at least one of the video shooting position and posture so that the main target body part is included in the exercise video.
[0090] Furthermore, the control unit (130) can use the prompt generation unit (131) to generate a prompt that performs an analysis of whether the user's exercise motion satisfies pre-set execution conditions based on user motion information and reference motion information.
[0091] The control unit (130) can process the generated prompt as input to the vision language model (132) to generate performance evaluation information regarding the user's exercise movements through the vision language model (132). At this time, although the present invention describes the vision language model (132) being included and operated in the control unit (130), it is not limited thereto and can also be understood as using a vision language model (132) included in a VLM server (140) that exists separately from the performance rate evaluation system (100). The VLM server (140) may include at least one vision language model (132) among Flamingo, BLIP (Bootstrapping Language-Image Pretraining), BLIP-2, GPT-4 with Vision, Kosmos-1, and Open Flamingo. At this time, the vision language model (132) included in the VLM server (140) according to the present invention is not limited to the model described above, and the VLM server (140) according to the present invention may further include a model having the same function as the vision language model (132).
[0092] In the following description, for convenience of explanation, the VLM server (140) and the vision language model (132) are used interchangeably, and it can be understood that the control unit (130) uses the vision language model (132) as using at least one vision language model included in the VLM server (140). Furthermore, at least one of the vision language model (132) and the VLM server (140) according to the present invention can be linked with a separate external server (e.g., motion analysis server) to evaluate the exercise performance rate of the user's exercise movements.
[0093] The method for evaluating the exercise performance rate for a user's exercise movement according to the present invention may be very diverse, and the present specification is not limited to any type or method as long as it is a module capable of evaluating the exercise performance rate for a user's exercise movement. For example, the control unit (130) can evaluate the exercise performance rate for a user's exercise movement by analyzing the movement pattern of a body part. For example, the vision language model (132) can recognize movement patterns along the time axis based on 3D ConvNet and Spatiotemporal models such as I3D (Inflated 3D ConvNet), SlowFast, X3D, etc., or can identify semantic-based exercise performance states such as 'raising arms' or 'sitting' without joint coordinates by using a Transformer-based video encoder such as VideoMAE or TimeSformer. Furthermore, the vision language model (132) can evaluate the exercise performance rate for a user's exercise movement by utilizing a vision-language model such as CLIP to extract natural language descriptions and motion labels corresponding to video sequences.
[0094] As another example, the vision language model (132) can perform joint position extraction and motion tracking for analyzing the user's exercise movements based on a reference exercise movement from an exercise video containing exercise movements. To this end, the vision language model (132) may include a preprocessing algorithm that divides the input exercise video into frames and extracts position information of human joints from each frame, and can evaluate the exercise performance rate of the user's exercise movements based on the joint position information extracted from the exercise video.
[0095] The method for evaluating the exercise performance rate of a user’s exercise movement based on joint position information extracted from an exercise video according to the present invention may be very diverse, and the present specification does not limit the method for evaluating the exercise performance rate of a user’s exercise movement based on joint position information extracted from an exercise video. In the present invention, the vision language model (132) is not limited in type or method as long as it is a model capable of evaluating the exercise performance rate of a user’s exercise movement based on joint position information extracted from an exercise video. The control unit (130) according to the present invention may further include at least one of a module and an algorithm that perform the same function as the vision language model (132). For example, the vision language model (132) may include a pre-trained motion analysis model, and through the pre-trained motion analysis model, the user's movement can be evaluated based on joint position information extracted from an exercise video.
[0096] The motion analysis model according to the present invention is an artificial intelligence model trained using a learning data set that includes position information for joint points, and can analyze the exercise posture of a user (U) from exercise video data to be analyzed. In the present invention, the vision language model (132) is described as including a pre-trained motion analysis model to generate an analysis result for the user's exercise motion, but is not limited thereto. For example, the control unit (130) may train the vision language model (132) to analyze the user's exercise motion by linking with a motion analysis server (not shown) that includes a motion analysis model. That is, the control unit (130) can evaluate the exercise performance rate of the user's exercise motion based on joint position information extracted from the exercise video using the vision language model (132) that includes a pre-trained motion analysis model, and the vision language model (132) may also evaluate the exercise performance rate using a motion analysis model included in a separate motion analysis server.
[0097] Furthermore, the control unit (130) may request the re-performance of a specific exercise movement based on the exercise performance rate of the user's exercise movement. For example, the control unit (130) may provide a notification message to the electronic device (10) requesting the re-performance of the specific exercise movement based on the fact that the exercise performance rate for the specific exercise movement is below a preset standard.
[0098] Additionally, the control unit (130) can generate feedback information for a specific exercise movement based on the exercise performance rate of the user's exercise movement. For example, the control unit (130) can provide at least one survey to the electronic device (10) to collect feedback information for the specific exercise movement based on the fact that the exercise performance rate for the specific exercise movement is below a preset standard.
[0099] Specifically, the control unit (130) may provide at least one user survey related to the exercise performance rate of the user's exercise movements on a service page provided to the electronic device (10). Furthermore, the control unit (130) may receive survey response data for at least one user survey received through the electronic device (10).
[0100] In another example, the control unit (130) may receive voice data of the user regarding the exercise performance rate of the user's exercise movements by using a microphone (13) provided in the electronic device (10). Furthermore, the control unit (130) may use a prompt generation unit (131) to generate a feedback prompt that updates the exercise program assigned to the user account based on at least one of the voice data and the survey response data. Here, “updating the exercise program” may mean changing at least one of the exercise movements, the number of repetitions of the movements, and the total exercise time that constitute the exercise program.
[0101] Meanwhile, the performance evaluation system (100) may include one or more processors, and such processors may include one or more general-purpose processors and / or one or more special-purpose processors (e.g., digital signal processors, tensor processing units (TPUs), graphics processing units (GPUs), neural network processing units (NPUs), application integrated circuits, application semiconductors (ASICs), etc.). One or more processors may be configured to execute instructions, computer-readable instructions, and / or other instructions described herein that are stored (or included) in the storage unit (120). Such a performance evaluation system (100) may perform data processing described below in cooperation with memory and at least one processor. The processor may perform a series of operations and data processing using data and information stored in memory. Here, “memory” may be a component of the storage unit (120), and “processor” may be used interchangeably with the control unit (130).
[0102] In the above description, the performance rate evaluation system (100) of the present invention has been described, and it can be implemented based on the exercise performance rate evaluation method using a vision language model described below.
[0103] Hereinafter, with reference to FIG. 2 together with FIG. 3, FIG. 4a, FIG. 4b, FIG. 5a, FIG. 5b, and FIG. 6, a method for evaluating exercise performance using a vision language model according to the present invention will be described in more detail. FIG. 2 is a flowchart for explaining a method for evaluating exercise performance using a vision language model according to the present invention, and FIG. 3 is a flowchart for explaining a process of collecting exercise video related to a user's exercise movements from an electronic device according to the present invention. FIG. 4a and FIG. 4b are conceptual diagrams for explaining a process of extracting user movement information from an exercise video according to the present invention, FIG. 5a and FIG. 5b are conceptual diagrams for explaining a process of generating a prompt according to the present invention, and FIG. 6 is a conceptual diagram for explaining a process of evaluating an exercise performance rate for a user's exercise movements using a vision language model according to the present invention.
[0104] In the present invention, a process of requesting an electronic device to perform a specific exercise motion may be carried out (S210, see FIG. 2).
[0105] The control unit (130) can verify the user account logged into the electronic device (10). Furthermore, the control unit (130) can collect user information corresponding to the user account from at least one of the storage unit (120), the database (200), and an external server (e.g., a medical staff server). Here, “user information” may include at least one of medical information, exercise history information, feedback information, and exercise program information related to the user. The method for collecting user information corresponding to the user account according to the present invention may be very diverse, and the present specification is not limited to the type and method as long as it is a module capable of collecting user information corresponding to the user account.
[0106] For example, the control unit (130) can verify user identification information corresponding to a user account in order to collect user information. At this time, the user identification information may include various identification information such as a unique identifier (ID) of the user account, a session token, and a login token, and the control unit (130) can identify (or verify) the user account logged into the user terminal using the user identification information.
[0107] Based on the identification of a user account, the control unit (130) can retrieve user information corresponding to the identified user account from at least one of a storage unit (120) that stores user information corresponding to the identified user account information, a database (200), and an associated external server. Furthermore, the control unit (130) can collect the retrieved user information as user information corresponding to the user account.
[0108] As an example, the control unit (130) may collect medical information including at least one of the user's age, the user's gender, the user's medical history, prescription information, and treatment plan. Additionally, the control unit (130) may collect exercise history information including at least one of the user's exercise performance date, the composition of exercise movements included in the exercise program performed, motion analysis results for each exercise movement, and exercise history analysis results.
[0109] Furthermore, the control unit (130) can collect feedback information including at least one of the user's voice data received from a microphone provided in the electronic device (10) and at least one response data to at least one survey provided to the electronic device (10) in relation to the exercise motion performed by the user.
[0110] Additionally, as illustrated in FIG. 3, the control unit (130) may collect exercise program information (310) for an exercise program set in a user account from a database (200, or storage unit (120)). At this time, the exercise program information of the present invention may include exercise movement information corresponding to each of a plurality of exercise movements. Furthermore, the exercise movement information may include reference movement information for the exercise movements. For example, the control unit (130) may collect first reference movement information (321) for a first exercise movement (e.g., squat, 320)) included in the exercise program and second reference movement information (331) for a second exercise movement (e.g., lunge, 330).
[0111] In this case, reference motion information related to the exercise motion may include information regarding a predefined target body part (or target body part information) in relation to the main motion of a specific exercise motion. Specifically, a target body part may be defined and exist for a specific exercise motion according to the main motion of the specific exercise motion.
[0112] In addition, the reference motion information according to the present invention may include target body part information for at least one of a main target body part and an auxiliary target body part, depending on the degree of relevance to the main motion of the exercise motion. For example, the reference motion information according to the present invention may include target body part information for at least one of a first main target body part (e.g., leg, knee, etc.) and a first auxiliary target body part (e.g., arm, waist, etc.), depending on the degree of relevance to the main motion (e.g., leg bending) of a first exercise motion (e.g., squat). Here, “relevance to the main motion” may refer to the degree of contribution of a body part that is critically involved in performing a specific exercise motion in the correct posture and securing the exercise effect of the specific exercise motion; such relevance to the main motion may be set in advance by medical personnel (e.g., medical staff, physical therapist, etc.) or may be set by an artificial intelligence model that has learned the relevance to the main motion based on a training data set for the specific exercise motion.
[0113] Additionally, reference motion information related to the exercise motion may include performance condition information for each of the predefined target body parts. For example, it may include performance condition information for each of the multiple predefined body parts (e.g., both feet, knees, pelvis, upper body, both arms, etc.) related to the first exercise motion (e.g., squat). In this case, the performance condition may refer to a condition set for a specific exercise motion based on movement characteristics regarding at least one of the position, posture, direction of movement, angle, alignment state, and whether the target body part is stationary.
[0114] Furthermore, the control unit (130) may request the electronic device (10), to which a user account is logged in, to perform an exercise for a specific exercise movement based on collected user information. For example, the control unit (130) may provide the electronic device (10) with an exercise guide video that matches a specific exercise movement included in an exercise program set (or assigned) to the user account. Specifically, the control unit (130) may control the electronic device (10) so that an exercise guide video corresponding to at least one exercise item is played sequentially according to the exercise program set in the user account.
[0115] The control unit (130) may request the user to perform an exercise by following an exercise guide video according to at least one exercise item. As an example, the control unit (130) may request the user to perform an exercise for a specific exercise through at least one of an exercise request message for an exercise (e.g., a pop-up message), a visual graphic related to the exercise request, a voice-based exercise start guidance signal, and a user interface screen that induces the exercise performance.
[0116] Next, in the present invention, a process of receiving a user’s exercise video related to a specific exercise movement from an electronic device may be performed (S220, see FIG. 2).
[0117] The control unit (130) can activate a sensor unit (11) including a microphone (13), a camera (12), and multiple different sensors provided in the electronic device (10) to receive an exercise video of the user's exercise movements. Furthermore, the control unit (130) can receive sensing information including at least one of video data and voice data based on the activation of at least one of the sensor unit (11) including a microphone (13), a camera (12), and multiple different sensors provided in the electronic device (10). For example, the control unit (130) can activate a camera provided in the electronic device (10) to receive an exercise video (310) including the user's exercise movements from the camera. Here, "exercise video (340)" may refer to video data captured in the process of the user performing a specific exercise movement in accordance with an exercise performance request for an exercise movement included in an exercise program pre-set in the user account. At this time, the explanation assumes that the user account is logged in to the application in advance.
[0118] At this time, the present invention describes receiving an exercise video (340) including the user's exercise movements using a camera (13) provided in an electronic device (10), but is not limited thereto, and may also receive an exercise video from a separate video equipment connected to the electronic device in at least one of wireless communication and wired communication. The method for receiving an exercise video (340) including the user's exercise movements according to the present invention may be very diverse, and in this specification, any module capable of receiving an exercise video (340) including the user's exercise movements is not limited to any type or method.
[0119] Furthermore, the control unit (130) can output an exercise video (340) received through a camera to the display unit of the electronic device (10) in real time. Specifically, based on camera activation, the control unit (130) can output an exercise video (340) in which a user performs an exercise movement according to an exercise guide video output to the display unit of the electronic device (10). Furthermore, the control unit (130) can store the exercise video (340) received from the electronic device (10) in a database (200).
[0120] Next, in the present invention, a process of extracting user motion information regarding a user’s motion corresponding to a specific motion from a motion video may be performed (S230, see FIG. 2).
[0121] The control unit (130) can extract user motion information from a user’s exercise video received from an electronic device (10) according to the performance of an exercise for a specific exercise motion. Here, the user motion information may include target body part information related to the user’s body part included in the user’s exercise video, a position corresponding to the target body part, the direction and speed of movement of the target body part, and motion performance information corresponding to at least one of the position change of the target body part. The target body part of the present invention may be defined in advance by medical personnel (e.g., medical professionals, physical therapists, etc.) or set by an artificial intelligence model trained to define the target body part based on a training data set for a specific exercise motion.
[0122] The method for extracting user motion information from a user’s exercise video according to the present invention may be very diverse, and the present specification is not limited to any type or method as long as it is a module capable of extracting user motion information from a user’s exercise video. The control unit (130) according to the present invention may further include at least one of a module and an algorithm that perform the same function as the motion information extraction module. At this time, the motion information extraction module can extract user motion information from the user’s exercise video through at least one algorithm.
[0123] As illustrated in FIG. 4a, the control unit (130) can extract user motion information (450) for a user’s exercise motion corresponding to a specific exercise motion from an exercise video using a motion information extraction module (410). Specifically, the motion information extraction module (410) can identify a predefined target body part for a specific exercise motion from an exercise video based on at least one of an exercise video (340) and exercise program information (310).
[0124] The method for identifying a predefined target body part for a specific exercise motion from an exercise video according to the present invention may be very diverse, and in this specification, any module capable of identifying a predefined target body part for a specific exercise motion from an exercise video is not limited to any specific type or method. The control unit (130) according to the present invention may further include at least one of a module and an algorithm that perform the same function as the motion information extraction module (410). At this time, the motion information extraction module (410) can identify a predefined target body part for a specific exercise motion from an exercise video through at least one algorithm.
[0125] For example, the control unit (130) can identify a predefined target body part for a specific exercise motion from an exercise video by using at least one of an object detection model and a body part segmentation algorithm included in the motion information extraction module (410). Specifically, the motion information extraction module (410) can detect a body part region in each frame constituting the exercise video by using a pre-trained object detection model. Furthermore, the motion information extraction module (410) can identify a body part region corresponding to a predefined target body part by analyzing the location, size, and relative positional relationship with other body part regions within the exercise video of the detected body part region.
[0126] The control unit (130) can analyze the movement of the target body part in the exercise video and extract user motion information for a specific exercise movement. Specifically, the control unit (130) can determine whether the main target body part is identifiable in the exercise video. As previously described, the main target body part may refer to a body part where a core movement takes place in relation to the main movement of a specific exercise movement. In the present invention, the main target body part may be predefined for the main movement of a specific exercise movement, or it may be set by medical personnel. In the present invention, the main target body part may be set in various ways depending on the exercise movement, and the entity setting the main target body part is not limited.
[0127] The control unit (130) can use the motion information extraction module (410) to calculate target body part information including whether a target body part is identified from the motion video (340), and based on the target body part information, determine whether a main target body part is identified.
[0128] Referring to FIG. 4a, the control unit (130) can identify whether the main target body part (433) and auxiliary target body parts (431, 432) among the target body parts (430) defined in the exercise video (340) are included for the main movement of a specific exercise movement. Specifically, the control unit (130) can use an exercise information extraction module to extract target body part information (420) that includes information on at least one of the target body part (330) and whether the target body part is included in the video (440).
[0129] For example, the control unit (130) can use target body part information (420) to identify whether a predefined main target body part (e.g., left leg, right leg, 433) is included in the exercise video for the main movement of a specific exercise movement (e.g., squat). At this time, the control unit (130) can perform an exercise performance rate evaluation according to the present invention based on the target body part information (420) and based on whether the main target body part (433) is included in the exercise video (340).
[0130] Specifically, the control unit (130) determines whether to identify a main target body part in the exercise video (340), and, based on the determination result, determines whether to extract user motion information for the target body part. More specifically, the control unit (130) determines whether to extract user motion information for at least one of the main target body part and the auxiliary target body part based on the determination result.
[0131] The control unit (130) can extract user motion information from the exercise video (340) based on the target body part information (420) and based on the fact that the main target body part (433) is included in the exercise video. Specifically, the control unit (130) can extract user motion information for a specific exercise motion by analyzing the movement of the target body part in the exercise video using a motion information extraction module. For example, the control unit (130) can extract at least one of the motion performance information (460) according to the movement of the target body part included in the target body part information (420) and the exercise video (340) using a motion information extraction module.
[0132] As previously explained, the motion performance information according to the present invention may include information corresponding to at least one of the target body part information related to the user's body part included in the user's exercise video, the position corresponding to the target body part, the direction and speed of the movement of the target body part, and the position change of the target body part. For example, the motion information extraction module (410) may extract motion performance information (460) based on the movement of the target body part included in the exercise video (340) using a pre-trained motion analysis model. Here, the pre-trained motion analysis model may refer to an artificial intelligence model trained to analyze the user's exercise motion by learning the temporal movement pattern of the target body part from the input exercise video and inferring at least one of the position change, speed, acceleration, and directionality of the target body part. Specifically, the motion information extraction module (410) may extract motion performance information (460) based on the position coordinate change, inter-frame movement vector, and speed and direction information based on the movement of the target body part included in the exercise video (340) using a pre-trained motion analysis model.
[0133] As another example, a pre-trained motion analysis model may refer to an artificial intelligence model capable of analyzing a user's exercise movements from an exercise video to be analyzed, which is trained based on a training dataset containing positional information for joint points.
[0134] In the present invention, a motion information extraction module (410) including a previously learned motion analysis model motion is described as extracting motion performance information (460) based on the movement of a target body part included in a motion video (340), but it is not limited thereto. It may also be configured to link with a separately provided motion analysis server to call the function of the motion analysis model included in the motion analysis server or to extract motion performance information (460) based on the movement of a target body part included in the motion video (340). In other words, the present invention may include both an implementation form in which the motion information extraction module (410) directly performs the function of the motion analysis model by incorporating it, and various implementation forms in which it performs in cooperation with an external server or an external analysis module.
[0135] Meanwhile, the motion information extraction module (410) according to the present invention may include a motion generation model. Here, the “motion generation model” may refer to a model that learns the movement of a user and generates a vector corresponding to the user’s movement (or motion). The motion generation model may analyze the user’s movement in a video to extract time-series data, and use the extracted time-series data to generate vector data including at least one of a joint position, joint angle, velocity, and acceleration corresponding to the user’s skeletal structure.
[0136] Furthermore, the motion generation model can generate motion data in which the user's movement over a specific period of time is vectorized based on the generated vector data, and can visually output the motion data in at least one of two-dimensional and three-dimensional spaces. For example, the motion generation model may include at least one of an artificial intelligence model based on an RNN (Recurrent Neural Network), LSTM (Long Short-Term Memory), GRU (Gated Recurrent Unit), CNN (Convolutional Neural Network), Transformer, or GNN (Graph Neural Network). In this way, the motion information extraction module (410) according to the present invention can extract motion performance information (460) based on the movement of a target body part included in the motion video (340) using the motion generation model.
[0137] Hereinafter, an example of an embodiment in which a motion information extraction module (410) is linked with a motion analysis server to extract motion performance information (460) according to the movement of a target body part included in a motion video (340) is described in detail. As previously described, the motion information extraction module (410) is trained to implement the motion of a motion analysis model included in the motion analysis server, and can extract motion performance information (460) according to the movement of a target body part included in a motion video (340) independently. Here, the “motion analysis model” is a motion analysis model trained using a learning data set that includes position information for joint points, and can extract motion performance information (460) from a motion video.
[0138] Referring to FIG. 4b, the motion analysis server (50) according to the present invention may refer to a cloud server that performs motion analysis of a user from a motion video (400) that captures the user's motion. The motion analysis server (50) can analyze the relative positional relationship between key points (P1, P2) corresponding to a plurality of joint points of the user (U) extracted from the motion video (400) through a motion analysis model learned using learning data related to joint points. Here, "joint point" may refer to a plurality of joints of the user (U) (or a part of the user (U)'s body including joints). And, "key point" may refer to an area corresponding to each of the plurality of joint points of the user (U) in the motion video (400). Accordingly, in the present invention, "joint point" and "key point" may be used interchangeably, and the same reference numeral "P2" may be assigned to each joint point and key point to explain them.
[0139] The control unit (130) can use a motion analysis model (52) to extract key points (P1, P2) corresponding to joint points from a user’s exercise video (400), and analyze the movement of a target body part based on an analysis of the positional relationship between the extracted key points (P1, P2). In the present invention, a series of processes for analyzing the movement of a target body part from an exercise video (400) using key points extracted through the motion analysis model (52) can be named the “exercise motion analysis process.”
[0140] In the present invention, the physical space and subject where the motion analysis process takes place are not separately distinguished, and it can be described as taking place in the performance rate evaluation system (100). The motion analysis process can be performed using key points extracted from the motion analysis model (52). As previously described, the control unit (130) can use the motion information extraction module (410) to extract key points corresponding to each of a plurality of pre-set joint points from a specific object corresponding to the user included in the motion video.
[0141] More specifically, the motion information extraction module (410) can analyze video data received from a camera on a frame-by-frame basis to extract key points corresponding to joint points corresponding to the user's movement. For example, the motion information extraction module (410) can use various object detection algorithms. For example, the motion information extraction module (410) can use an algorithm that ensembles multiple bounding boxes (Weighted Box Fusion, WBF). However, it is obvious that the motion information extraction module (410) is not limited to the object detection algorithm described above, but can use various object detection algorithms capable of detecting objects corresponding to the user (U) from video data. As another example, the motion analysis model (52) included in the motion analysis server (50) can identify or estimate the user's joint points from the movement video (400) through learning on training data specialized for joint points, and extract key points corresponding thereto.
[0142] In the present invention, the training data for which the motion analysis model (52) performs training may be stored in a motion analysis database (60), and such a motion analysis database (60) may also be named a “training data DB.” Further details regarding the training data will be described later.
[0143] According to the present invention, the motion analysis server (50) may include at least one of a learning unit (51) and a motion analysis model (52). The motion analysis server (50) may be provided inside the performance rate evaluation system (100) according to the present invention or may be an external server. That is, the motion analysis server (50) according to the present invention performs the function of learning the movement analysis of a target body part in conjunction with the motion information extraction module (410), and it can be understood that there are no physical spatial constraints. Detailed information regarding the motion analysis server (50) will be described later along with the learning data.
[0144] The motion analysis database (60) is a storage facility where a learning data set is stored, and may be provided within the performance rate evaluation system (100) itself according to the present invention or may be an external storage facility (or external DB). It can be understood that the motion analysis database (60) according to the present invention is sufficient as long as it is a space where the learning data set is stored, and there are no restrictions on the physical space.
[0145] Meanwhile, the “exercise video” described in the present invention may include at least one of “exercise video data to be analyzed” and “exercise video data to be learned.” The “exercise video data to be analyzed” is exercise video data that is subject to movement analysis of a target body part, and the “exercise video data to be learned” can be understood as an exercise video (400) that is subject to machine learning for a motion analysis model.
[0146] The learning unit (51) may be configured to perform learning for a motion analysis model (52) based on the exercise video data to be learned. The learning unit (51) may train the motion analysis model (52) using the learning data. The learning unit (51) may detect a user (U) in the exercise video data to be learned and extract various learning data used for estimating exercise posture from the detected user (U). Such learning data may be used interchangeably with “information,” “data,” “data value,” or “data value.”
[0147] Meanwhile, the extraction of training data may be performed by means other than the training unit (51). The training unit (51) may use various object detection algorithms to detect the user (U) from the training target exercise video data. For example, the training unit (51) may use an algorithm that ensembles multiple bounding boxes (Weighted Box Fusion, WBF). However, it is obvious that the training unit (51) is not limited to the object detection algorithm described above and may use various object detection algorithms capable of detecting the movement of a target body part from the training target exercise video data.
[0148] Furthermore, the learning unit (51) can perform learning for the motion analysis model (52) based on the learning data set existing in the motion analysis database (60). As previously explained, the learning data set may include location information of joint points. The motion analysis model (52) is a motion analysis model learned using the learning data set containing location information for joint points, and can analyze the movement of a target body part from the motion video data to be analyzed.
[0149] Meanwhile, the motion analysis model (52) can extract key points corresponding to the user's joint points from the motion video (400) using the learning data set generated by the learning unit (51). The motion analysis model (52) can analyze the movement of the target body part in the motion video (400) using the extracted key points. Specifically, the motion analysis model (52) can analyze the relative positional relationship between key points and perform an analysis of the movement of the target body part based on the analysis of the positional relationship. For example, it can estimate and analyze information regarding at least one of i) the position of the joint point, ii) the range of motion of the joint point, iii) the movement path of the joint point, iv) the connection relationship between the joint points, and v) the symmetry relationship of the joint point for the user (U).
[0150] Furthermore, the motion analysis model (52) can perform analysis on at least one of the range of motion of the joint, the speed (or acceleration) of the movement of the joint, the change in position of the target body part, the speed, the acceleration, and the directionality. In the present invention, the motion analysis model (52) may also be configured to include a learning unit (51). Furthermore, conversely, the learning unit (51) may include the motion analysis model (52), and in this case, the learning unit (51) can train the motion analysis model (52) to perform the motion analysis function of the target body part. Accordingly, in the present invention, the function performed by the motion analysis model (52) may be described interchangeably as being performed by the learning unit (51).
[0151] That is, the method for the motion information extraction module (410) according to the present invention to extract motion performance information (460) based on the movement of a target body part included in the exercise video (340) can be very diverse, and in this specification, any artificial intelligence model capable of extracting motion performance information (460) based on the movement of a target body part included in the exercise video (340) is not limited to any specific type or method.
[0152] As described in the example above, the control unit (130) determines whether the main target body part is identified in the exercise video, and based on whether the main target body part is included in the exercise video, can extract user motion information for at least one of the main target body part and the auxiliary target body part.
[0153] In a different example, if the main target body part is not identified in the exercise video, the control unit (130) may request the electronic device to include the main target body part in the exercise video. Specifically, the control unit (130) may output a guidance message to the electronic device instructing it to adjust at least one of the video shooting position and posture so that the main target body part is included in the exercise video.
[0154] For example, the control unit (130) can output a guidance message (e.g., “Please ensure the right leg and left leg are included in the screen”) on the electronic device (10) so that the main target body part is included within a specific area of the exercise video (or the display of the electronic device) in order to identify the main target body part from the exercise video.
[0155] That is, the control unit (130) can extract user motion information (450) including at least one of target body part information and motion performance information for a specific motion from a motion video through different processes, depending on whether the main target body part is included in the motion video.
[0156] Meanwhile, in the present invention, a process of generating a prompt to perform an analysis of whether the user's exercise motion satisfies a preset execution condition using user motion information and reference motion information corresponding to a specific exercise motion may be carried out (S240, see FIG. 2).
[0157] As illustrated in FIG. 5a, the control unit (130) can generate a prompt to perform an analysis of whether the user's exercise motion satisfies a preset execution condition based on at least one of user motion information (450) extracted from an exercise video (340) and reference motion information (510) extracted from a database (200) using a prompt generation unit (131).
[0158] The prompt generation unit (131) according to the present invention can utilize natural language processing (NLP) technology to receive multiple different data inputs, extract prompt information, and generate a prompt to be input into a vision language model (132). Here, natural language processing (NLP) technology may refer to technology capable of understanding, interpreting, and generating human language using artificial intelligence technology (e.g., deep learning).
[0159] Specifically, the prompt generation unit (131) can generate a prompt based on prompt engineering by using at least one of user action information and reference action information. Here, prompt engineering may refer to a technique for designing and optimizing input text (or prompt) to effectively utilize a natural language processing (NLP) model.
[0160] As previously explained, the database (200) according to the present invention may store reference motion information (510) corresponding to a specific exercise motion. At this time, the reference motion information corresponding to the specific exercise motion may store performance condition information for a target body part according to the main motion of the specific exercise motion. At this time, the “performance condition information” according to the present invention is condition information referenced to evaluate the exercise performance rate for a specific exercise motion, and may include performance condition information set for a target body part in relation to the specific exercise motion. At this time, the performance condition may refer to a condition set for a specific exercise motion based on movement characteristics regarding at least one of the position, posture, direction of movement, angle, alignment state, and whether or not the target body part is stopped.
[0161] As illustrated in FIG. 5b, the performance condition information (520) of the present invention may include at least one of whether the target body part (430) is included in the video according to the main movement of a specific exercise movement (440), whether the performance condition is applied (531), and the performance condition (540). At this time, it may be understood that in the present invention, the pre-set performance condition (540) is applied to the target body part (432) included in the exercise video, and the pre-set performance condition (450) is not applied to the target body part (431) that is not included in the exercise video.
[0162] The prompt generation unit (131) according to the present invention can generate a prompt (500) that performs an analysis of whether the user's exercise movement satisfies pre-set performance conditions by using user motion information (450) and reference motion information (510). Specifically, the prompt generation unit (131) can generate a prompt (500) that includes at least one of motion performance information (461) extracted from the user's exercise video and target body part information (420) according to the main motion of a specific exercise movement. Additionally, the prompt generation unit (131) can generate a prompt (500) that includes performance condition information (520) for each target body part according to the main motion of a specific exercise movement.
[0163] That is, the control unit (130) can use the prompt generation unit (131) to generate a prompt (500) that analyzes whether the user's exercise movement satisfies a preset execution condition based on the extracted action execution information (461) and the target body part information (420), and the execution condition information (520) for the target body part.
[0164] Meanwhile, in the present invention, a process may be performed in which a generated prompt and a movement video are input into a pre-trained vision language model to obtain performance evaluation information regarding the user's movement from the pre-trained vision language model (S250, see FIG. 2).
[0165] As illustrated in FIG. 6, the control unit (130) can input a motion video (340) received from the electronic device (10) and a prompt (500) to perform an analysis of whether the user's motion satisfies pre-set performance conditions into a pre-trained vision language model. As previously described, the vision language model (132) according to the present invention may refer to a multimodal neural network-based vision language model that is trained to receive both image data and natural language text data together and to infer whether the user's motion satisfies pre-set performance conditions by analyzing the semantic relationship between the image data and the text data.
[0166] For example, the pre-trained vision language model (132) may refer to a vision language model that has been pre-trained using training data that includes a movement video containing body movements according to multiple movement actions and a natural language description of a pre-set performance condition for each of the multiple movement actions. Alternatively, the pre-trained vision language model may refer to a vision language model trained by the control unit (130) to infer whether a user's movement action satisfies a pre-set performance condition using training data that includes a movement video for training, movement program information for training, and correct answer performance evaluation information.
[0167] The control unit (130) can use a pre-trained vision language model (132) to analyze whether a pre-set performance condition is satisfied in a target body part for a specific exercise movement for a pre-set number of times. Specifically, the vision language model (132) can analyze whether a pre-set performance condition is satisfied in a target body part for a specific exercise movement according to a pre-set number of times. For example, the vision language model (132) can analyze whether a pre-set performance condition is satisfied in a target body part for each of the pre-set number of times the specific exercise movement is performed. Alternatively, the vision language model (132) can analyze whether a pre-set performance condition is satisfied in a target body part for each of the pre-set number of times the specific exercise movement is performed.
[0168] In the present invention, when an action is performed a predetermined number of times, it is described that the determination of whether a predetermined execution condition is satisfied in a target body part is analyzed for each of the predetermined number of times, but it is not limited thereto, and depending on the predetermined number of times, the determination of whether a predetermined execution condition is satisfied in a target body part while performing a specific exercise action may also be analyzed.
[0169] The method for analyzing whether a pre-set performance condition is satisfied in a target body part according to the present invention may be very diverse, and in this specification, any module capable of analyzing whether a pre-set performance condition is satisfied in a target body part is not limited to its type or method. The control unit (130) according to the present invention may further include at least one of a module and an algorithm that perform the same function as the vision language model (132). At this time, the vision language model (132) can analyze whether a pre-set performance condition is satisfied in a target body part through at least one algorithm. For example, the vision language model (132) can analyze whether a pre-set performance condition is satisfied in a target body part based on the movement video and the prompt by inferring whether there is a semantic match between the movement of the target body part included in the input exercise video and the natural language description of the performance condition included in the prompt.
[0170] At this time, it can be understood that the previously trained vision language model (132) is trained to perform the actions of the motion information extraction model (410) and can analyze the movement of a target body part from a motion video. Since the method of the motion information extraction model (410) according to the present invention identifying a target body part from a motion video and analyzing the movement of the target body part has already been explained above, a redundant explanation regarding the specific method of the previously trained vision language model (132) analyzing the movement of the target body part from a motion video is omitted below.
[0171] The control unit (130) according to the present invention can generate performance evaluation information (600) for a user’s exercise movement corresponding to a specific exercise movement by using a previously learned vision language model (132). Specifically, the vision language model (132) can generate performance evaluation information (600) for a user’s exercise movement corresponding to a specific exercise movement by using motion performance information of a target body part, target body part information, and reference motion information (or performance condition information) extracted from at least one of an exercise video (340) and a prompt (500).
[0172] For example, the vision language model (132) can generate performance evaluation information (600) including evaluation results (620) corresponding to each of the pre-set number of performances (or, number of performances, set of performances, 610) for a specific exercise movement. Specifically, the vision language model (132) can analyze a frame sequence corresponding to a specific number of performances (or, number of performances, set of performances) in an exercise video (340) and determine whether a target body part has satisfied a pre-set performance condition at a specific number of performances according to the performance condition information included in the prompt (500). For example, if the vision language model (132) has pre-set conditions for each target body part according to the main movement (e.g., leg bending) of a specific exercise movement (e.g., squat), the vision language model (132) can analyze whether the pre-set conditions are satisfied for each number of times the exercise is performed, on a frame-by-frame basis.
[0173] There may be a wide variety of methods for analyzing whether the performance conditions set for the target body part according to the present invention are satisfied, and the present specification is not limited to any type or method as long as it is a module capable of analyzing whether the performance conditions set for the target body part are satisfied. As an example, a vision language model (132) can analyze whether the performance conditions set for the main target body part according to the main movement of a specific exercise movement are satisfied. Specifically, if the performance conditions for the main target body part are not satisfied, the vision language model (132) can generate a first evaluation result (e.g., 'Unsatisfied', 622) for the user's exercise movement.
[0174] As another example, the vision language model (132) can analyze whether the performance conditions for the main target body part corresponding to the specific exercise movement are satisfied, and whether the performance conditions for the auxiliary target body part corresponding to the specific exercise movement are satisfied. Specifically, the vision language model (132) can generate a first evaluation result (e.g., 'not satisfied', 622) for the user's exercise movement if the number of auxiliary target body parts that do not satisfy the performance conditions for the main target body part corresponding to the specific exercise movement is greater than a certain threshold (e.g., 2 or more) even if the performance conditions for the main target body part corresponding to the specific exercise movement are satisfied.
[0175] In contrast, the vision language model (132) can generate a second evaluation result (e.g., 'satisfied', 621) if the performance conditions set for the main target body part according to the main movement of a specific exercise movement are satisfied, and even if some of the auxiliary target body parts fail to meet the evaluation, the number of auxiliary target body parts that fail to meet the evaluation is less than a specific threshold. At this time, the specific threshold can be set in various ways by the user, medical staff, and the system.
[0176] Furthermore, the control unit (130) may obtain performance evaluation information regarding the user's exercise movements from a pre-trained vision language model (132) based on the analysis result regarding the satisfaction status. Specifically, the vision language model (132) may generate performance evaluation information for the user's exercise movements, including at least one of information on the number of times (or number of times, sets) of a specific exercise movement that is pre-set, an evaluation result (620) for each of the number of times of the exercise, and information on the reason for evaluation (630) for each of the evaluation results.
[0177] Next, in the present invention, a process of evaluating the user's exercise performance rate based on performance evaluation information may be carried out (S260, see FIG. 2).
[0178] The control unit (130) can evaluate (or calculate) the user's exercise performance rate using performance evaluation information regarding the user's exercise movement according to a specific exercise movement requested for execution. Here, the exercise performance rate may refer to a quantitative indicator calculated based on whether the user's exercise movement satisfies the performance conditions with a reference movement corresponding to the specific exercise movement.
[0179] The method for evaluating the user’s exercise performance rate according to the present invention may be very diverse, and the control unit (130) may evaluate the user’s exercise performance rate through various methods. For example, the control unit (130) may calculate the exercise performance rate (640) in a weighted average manner by applying a predefined weight based on whether the performance conditions of each main target body part and auxiliary target body part included in the performance evaluation information (600) are satisfied. As an example, the control unit (130) may calculate the exercise performance rate by applying a first weight (e.g., 0.7) to the main target body part and a second weight (e.g., 0.3) to the auxiliary target body part, and then multiplying and summing the performance condition satisfaction rates for each target body part. Here, “performance condition satisfaction rate” may refer to a quantitative indicator calculated based on the number (or ratio) of performance conditions actually satisfied among one or more performance conditions pre-set for a specific exercise movement.
[0180] As another example, the control unit (130) can calculate the exercise performance rate based on the ratio of frames satisfying the performance conditions among all frames by accumulating whether the performance conditions are satisfied on a frame-by-frame basis during the performance of a specific exercise movement. For example, if the ratio of frames satisfying the performance conditions among all frames constituting the exercise video is 85%, the control unit (130) can evaluate the user's exercise performance rate as 85 points (or 85%).
[0181] Additionally, the control unit (130) can calculate the exercise performance rate by calculating the ratio of the number of times a specific exercise movement satisfies the performance conditions among the number of times a specific exercise movement is pre-set, based on the evaluation results corresponding to the number of times (or number of times, sets) of performance included in the performance evaluation information. Specifically, the control unit (130) can calculate the ratio of the number of times a specific exercise movement satisfies the performance conditions pre-set for a target body part among the pre-set number of times a specific exercise movement. Furthermore, the control unit (130) can evaluate the user's exercise performance rate as at least one of a performance score and a percentile based on the ratio of the calculated number of times a specific exercise movement is pre-set as “5 times.” For example, if the number of times a specific exercise movement is pre-set as “5 times,” the control unit (130) can evaluate the user's exercise performance rate (640) as 80 points (or 80%) based on the fact that among the “5 times,” the number of times a specific exercise movement is a second evaluation result (e.g., “satisfied”) is “4 times.”
[0182] That is, the exercise performance rate according to the present invention can be evaluated (or calculated) in various ways, and the present invention is not limited to the methods exemplified above, and may also include other evaluation methods capable of calculating the user's exercise performance rate based on the results of the performance evaluation.
[0183] In the foregoing, a method for evaluating exercise performance using a vision language model according to the present invention has been described in detail. Below, with reference to FIGS. 7a to 7c, a process for requesting the re-execution of an exercise movement and updating an exercise program based on the user's exercise performance rate will be described in detail according to the method for evaluating exercise performance using a vision language model. FIGS. 7a to 7c are conceptual diagrams for explaining an embodiment in which a request for the re-execution of an exercise movement is made based on the exercise performance rate according to the present invention, and an exercise program set in a user account is updated using user feedback information.
[0184] Meanwhile, in the present invention, a process of requesting the user to perform a specific exercise movement corresponding to the user's exercise movement again can be carried out based on the exercise performance rate calculated for the user's exercise movement.
[0185] As illustrated in FIG. 7a (a), the control unit (130) may provide a request message to the electronic device requesting the re-execution of a specific exercise movement based on the fact that the evaluated user's exercise performance rate does not satisfy a preset standard condition. Here, the preset standard condition may refer to a condition satisfied by the evaluated user's exercise performance rate being greater than or equal to a preset target value. Specifically, the control unit (130) may determine whether the user's exercise performance rate satisfies the preset standard condition based on the preset target value. At this time, the preset target value in the present invention may be set in various ways, and the present specification does not limit the method of setting the preset target value. For example, the control unit (130) may be set according to user information corresponding to a user account logged into the electronic device. Specifically, the control unit (130) may be set by at least one of the user, medical staff, and system, taking into account the difficulty level of the exercise program set in the user account, the type of specific exercise movement, the user's age, gender, and exercise history information.
[0186] For example, the control unit (130) may output (or provide) a request message (710) to the service page (700) requesting the user to perform a specific exercise action again based on the user's exercise performance rate being a first performance rate (e.g., 60%) when the user's exercise performance rate is a specific target value (e.g., 80%).
[0187] Specifically, the control unit (130) may receive a first user input for a request icon (711) displayed in a portion of a request message (710) from the electronic device (10). Here, the first user input may refer to user input for a request icon (711) displayed in a portion of a request message (710), and such first user input may be performed in at least one of a tap, double tap, long press, click, swipe, drag, pinch in, pinch out, and rotate on the icon. In the present invention, "receiving user input" may mean receiving an input signal (or selection signal) corresponding to the user input based on input being made by a user through an input unit configuration provided in the electronic device (10).
[0188] In addition, in the present invention, the input unit does not necessarily refer to a hardware means, but can be understood as a channel for receiving input from a user. The input unit may also be referred to as a user interface module. The input unit may include a touch screen, touch input means (e.g., virtual key, soft key, visual key, touch key, etc.), a computer mouse, a keyboard, a keypad, a touchpad, a trackball, a joystick, a dome switch, a jog wheel, a jog switch, a voice recognition module, or other similar devices. However, the present invention does not limit the type of input unit.
[0189] The control unit (130) can re-evaluate the exercise performance rate of a user’s exercise corresponding to a specific exercise motion based on the occurrence of an activation event for the request icon (711) according to the first user input. Specifically, when an activation event for the request icon (711) occurs, the control unit (130) can receive the user’s exercise video performed after the first user input again. Furthermore, the control unit (130) can analyze whether the performance conditions for a target body part are satisfied by using the new user motion information extracted from the exercise video and the reference motion information corresponding to the specific exercise motion. The control unit (130) can generate updated performance evaluation information for the specific exercise motion according to the analysis results using a vision language model (132) and recalculate the exercise performance rate based thereon.
[0190] Alternatively, the control unit (130) may receive user input for a progress icon (712) displayed in a portion of a request message (710) from the electronic device (10). Based on the received second user input, the control unit (130) may perform a process of evaluating the exercise performance rate for the next exercise movement without re-evaluating the exercise performance rate for the user's exercise movement corresponding to a specific exercise movement, based on the occurrence of an activation event for the progress icon (712). Here, the second user input may be performed in at least one of the following methods: tap, double tap, long press, click, swipe, drag, pinch in, pinch out, and rotate on the icon.
[0191] At this time, the present invention describes a process of requesting the re-execution of a specific exercise action corresponding to the user's exercise action based on user input regarding an icon included in a request message, but it is not limited thereto. The present invention may also perform a process of requesting the re-execution of a specific exercise action corresponding to the user's exercise action through various methods such as user voice commands, gesture recognition, detection of the status of external sensors, and automatic repetition based on user settings.
[0192] Meanwhile, in the present invention, a process of updating the exercise program set in the user's user account can be performed based on the exercise performance rate calculated for the user's exercise movements.
[0193] The control unit (130) can generate feedback information related to the user's exercise movements in order to update the exercise program set in the user account. Here, the “feedback information” may include at least one of the user’s voice data received from a microphone provided in the electronic device (10) and response data to at least one survey provided to the electronic device (10) in relation to the exercise movements performed by the user.
[0194] As illustrated in (b) of FIG. 7a, the control unit (130) may provide a feedback message related to the evaluated user’s exercise performance rate to an electronic device. Here, the feedback message may include a feedback icon for receiving the user’s voice data and survey response data in relation to the user’s exercise performance rate. As an example, the control unit (130) may output (or provide) a feedback message related to the user’s exercise performance rate to a service page (700) based on the fact that the evaluated user’s exercise performance rate did not satisfy a preset standard condition.
[0195] Specifically, the control unit (130) may receive a third user input for a voice feedback icon (721) output in a portion of a feedback message (720) from the electronic device (10). This third user input may be performed in at least one of the following ways: tap, double tap, long press, click, swipe, drag, pinch in, pinch out, and rotate on the icon.
[0196] Furthermore, the control unit (130) can receive voice data corresponding to voice received through a microphone equipped in an electronic device based on the occurrence of a voice feedback activation event according to a third user input.
[0197] As illustrated in FIG. 7b, the control unit (130) can receive voice data (730) corresponding to voice received through a microphone (13) provided in the electronic device (10). Furthermore, the control unit (130) can analyze the voice data (730) to generate voice analysis data. More specifically, the control unit (130) according to the present invention may include a voice recognition model capable of analyzing voice data (710) corresponding to voice received through a microphone (13) provided in the electronic device. There may be a wide variety of methods for analyzing voice data (730) corresponding to voice received through a microphone provided in the electronic device according to the present invention, and the present specification is not limited to the type and method of any module or model capable of analyzing voice data (730) corresponding to voice received through a microphone provided in the electronic device. The control unit (130) according to the present invention may further include at least one of a module and an algorithm that perform the same function as the voice recognition model. For example, a speech recognition model can generate speech analysis data by analyzing speech data (730) based on at least one of a STT (Speech-to-Text) algorithm, HMM (Hidden Markov Model), HMM-GMM (Hidden Markov Model-Gaussian Mixture Model), CTC (Connectionist Temporal Classification) based model, Beam Search based model, DNN (Deep Neural Network) based model, RNN (Recurrent Neural Network) based model, Seq2Seq (Sequence-to-Sequence) model, and Transformer based model.
[0198] Furthermore, the control unit (130) may receive a fourth user input for a survey feedback icon (722) displayed in a portion of the feedback message (720) from the electronic device (10). This fourth user input may be performed in at least one of the following ways: tap, double tap, long press, click, swipe, drag, pinch in, pinch out, and rotate on the icon.
[0199] At this time, the present invention describes a process of updating an exercise program set in a user account based on user input regarding icons (721, 722) included in a feedback message, but is not limited thereto. The present invention may also perform a process of updating an exercise program set in a user account through various methods such as user voice commands, gesture recognition, detection of the status of external sensors, and automatic repetition based on user settings.
[0200] Furthermore, the control unit (130) may provide at least one user survey related to the user's exercise movements on a service page output to the electronic device (10) based on the occurrence of a survey feedback activation event according to the fourth user input. Here, the at least one user survey (740) may include a multiple-choice survey composed of at least one multiple-choice question item related to the user's exercise movements and a plurality of selection items corresponding to the multiple-choice question item.
[0201] The control unit (130) may provide a survey page (or service page) containing a plurality of surveys on the electronic device (10). In this case, the page containing the plurality of surveys may be output through the touch screen (or display) of the electronic device (10). Here, the survey (or question, or problem, or item, or test) provided on the page may include a survey related to the user's exercise performance rate. In the present invention, the survey related to the exercise performance rate may be diverse. For example, the survey related to the exercise performance rate may include various factors related to the exercise performance rate, such as the presence or absence of pain, the duration of pain, mental health, and physical health.
[0202] At this time, the survey related to the exercise performance rate may include a credible survey actually used in the Department of Psychiatry to diagnose the user's condition. In addition, the control unit (130) may periodically update the at least one survey by additionally collecting the survey related to the exercise performance rate through a central server, an external server, a website, etc.
[0203] Furthermore, the control unit (130) may receive user survey response data (740) for at least one user survey provided to the electronic device. For example, the control unit (130) may receive objective survey response data (741) corresponding to an objective survey included in at least one user survey provided to the electronic device (10). Specifically, the control unit (130) may receive response data for an objective survey (or objective survey response data) from the electronic device (10) that includes a response matched to an item selected by user input among a plurality of selection items. At this time, the objective survey response data may include natural language response data corresponding to a specific selection item selected by user input among a plurality of selection items.
[0204] Meanwhile, the control unit (130) may provide at least one open-ended survey (742) regarding the user's exercise performance rate on a service page output to the electronic device (10). Specifically, the at least one user survey provided on the service page may further include an open-ended survey capable of receiving natural language input for at least one open-ended question item from the user terminal regarding the exercise performance rate.
[0205] Furthermore, the control unit (130) can receive natural language input for a subjective question item as subjective survey response data (742). At this time, the control unit (130) can receive natural language input entered into the electronic device (10) as subjective survey response data (742) for the user's subjective survey. At this time, if the user's voice is received by the electronic device (10) as subjective survey response data (742), the control unit (130) can receive voice data corresponding to the voice received from the electronic device (10) as subjective survey response data (742) for the subjective survey.
[0206] The control unit (130) can generate feedback information (750) using at least one of voice data and survey response data. Furthermore, the control unit (130) can generate a feedback prompt (760) for updating an exercise program using the feedback information (750). The control unit (130) can update the exercise program from the pre-trained vision language model by processing the feedback prompt as input to the pre-trained vision language model. Here, “exercise program update” may mean changing at least one of the exercise movements, the number of repetitions of the movements, and the total exercise time that constitute the exercise program. Below, an exercise program update that changes the exercise movements is described as an example, but is not limited thereto, and at least one of the number of repetitions and the total exercise time may also be changed.
[0207] As illustrated in FIG. 7c, the control unit (130) can extract text (751, 752) of a pre-set topic related to an exercise program update from feedback information (750) and include the text of the pre-set topic in a feedback prompt (760). Here, “pre-set topic” may mean at least one of adjusting the difficulty of the exercise, selecting the type of exercise, and changing the type of exercise.
[0208] For example, the control unit (130) can generate a feedback prompt (760) to change the “squat movement” to another exercise item based on receiving voice data such as “I felt a lot of pain in my left knee during the squat movement!” through a microphone equipped in the electronic device. Furthermore, the control unit (130) can update the exercise program (770) set in the user account by inputting the feedback prompt (760) into the vision language model (132). At this time, the exercise program (770) set in the user account may be stored in the database (200), and the exercise program (770) may include multiple different exercise movements (771 to 773).
[0209] Furthermore, the control unit (130) can change a specific exercise action (771) among the plurality of exercise actions to another action (781) according to the feedback prompt (760) using the vision language model (132). Specifically, the vision language model (132) can change the exercise actions constituting the exercise program set in the user account according to the feedback prompt (760). Similar exercise actions related to the specific exercise action (771) may be matched and stored in the storage unit (120). The vision language model (132) can change the exercise items constituting the exercise program (770) set in the user account by using the similar exercise actions matched to the specific action. For example, the vision language model (132) can receive the feedback prompt and update the exercise program to provide a specific exercise action (e.g., squat) by changing it to a similar exercise action related to the specific exercise action (e.g., “hip extension exercise”). The control unit (130) can obtain an updated exercise program (780) from a vision language model (132) and provide the updated exercise program (780) to an electronic device (10).
[0210] Furthermore, the vision language model (132) can update the exercise program (770) by changing the exercise movements assigned after the next day of the specific exercise period (e.g., “Day 1 of Week 2”) based on feedback information (750) regarding the exercise movements assigned during the specific exercise period (e.g., “Day 6 of Week 1”). The control unit (130) can provide the updated exercise program (780) on the electronic device (10) from the day after the specific day.
[0211] As described above, the method and system for evaluating exercise performance rate using a vision language model according to the present invention can determine whether the user actually performs the exercise and prevent perfunctory exercise performance by performing an evaluation of the user's exercise movements included in an exercise video and calculating the exercise performance rate.
[0212] Furthermore, the method and system for evaluating exercise performance using a vision language model according to the present invention can contribute to determining whether a key movement is accurately performed and maximizing exercise effects by identifying a target body part corresponding to a key movement of a specific exercise motion from an exercise video and quantitatively evaluating whether the performance conditions set for the target body part are satisfied.
[0213] Furthermore, the method and system for evaluating exercise performance rate using a vision language model according to the present invention can contribute to preventing the user from abandoning exercise performance and inducing continuous exercise performance by inducing the user to re-perform the same exercise according to the user's exercise performance rate.
[0214] Furthermore, the method and system for evaluating exercise performance using a vision language model according to the present invention can generate a customized exercise program with adjusted difficulty, intensity, and number of repetitions based on the exercise performance rate. Through this, the exercise effect can be maximized by providing an individualized exercise program suitable for the user's physical condition, exercise performance ability, and recovery progress.
[0215] Furthermore, the exercise performance rate evaluation system (100) using a vision language model according to the present invention can be implemented through a computing device described below and can perform data processing related to the exercise performance rate evaluation method using the vision language model described above.
[0216] Meanwhile, FIG. 8 illustrates an example of a block diagram of a computing system in which the present invention can be implemented.
[0217] Referring to FIG. 8, a computing system (10000) that performs a method for evaluating exercise performance using a vision language model according to one embodiment of the present invention may include at least one computing device. At this time, the at least one computing device may be a single processor or a multiprocessor computing device.
[0218] The components of at least one computing device of the present invention may include various hardware components such as one or more processors, memory, other hardware, and a system bus (not shown) that connects various system components so that they can transmit and receive data to and from each other (e.g., telecommutatively connected, physically connected, electrically connected), and the components of at least one computing device are not limited thereto and may be very diverse.
[0219] Meanwhile, at least one computing device included in a computing system (10000) that performs a method for evaluating exercise performance using a vision language model may be connected to communicate via a network (1070). For example, at least one computing device included in the computing system (10000) may be clustered or part of a local area network (LAN). Additionally, at least one computing device may be part of a wide area network (WAN) or connected to at least one of a client-server network and a peer-to-peer network within the cloud.
[0220] Meanwhile, when at least one computing device is used in at least one of a network environment and a cloud computing environment, the at least one computing device may be connected to at least one of a public and private network through a network interface or adapter. In one embodiment, other communication connection devices, such as a modem, may be used to establish communication through the network. The modem may be at least one of an internal modem and an external modem, and may be connected to a system bus through a network interface or a specific mechanism, etc. A wireless network component consisting of an interface and an antenna may be coupled to the network through a device such as an access point, a peer computer, etc. In the present invention, the method of connecting at least one computing device to communicate through the network (1070) is not limited, and it may be connected to communicate in a manner different from the described example.
[0221] Furthermore, other computer-type devices and / or systems not shown in FIG. 8 may also interact technically with at least one computing device or other system through one or more connections to the network (1070) via a network interface. Here, the network interface may include network interface equipment such as a physical network interface controller (NIC) or a virtual network interface (VIF).
[0222] The network (1070) of the present invention may include various forms such as the Internet, WLAN (Wireless LAN), Wi-Fi (Wireless-Fidelity), Wi-Fi (Wireless Fidelity) Direct, DLNA (Digital Living Network Alliance), WiBro (Wireless Broadband), WiMAX (World Interoperability for Microwave Access), HSDPA (High Speed Downlink Packet Access), HSUPA (High Speed Uplink Packet Access), LTE (Long Term Evolution), LTE-A (Long Term Evolution-Advanced), 5G (5th Generation Mobile Telecommunication), Bluetooth (Bluetooth™ Frequency Identification), Infrared Data Association (IrDA), UWB (Ultra-Wideband), ZigBee, NFC (Near Field Communication), Wi-Fi Direct, Wireless USB (Wireless Universal Serial Bus), etc., and in the present invention, data transmission may be performed based on standard communication protocols such as TCP / IP, HTTP, SSL, etc.
[0223] A computing system (10000) that performs a method for evaluating exercise performance using a vision language model according to the present invention may include at least one of a user computing device (1010, or system), a training computing system (1050, or device), and a server computing system (1030, or device).
[0224] A user computing device (1010) according to the present invention may be understood as a computing device comprising at least one processor (1011) and a memory (1012) for performing a method of evaluating exercise performance using a vision language model. For example, the user computing device (1010) may include at least one computing device among a smartphone, a smart TV, a laptop computer, a desktop computer, a digital broadcasting terminal, a PDA (personal digital assistants), a PMP (portable multimedia player), a navigation device, a slate PC, a tablet PC, an ultrabook, a wearable device (e.g., a smartwatch, a smart glass, and a head-mounted display).
[0225] At least one processor (1011) constituting the user computing device (1010) may include one or more general-purpose processors and / or one or more special-purpose processors. For example, at least one processor (1011) constituting the user computing device (1010) may be composed of at least one or a plurality of electrically connected processors among a central processing unit (CPU), a graphics processing unit (GPU), a tensor processing unit (TPU), a neural network processing unit (NPU), an arithmetic logic unit (ALU), a floating-point arithmetic unit (FPU), an application integrated circuit, an application semiconductor (ASIC), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, and / or other electrical units for performing functions.
[0226] Furthermore, at least one processor (1011) may be configured to execute computer-readable instructions contained in memory (1012) and / or other instructions described herein. Memory (1012) constituting a user computing system (1010) according to the present invention may include volatile memory, non-volatile memory, fixed media, removable media, magnetic media, optical media, semiconductor media and / or other types of physically durable storage media. For example, memory (1012) may include one or more non-transient / transient computer-readable storage media such as RAM, ROM, HDD (Hard Disk Drive), SSD (Solid State Disk), SSD (Silicon Disk Drive), EEPROM, EPROM, flash memory device, magnetic disk, and combinations thereof, and may include web storage of a server performing memory storage functions over the internet. This memory (1012) can store data and instructions necessary for the at least one processor (1011) to perform the operation of an application for evaluating exercise performance using a vision language model.
[0227] A user computing device (1010) may include one or more user input components (1021) that detect user input. For example, the user input component (1021) may also be referred to as a user interface module. The user input component (1021) may include a touch screen, a computer mouse, a keyboard, a keypad, a touchpad, a trackball, a joystick, a voice recognition module, or other similar devices. However, the present invention does not limit the type of user input component (1021). In this case, the user input component (1021) in the present invention does not necessarily mean a hardware means, but can be understood as a channel for receiving input from a user. Meanwhile, the user of the present invention may refer to an automated agent, script, playback software, etc., that operates on behalf of one or more people.
[0228] A user can interact with a computing system (10000) including at least one computing device through input text, touch, voice, movement, computer vision, gestures and / or other forms of input / output using a user input component (1021). For example, the user input component (1021) may include one or more of a command line interface (CLI), a graphical user interface (GUI), a natural user interface (NUI), a voice command interface and / or other user interface (UI) representations.
[0229] Between the user input component (1021) and the user computing device (1010), one or more application programming interface (API) calls may be made based on user input received from a user interface and / or a network. Here, the expression “based on” may be interpreted to include cases where it is based on the use of a specific configuration, modified from, derived from, influenced by, dependent on, or otherwise derived from a specific configuration. In some embodiments, an API call may be configured for a specific API, which may be interpreted or converted into an API call configured for another API. Here, an API may mean a defined interface or connection between computers or between computer programs.
[0230] In one embodiment, the user computing device (1010) may store at least one machine learning model (1020). For example, the user computing device (1010) may be various machine learning models, such as a plurality of neural networks (e.g., deep neural networks) that analyze the user's exercise movements based on exercise video and exercise program information and evaluate exercise performance using a vision language model, or other types of machine learning models including non-linear models and / or linear models, and may be composed of a combination thereof.
[0231] According to an embodiment of the present invention, a user computing device (1010) may perform an exercise performance evaluation method using a vision language model by using a local or / and external machine learning model (1020). Alternatively, the user computing device (1010) may perform an exercise performance evaluation method using a vision language model by using a machine learning model (1040) provided by a server.
[0232] In addition, according to another embodiment of the present invention, a server computing system (1030) communicating with a user computing device (1010) may provide performance evaluation information regarding a user's exercise movements to the user computing device (1010) via an application or / and the web in accordance with a request from a user received through the user computing device (1010).
[0233] In addition, according to another embodiment of the present invention, at least a part of the user computing device (1010) and the server computing system (1030) are linked together to perform an exercise performance evaluation method using a vision language model, thereby providing the user with performance evaluation information regarding the user's exercise movements.
[0234] Additionally, according to various embodiments of the present invention, a user computing device (1010) and / or a server computing system (1030) can learn machine learning models (1020, 1040) performed in a method for evaluating exercise performance using a vision language model through interaction with a training computing system (1050) that is communicatedly connected via a network (1070). In this case, the training computing system (1050) may be a computing system separate from the server computing system (1030). Alternatively, in some embodiments, the training computing system (1050) may be part of the server computing system (1030) or part of the user computing device (1010).
[0235] Meanwhile, the server computing system (1030) may include at least one processor (1031) and memory (1032). Here, the processor (1031) may be composed of at least one or a plurality of electrically connected processors among a central processing unit (CPU), a graphics processing unit (GPU), a tensor processing unit (TPU), a neural network processing unit (NPU), an application integrated circuit, an application semiconductor (ASIC), an arithmetic logic unit (ALU), a floating-point arithmetic unit (FPU), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, and / or other electrical units for performing functions. For example, at least one processor (1031) may include a circuit and a transistor configured to execute instructions from memory (1032).
[0236] The memory (1032) constituting the server computing system (1030) according to the present invention may include volatile memory, non-volatile memory, fixed media, removable media, magnetic media, optical media, semiconductor media, and / or other types of physically durable storage media. For example, the memory (1032) may include one or more non-transient / transient computer-readable storage media such as RAM, ROM, HDD (Hard Disk Drive), SSD (Solid State Disk), SSD (Silicon Disk Drive), EEPROM, EPROM, flash memory device, magnetic disk, etc., and combinations thereof, and may include web storage of a server that performs the storage function of memory over the internet. Additionally, the server computing system (1030) may further include a data storage (data store). For example, the data storage may be composed of at least one of a relational database, a NoSQL database, a data warehouse, and a local file system.
[0237] In the memory (1032) constituting the server computing system (1030) according to the present invention, data and instructions necessary for the at least one processor (1031) to perform the operation of an application for evaluating exercise performance using a vision language model may be stored.
[0238] In one embodiment, the server computing system (1030) may be composed of a single device or a plurality of computing devices, and these may be configured to operate according to a sequential or parallel computing architecture. Additionally, a distributed processing system may be configured with a plurality of networked devices.
[0239] Meanwhile, the training computing system (1050) may include at least one processor (1051) and memory (1052). The model trainer (1060) is a logical component that executes the training of at least one machine learning model (1020, 1040) and may be implemented in the form of hardware, firmware, or software. For example, the model trainer (1060) may be executed by the processor (1051) after loading training data (1061) stored in a storage device into memory (1052). For example, the model trainer (1060) may be configured to execute one or more operations (e.g., model training, model reconstruction, model validation, model testing) on at least one machine learning model.
[0240] The machine learning model of the present invention may include at least one of a statistical model, an algorithm, a neural network (NN), a convolutional neural network (CNN), a generative neural network (GNN), a Word2Vec model, a Bag of Words model, a TF-IDF (document frequency-inverse document frequency) model, a GPT (Generative Pre-trained Transformer) model (or other autoregressive models), a PPO (Proximal Policy Optimization) model, a nearest neighbor model (e.g., a k-nearest neighbor model), a linear regression model, a K-means clustering model, a Q-learning model, a TD (Temporal Difference) model, a Deep Adversarial Network model, and all other types of models further described herein.
[0241] Specifically, the model trainer (1060) may execute operations to train a machine learning model, and said operations may include at least one of adding, removing, and modifying model parameters. At this time, the training of the machine learning model may be at least one of supervised learning, semi-supervised learning, and unsupervised learning. In one embodiment, the training of the machine learning model may include the step of repeatedly inputting training data (1061) based on epochs and repeatedly performing the machine learning model training process configured in this way. Here, an epoch may refer to a unit in which the entire set of training data (1061) undergoes forward and backpropagation processing once. In some implementations, different levels of training methods (e.g., supervised learning, semi-supervised learning, unsupervised learning) may be used for different epochs.
[0242] The training data (1061) of the present invention may include input data and / or data previously output from at least one machine learning model (e.g., recursive learning feedback). The parameters of at least one machine learning model may include at least one of a seed value, a model node, a model layer, an algorithm, a function, connections between different machine learning models, connections between parameters, machine learning model constraints, and other digital components that influence the output of the machine learning model. In this case, the model connections between different machine learning models may include or represent relationships between model parameters and / or models, which may be dependent or interdependent, hierarchical, and / or static or dynamic. The combinations and configurations of model parameters described herein may be too complex to be maintained or used by human cognitive abilities.
[0243] In the present invention, the machine learning parameters described according to the embodiments are not limited, and a single machine learning model may further include a plurality of model parameters.
[0244] Meanwhile, FIG. 9 illustrates an example of a block diagram of a computing device (1100) that may be included in a user computing device (1010), a server computing system (1030), and a training computing system (1050), as an embodiment of a computing system (10000) in which the present invention can be implemented.
[0245] As illustrated in FIG. 9, the computing device (1100) may include at least one application (e.g., Application 1 to Application N), and each of the at least one application may include a machine learning library and a model execution environment for performing a method of evaluating exercise performance using a machine learning-based vision language model. The at least one application included in the computing device (1100) may communicate with the sensor, context manager, device state manager, or additional component(s) within the computing device (1100) via an Application Programming Interface (API). In one embodiment, the at least one application may interface with device components, such as receiving sensor data or state data or transmitting prediction results to an output device via a public or private API.
[0246] Meanwhile, FIG. 10 illustrates an example of a block diagram in another aspect of a computing device (1200), which is one of the components of a computing system (10000) that performs a method for evaluating exercise performance using a vision language model according to an embodiment of the present invention.
[0247] A computing device (1200) according to the present invention may include at least one application (e.g., Application 1 to Application N), and at least one application may communicate with a central intelligence layer (1210). Each application may interact with a shared model within the central intelligence layer (1210) through an API (e.g., a common API).
[0248] The central intelligence layer (1210) includes one or more machine learning models and may share them among multiple applications or provide them independently to each. In one embodiment, the central intelligence layer (1210) may be integrated as part of an operating system or implemented as a separate logical layer.
[0249] Additionally, the central intelligence layer (1210) can communicate with the central device data layer (1220). The central device data layer (1220) can integrate and store exercise videos, user information, and performance evaluation information stored within the computing device (1200), and provide this as input data required for evaluating exercise performance rates using a vision language model. Each device component (e.g., sensor, state manager, etc.) can communicate with the central device data layer (1220) through a private API, etc.
[0250] The technology described in this specification may be composed of a single or multiple computing devices, and a machine learning model that performs a method for evaluating exercise performance using a vision language model may be executed sequentially or in parallel on one component or multiple distributed components. Data storage, machine learning models, and applications may be distributed and operated locally or over a network, and these configurations can be flexibly applied to various system architectures.
[0251] Meanwhile, computer-readable media include all types of recording devices in which data that can be read by a computer system is stored. Examples of computer-readable media include HDD (Hard Disk Drive), SSD (Solid State Disk), SSD (Silicon Disk Drive), ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, etc.
[0252] Furthermore, the computer-readable medium may be a server or cloud storage that includes a storage and is accessible to an electronic device via communication. In this case, the computer may download the program according to the present invention from the server or cloud storage via wired or wireless communication.
[0253] Furthermore, in the present invention, the computer described above is an electronic device equipped with a processor, namely a CPU (Central Processing Unit), and no special limitations are placed on its type.
[0254] Meanwhile, the above detailed description should not be interpreted restrictively in all respects but should be considered exemplary. The scope of the invention shall be determined by a reasonable interpretation of the appended claims, and all modifications within the equivalent scope of the invention are included within the scope of the invention.
Claims
1. A step of requesting an electronic device to perform an exercise for a specific movement; A step of receiving a user’s exercise video related to the specific exercise movement from the electronic device; A step of extracting user motion information regarding the user’s exercise motion corresponding to the specific exercise motion from the exercise video; A step of generating a prompt to perform an analysis of whether the user's exercise motion satisfies a preset execution condition using the user action information and reference action information corresponding to the specific exercise motion; A step of inputting the generated prompt and the exercise video into a pre-trained vision language model to obtain performance evaluation information regarding the user's exercise movements from the pre-trained vision language model; and A method for evaluating exercise performance using a vision language model, characterized by including a step of evaluating the user's exercise performance rate based on the above-mentioned performance evaluation information.
2. In Paragraph 1, The specific exercise movements mentioned above include, According to the main movements of the specific exercise motion mentioned above, a target body part is defined and exists, and The above target body part is, A method for evaluating exercise performance using a vision language model, characterized by including at least one of a main target body part and an auxiliary target body part, depending on the degree of correlation with the above-mentioned main movement.
3. In Paragraph 2, The step of extracting the above user action information is, A step of identifying the target body part, which is predefined for the specific exercise movement, from the exercise video; and A method for evaluating exercise performance using a vision language model, characterized by including the step of analyzing the movement of the target body part in the exercise video and extracting user motion information for the specific exercise movement.
4. In Paragraph 3, The step of extracting the above user action information is, A step of determining whether the main target body part is identifiable in the above exercise video; and A method for evaluating exercise performance using a vision language model, characterized by including a step of determining whether to extract user motion information for at least one of the main target body part and the auxiliary target body part according to the above judgment result.
5. In Paragraph 4, The step of determining whether the above main target body part is identified is: If the main target body part is not identified in the exercise video, the method further includes the step of requesting the electronic device to include the main target body part in the exercise video. In the step of requesting that the above main target body part be included in the exercise video, A method for evaluating exercise performance using a vision language model, characterized by outputting a guidance message to an electronic device that guides the adjustment of at least one of the video recording position and posture so that the main target body part is included in the exercise video.
6. In Paragraph 2, The step of obtaining the performance evaluation information regarding the exercise movements of the user is, A step of analyzing whether the performance condition set for the target body part is satisfied for a predetermined number of times for the specific exercise movement using the above-mentioned previously trained vision language model; and A method for evaluating exercise performance using a vision language model, characterized by including the step of obtaining performance evaluation information regarding the user's exercise movements from the previously learned vision language model based on the analysis result of the above satisfaction.
7. In Paragraph 6, The step of evaluating the above exercise performance rate is, A step of calculating the ratio of the number of executions satisfying the execution conditions set for the target body part among the aforementioned preset execution counts; and A method for evaluating exercise performance using a vision language model, characterized by including the step of evaluating the user’s exercise performance rate as at least one of a performance score and a percentile according to the ratio of the calculated number of times the exercise is performed.
8. In Paragraph 1, The method further includes the step of requesting the user to perform a specific exercise movement corresponding to the user's exercise movement again, based on the exercise performance rate calculated for the user's exercise movement. The step of requesting the re-performance of the specific exercise movement described above is, A step of providing a request message to the electronic device requesting the specific exercise movement to be performed again, based on the fact that the exercise performance rate does not satisfy a preset standard condition; A step of receiving user input for a request icon displayed in a portion of the request message from the electronic device; and A method for evaluating exercise performance using a vision language model, characterized by including the step of re-evaluating the exercise performance rate for the user’s exercise movement corresponding to the specific exercise movement based on the occurrence of an activation event for the request icon according to the user input.
9. In Paragraph 1, The method further includes the step of updating the exercise program set in the user account of the user based on the exercise performance rate calculated for the exercise movements of the user. The step of updating the above exercise program is, A step of generating feedback information related to the exercise movements of the above user; A step of generating a feedback prompt for updating the exercise program using the above feedback information; and A method for evaluating exercise performance using a vision language model, characterized by including the step of processing the above feedback prompt as input to the above-mentioned vision language model and updating the exercise program from the above-mentioned vision language model.
10. In Paragraph 9, The step of generating the above feedback information is, A step of receiving voice data corresponding to voice received through a microphone provided in the electronic device; A step of receiving user survey response data for at least one user survey provided to the electronic device; and A method for evaluating exercise performance using a vision language model, characterized by including the step of generating feedback information using at least one of the voice data and the survey response data.
11. A communication unit that receives a user’s exercise video related to a specific exercise movement from an electronic device, and a control unit that requests the electronic device to perform the exercise for the specific exercise movement, The above control unit is, Extract user motion information regarding the user's exercise motion corresponding to the specific exercise motion from the exercise video, and Using the above user motion information and the reference motion information corresponding to the above specific motion, a prompt is generated to perform an analysis of whether the user's motion satisfies a preset execution condition, and The generated prompt and the exercise video are input into a pre-trained vision language model to obtain performance evaluation information regarding the user's exercise movements from the pre-trained vision language model, and An exercise performance rate evaluation system using a vision language model characterized by evaluating the user's exercise performance rate based on the above-mentioned performance evaluation information.
12. A program that is executed by one or more processes in an electronic device and stored on a computer-readable recording medium, The above program is, A step of requesting an electronic device to perform an exercise for a specific movement; A step of receiving a user’s exercise video related to the specific exercise movement from the electronic device; A step of extracting user motion information regarding the user’s exercise motion corresponding to the specific exercise motion from the exercise video; A step of generating a prompt to perform an analysis of whether the user's exercise motion satisfies a preset execution condition using the user action information and reference action information corresponding to the specific exercise motion; A step of inputting the generated prompt and the exercise video into a pre-trained vision language model to obtain performance evaluation information regarding the user's exercise movements from the pre-trained vision language model; and A program stored on a computer-readable recording medium characterized by including instructions for performing a step of evaluating the exercise performance rate of the user based on the above performance evaluation information.
Citation Information
Patent Citations
Refill type lipstick container
KR1020250037876A
Primer Pairs for Identification of Bacillus velezensis K10 and Usage Thereof
KR1020250041682A
Bio-packaging box m of distribution and circulation for managing freshness easily
KR102641220B1
Instance level scene recognition with a vision language model
US11978271B1
KR20240093317A