Model training method and virtual coach guidance strategy generation method
By acquiring multi-dimensional user data and using the DQN model to generate personalized training guidance strategies, the problem of the lack of adaptability in traditional virtual coaching systems is solved, and higher quality training guidance is achieved.
Patent Information
- Application Number
- CN202510936014.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-07-07
AI Technical Summary
Traditional virtual coaching systems lack adaptability to individual states and dynamic changes, resulting in a lack of targeted training guidance and affecting the quality of user training.
By acquiring multi-dimensional data such as user action quality, fatigue level, self-satisfaction, and comprehension level, the DQN model is used to generate personalized training guidance strategies, which are then adjusted in real time to adapt to changes in user status.
It improves the personalization and accuracy of training guidance, prevents ineffective training or injuries, and enhances the guidance capabilities of virtual coaches.
Smart Images

Figure CN120873589A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a model training method and a method for generating virtual coach guidance strategies. Background Technology
[0002] In sports training and fitness, scientific and personalized training guidance is crucial for improving results, preventing injuries, and maintaining motivation. Traditional face-to-face coaching models are limited by time, location, cost, and coaching resources, making it difficult to meet the growing demand for personalized and convenient training. In recent years, virtual coaching systems based on artificial intelligence, sensor technology, and human-computer interaction have developed rapidly, providing users with an accessible, scalable, and cost-effective alternative or supplementary solution.
[0003] Existing methods generally rely on pre-defined, universal training rule bases to provide training guidance to users. Users typically select a fixed set of rules based on their goals (such as fat loss, muscle gain, or improved endurance) or experience level. These rules lack adaptability to individual conditions and dynamic changes, thus affecting the quality of training. Summary of the Invention
[0004] In view of this, embodiments of this application provide a model training method and a method for generating virtual coaching guidance strategies, which can provide users with more targeted training guidance based on their specific training situation and improve the quality of their training.
[0005] In a first aspect, embodiments of this application provide a model training method, including:
[0006] Obtain the current state of the user after performing the current training action; wherein, the current state includes: action quality score, fatigue level score, self-satisfaction score, comprehension level score, and any multiple of the following: action quality score, fatigue level score, self-satisfaction score, comprehension level score, and sensor data from the exercise device. The action quality score is used to characterize the execution quality of the current training action, the fatigue level score is used to characterize the user's fatigue level, the self-satisfaction score is used to characterize the user's satisfaction with performing the current training action, and the comprehension level score is used to characterize the user's comprehension of the current training action.
[0007] The current state is input into the DQN (DeepQ-Network) model to obtain the Q-value of each strategy in the preset action space; wherein, the action space includes: any number of strategies used by the virtual coach to guide the user's training, such as outputting preset encouraging feedback, outputting corrective instructions, outputting rest suggestions, outputting questions, and no output.
[0008] Based on the Q value of each of the aforementioned strategies, the current strategy is selected from the action space;
[0009] Execute the current policy;
[0010] After the user responds to the current strategy and performs the next training action, the user's next state is obtained;
[0011] The parameters of the DQN model are adjusted based on the current state, the next state, and the current strategy.
[0012] Secondly, embodiments of this application provide a method for generating virtual coaching guidance strategies, including:
[0013] Based on the DQN model trained using the method described in the above embodiments, a strategy is generated for virtual coaches to guide the training of target users.
[0014] Thirdly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the method described in any of the above embodiments.
[0015] One embodiment of the above invention has the following advantages or beneficial effects: The DQN model is trained based on data from multiple dimensions such as movement quality and fatigue level, enabling it to learn optimal response strategies for multiple user dimensions. It can adjust guidance strategies in a timely manner according to changes in the user's state, providing more personalized training guidance. For example, if the user easily completes the set task, the DQN model can automatically increase the difficulty by adjusting the strategy to maintain effective stimulation; if the user is clearly struggling or their movements are distorted, it can promptly reduce the difficulty or suggest rest to prevent ineffective training or injury. Furthermore, the training process also considers the user's internal state, such as fatigue level, self-satisfaction level, and comprehension level, further improving the guidance quality of the virtual coach. The data and experience accumulated by the DQN model will continuously optimize its decision-making strategy, making the virtual coach's guidance capabilities increasingly stronger and more accurate.
[0016] The further effects of the aforementioned unconventional alternative methods will be explained below in conjunction with specific implementation methods. Attached Figure Description
[0017] The accompanying drawings are provided to better understand this application and do not constitute an undue limitation thereof. Wherein:
[0018] Figure 1 This is a flowchart of a model training method provided in one embodiment of this application;
[0019] Figure 2 This is a schematic diagram of a model training device provided in one embodiment of this application. Detailed Implementation
[0020] The following description, in conjunction with the accompanying drawings, illustrates exemplary embodiments of this application, including various details to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this application. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0021] like Figure 1 As shown in the embodiments of this application, a model training method is provided, including:
[0022] Step 101: Obtain the current state of the user after performing the current training action.
[0023] This method is applied to a virtual coaching system. The current state includes any combination of the following: movement quality score, fatigue level score, self-satisfaction score, comprehension score, and sensor data from the exercise equipment. The movement quality score characterizes the execution quality of the current training movement, the fatigue level score characterizes the user's fatigue level, the self-satisfaction score characterizes the user's satisfaction with performing the current training movement, and the comprehension score characterizes the user's understanding of the current training movement.
[0024] The user's state, including the current state and the next state, can be obtained through visual data, voice data, sensor data from motion devices, and interface interaction data. Visual data is obtained by capturing the user's motion video stream through a camera; voice data refers to the user's voice interaction with the virtual coach, including answers to specific questions and spontaneous speech; sensor data refers to kinematic or dynamic data collected by wearable devices or training equipment; and interface interaction data refers to data generated by the user operating the software interface provided by the virtual coaching system.
[0025] The exercise quality score can include scores for one or more exercise dimensions. For example, for squats, there are five working dimensions: squat depth, knee trajectory, back posture, hip angle, and movement speed, which evaluate the squat exercise from different dimensions.
[0026] For example, there are three scores for the action dimension: 0 for poor quality or needing correction, 1 for medium or acceptable quality, and 2 for excellent or standard quality. The quality or score of the action dimension can be determined by comparing it with a preset standard action; the calculation method for the action quality score is not limited here. In practical applications, this method can directly obtain the scores for each item in the current state, or it can calculate the scores using acquired video data, etc.
[0027] Fatigue level scores can be obtained through voice data and interface interaction data between the user and the virtual coach. For example, the virtual coach can ask the user whether their current fatigue level is high, medium, or low. If it is low, the fatigue level score is 0; if it is medium, the fatigue level score is 1; and if it is high, the fatigue level score is 2.
[0028] Self-satisfaction scores can be obtained through voice data and interface interaction data between the user and the virtual coach. For example, the user actively expresses their satisfaction with performing the current training action. If they are not satisfied, the self-satisfaction score is 0; if they are neutral, the self-satisfaction score is 1; and if they are satisfied, the self-satisfaction score is 2.
[0029] Comprehension score can be obtained through voice data and action quality scores between the user and the virtual coach. For example, if the user expresses misunderstanding or repeatedly makes mistakes in the action points, the comprehension level is low and the comprehension score is 0. If the user has some uncertainty or occasionally makes mistakes in the action points, the comprehension level is medium and the comprehension score is 1. If there are no mistakes in the action points or the user does not express any doubts, the comprehension level is high and the comprehension score is 2.
[0030] Among them, sports equipment can be wearable devices or sports equipment, and the corresponding sensor data includes multiple continuous values, such as peak force, which is the peak value of force, as well as average angular velocity, maximum range of motion, etc.
[0031] The current training exercise can be a single movement or a movement unit that repeats the same movement, such as a movement unit consisting of 3 squats.
[0032] Step 102: Input the current state into the DQN model to obtain the Q-values of each policy in the preset action space.
[0033] The action space includes: outputting preset encouraging feedback, outputting corrective instructions, outputting rest suggestions, outputting questions, and any number of strategies used by the virtual coach to guide user training.
[0034] The current state is input into the DQN model in vector form. If the current state is not in vector form, it needs to be preprocessed. The specific process will be further explained in subsequent embodiments.
[0035] Encouraging feedback is used to encourage users to maintain the quality of their actions, such as outputting encouraging words like "Well done" or "Keep it up."
[0036] Corrective instructions are used to correct training movements, such as outputting voice correction instructions like "Be careful not to let your knees buckle inward" and "Keep your core engaged and your back straight."
[0037] Rest suggestions are used to advise users to take a short break, such as "Please take a short break".
[0038] Asking questions is used to inquire about the user's feelings or level of understanding, such as "How are you feeling now?" or "Did you understand this action clearly?"
[0039] Encouraging feedback, corrective instructions, rest suggestions, and questions can be output in the form of text or voice.
[0040] "No output" means that the virtual coach does not need to intervene at present and can remain silent to continue observing the user's actions.
[0041] Step 103: Select the current policy from the action space based on the Q value of each policy.
[0042] The embodiments of this application adopt the ε-greedy strategy, which is a strategy for selecting the Q value based on the current network output in the DQN model.
[0043] Step 104: Execute the current policy.
[0044] The virtual coach, acting as an intelligent agent, can execute the current strategy, and the user can then execute the next training action based on that strategy.
[0045] For example, a user performs 5 consecutive squats, and the virtual coach gives the corrective instruction "straighten your back". The user adjusts their posture and then performs 5 more squats. The 5 squats constitute one action unit.
[0046] Step 105: After the user responds to the current strategy and executes the next training action, obtain the user's next state.
[0047] Step 106: Adjust the parameters of the DQN model based on the current state, the next state, and the current policy.
[0048] This application's embodiments train a DQN model based on data from multiple dimensions, including movement quality and fatigue level. This enables the model to learn optimal response strategies for various user dimensions and adjust guidance strategies in a timely manner according to changes in the user's state, providing more personalized training guidance. For example, if the user easily completes the set task, the DQN model can automatically increase the difficulty by adjusting the strategy to maintain effective stimulation; if the user is clearly struggling or their movements are distorted, it can promptly reduce the difficulty or suggest rest to prevent ineffective training or injury. Furthermore, the training process also considers the user's internal state, such as fatigue level, self-satisfaction level, and comprehension level, further improving the quality of virtual coaching. The data and experience accumulated by the DQN model continuously optimize its decision-making strategies, making the virtual coach's guidance capabilities increasingly stronger and more precise.
[0049] In one embodiment of this application, the method further includes:
[0050] Since the current item is a discrete value, the discrete value is converted into a numerical vector through one-hot encoding or an embedding layer; among them, the action quality score, fatigue score, self-satisfaction score, and comprehension score are discrete values, while the sensor data of the motion device are continuous values.
[0051] In response to the current item being a continuous value, normalize the current item.
[0052] Input the current state into the DQN model, including:
[0053] The vectors corresponding to each item in the current state are concatenated and then input into the DQN model.
[0054] For example, the current state includes motion quality score, fatigue score, and sensor data from the motion equipment. The motion quality score and fatigue score are converted into numerical vectors, the sensor data from the motion equipment is normalized, and the processed items are concatenated and input into the DQN model.
[0055] The current state s obtained after preprocessing can be a mixed-type vector s = [q1, ... q2] . M [f,sat,u], where q1...q M There are M continuous values, representing the scores for the same training action on M different action dimensions. For example, after normalization to the [0, 1] interval, higher values indicate better quality. F is used to characterize the fatigue level score and is a discrete value, sat is used to characterize the self-satisfaction score and is a discrete value, and u is used to characterize the comprehension level score and is a discrete value.
[0056] In one embodiment of this application, adjusting the parameters of the DQN model based on the current state, the next state, and the current policy includes:
[0057] Calculate the immediate reward based on the current state, the next state, and the current policy;
[0058] Store the experience consisting of the current state, current strategy, immediate reward, and next state into the experience replay buffer.
[0059] A preset number of target experiences are sampled from the experience replay buffer;
[0060] The objective Q-values of each objective are calculated based on the objective network in the DQN model;
[0061] The predicted Q-values for each objective are calculated based on the current network in the DQN model.
[0062] Calculate the value of the loss function based on the target Q-value and predicted Q-value of each objective based on the experience of each objective.
[0063] The parameters of the DQN model are adjusted based on the value of the loss function.
[0064] The calculation process for the target Q-value and the predicted Q-value can refer to the training process of existing DQN models, and will not be repeated here. In this embodiment, mean squared error can be used for calculation.
[0065] The embodiments of this application construct an action space based on the user's training state and the guidance strategy of a virtual coach, and train a DQN model, which enables the model to learn personalized strategies, thereby providing users with high-quality training guidance.
[0066] In one embodiment of this application, calculating the immediate reward based on the current state, the next state, and the current policy includes:
[0067] Calculate the action quality reward based on the action quality score in the current state and the action quality score in the next state;
[0068] The user's state reward is calculated based on the self-satisfaction score, fatigue score, and comprehension score in the next state, as well as the comprehension score in the current state.
[0069] Calculate training progress rewards based on sensor data from the motion device in the next state;
[0070] Calculate the policy cost penalty based on the current policy;
[0071] The immediate reward is obtained by weighting and summing the action quality reward, user status reward, training progress reward, and policy cost penalty.
[0072] For example, Equation (1) is used to calculate instant rewards.
[0073] r t =w q ×R quality +w us ×R user_state +w p ×R progression +w c ×R action_cost (1)
[0074] Where, r t Used to represent immediate rewards, R quality Used to characterize the reward for action quality, R user_state R is used to characterize user status rewards. progression R is used to represent training progress rewards. action_cost Used to characterize the policy cost penalty, w q w us w pw c These are the corresponding weights.
[0075] This application's embodiments calculate the immediate reward generated by the virtual coach using the current strategy from multiple dimensions, such as action quality and user status. Because this immediate reward considers multiple different dimensions, it can more comprehensively and accurately measure the current strategy, select a strategy more suitable for the user, and thus improve the quality of training guidance. In practical application scenarios, the calculation method of the immediate reward can also be adjusted according to business needs, such as calculating the immediate reward based on action quality rewards and user status rewards.
[0076] Furthermore, the embodiments of this application employ a weighted summation method to calculate immediate rewards, which further reflects the differences in importance among various dimensions and improves the accuracy of immediate reward calculation. Rewards or penalties, such as action quality rewards, can also be calculated in different ways; only one such calculation method will be described in detail below.
[0077] In one embodiment of this application, the action quality score in the current state and the action quality score in the next state both include scores for multiple action dimensions.
[0078] Based on the action quality score in the current state and the action quality score in the next state, calculate the action quality reward, including:
[0079] Calculate the basic action quality reward based on the scores of multiple action dimensions in the next state;
[0080] The increment of action quality reward is calculated based on the score difference between the current state and the previous state in each action dimension.
[0081] The action quality reward is calculated based on the basic action quality reward and the action quality reward increment.
[0082] The action quality reward can be calculated using equation (2).
[0083]
[0084] Where, q t+1,i q is the score for action dimension i in the next training action. t,i w represents the score of action dimension i in the current training action. q1 The weights used to characterize the rewards for the quality of basic movements, w q2 The weights used to characterize the incremental reward for action quality.
[0085] This embodiment of the application encourages users to improve the quality of their training movements through action quality rewards.
[0086] It not only considers the quality of the next state, but also the changes of the next state relative to the current state, which makes the action quality reward more accurately reflect the quality of the current strategy.
[0087] In one embodiment of this application, a training progress reward is calculated based on sensor data of the motion device in the next state, including:
[0088] Based on the sensor data of the motion device in the next state, determine whether the user has completed the next training action;
[0089] In response to completing the next training action, the training progress reward is set to a preset first threshold.
[0090] In response to the failure to complete the next training action, the training progress reward is set to a preset second threshold.
[0091] Among them, the first threshold is greater than the second threshold, and the first threshold is a positive number.
[0092] To encourage users to complete training exercises, this application provides greater rewards for completed exercises, such as R for completing the next exercise. progression =C prog , where C prog This is the first threshold. In practical applications, failing to complete the next training action can discourage or penalize the user; for example, the second threshold can be 0 or a negative value.
[0093] In one embodiment of this application, when the action quality reward, user status reward, and training progress reward are fixed, the immediate reward corresponding to any one of the following strategies—outputting preset encouraging feedback, outputting correction instructions, outputting rest suggestions, and outputting questioning—is less than the immediate reward corresponding to no output.
[0094] No output indicates that no virtual coach intervention is needed. To avoid excessive intervention, this application encourages the no-output strategy.
[0095] For example, R can output pre-defined encouraging feedback, corrective instructions, rest suggestions, and questions. action_cost =-C cost For no output, R action_cost =0, while C cost It is used to characterize the fixed cost attached to each proactive intervention by a virtual coach.
[0096] In one embodiment of this application, a user state reward is calculated based on the self-satisfaction score, fatigue score, and comprehension score in the next state, as well as the comprehension score in the current state, including:
[0097] Determine the self-satisfaction reward based on the self-satisfaction score in the next state;
[0098] Based on the fatigue score in the next state, determine the fatigue reward;
[0099] Based on the comprehension score in the next state, determine the comprehension reward;
[0100] The increment of the comprehension reward is determined based on the difference between the comprehension score in the next state and the comprehension score in the current state;
[0101] The user state reward is calculated based on self-satisfaction reward, fatigue level reward, comprehension level reward and comprehension level reward increment.
[0102] Taking self-satisfaction reward as an example, if the self-satisfaction score is 2, then the self-satisfaction reward is C. sat Otherwise, the self-satisfaction reward is 0.
[0103] In this embodiment of the application, the user status reward can also be directly calculated using equation (3).
[0104] R user_state =C sat ×Π(sat t+1 =High)-P fat ×Π(f t+1 =High)-P und ×Π(u t+1 =Low)+C und_inc ×Π(u t+1 >U t ) (3)
[0106] Here, Π(·) is an indicator function that outputs 1 when the condition within the parentheses is true, and 0 otherwise; while C sat P fat P und C und_inc These are preset positive numbers, representing rewards for users reaching high satisfaction, high fatigue, low comprehension, and improved comprehension compared to their current state, respectively. Negative values represent penalties. t+1 f is used to characterize self-satisfaction in the next state. t+1 Used to characterize the degree of fatigue in the next state, u t+1 U is used to characterize the level of understanding of the next state. t Used to characterize the level of understanding of the current state, High and Low correspond to high and low levels of understanding, respectively. They can also be replaced by scores of 2 and 0.
[0107] The embodiments of this application can encourage positive user states, such as high comprehension, and punish negative states, such as high fatigue. By rewarding positive changes in states through comprehension increments, the training quality of users can be improved.
[0108] This application provides a method for generating virtual coaching guidance strategies, including:
[0109] Based on the DQN model trained using any of the above embodiments, a strategy is generated for virtual coaches to guide the training of target users.
[0110] Specifically, the state of the target user after performing the target training action is obtained, and this state is input into the trained DQN model to obtain a strategy to guide the training of the target user.
[0111] like Figure 2 As shown in the figure, this application embodiment provides a model training apparatus, including:
[0112] The acquisition module 201 is configured to acquire the current state of the user after performing the current training action; wherein, the current state includes: action quality score, fatigue level score, self-satisfaction score, comprehension level score and any number of sensor data from the motion device. The action quality score is used to characterize the execution quality of the current training action, the fatigue level score is used to characterize the user's fatigue level, the self-satisfaction score is used to characterize the user's satisfaction with performing the current training action, and the comprehension level score is used to characterize the user's comprehension of the current training action.
[0113] The input module 202 is configured to input the current state into the deep Q-network (DQN) model to obtain the Q-values of each strategy in the preset action space. The action space includes any number of strategies used by the virtual coach to guide user training, such as outputting preset encouraging feedback, outputting corrective instructions, outputting rest suggestions, outputting questions, and no output.
[0114] Processing module 203 is configured to select the current policy from the action space based on the Q value of each policy; and execute the current policy.
[0115] Adjust module 204, configured to obtain the user's next state after the user responds to the current policy and executes the next training action; adjust the parameters of the DQN model based on the current state, the next state, and the current policy.
[0116] In one embodiment of this application, the acquisition module 201 is configured to convert the discrete value into a numerical vector through one-hot encoding or an embedding layer in response to the current item being a discrete value; wherein, the action quality score, fatigue score, self-satisfaction score, and comprehension score are discrete values, and the sensor data of the motion device are continuous values; in response to the current item being a continuous value, the current item is normalized;
[0117] Input module 202 is configured to concatenate the vectors corresponding to each item in the current state and then input them into the DQN model.
[0118] In one embodiment of this application, the adjustment module 204 is configured to: calculate an immediate reward based on the current state, the next state, and the current policy; store the experience consisting of the current state, the current policy, the immediate reward, and the next state into an experience replay buffer; sample a preset number of target experiences from the experience replay buffer; calculate the target Q-value of each target experience based on the target network in the DQN model; calculate the predicted Q-value of each target experience based on the current network in the DQN model; calculate the value of the loss function based on the target Q-value and the predicted Q-value of each target experience; and adjust the parameters of the DQN model based on the value of the loss function.
[0119] In one embodiment of this application, the adjustment module 204 is configured to calculate an action quality reward based on the action quality score in the current state and the action quality score in the next state; calculate a user state reward based on the self-satisfaction score, fatigue score, and comprehension score in the next state, and the comprehension score in the current state; calculate a training progress reward based on the sensor data of the motion device in the next state; calculate a strategy cost penalty based on the current strategy; and obtain an immediate reward by weighted summing of the action quality reward, user state reward, training progress reward, and strategy cost penalty.
[0120] In one embodiment of this application, the adjustment module 204 is configured to calculate a basic action quality reward based on the scores of multiple action dimensions in the next state; calculate an action quality reward increment based on the score difference between the next state and the current state in each action dimension; and calculate an action quality reward based on the basic action quality reward and the action quality reward increment; wherein the action quality score in the current state and the action quality score in the next state both include scores of multiple action dimensions.
[0121] In one embodiment of this application, the adjustment module 204 is configured to determine whether the user has completed the next training action based on the sensor data of the motion device in the next state; in response to completing the next training action, determine the training progress reward as a preset first threshold; in response to not completing the next training action, determine the training progress reward as a preset second threshold; wherein the first threshold is greater than the second threshold, and the first threshold is a positive number.
[0122] In one embodiment of this application, the adjustment module 204 is configured to determine a self-satisfaction reward based on the self-satisfaction score in the next state; determine a fatigue reward based on the fatigue score in the next state; determine a comprehension reward based on the comprehension score in the next state; determine a comprehension reward increment based on the difference between the comprehension score in the next state and the comprehension score in the current state; and calculate a user state reward based on the self-satisfaction reward, fatigue reward, comprehension reward, and comprehension reward increment.
[0123] This application provides a virtual coaching system, including:
[0124] The guidance module is configured to generate a strategy for virtual coaching of target users' training based on the DQN model trained using any of the methods described in the above embodiments.
[0125] Specifically, the guidance module is configured to obtain the state of the target user after performing the target training action, input the state into the trained DQN model, and obtain a strategy to guide the target user's training.
[0126] The virtual coaching system also includes a model training device. The guidance module generates strategies for the virtual coach to guide the target user's training based on the DQN model trained by the model training device.
[0127] This application provides an electronic device, including:
[0128] One or more processors;
[0129] Storage device for storing one or more programs.
[0130] When one or more programs are executed by one or more processors, the one or more processors implement the methods as described in any of the above embodiments.
[0131] This application provides a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements the method as described in any of the above embodiments.
[0132] It should be noted that the computer-readable medium shown in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable medium other than computer-readable storage media, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0133] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0134] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A model training method, characterized in that, include: Obtain the current state of the user after performing the current training action; wherein, the current state includes: action quality score, fatigue level score, self-satisfaction score, comprehension level score, and any multiple of the following: action quality score, fatigue level score, self-satisfaction score, comprehension level score, and sensor data from the exercise device. The action quality score is used to characterize the execution quality of the current training action, the fatigue level score is used to characterize the user's fatigue level, the self-satisfaction score is used to characterize the user's satisfaction with performing the current training action, and the comprehension level score is used to characterize the user's comprehension of the current training action. The current state is input into a deep Q-network (DQN) model to obtain the Q-values of each strategy in a preset action space; wherein, the action space includes: any number of strategies used by the virtual coach to guide user training, such as outputting preset encouraging feedback, outputting corrective instructions, outputting rest suggestions, outputting questions, and no output. Based on the Q value of each of the aforementioned strategies, the current strategy is selected from the action space; Execute the current policy; After the user responds to the current strategy and performs the next training action, the user's next state is obtained; The parameters of the DQN model are adjusted based on the current state, the next state, and the current strategy.
2. The method as described in claim 1, characterized in that, Further includes: In response to the current item being a discrete value, the discrete value is converted into a numerical vector through one-hot encoding or an embedding layer; wherein, the action quality score, the fatigue level score, the self-satisfaction score, and the comprehension level score are the discrete values, and the sensor data of the motion device are continuous values; In response to the fact that the current item is a continuous value, the current item is normalized; Inputting the current state into the DQN model includes: The vectors corresponding to each item in the current state are concatenated and then input into the DQN model.
3. The method as described in claim 1, characterized in that, Based on the current state, the next state, and the current policy, adjust the parameters of the DQN model, including: Calculate the immediate reward based on the current state, the next state, and the current strategy; The experience consisting of the current state, the current strategy, the immediate reward, and the next state is stored in the experience replay buffer. A preset number of target experiences are sampled from the experience playback buffer; Calculate the target Q value of each target experience based on the target network in the DQN model; The predicted Q-values for each of the target experiences are calculated based on the current network in the DQN model. Based on the target Q-value and predicted Q-value of each of the aforementioned target experiences, the value of the loss function is calculated; The parameters of the DQN model are adjusted based on the value of the loss function.
4. The method as described in claim 3, characterized in that, Based on the current state, the next state, and the current policy, calculate the immediate reward, including: Calculate the action quality reward based on the action quality score in the current state and the action quality score in the next state; The user's state reward is calculated based on the self-satisfaction score, fatigue score, and comprehension score in the next state, as well as the comprehension score in the current state. Based on the sensor data of the motion device in the next state, calculate the training progress reward; Calculate the policy cost penalty based on the current policy; The immediate reward is obtained by weighted summation of the action quality reward, the user status reward, the training progress reward, and the policy cost penalty.
5. The method as described in claim 4, Its features are, The action quality score in the current state and the action quality score in the next state both include scores for multiple action dimensions. Based on the action quality score in the current state and the action quality score in the next state, an action quality reward is calculated, including: Based on the scores of multiple action dimensions in the next state, calculate the basic action quality reward; The action quality reward increment is calculated based on the score difference between the previous state and the current state in each of the aforementioned action dimensions. The action quality reward is calculated based on the basic action quality reward and the action quality reward increment.
6. The method as described in claim 4, characterized in that, Based on the sensor data of the motion device in the next state, a training progress reward is calculated, including: Based on the sensor data of the motion device in the next state, it is determined whether the user has completed the next training action; In response to completing the next training action, the training progress reward is determined to be a preset first threshold; In response to the failure to complete the next training action, the training progress reward is determined to be a preset second threshold; Wherein, the first threshold is greater than the second threshold, and the first threshold is a positive number.
7. The method as described in claim 4, characterized in that, in, With the action quality reward, user status reward, and training progress reward fixed, the immediate reward corresponding to any one of the following strategies—the preset output encouragement feedback, the output correction instruction, the output rest suggestion, and the output question—is less than the immediate reward corresponding to no output.
8. The method as described in claim 4, characterized in that, Based on the self-satisfaction score, fatigue score, and comprehension score in the next state, and the comprehension score in the current state, the user's state reward is calculated, including: The self-satisfaction reward is determined based on the self-satisfaction score in the next state. Based on the fatigue level score in the next state, a fatigue level reward is determined; Based on the comprehension score in the next state, a comprehension reward is determined. The comprehension reward increment is determined based on the difference between the comprehension score in the next state and the comprehension score in the current state; The user state reward is calculated based on the self-satisfaction reward, the fatigue level reward, the comprehension level reward, and the comprehension level reward increment.
9. A method for generating a virtual coaching strategy, characterized in that, include: Based on the DQN model trained by any one of the methods described in claims 1-8, a strategy for virtual coaching to guide target user training is generated.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-8.
Citation Information
Patent Citations
Virtual object behavior strategy training method and device, electronic equipment and storage medium
CN111026272A
Model training and intervention strategy determination method and device, and electronic equipment
CN116844695A
Training action selection neural networks using a differentiable credit function
US20200175364A1
Off-line learning for robot control using a reward prediction model
US20230256593A1