Motion tracking method and device and electronic equipment
By employing a distributed perception and failure-first resampling training motion strategy, combined with motion increment and feasible region projection, the problem of long-term drift and error accumulation in robot motion tracking is solved, achieving a more stable motion tracking effect.
Patent Information
- Application Number
- CN202511381878.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-25
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2045-09-25
AI Technical Summary
In existing technologies, the systematic differences between the robot's actual dynamics and the reference motion are ignored when imitating or tracking robot movements. This leads to problems such as long-term drift and error accumulation, long data tails and inefficient learning, limitations of phase scalars, and gaps between simulation and reality.
By using a trained motion strategy, motion increments are determined based on distribution-aware resampling and failure-first resampling, and the target robot is controlled to perform motion tracking. This includes distribution-aware resampling and failure-first dynamic sampling, constructing a hybrid sampling distribution, selecting training segments, training the untrained motion strategy, and updating the motion strategy by combining feasible region projection and joint actuator execution.
It achieves more stable long-term high dynamic tracking, reduces drift and accumulated errors, improves training efficiency, and enhances the stability and accuracy of robot motion tracking.
Smart Images

Figure CN121004610A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of robots, and in particular to a motion tracking method and device and electronic equipment. BACKGROUND
[0002] In the prior art, robot motion imitation or tracking is mostly direct prediction of absolute joint commands, ignoring systematic differences between real robot dynamics and reference motions, resulting in problems such as long-term drift and error accumulation, long-tail data and inefficient learning, phase scalar limitations, and simulation-to-real discrepancies. Therefore, an effective solution is urgently needed to solve at least one of the above problems. SUMMARY
[0003] To solve the above problems, the present application provides a motion tracking method, device and electronic equipment.
[0004] The present application provides a motion tracking method, comprising: Determining a motion increment for executing a tracking motion based on the tracking motion and motion parameters of a target robot through a trained motion strategy, the motion strategy being obtained by training based on distribution perception and failure priority resampling; Controlling the target robot to perform motion tracking based on the motion increment and the tracking motion.
[0005] According to the motion tracking method provided by the present application, the training process of the motion strategy comprises: Performing distribution perception resampling and failure priority dynamic sampling on a reference motion sequence to obtain a hybrid sampling distribution; Selecting a training segment based on the hybrid sampling distribution; Training an untrained motion strategy based on the training segment to obtain the trained motion strategy.
[0006] According to the motion tracking method provided by the present application, the distribution perception resampling and failure priority dynamic sampling on the reference motion sequence to obtain the hybrid sampling distribution comprises: Dividing the reference motion sequence into at least two segments in chronological order; Performing distribution perception balanced sampling on all the segments using an occupation vector of a key joint degree of freedom to obtain a distribution matrix, the distribution matrix containing sampling distributions of occupation grids of each segment; For each segment, performing difficult example priority sampling based on a failure rate of the segment to obtain a sampling probability of the segment; Determining a hybrid sampling distribution of each segment based on the distribution matrix and the sampling probability corresponding to each segment.
[0007] According to the motion tracking method provided by the application, the distribution matrix is obtained by performing distribution-aware balanced sampling on all the segments based on the occupation vector of the key joint degree of freedom, comprising: The target segment corresponding to the key joint degree of freedom is discretized into a grid, and the occupation vector of each segment occupying the grid is calculated; The distribution matrix is obtained by taking the average of the occupation vector as the target.
[0008] According to the motion tracking method provided by the application, the sampling probability of the segment is obtained by performing difficult example priority sampling based on the failure rate of the segment, comprising: The local failure rate of the segment in the current batch training and the global failure rate of the segment in the overall training are obtained; The global failure rate is updated based on the local failure rate; The sampling probability of the segment is determined based on the updated global failure rate.
[0009] According to the motion tracking method provided by the application, the trained motion strategy is obtained by training the untrained motion strategy based on the training segment, comprising: Based on the interaction of each training segment, an observation balance condition is constructed, which represents that the overall observation is balanced with the perception action, and the perception action includes the body perception of the target robot and the reference action corresponding to each training segment; Based on the overall observation, the reference increment corresponding to the reference action is determined; Based on the reference increment, an action tracking target is constructed, which represents that the target action is the sum of the reference action and the reference increment; Based on the action tracking target, the feasible region projection and joint actuator execution of the target joint of the target robot are performed to update the motion strategy, and the trained motion strategy is obtained.
[0010] According to the motion tracking method provided by the application, the motion strategy is updated based on the action tracking target, comprising: Based on the action tracking target, the feasible region projection and joint actuator execution of the target joint of the target robot are performed to obtain the change amount of the state of the target robot and the environment after interaction; The reward is calculated based on the change amount; Based on the reward, the parameters of the motion strategy and the parameters of the overall strategy are updated, and the overall strategy is used to judge the state value.
[0011] The action tracking method provided by the application further comprises: In the process of training the motion strategy, the weights and thresholds in the tracking environment are adjusted according to the training progress; and the physical parameters of the tracking environment are randomized when the action tracking is reset.
[0012] The application further provides an action tracking device, comprising: A determination module is configured to determine, by using a trained motion strategy, a motion increment for performing a tracking action based on the motion parameters of the tracking action and a target robot, wherein the motion strategy is trained based on distribution perception and failure-priority resampling. A control module is configured to control the target robot to perform action tracking based on the motion increment and the tracking action.
[0013] The application further provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the action tracking method according to any one of the above when executing the computer program.
[0014] The application further provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program is executable by a processor to implement the action tracking method according to any one of the above.
[0015] The application further provides a computer program product comprising a computer program, wherein the computer program is executable by a processor to implement the action tracking method according to any one of the above.
[0016] The action tracking method, device and electronic device provided by the application can make the motion strategy converge faster, improve the training efficiency, reduce drift and accumulated error by combining the motion increment with the focus on the dynamics compensation, and can achieve more stable long-term high-dynamic tracking through the distribution perception and failure-priority resampling. BRIEF DESCRIPTION OF DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the application or prior art, the following will briefly introduce the drawings needed in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the application, and those skilled in the art can obtain other drawings according to these drawings without any creative effort.
[0018] Figure 1is a flowchart of a motion tracking method provided by the present application.
[0019] Figure 2 is a structural diagram of a motion tracking device provided by the present application.
[0020] Figure 3 is a structural diagram of an electronic device provided by the present application. DETAILED DESCRIPTION
[0021] In order to make the objectives, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below with reference to the drawings in the present application. Obviously, the described embodiments are some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0022] First, the related content involved in the present application is briefly described.
[0023] MDP is Markov Decision Process, and POMDP is Partially Observable Markov Decision Process.
[0024] s represents a system state, which is not directly observable; O represents an overall observation; p represents proprioception of the robot; and G represents a reference / motion library. In the present application, O=Ψ(P,G) is a fused observation of P and G.
[0025] Residual / Delta Action is an increment a of a motion strategy output relative to a reference joint / motion (qref) command, and a final joint target qtar=qref+a.
[0026] Goal-conditioned RL is goal-conditioned reinforcement learning, which takes a reference segment / local future as a “goal” encoding g, and a motion strategy π(a|O,g) outputs an action.
[0027] A PD (Proportion Differentiation) controller, i.e., a joint actuator, i.e., a joint space PD controller, is used to convert qtar into an execution torque τ, where τ=PD(qtar,q,q , ), where q represents a joint position, and q ,Reference position.
[0028] Feasible region projection ΠC is used to limit and project at least one of the position, velocity and torque of the joint, to ensure hardware safety.
[0029] Distribution-aware resampling is used to sample equally according to the joint occupancy histogram, to improve long-tail pose coverage.
[0030] Failure-aware priority is used to dynamically increase the sampling probability according to the exponential moving average failure rate of difficult examples.
[0031] Curriculum learning is used to gradually increase the difficulty from“lenient to strict”.
[0032] Domain randomization is used to randomize physical parameters such as friction, inertia, gain and delay to enhance robustness and transferability.
[0033] Selective residualization is used to output residuals for key / sensitive DOFs, and other joint DOFs (Degree of Freedom) are transmitted transparently.
[0034] The related art has the following problems due to neglecting the systematic difference between the real dynamics of the robot and the reference motion: long-term drift and error accumulation, such as small deviations being amplified step by step in long sequences (such as dance); data long tail and learning inefficiency, such as over-sampling of common poses and under-sampling of rarely occurring“key poses”; limitations of phase scalar, such as using a single“phase”to encode progress, which is not robust to rhythm changes, time warping and segment splicing; simulation to reality (sim2real) gap, such as friction, inertia, delay and other biases that make tracking unstable and chattering worse.
[0035] To solve the problems of long-time sequence motion tracking drift, data long-tail learning inefficiency and dynamics mismatch, the present application provides a motion tracking method, device and electronic equipment.
[0036] The motion tracking method, device and electronic equipment of the present application will be described below. Figures 1-3 The motion tracking method, device and electronic equipment of the present application will be described below.
[0037] Figure 1 The flowchart of the motion tracking method provided by the present application is shown in FIG. 1, which includes steps 101 and 102. Figure 1
[0038] Step 101: determining a motion increment for performing the tracking action based on the tracking action and the motion parameters of the target robot by using the trained motion strategy, wherein the motion strategy is trained based on distribution perception and failure-priority resampling.
[0039] Specifically, the tracking action is an action that needs to be tracked by the target robot. The motion increment can be a residual action, which is used to concentrate capacity on dynamics compensation and is more stable in long-term.
[0040] In actual application, the tracking action can be obtained first, and then the tracking action and the motion parameters of the target robot are processed by using the trained motion strategy to output the increment a of the tracking action, i.e., the motion increment.
[0041] Before using the trained motion strategy, the motion strategy can be trained by using distribution / failure perception resampling to improve training efficiency.
[0042] Step 102: controlling the target robot to perform action tracking based on the motion increment and the tracking action.
[0043] In actual application, on the basis of obtaining the motion increment, the motion increment and the tracking action are superimposed to obtain a target action, and further, the target robot is controlled to perform the target action to perform action tracking.
[0044] The action tracking method provided by the application determines a motion increment for performing the tracking action based on the tracking action and the motion parameters of the target robot by using the trained motion strategy, wherein the motion strategy is trained based on distribution perception and failure-priority resampling; and the target robot is controlled to perform action tracking based on the motion increment and the tracking action. The motion strategy can converge faster and improve training efficiency by using distribution perception and failure-priority resampling. In combination with the motion increment focusing on dynamics compensation, the drift and cumulative error are reduced, and more stable long-term high-dynamic tracking can be achieved.
[0045] In one or more optional embodiments of the application, the training process of the motion strategy comprises: performing distribution perception resampling and failure-priority dynamic sampling on the reference action sequence to obtain a hybrid sampling distribution; selecting a training segment based on the hybrid sampling distribution; training the untrained motion strategy based on the training segment to obtain the trained motion strategy.
[0046] In actual application, distribution-aware resampling can be performed to balance sampling with reference to the occupation histogram of the key joint DOF in the action sequence, and failure-aware dynamic sampling can be performed to construct a difficult case priority distribution with reference to the EMA (Exponential Moving Average) failure rate corresponding to the action sequence, so as to obtain a hybrid sampling distribution, wherein the difficult case refers to a case with a high failure rate, and the difficult case priority distribution refers to a distribution in which the case with a higher failure rate is more preferentially distributed.
[0047] Further, after obtaining the hybrid sampling distribution, a segment with a higher hybrid sampling distribution can be selected from the segments corresponding to the reference action sequence as a training segment to train the motion model, or the segments can be arranged in a sequence from high to low according to the hybrid sampling distribution to form training segments for training the motion model, and then a trained motion strategy is obtained.
[0048] In the embodiments of the present application, the distribution-aware resampling can improve long-tail pose coverage and reduce training bias, and the failure-aware dynamic sampling can adaptively focus on the difficult segments of the reference action sequence in the continuous reference action sequence and accelerate convergence.
[0049] In one or more optional embodiments of the present application, the distribution-aware resampling and the failure-aware dynamic sampling of the reference action sequence to obtain the hybrid sampling distribution include: The reference action sequence is divided into at least two segments in time sequence; The distribution-aware balanced sampling is performed on all the segments with the occupation vector of the key joint DOF to obtain a distribution matrix, and the distribution matrix contains the sampling distribution of the occupation grid of each segment; For each segment, difficult case priority sampling is performed based on the failure rate of the segment to obtain the sampling probability of the segment; The hybrid sampling distribution of each segment is determined based on the distribution matrix and the sampling probability corresponding to each segment.
[0050] In actual application, segment division and histogram estimation can be performed first: the reference action sequence is divided into S segments in time, wherein S is a positive integer greater than 1. Then, histogram balanced sampling, i.e., distribution-aware balanced sampling, is performed with the occupation vector of the key joint DOF, so as to obtain a distribution matrix ω, thereby improving long-tail pose coverage and reducing training bias.
[0051] Further, the EMA failure rate is used to construct a difficult case priority distribution to obtain the sampling probability P of the s-th segment. sAdaptively focus on the difficult segment to accelerate the convergence. Wherein s is an integer greater than or equal to 1 and less than or equal to S.
[0052] Then, the mixed sampling distribution P of the s-th segment is calculated according to the following formula mix .
[0053] P mix = (1-λ) softmax (ω) + λ softmax (P s ) Wherein, λ is an annealing parameter, which can be 0→0.5.
[0054] In the embodiment of the application, the mixed sampling distribution is determined based on the distribution matrix and the sampling probability corresponding to each segment, so as to train the motion strategy, which helps to improve the training efficiency.
[0055] In one or more optional embodiments of the application, the distribution-aware balanced sampling of all the segments is performed by using the occupation vector of the key joint DOF to obtain a distribution matrix, including: The target segment corresponding to the key joint DOF is discretized into a grid, and the occupation vector of each segment occupying the grid is calculated; The distribution matrix is obtained by taking the occupation vector as the target.
[0056] In practical applications, histogram estimation can be performed: the target segment corresponding to the key joint DOF (such as hip / knee pitch) is discretized into a grid, and the occupation vector occ s of each segment relative to the grid is calculated as an optimization problem, that is, the occupation vector is averaged, that is, Pω=u. Wherein ω is a matrix formed by the distribution of each segment occupying the grid, that is, a distribution matrix; P is a matrix formed by the sampling probability of each segment, that is, a probability matrix, and u is the target of the average distribution. In this way, the accuracy and efficiency of the distribution matrix can be ensured.
[0057] Exemplarily, the distribution-aware weight (probability) is solved in a static manner, that is, the following formula is solved to obtain the distribution matrix ω.
[0058] Wherein, ω is the distribution matrix; P is the probability matrix, and u is the target of the average distribution.
[0059] Or use the inverse frequency weight (probability) method to determine the distribution matrix ω, that is, the distribution ω of the s-th segment is determined according to the following formula s , and the distributions of the segments are arranged in the form of a matrix to obtain the distribution matrix ω.
[0060] ω s ∝1 / (occs +ε1) Wherein, ε1 is a minimum value, preventing the denominator from being zero.
[0061] In one or more optional embodiments of the present application, the difficulty example priority sampling based on the failure rate of the segment obtains a sampling probability of the segment, including: Obtaining a local failure rate of the segment in the current batch training and a global failure rate of the segment in the overall training; Updating the global failure rate based on the local failure rate; Determining the sampling probability of the segment based on the updated global failure rate.
[0062] In practical applications, the failure rate can be counted online, and the global failure rate is updated after one batch training is completed. The formula for updating the global failure rate is as follows: r s ←αf s +(1-α)r s Wherein, r s is the global failure rate corresponding to the s th segment, f s is the local failure rate corresponding to the s th segment, and α is a parameter based on a sliding average.
[0063] After updating the global failure rate, the updated global failure rate is processed according to the following formula to obtain the sampling probability P s of the s th segment, which is used for difficulty example priority.
[0064] P s ∝(r s +ε2) β Wherein, β is a sampling temperature, and ε2 is a basic sampling probability (minimum value).
[0065] In the embodiments of the present application, the global failure rate is updated based on the local failure rate, and the difficulty example priority is performed based on the updated global failure rate, which can improve the accuracy of the sampling probability.
[0066] In one or more optional embodiments of the present application, the training of the untrained motion strategy based on the training segment obtains the trained motion strategy, including: Based on the expansion interaction of each training segment, an observation balance condition is constructed, and the observation balance condition represents that the overall observation is balanced with the perception action, and the perception action includes the body perception of the target robot and the reference action corresponding to each training segment; Based on the overall observation, a reference increment corresponding to the reference action is determined; constructing a motion tracking target based on the reference increment, the motion tracking target representing the target motion as a sum of the reference motion and the reference increment; projecting and executing the target joint of the target robot based on the motion tracking target to update the motion strategy, to obtain the trained motion strategy.
[0067] In practical applications, rolling and strategy updating can be performed. Specifically, the sampling segment is unfolded to interact, and an observation balance condition O t =[p t ;G t ], wherein t represents the current training batch t, O is the overall observation, p is the body perception, and G is the reference motion.
[0068] Then, selective residualization and safety projection are performed: based on the overall observation, a neural network (NN, Neural Network) (built-in the robot or built-in the motion strategy) is used to determine the reference increment corresponding to the reference motion, i.e., a=NN(O t ), wherein NN represents the neural network; based on the reference increment, a motion tracking target qtar=qref+a is constructed, wherein qref is the reference motion, and qtar is the target motion. In this way, the reference condition that is independent of the phase, i.e., qref, is used to replace a single phase, which can improve the robustness of rhythm changes / segment splicing.
[0069] Then, based on the motion tracking target, projection and smoothing regularization are performed through ΠC (feasible region projection) and PD (joint actuator) execution to suppress chattering and out-of-bound, to update the motion strategy.
[0070] In one or more optional embodiments of the present application, the projecting and executing the target joint of the target robot based on the motion tracking target to update the motion strategy comprises: projecting and executing the target joint of the target robot based on the motion tracking target to obtain a change amount of the state of the target robot after interacting with the environment; calculating a reward based on the change amount; updating parameters of the motion strategy and parameters of an overall strategy based on the reward, the overall strategy being used to judge the state value.
[0071] In practical applications, based on the motion tracking target, the change amount of the state of the target robot after interacting with the environment can be obtained through ΠC (feasible region projection) and PD (joint actuator) execution, and then a reward R is calculated based on the change amount (including Rimit and Rphys). See the following formula: R = wimit·Rimit + wphys·Rphys.
[0072] where wimit is an imitation reward weight; Rimit is joint / speed / root pose / end / timing of contact; wphys is a reward weight of physical constraints; Rphys is at least one of torque, joint limits, foot slip, and Δa, and Δa is a change amount of a reference increment.
[0073] Further, a proximal policy optimization algorithm (PPO), a soft actor-critic algorithm (SAC), a twin delayed deep deterministic policy gradient algorithm (TD3), or an importance weighted actor-learner architecture (IMPALA) can be used to update the reward based on the reward, wherein πθ is a parameter of the motion policy, and Vϕ is a parameter of the overall policy.
[0074] In the embodiment of the present application, the reward is calculated by determining the change amount, and the parameter is updated based on the reward, which can improve the robustness.
[0075] In one or more optional embodiments of the present application, the method further comprises: During training of the motion policy, the weights and thresholds in the tracking environment are adjusted according to the training progress; and the physical parameters of the tracking environment are randomized when the action tracking is reset.
[0076] In practical applications, curriculum / randomization can be performed, i.e., the weights and thresholds are adjusted according to the training progress; and the physical parameters are randomized when the action tracking (episode) is reset. In this way, the difficulty progression and the randomization of the physical parameters can take into account the accuracy, stability, and transferability.
[0077] It should be noted that the action space can be a joint position residual, a joint torque / speed residual, or an end space residual.
[0078] The condition input of the overall observation can be a short look-ahead window Or a shape embedding (mass, inertia, joint limits, etc.).
[0079] In addition, the phase signal of the overall observation O and the reference motion / action library G: in a specific rhythm task, the phase can be provided in parallel with the reference joint (as a redundant robustness). In online adaptation, meta-learning / adaptive gain can be introduced, or the sampling temperature β and the mixing coefficient λ can be updated online.
[0080] Exemplarily, the motion tracking method is applied to a motion tracking system, which comprises: A data / reference library module (G): storing a reference joint sequence {qref}, a segment index, and a small amount of a look-ahead window; A body perception module (P): collecting robot joint positions / speeds, root speeds / angles, contact flags, historical residuals, and the like; An observation fusion module (O): Ot=Ψ(Pt,Gt) (supporting splicing or encoder fusion); A resampling module: two-level samplers of distribution perception (static) and failure priority (online); A residual strategy and training module: πθ(a t ∣O t ,G t ) outputs a residual (reference increment) a t , and a PPO or the like optimizer updates parameters, where t represents a current training batch t; A safety and control execution module: qtar=qref+ a → ΠC → joint actuator output τ; A course scheduling and field randomization module: weight / threshold annealing; friction / inertia / gain / delay randomization; A log and evaluation module: recording at least one index of a root mean square error (RMSE), a foot slip time, a root mean square (RMS) of a torque, and a success rate.
[0081] A pseudo code of the motion tracking method is as follows: Input: reference library G, control frequency fc, PPO parameters (γ,λGAE,clip); Output: policy parameters θ.
[0082] Initialize w ← balance_by_hist(P), r s ← 0 for iter = 1..T: Calculate P s ∝(r s +ε2) β , P mix =(1-λ)softmax(ω)+λsoftmax(P s ) Sample a segment s ~ P mix , collect a trajectory: O=[P;G], a~πθ(a|o,g), qtar=qref+a (qtar, t) ← Π_C(qtar, PD(qtar, q, q , ); compute reward and advantage PPO update (q, V); update r s EMA failure rate of Course / randomization scheduling (weight and threshold annealing; physical parameter randomization) Exemplarily (Unitree G1): reference dance action library, control frequency 50 Hz; adopt selective residual (output residual of 12 DOFs of two legs); realize stable tracking of long-term dance segments after 100,000 steps of training.
[0083] Exemplarily (H1-2): use the reward model and various hyperparameters, reward functions and observation space settings to realize cross-platform migration.
[0084] It should be noted that the hardware implementation of the motion tracking method includes: NV Orin AGX; joint driver supports position / speed / torque mode and PD; the software implementation of the motion tracking method includes: Linux + Python / C++; RL framework (PyTorch), simulation (MuJoCo / IsaacGym), real-time execution (LCM kit).
[0085] The motion tracking method provided by the application focuses on dynamics compensation through residual learning, reduces drift and cumulative error, that is, long-term stability; through distribution balance + difficult instance priority, key postures are covered and sample efficiency is improved, that is, training efficiency; through phase-independent reference + domain randomization, rhythm changes and parameter uncertainties are considered, that is, robust generalization; through selective residualization and feasible region projection, high-frequency chattering and hardware out-of-bound are suppressed, that is, safety and controllability; through reuse or small sample adaptation of the same motion strategy on G1 / H1-2 / PND Adam platforms, cross-platform migration is realized.
[0086] The motion tracking device provided by the application is described below, and the motion tracking device described below can be correspondingly referred to the motion tracking method described above.
[0087] Figure 2 is a structural schematic diagram of the motion tracking device provided by the application, as Figure 2 shown, the device comprises: A determination module 201 configured to determine a motion increment for executing a tracking action based on the motion parameters of the tracking action and a target robot through a trained motion strategy, wherein the motion strategy is obtained by training based on distribution perception and failure priority resampling. The control module 202 is configured to control the target robot to perform action tracking based on the motion increment and the tracking action.
[0088] The action tracking device provided by the application determines a motion increment for performing the tracking action based on the tracking action and the motion parameters of the target robot through the trained motion strategy, and the motion strategy is obtained through training based on distribution perception and failure-priority resampling; the target robot is controlled to perform action tracking based on the motion increment and the tracking action. Through distribution perception and failure-priority resampling, the motion strategy can converge faster, the training efficiency is improved, the drift and cumulative error are reduced through the motion increment focusing on dynamics compensation, and more stable long-term high-dynamic tracking can be realized.
[0089] In one or more optional embodiments of the application, the device further comprises a training module configured to: perform distribution perception resampling and failure-priority dynamic sampling on the reference action sequence to obtain a hybrid sampling distribution; select a training segment based on the hybrid sampling distribution; train an untrained motion strategy based on the training segment to obtain the trained motion strategy.
[0090] In one or more optional embodiments of the application, the training module is further configured to: divide the reference action sequence into at least two segments in chronological order; perform distribution perception balanced sampling on all the segments with the occupation vector of the key joint degree of freedom to obtain a distribution matrix, and the distribution matrix contains the sampling distribution of the occupation grid of each segment; for each segment, perform difficult case priority sampling based on the failure rate of the segment to obtain a sampling probability of the segment; determine a hybrid sampling distribution of each segment based on the distribution matrix and the sampling probability corresponding to each segment.
[0091] In one or more optional embodiments of the application, the training module is further configured to: select a target subsegment corresponding to the key joint degree of freedom to discretize into a grid, and calculate the occupation vector of each segment occupying the grid; obtain the distribution matrix with the average of the occupation vectors as the target.
[0092] In one or more optional embodiments of the application, the training module is further configured to: obtain the local failure rate of the segment in the current batch training and the global failure rate of the segment in the overall training. updating the global failure rate based on the local failure rate; determining a sampling probability of the segment based on the updated global failure rate.
[0093] In one or more optional embodiments of the present application, the training module is further configured to: construct an observation balance condition based on the expanded interaction of each training segment, the observation balance condition representing that the overall observation is balanced with a perception action, the perception action including body perception of the target robot and a reference action corresponding to each training segment; determine a reference increment corresponding to the reference action based on the overall observation; construct an action tracking target based on the reference increment, the action tracking target representing that the target action is the sum of the reference action and the reference increment; perform feasible region projection and joint actuator execution on the target joint of the target robot based on the action tracking target to update the motion strategy, thereby obtaining the trained motion strategy.
[0094] In one or more optional embodiments of the present application, the training module is further configured to: perform feasible region projection and joint actuator execution on the target joint of the target robot based on the action tracking target to obtain a change amount of the state of the target robot and the environment after interaction; calculate a reward based on the change amount; update parameters of the motion strategy and parameters of an overall strategy based on the reward, the overall strategy being used to judge the value of the state.
[0095] In one or more optional embodiments of the present application, the training module is further configured to: adjust the weight and the threshold value in the tracking environment according to the training progress during training of the motion strategy; and randomize the physical parameters of the tracking environment when the action tracking is reset.
[0096] Figure 3 is a structural schematic diagram of an electronic device provided by the present application, as Figure 3As shown, the electronic device can include a processor 310, a communications interface 320, a memory 330, and a communications bus 340, wherein the processor 310, the communications interface 320, and the memory 330 communicate with each other through the communications bus 340. The processor 310 can invoke a logical instruction in the memory 330 to execute an action tracking method, which includes determining a motion increment for performing a tracking action based on a motion parameter of the tracking action and a target robot through a trained motion strategy, the motion strategy being trained based on distribution perception and failure-first resampling; and controlling the target robot to perform action tracking based on the motion increment and the tracking action.
[0097] In addition, the logical instruction in the memory 330 described above can be implemented in the form of a software functional unit and sold or used as an independent product, which can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the methods described in the embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various program code storage media.
[0098] On the other hand, the present application also provides a computer program product, which includes a computer program, the computer program can be stored on a non-transitory computer readable storage medium, and the computer program can be executed by a processor to enable a computer to execute the action tracking method provided by the above-mentioned methods, which includes determining a motion increment for performing a tracking action based on a motion parameter of the tracking action and a target robot through a trained motion strategy, the motion strategy being trained based on distribution perception and failure-first resampling; and controlling the target robot to perform action tracking based on the motion increment and the tracking action.
[0099] In yet another aspect, the present application also provides a non-transitory computer-readable storage medium having stored thereon a computer program, which, when executed by a processor, implements the action tracking method provided by the above method, and the method comprises: determining a motion increment of performing a tracking action based on a tracking action and a motion parameter of a target robot by a trained motion strategy, the motion strategy being trained based on distribution perception and failure priority-based resampling; and controlling the target robot to perform action tracking based on the motion increment and the tracking action.
[0100] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e., can be located in one place or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiment scheme according to actual needs. Those skilled in the art can understand and implement without creative labor.
[0101] From the above description of the embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software plus necessary universal hardware platforms, and of course can also be realized by hardware. Based on such understanding, the above technical solutions, essentially or in other words, the part that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.
[0102] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements to some technical features; and these modifications or replacements do not make the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A motion tracking method, characterized in that, include: Using a trained motion strategy, the motion increment for performing the tracking action is determined based on the tracking action and the motion parameters of the target robot. The motion strategy is trained based on distributed perception and failure-first resampling. Based on the motion increment and the tracking action, the target robot is controlled to perform motion tracking.
2. The motion tracking method according to claim 1, characterized in that, The training process for the aforementioned movement strategy includes: The reference action sequence is subjected to distribution-aware resampling and failure-priority dynamic sampling to obtain a hybrid sampling distribution; Based on the aforementioned mixed sampling distribution, training segments are selected; The untrained movement strategy is trained based on the training segments to obtain the trained movement strategy.
3. The motion tracking method according to claim 2, characterized in that, The process of performing distribution-aware resampling and failure-priority dynamic sampling on the reference action sequence to obtain a hybrid sampling distribution includes: The reference action sequence is divided into at least two segments according to chronological order; Using the occupancy vector of the key joint degrees of freedom, a distribution-aware equalization sampling is performed on all the segments to obtain a distribution matrix, which contains the sampling distribution of the grid occupied by each segment; For each segment, hard cases are sampled first based on the failure rate of the segment to obtain the sampling probability of the segment; Based on the distribution matrix and the sampling probability corresponding to each segment, the mixed sampling distribution of each segment is determined.
4. The motion tracking method according to claim 3, characterized in that, The distribution matrix is obtained by performing distribution-aware equalization sampling on all segments using the occupancy vectors of key joint degrees of freedom, including: The target segments corresponding to the key joint degrees of freedom are selected and discretized into a mesh, and the occupancy vector of each segment in the mesh is calculated. The distribution matrix is obtained by taking the average of the occupancy vectors as the objective.
5. The motion tracking method according to claim 3, characterized in that, The step of performing hard-case priority sampling based on the failure rate of the fragment to obtain the sampling probability of the fragment includes: Obtain the local failure rate of the segment in the current batch training and the global failure rate of the segment in the overall training; The global failure rate is updated based on the local failure rate. The sampling probability of the segment is determined based on the updated global failure rate.
6. The motion tracking method according to any one of claims 2-5, characterized in that, The step of training the untrained motion strategy based on the training fragment to obtain the trained motion strategy includes: Based on the interaction of each training segment, an observation balance condition is constructed. The observation balance condition represents the balance between overall observation and perception action. The perception action includes the body perception of the target robot and the reference action corresponding to each training segment. Based on the overall observation, the reference increment corresponding to the reference action is determined; Based on the reference increment, an action tracking target is constructed, wherein the action tracking target represents the target action as the sum of the reference action and the reference increment; Based on the motion tracking target, feasible domain projection and joint actuator execution are performed on the target joints of the target robot to update the motion strategy and obtain the trained motion strategy.
7. The motion tracking method according to claim 6, characterized in that, The step of updating the motion strategy by projecting feasible regions and executing joint actuators on the target joints of the target robot based on the motion tracking target includes: Based on the motion tracking target, feasible domain projection and joint actuator execution are performed on the target joint of the target robot to obtain the change in state of the target robot after interaction with the environment. Calculate the reward based on the change; Based on the reward, the parameters of the motion strategy and the parameters of the overall strategy are updated, and the overall strategy is used to determine the state value.
8. The motion tracking method according to claim 1, characterized in that, The method further includes: During the training of the motion strategy, the weights and thresholds in the tracking environment are adjusted according to the training progress; when the motion tracking is reset, the physical parameters of the tracking environment are randomized.
9. A motion tracking device, characterized in that, include: The determination module is configured to determine the motion increment for performing the tracking action based on the tracking action and the motion parameters of the target robot using a trained motion strategy, wherein the motion strategy is trained based on distribution perception and failure-first resampling. The control module is configured to control the target robot to perform motion tracking based on the motion increment and the tracking action.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the motion tracking method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Quadruped robot motion control method and system based on reinforcement learning action simulation
CN118012077A
Industrial robot digital twin model self-updating method based on reinforcement learning algorithm
CN119670841A
Whole-body control method and control device of robot, storage medium and robot
CN120395889A
Action prediction networks for robotic grasping
US20200086483A1
Machine learning system for marker-based motion capture
WO2025056928A2