Quadruped machine gait generation and control method, device, equipment and medium

By constructing gait animation data of virtual horses and reinforcement learning in a physical simulation environment, the problem of low efficiency in generating gait of quadrupedal robotic horses was solved, and efficient control and flexible adjustment of the movement pattern of quadrupedal robotic horses were achieved.

CN121477655BActive Publication Date: 2026-05-12HANGZHOU YUNSHENCHU TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HANGZHOU YUNSHENCHU TECH CO LTD
Filing Date
2026-01-09
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

In the existing technology, the gait generation efficiency of quadrupedal robotic horses is low, and it is difficult to effectively control the movement pattern of robotic horses.

Method used

By constructing gait animation data of a virtual horse, establishing a mapping relationship between the animated skeletal joints and the robotic horse joints, redirecting it to the kinematic model, constructing a target dataset for adversarial imitation learning, expanding the parameterized trajectory, and performing reinforcement learning in a physical simulation environment, a unified gait generation strategy is obtained, achieving efficient control of the quadrupedal robotic horse.

Benefits of technology

It enables continuous adjustment of the gait type, stride frequency, speed and behavior pattern of the quadrupedal robotic horse, improving the control efficiency and stability of the movement pattern, and is able to reproduce flexible and smooth movement in real environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121477655B_ABST
    Figure CN121477655B_ABST
Patent Text Reader

Abstract

The present disclosure provides a quadruped robot horse gait generation and control method, device, equipment and medium, comprising: constructing gait animation data of a virtual horse; according to a mapping relationship, the gait animation data of the virtual horse is redirected to a kinematics model of a quadruped robot horse to obtain reference trajectory data; according to the reference trajectory data, a target data set is constructed; the target data set is parameterized trajectory augmented through a predefined continuous adjustable parameter; through adversarial imitation learning, the robot horse control strategy is pre-trained according to the target data set; through a predefined high-level parameterized instruction and a task reward mechanism, the robot horse control strategy is reinforced learning joint training in a physical simulation environment to obtain a unified gait generation strategy; the quadruped robot horse is controlled to perform corresponding actions through the unified gait generation strategy. Thus, efficient control of the quadruped robot horse form in a real environment is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments disclosed herein relate to the field of artificial intelligence technology, and more specifically, to a method, apparatus, device, and medium suitable for generating and controlling the gait of a quadrupedal robotic horse. Background Technology

[0002] The quadrupedal robotic horse is a biomimetic robot with four leg actuators. Its body structure includes a torso, head, limbs and tail. It can use high-torque servo motors or hydraulic / pneumatic actuators, combined with multi-degree-of-freedom joint design, to simulate the biological movement characteristics of equines.

[0003] The mechanical gait of the quadrupedal robotic horse can highly biomimeticly simulate various typical movement patterns of real horses, such as walking, trotting, slow walking, and sprinting. By collecting and analyzing equine kinematic data, a precise gait parameter model can be constructed to achieve dynamic adjustment of key movement parameters such as stride length, stride frequency, leg lift height, and landing angle. However, current technologies suffer from low gait generation efficiency, making it difficult to effectively control the robotic horse's movement pattern. Summary of the Invention

[0004] The embodiments described herein provide a method, apparatus, device, and medium for generating and controlling the gait of a quadrupedal robotic horse, overcoming the aforementioned problems.

[0005] Firstly, according to the content of this disclosure, a method for generating and controlling the gait of a quadrupedal robotic horse is provided, including:

[0006] Construct gait animation data for a virtual horse; the gait animation data for the virtual horse includes: keyframe animations corresponding to a preset specific gait and keyframe animations corresponding to a preset specific behavior;

[0007] Align the skeletal model of the virtual horse with the joint topology of the quadrupedal robotic horse to establish a mapping relationship between the animation skeletal joints and the robotic horse joints.

[0008] Based on the mapping relationship between the animated skeletal joints and the robotic horse joints, the gait animation data of the virtual horse is redirected to the kinematic model of the quadrupedal robotic horse to obtain the reference trajectory data of the quadrupedal robotic horse that can be executed within a preset joint range.

[0009] Construct a target dataset for adversarial imitation learning based on the reference trajectory data;

[0010] Construct predefined, continuously adjustable parameters;

[0011] The target dataset is augmented with parameterized trajectories using predefined, continuously adjustable parameters.

[0012] The control strategy for the machine horse is pre-trained based on the target dataset through adversarial imitation learning.

[0013] By using predefined high-level parameterized instructions and task reward mechanisms, the control strategy of the robotic horse is jointly trained through reinforcement learning in a physical simulation environment to obtain a unified gait generation strategy for continuously adjusting gait type, gait frequency, speed, and behavior pattern.

[0014] The quadrupedal robotic horse is controlled to perform corresponding actions by the unified gait generation strategy, which continuously adjusts gait type, gait frequency, speed, and behavior pattern.

[0015] Secondly, according to the present disclosure, a quadrupedal robotic horse gait generation and control device is provided, comprising:

[0016] The first construction module is used to construct the gait animation data of the virtual horse; the gait animation data of the virtual horse includes: keyframe animations corresponding to a preset specific gait of the virtual horse and keyframe animations corresponding to a preset specific behavior;

[0017] A module is established to align the skeletal model of the virtual horse with the joint topology of the quadrupedal robotic horse, and to establish a mapping relationship between the animation skeletal joints and the robotic horse joints.

[0018] The redirection module is used to redirect the gait animation data of the virtual horse to the kinematic model of the quadrupedal ...

[0019] The second construction module is used to construct a target dataset for adversarial imitation learning based on the reference trajectory data;

[0020] The third building module is used to build predefined continuously adjustable parameters;

[0021] An expansion module is used to expand the target dataset with parameterized trajectories using predefined continuously adjustable parameters;

[0022] The first training module is used to pre-train the control strategy of the machine horse based on the target dataset through adversarial imitation learning.

[0023] The second training module is used to perform reinforcement learning joint training on the control strategy of the robotic horse in a physical simulation environment through predefined high-level parameterized instructions and task reward mechanism, so as to obtain a unified gait generation strategy for continuously adjusting gait type, gait frequency, speed and behavior pattern.

[0024] The control module is used to control the quadrupedal robotic horse to perform corresponding actions through the unified gait generation strategy for continuously adjusting gait type, gait frequency, speed, and behavior pattern.

[0025] Thirdly, a computer device is provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the quadrupedal mechanical horse gait generation and control method as described in any of the above embodiments.

[0026] Fourthly, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the quadrupedal horse gait generation and control method as described in any of the above embodiments.

[0027] The quadrupedal robotic horse gait generation and control method provided in this application constructs gait animation data of a virtual horse. The virtual horse gait animation data includes: keyframe animations corresponding to a preset specific gait and keyframe animations corresponding to a preset specific behavior. The skeletal model of the virtual horse is aligned with the joint topology of the quadrupedal robotic horse to establish a mapping relationship between the animation skeletal joints and the robotic horse joints. Based on the mapping relationship between the animation skeletal joints and the robotic horse joints, the gait animation data of the virtual horse is redirected to the kinematic model of the quadrupedal robotic horse to obtain reference trajectory data executable within a preset joint range. Based on the reference trajectory data... The process involves constructing a target dataset for adversarial imitation learning; building predefined continuously adjustable parameters; expanding the target dataset with parameterized trajectories using these parameters; pre-training a robotic horse control strategy based on the target dataset using adversarial imitation learning; and then jointly training the robotic horse control strategy with reinforcement learning in a physical simulation environment using predefined high-level parameterized instructions and task reward mechanisms to obtain a unified gait generation strategy for continuously adjusting gait type, gait frequency, speed, and behavior patterns. This unified gait generation strategy is then used to control the quadrupedal robotic horse to perform corresponding actions. In this way, a controllable and editable virtual motion capture dataset is constructed using animation data of a virtual horse for pre-training of the control strategy, followed by post-training in a physical simulation environment. This results in a unified gait generation strategy capable of continuously adjusting gait type, gait frequency, speed, and behavior patterns, thereby achieving efficient control of the quadrupedal robotic horse's morphology in a real-world environment.

[0028] The above description is merely an overview of the technical solutions of the embodiments of this application. In order to better understand the technical means of the embodiments of this application and to implement them in accordance with the contents of the specification, and to make the above and other objects, features and advantages of the embodiments of this application more obvious and understandable, specific implementation methods of this application are described below. Attached Figure Description

[0029] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings of the embodiments will be briefly described below. It should be understood that the drawings described below only relate to some embodiments of this disclosure and are not intended to limit this disclosure, wherein:

[0030] Figure 1 This is a flowchart illustrating a method for generating and controlling the gait of a quadrupedal robotic horse, as disclosed in this publication.

[0031] Figure 2 This is a schematic diagram of the structure of a quadrupedal robotic horse gait generation and control device provided in this disclosure.

[0032] Figure 3 This is a schematic diagram of the structure of a computer device provided in this disclosure.

[0033] It should be noted that the elements in the attached diagram are schematic and not drawn to scale. Detailed Implementation

[0034] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the described embodiments of this disclosure without creative effort are also within the scope of protection of this disclosure.

[0035] Unless otherwise defined, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which this subject matter pertains. It will be further understood that terms such as those defined in commonly used dictionaries shall be interpreted as having the meaning consistent with their meaning in the context of the specification and in the relevant art, and shall not be interpreted in an idealized or overly formal form unless otherwise explicitly defined herein. As used herein, the statement of “connecting” or “coupling” two or more parts together shall mean that these parts are directly joined together or joined through one or more intermediate components.

[0036] The term "embodiment" as used herein means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of the phrase "embodiment" in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0037] In this document, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists, A and B exist simultaneously, or B exists. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship. Terms such as "first" and "second" are only used to distinguish one component (or part of a component) from another component (or another part of a component).

[0038] In the description of this application, unless otherwise stated, "multiple" means two or more (including two), and similarly, "multiple groups" means two or more (including two groups).

[0039] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0040] Figure 1 This is a flowchart illustrating a method for generating and controlling the gait of a quadrupedal robotic horse according to an embodiment of this disclosure, as shown below. Figure 1 As shown, the specific process of the quadrupedal robotic horse gait generation and control method includes:

[0041] S110. Construct the gait animation data of the virtual horse; align the skeletal model of the virtual horse with the joint topology of the quadrupedal robotic horse, and establish the mapping relationship between the animation skeletal joints and the robotic horse joints.

[0042] The virtual horse's skeletal and skinned models can be created using 3D animation software, and keyframe animations of various basic gaits and performance movements can be drawn manually or semi-automatically. The virtual horse's gait animation data includes keyframe animations corresponding to preset specific gaits and preset specific behaviors. Metadata such as gait type, cadence, rhythm information, and movement labels can be added to each animation data point.

[0043] Preset specific gait patterns can include, but are not limited to, typical gait patterns at different movement speeds such as walk, trot, trot, run, and backward movement. Each gait pattern contains skeletal joint posture data of the virtual horse at key time points within a complete movement cycle. Preset specific behaviors can include, but are not limited to, non-periodic or scenario-specific actions such as leg lifting, stepping, turning, and tail wagging. Keyframe animations record the skeletal joint movement states of the virtual horse at key nodes from the initial posture to the final posture when performing these behaviors.

[0044] By analyzing the degrees of freedom (such as rotation axis direction and rotation range) of each joint in the virtual horse skeletal model and the driving parameters (such as servo motor rotation angle range and motor output torque limit) of the physical joints of the quadrupedal robotic horse, a joint kinematic mapping matrix can be constructed to represent the mapping relationship between the animated skeletal joints and the robotic horse joints. Specifically, for key motion joints such as the hip, knee, and ankle joints of the virtual horse, their rotation angles and motion trajectories in the animation data are converted into control command values ​​for the corresponding joints of the quadrupedal robotic horse through coordinate transformation and scaling. For example, the rotation angle of the virtual horse's hip joint around the X-axis can be... α Based on the actual installation position and transmission ratio of the machine horse's hip joint, the target rotation angle of the servo motor is converted. β=α× The transmission coefficient plus zero-point compensation value ensures that the joint angle after conversion does not exceed the physical limit range of the quadrupedal mechanical horse.

[0045] S120. Based on the mapping relationship between the animated skeletal joints and the robot horse joints, the gait animation data of the virtual horse is redirected to the kinematic model of the four-legged robot horse to obtain the reference trajectory data of the four-legged robot horse that can be executed within the preset joint range.

[0046] Redirection is used to convert and map the skeletal / joint motion data of the virtual horse onto the kinematic model of the quadrupedal robotic horse itself, so that even if the two sides have different numbers of joints, joint axes, and limb lengths, they can still obtain the joint trajectories or foot trajectories that the quadrupedal robotic horse can execute.

[0047] In some embodiments, based on the mapping relationship between the animated skeletal joints and the robot horse joints, the gait animation data of the virtual horse is redirected to the kinematic model of the quadrupedal robot horse to obtain reference trajectory data executable within a preset joint range for the quadrupedal robot horse. This includes: obtaining gait-related data from the gait animation data of the virtual horse based on the mapping relationship between the animated skeletal joints and the robot horse joints; redirecting the gait-related data to the kinematic model of the quadrupedal robot horse to generate initial trajectory data executable within a preset joint range for the quadrupedal robot horse; and integrating the initial trajectory data according to time series requirements to obtain reference trajectory data executable within a preset joint range for the quadrupedal robot horse.

[0048] This includes gait-related data such as joint angles, root node pose, and foot trajectory. After generating initial trajectory data for the quadrupedal robot horse within a preset joint range, time-series requirements can be set according to the reinforcement learning environment. This requires a sequence of joint positions, joint velocities, trunk position and orientation, foot contact markers, and gait phases at each moment, integrating the initial trajectory data. This facilitates effective training and optimization of the quadrupedal robot horse's gait control strategy in a simulation environment, enabling the quadrupedal robot horse to more accurately reproduce the reference trajectory and possess better dynamic stability and adaptability during actual movement.

[0049] S130. Construct a target dataset for adversarial imitation learning based on the reference trajectory data.

[0050] The target dataset contains reference trajectory data samples of a quadrupedal robotic horse in different gaits (such as walking, trotting, jumping, etc.). Each sample covers key kinematic information such as joint position, joint velocity, trunk posture (position and orientation), foot trajectory, and foot-ground contact state in a time series.

[0051] In some embodiments, constructing a target dataset for adversarial imitation learning based on reference trajectory data includes: performing morphological transformation on the reference trajectory data to obtain an initial dataset for adversarial imitation learning; extracting gait style-related features from the initial dataset and performing physical simulation of the gait style-related features using an environmental physics engine to obtain the target dataset for adversarial imitation learning.

[0052] This involves constructing an initial dataset suitable for adversarial imitation learning by converting reference trajectory data into state-action or state-next-state formats; and extracting gait style-related features such as gait phase, stride frequency, contact sequence, stride length, and joint posture patterns. Stride frequency is the number of gait cycles a quadrupedal robotic horse completes within a given time, i.e., the number of "steps" or gait cycles completed per unit time, reflecting the speed of movement; a higher stride frequency indicates more frequent and faster steps. Stride length is the displacement distance of the quadrupedal robotic horse's body or foot in the forward direction between two consecutive landings (on the same foot), describing the size of each step. Gait phase is the relative progress within a complete gait cycle, describing the current position within the gait cycle, such as lifting the foot, swinging, landing, or supporting the weight. Combined with an environmental physics engine (gravity, contact, friction, etc.), the feasibility of the animated trajectory in physical simulation is ensured, providing realistic physical feedback for subsequent reinforcement learning.

[0053] S140. Construct predefined continuously adjustable parameters; perform parameterized trajectory expansion on the target dataset using the predefined continuously adjustable parameters.

[0054] The predefined continuously adjustable parameters include: time scaling factor, gait frequency factor, stride scaling factor, and movement style factor. The time scaling factor controls the overall movement rate; the gait frequency factor controls the single-step cycle and swing-support ratio; the stride scaling factor controls the displacement of each forward or backward step; and the movement style factor controls the movement style, such as more "exaggerated" or more "gentle".

[0055] In some embodiments, parameterized trajectory augmentation of the target dataset is performed using predefined continuously adjustable parameters, including: elastically deforming the animation timeline of the target dataset using time scaling factors, gait frequency factors, and stride scaling factors, respectively, to generate reference trajectory variants with different rates, different cadence frequencies, and different amplitudes; adjusting the animation style of the target dataset using motion style factors to generate reference trajectory variants with different styles; and augmenting the reference trajectory variants with different rates, different cadence frequencies, different amplitudes, and different styles into the target dataset.

[0056] Specifically, when the animation timeline is elastically deformed using a time scaling factor, the original animation's time sequence can be stretched or compressed as a whole. For example, when the time scaling factor is greater than 1, the animation playback speed is increased, and the overall motion time is shortened; when the time scaling factor is less than 1, the animation playback speed is slowed down, and the overall motion time is prolonged, thereby generating reference trajectory variants with different speeds.

[0057] The gait frequency coefficient can be adjusted by modifying the duration of the single-step cycle and the proportion of the swing phase and support phase within the single-step cycle to deform the animation timeline. Increasing the gait frequency coefficient shortens the single-step cycle and may change the swing-support ratio, such as decreasing the proportion of the support phase and increasing the proportion of the swing phase, thus increasing the number of steps per unit time and forming a high-frequency reference trajectory variant; conversely, decreasing the gait frequency coefficient generates a low-frequency reference trajectory variant.

[0058] The stride scaling factor can linearly or non-linearly scale the displacement of each step while keeping the gait cycle and motion pattern basically unchanged. When the stride scaling factor increases, the displacement distance of each step when the quadruped robot moves forward or backward increases, and the stride becomes larger; when the stride scaling factor decreases, the displacement distance of each step decreases, and the stride becomes smaller, thus obtaining reference trajectory variants with different strides.

[0059] For the motion style coefficient, different styles of reference trajectory variants can be generated by modifying details such as joint angles, limb movement amplitude, and speed change curves in the animation trajectory. For example, when the motion style coefficient is set to "exaggerated," the contrast between the limb movement amplitude and speed change will be increased, making the action appear more powerful and dramatic; when set to "soft," the movement amplitude will be reduced and speed changes will be smoothed, giving the action a more relaxed and coherent style, thus generating different styles of reference trajectory variants.

[0060] After generating reference trajectory variants with different rates, different frequencies, different amplitudes, and different styles, physical constraints can be applied to the corresponding trajectories to ensure that the transformed trajectories still meet the constraints (such as not excessively exceeding joint angle and velocity limits), and the corresponding gait phase, contact markers, and other information can be updated.

[0061] Therefore, by multiplying the dataset through programmed parameter transformation, the control strategy's ability to continuously adjust step frequency, speed, and stride can be significantly enhanced, while avoiding repeated motion capture and step-by-step training.

[0062] S150. Through adversarial imitation learning, the control strategy of the robotic horse is pre-trained based on the target dataset. Through predefined high-level parameterized instructions and task reward mechanism, the control strategy of the robotic horse is jointly trained by reinforcement learning in a physical simulation environment to obtain a unified gait generation strategy for continuously adjusting gait type, gait frequency, speed and behavior pattern.

[0063] In adversarial imitation learning, a "policy network" generates the robot's motion trajectory in physical simulation; a "discriminator network" distinguishes whether the trajectory comes from expert data or policy generation; adversarial training enables the policy-generated trajectory to closely resemble the expert trajectory in style and distribution. The "policy network" is a function or neural network that outputs control commands (such as joint positions, joint velocities, and joint torques) based on current observations (e.g., robot state, gait phase, high-level commands), used to achieve the mapping from "state-command" to "action." Reinforcement learning is used to further optimize the control strategy of the quadrupedal robot in physical simulation, enabling it to achieve task objectives such as speed tracking and stability while satisfying gait style requirements.

[0064] During pre-training, the observation space of the robotic horse control strategy includes: the current state of the robotic horse (joint angles, joint velocities, trunk linear and angular velocities, foot position, etc.); and gait-related information (current gait type label, current gait phase, expected stride frequency, expected stride length, etc.). The discriminator is input to expert trajectory features from the animation (i.e., the expected trajectory) and trajectory features of the robot after executing the strategy in the physical simulation (i.e., the reference trajectory in the target dataset). Through adversarial training, the actions generated by the strategy are made to stylistically approximate the animation demonstration. The pre-trained robotic horse control strategy internally encodes the basic motion styles and rhythmic features of multiple gaits. Thus, by learning multiple gaits and rhythms simultaneously in a single network, the engineering complexity is greatly reduced.

[0065] In some embodiments, the method further includes: abstracting user requirements into high-level parameterized instructions; the high-level parameterized instructions include: desired forward velocity, desired lateral velocity, desired angular velocity, desired gait type, target cadence, target stride length, and target behavior pattern.

[0066] Among them, gait types include walking, fast walking, trotting, running, and walking backward; target behavior patterns include stationary stepping performances, leg-raising shows, and spinning dance steps. By feeding high-level instructions as conditional inputs and the robot's state into the policy network, the same policy can generate movements with different styles and rhythms under different instructions.

[0067] In some embodiments, the task reward mechanism is used to describe tracking rewards, task rewards, gait phase and gait frequency consistency rewards, and stability and safety rewards; the tracking reward is used to constrain the quadrupedal robot horse to follow the reference gait generated by the integration of animation data and high-level instructions in the physical simulation; the task reward is used to constrain the quadrupedal robot horse to successfully complete the specified displacement task under the premise of satisfying the preset behavioral style; the gait phase and gait frequency consistency reward is used to align the timing of the quadrupedal robot horse's foot contact / swing pattern with the given gait frequency and phase progress; the stability and safety reward is used to constrain the quadrupedal robot horse's limb stability when performing the corresponding action.

[0068] In the physical simulation, the quadrupedal robotic horse follows a reference gait generated from animation data and high-level commands, including foot landing sequence, average stride length, and rhythm. Tracking rewards ensure the movement style and rhythm meet expectations. Task rewards are designed to match the desired speed, direction of travel, and turning angular velocity from high-level commands, allowing the robotic horse to complete specified displacement tasks while maintaining its performance style. Gait phase and gait frequency consistency rewards encourage foot contact / swing patterns to align temporally with a given gait frequency and phase progression, enabling the strategy to walk stably or stand still at a specified gait frequency, effectively solving the problem of continuous gait frequency control. Stability and safety rewards include body posture stability, non-slipping feet, joint non-over-limitation, and energy consumption penalties, ensuring the generated movements are executable on a real robot and less prone to instability.

[0069] In some embodiments, a unified gait generation strategy for continuously adjusting gait type, gait frequency, speed, and behavior pattern is obtained by jointly training the robotic horse control strategy using reinforcement learning in a physical simulation environment through predefined high-level parameterized instructions and task reward mechanisms. This includes: randomly sampling different high-level parameterized instructions and executing the robotic horse control strategy in a physical simulation environment; summarizing the reinforcement learning rewards obtained according to the task reward mechanism; and optimizing the robotic horse control strategy based on the reinforcement learning rewards using a reinforcement learning algorithm to obtain a unified gait generation strategy for continuously adjusting gait type, gait frequency, speed, and behavior pattern.

[0070] In this process, different high-level instructions (different gaits, cadences, speeds, and performance modes) are randomly sampled in a physical simulation environment. The control strategy of the robotic horse is executed in the simulation, and the reward for reinforcement learning is obtained by summarizing the unified reward. The control strategy of the robotic horse is updated using a reinforcement learning algorithm, while retaining an appropriate amount of style constraints from adversarial imitation learning to prevent the loss of the original aesthetic beauty of the horse's gait during training.

[0071] Thus, through reinforcement learning with high-level instructions and uniform rewards, it can not only imitate existing animations, but also flexibly extrapolate and combine actions in a continuous instruction space, such as smoothly transitioning from a slow walk to a high-frequency stepping in place, and then connecting a backward jogging performance action; without having to recollect data and train new strategies for each combination.

[0072] S160, Control the quadrupedal robotic horse to perform corresponding actions by using a unified gait generation strategy for continuously adjusting gait type, gait frequency, speed, and behavior pattern.

[0073] Specifically, a unified gait generation strategy, used for continuous adjustment of gait type, gait frequency, speed, and behavior pattern, can control the corresponding movement patterns of the quadrupedal robotic horse in a real environment. The unified gait generation strategy receives continuous high-level command parameters, such as setting the gait type to "walk," the gait frequency to 0.8Hz, and the forward speed to 0.3m / s, or setting the behavior pattern to "performance mode - spinning in place" and specifying a rotational angular velocity of 1rad / s. Through the parsing and mapping of these continuous parameters, the strategy drives the various joints of the quadrupedal robotic horse to move collaboratively along an optimized trajectory, achieving precise conversion from commands to specific actions. During the driving process, the unified gait generation strategy can automatically adapt to the dynamic balance requirements under different command combinations, ensuring that the quadrupedal robotic horse maintains a stable trunk posture during the speed increase from trot to trot; when performing complex behavioral patterns such as backward trot, the stepping sequence and landing timing of the limbs conform to the kinematic laws, avoiding phenomena such as slipping or tipping over, thereby reproducing the flexible, smooth and aesthetically pleasing movement ability obtained in simulation training in the real physical world.

[0074] In this embodiment, gait animation data of a virtual horse is constructed. This data includes keyframe animations corresponding to preset specific gaits and preset specific behaviors. The skeletal model of the virtual horse is aligned with the joint topology of a quadrupedal robotic horse to establish a mapping relationship between the animated skeletal joints and the robotic horse joints. Based on this mapping relationship, the gait animation data of the virtual horse is redirected to the kinematic model of the quadrupedal robotic horse to obtain reference trajectory data executable within preset joint ranges. Based on this reference trajectory data, adversarial mimicry is constructed. The learning process involves: constructing a target dataset; building predefined, continuously adjustable parameters; expanding the target dataset with parameterized trajectories using these parameters; pre-training the robotic horse control strategy using adversarial imitation learning based on the target dataset; jointly training the robotic horse control strategy in a physical simulation environment using reinforcement learning through predefined high-level parameterized instructions and task reward mechanisms, resulting in a unified gait generation strategy for continuously adjusting gait type, gait frequency, speed, and behavior patterns; and controlling the quadrupedal robotic horse to perform corresponding actions using this unified gait generation strategy. In this way, a controllable and editable virtual motion capture dataset is constructed using the animation data of a virtual horse for pre-training of the control strategy, and then post-training the control strategy in a physical simulation environment, resulting in a unified gait generation strategy capable of continuously adjusting gait type, gait frequency, speed, and behavior patterns, thereby achieving efficient control of the quadrupedal robotic horse's morphology in a real-world environment.

[0075] In summary, this embodiment uses 3D animation software to draw and adjust various gaits and performance movements of horses, constructing a controllable and editable virtual motion capture dataset; it unifies the animation data into training data suitable for imitation learning and reinforcement learning, learning the movement styles and dynamic characteristics of various horse gaits in a unified policy network; it automatically expands the dataset by parametrically transforming the animation data's speed, stride frequency, stride length, etc., enabling the policy to generate new variant movements in a continuous parameter space; it introduces high-level parametric instructions and a reward design combined with gait phase and contact state to achieve real-time adjustment and smooth switching of the quadrupedal robotic horse's gait, stride frequency, and rhythm, allowing it to naturally complete multi-gait and multi-movement combinations and transitions in stage performance scenarios.

[0076] Figure 2 This is a schematic diagram of the structure of a quadrupedal robotic horse gait generation and control device provided in this embodiment. The quadrupedal robotic horse gait generation and control device may include:

[0077] The first construction module 210 is used to construct the gait animation data of the virtual horse; the gait animation data of the virtual horse includes: keyframe animations corresponding to the preset specific gait of the virtual horse and keyframe animations corresponding to the preset specific behavior.

[0078] Module 220 is used to align the skeletal model of the virtual horse with the joint topology of the quadrupedal robotic horse, and to establish a mapping relationship between the animated skeletal joints and the robotic horse joints.

[0079] The redirection module 230 is used to redirect the gait animation data of the virtual horse to the kinematic model of the quadrupedal ...

[0080] The second building module 240 is used to build a target dataset for adversarial imitation learning based on reference trajectory data.

[0081] The third building module 250 is used to build predefined continuously adjustable parameters.

[0082] The extension module 260 is used to extend the target dataset with parameterized trajectories using predefined continuously adjustable parameters.

[0083] The first training module 270 is used to pre-train the control strategy of the machine horse based on the target dataset through adversarial imitation learning.

[0084] The second training module 280 is used to perform reinforcement learning joint training on the control strategy of the robotic horse in a physical simulation environment through predefined high-level parameterized instructions and task reward mechanism, so as to obtain a unified gait generation strategy for continuously adjusting gait type, gait frequency, speed and behavior pattern.

[0085] The control module 290 is used to control the quadrupedal robotic horse to perform corresponding actions through a unified gait generation strategy for continuously adjusting gait type, gait frequency, speed, and behavior pattern.

[0086] In this embodiment, optionally, the redirection module 230 is specifically used for:

[0087] Based on the mapping relationship between the skeletal joints of the animation and the joints of the robotic horse, gait-related data is obtained from the gait animation data of the virtual horse; the gait-related data is redirected to the kinematic model of the quadrupedal robotic horse to generate initial trajectory data that can be executed within the preset joint range of the quadrupedal robotic horse; the initial trajectory data is integrated according to the time series requirements to obtain reference trajectory data that can be executed within the preset joint range of the quadrupedal robotic horse.

[0088] In this embodiment, optionally, the second building module 240 is specifically used for:

[0089] The reference trajectory data is morphologically transformed to obtain the initial dataset for adversarial imitation learning; gait style-related features are extracted from the initial dataset, and the gait style-related features are physically simulated using an environmental physics engine to obtain the target dataset for adversarial imitation learning.

[0090] In this embodiment, optionally, the expansion module 260 is specifically used for:

[0091] The animation timeline of the target dataset is elastically deformed using time scaling factor, gait frequency factor, and stride scaling factor, respectively, to generate reference trajectory variants with different rates, different gait frequencies, and different gait amplitudes. The animation style of the target dataset is adjusted using motion style factor to generate reference trajectory variants with different styles. The reference trajectory variants with different rates, different gait frequencies, different gait amplitudes, and different styles are then expanded into the target dataset.

[0092] In this embodiment, optionally, the second training module 280 is specifically used for:

[0093] Different high-level parameterized instructions are randomly sampled and the control strategy of the robotic horse is executed in a physical simulation environment. The reinforcement learning rewards obtained are summarized according to the task reward mechanism. The control strategy of the robotic horse is optimized by reinforcement learning algorithm based on the reinforcement learning rewards to obtain a unified gait generation strategy for continuously adjusting gait type, gait frequency, speed and behavior pattern.

[0094] In this embodiment, optionally, an abstract module may also be included.

[0095] The abstract module is used to abstract user requirements into high-level parameterized instructions. The high-level parameterized instructions include: desired forward velocity, desired lateral velocity, desired angular velocity, desired gait type, target cadence, target stride length, and target behavior pattern.

[0096] In this embodiment, optionally, the task reward mechanism is used to describe tracking rewards, task rewards, gait phase and gait frequency consistency rewards, and stability and safety rewards; the tracking reward is used to constrain the quadrupedal robotic horse to follow the reference gait generated by the integration of animation data and high-level instructions in the physical simulation; the task reward is used to constrain the quadrupedal robotic horse to successfully complete the specified displacement task under the premise of satisfying the preset behavioral style; the gait phase and gait frequency consistency reward is used to align the timing of the quadrupedal robotic horse's foot contact / swing pattern with the given gait frequency and phase progress; the stability and safety reward is used to constrain the quadrupedal robotic horse's limb stability when performing the corresponding action.

[0097] The quadrupedal robotic horse gait generation and control device provided in this disclosure can execute the above-described method embodiments. Its specific implementation principle and technical effects can be found in the above-described method embodiments, and will not be repeated here.

[0098] This application also provides a computer device. Please refer to the following for details. Figure 3 , Figure 3 This is a basic structural block diagram of the computer device in this embodiment.

[0099] The computer device includes a memory 310 and a processor 320 that are interconnected via a system bus. It should be noted that only a computer device with memory 310 and processor 320 is shown in the figure; however, it should be understood that it is not required to implement all the components shown, and more or fewer components may be implemented alternatively. Those skilled in the art will understand that the computer device described herein is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0100] Computer devices can include desktop computers, laptops, handheld computers, and cloud servers. These devices allow for human-computer interaction with users through keyboards, mice, remote controls, touchpads, or voice-activated devices.

[0101] The memory 310 includes at least one type of readable storage medium, including non-volatile memory or volatile memory, such as flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. RAM may include static RAM or dynamic RAM. In some embodiments, the memory 310 may be an internal storage unit of a computer device, such as the hard disk or memory of the computer device. In other embodiments, the memory 310 may also be an external storage device of the computer device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, or flash card equipped on the computer device. Of course, the memory 310 may include both internal storage units and external storage devices of the computer device. In this embodiment, the memory 310 is typically used to store the operating system and various application software installed on the computer device, such as the program code of the method described above. In addition, the memory 310 can also be used to temporarily store various types of data that have been output or will be output.

[0102] Processor 320 is typically used to perform overall operations of a computer device. In this embodiment, memory 310 is used to store program code or instructions, including computer operation instructions, and processor 320 is used to execute the program code or instructions stored in memory 310 or process data, such as program code that runs the methods described above.

[0103] In this article, the bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. This bus system can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.

[0104] Another embodiment of this application also provides a computer-readable medium, which may be a computer-readable signal medium or a computer-readable medium. A processor in a computer reads computer-readable program code stored in the computer-readable medium, enabling the processor to execute the functional actions specified in each step or combination of steps in the above method; and to generate means for implementing the functional actions specified in each block or combination of blocks in the block diagram.

[0105] Computer-readable media include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared memory or semiconductor systems, devices or apparatuses, or any suitable combination thereof, wherein the memory is used to store program code or instructions, the program code including computer operation instructions, and the processor is used to execute the program code or instructions of the above-described methods stored in the memory.

[0106] The definitions of memory and processor can be found in the description of the foregoing computer device embodiments, and will not be repeated here.

[0107] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0108] In the various embodiments of this application, the functional units or modules can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0109] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0110] In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" as described in this application does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. This application can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In the unit claims listing several means, several units of these means may be embodied by the same item of hardware. The use of "first," "second," and "third," etc., does not indicate any order and these words should be interpreted as names. Unless otherwise specified, the steps in the above embodiments should not be construed as limiting the order of execution.

[0111] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A method for generating and controlling the gait of a quadrupedal robotic horse, characterized in that, include: Construct gait animation data for a virtual horse; The gait animation data of the virtual horse includes: keyframe animations corresponding to a preset specific gait of the virtual horse and keyframe animations corresponding to a preset specific behavior; Align the skeletal model of the virtual horse with the joint topology of the quadrupedal robotic horse to establish a mapping relationship between the animation skeletal joints and the robotic horse joints. Based on the mapping relationship between the animated skeletal joints and the robotic horse joints, the gait animation data of the virtual horse is redirected to the kinematic model of the quadrupedal robotic horse to obtain the reference trajectory data of the quadrupedal robotic horse that can be executed within a preset joint range. Construct a target dataset for adversarial imitation learning based on the reference trajectory data; Construct predefined, continuously adjustable parameters; Parametric trajectory expansion of the target dataset is performed using predefined continuously adjustable parameters. This includes: elastically deforming the animation timeline of the target dataset using time scaling factors, gait frequency factors, and stride scaling factors to generate reference trajectory variants with different rates, different cadences, and different amplitudes; adjusting the animation style of the target dataset using motion style factors to generate reference trajectory variants with different styles; and expanding the reference trajectory variants with different rates, different cadences, different amplitudes, and different styles into the target dataset. The control strategy for the machine horse is pre-trained based on the target dataset through adversarial imitation learning. By using predefined high-level parameterized instructions and task reward mechanisms, the control strategy of the robotic horse is jointly trained through reinforcement learning in a physical simulation environment to obtain a unified gait generation strategy for continuously adjusting gait type, gait frequency, speed, and behavior pattern. The quadrupedal robotic horse is controlled to perform corresponding actions by the unified gait generation strategy, which continuously adjusts gait type, gait frequency, speed, and behavior pattern.

2. The method according to claim 1, characterized in that, Based on the mapping relationship between the animated skeletal joints and the robotic horse joints, the gait animation data of the virtual horse is redirected to the kinematic model of the quadrupedal robotic horse to obtain reference trajectory data of the quadrupedal robotic horse that can be executed within a preset joint range, including: Based on the mapping relationship between the animated skeletal joints and the robotic horse joints, gait-related data are obtained from the gait animation data of the virtual horse; The gait-related data is redirected to the kinematic model of the quadrupedal robotic horse to generate initial trajectory data for the quadrupedal robotic horse that can be executed within a preset joint range; The initial trajectory data is integrated according to the time series requirements to obtain the reference trajectory data of the quadrupedal robotic horse that can be executed within a preset joint range.

3. The method according to claim 1, characterized in that, Based on the reference trajectory data, a target dataset for adversarial imitation learning is constructed, including: The reference trajectory data is morphologically transformed to obtain an initial dataset for adversarial imitation learning; Gait style-related features are extracted from the initial dataset, and the gait style-related features are physically simulated using an environmental physics engine to obtain a target dataset for adversarial imitation learning.

4. The method according to claim 1, characterized in that, By using predefined high-level parameterized instructions and task reward mechanisms, the robotic horse control strategy is jointly trained through reinforcement learning in a physical simulation environment to obtain a unified gait generation strategy for continuously adjusting gait type, gait frequency, speed, and behavior patterns, including: Different high-level parameterized instructions are randomly sampled, and the machine horse control strategy is executed in a physical simulation environment; the reinforcement learning rewards obtained are then summarized according to the task reward mechanism. By using a reinforcement learning algorithm, the control strategy of the robotic horse is optimized based on the reinforcement learning reward, resulting in a unified gait generation strategy for continuously adjusting gait type, gait frequency, speed, and behavior pattern.

5. The method according to claim 1, characterized in that, Also includes: User requirements are abstracted into the high-level parameterized instructions; the high-level parameterized instructions include: desired forward velocity, desired lateral velocity, desired turning angular velocity, desired gait type, target cadence, target stride length, and target behavior pattern.

6. The method according to claim 4, characterized in that, The task reward mechanism describes tracking rewards, task rewards, gait phase and gait frequency consistency rewards, and stability and safety rewards. The tracking reward is used to constrain the quadrupedal robotic horse to follow a reference gait generated by combining animation data and high-level instructions in physical simulation. The task reward is used to constrain the quadrupedal robotic horse to successfully complete a specified displacement task while meeting a preset behavioral style. The gait phase and gait frequency consistency reward is used to align the timing of the quadrupedal robotic horse's foot contact / swing pattern with a given gait frequency and phase progress. The stability and safety reward is used to constrain the quadrupedal robotic horse's limb stability when performing corresponding actions.

7. A device for generating and controlling the gait of a quadrupedal robotic horse, characterized in that, include: The first building module is used to construct the gait animation data of the virtual horse; The gait animation data of the virtual horse includes: keyframe animations corresponding to a preset specific gait of the virtual horse and keyframe animations corresponding to a preset specific behavior; A module is established to align the skeletal model of the virtual horse with the joint topology of the quadrupedal robotic horse, and to establish a mapping relationship between the animation skeletal joints and the robotic horse joints. The redirection module is used to redirect the gait animation data of the virtual horse to the kinematic model of the quadrupedal ... The second construction module is used to construct a target dataset for adversarial imitation learning based on the reference trajectory data; The third building module is used to build predefined continuously adjustable parameters; An expansion module is used to parametrically expand the target dataset using predefined continuously adjustable parameters. Specifically, the expansion module is used to: elastically deform the animation timeline of the target dataset using time scaling factors, gait frequency factors, and stride scaling factors, respectively, to generate reference trajectory variants with different rates, different cadence frequencies, and different amplitudes; adjust the animation style of the target dataset using motion style factors, generating reference trajectory variants with different styles; and expand the reference trajectory variants with different rates, different cadence frequencies, different amplitudes, and different styles into the target dataset. The first training module is used to pre-train the control strategy of the machine horse based on the target dataset through adversarial imitation learning. The second training module is used to perform reinforcement learning joint training on the control strategy of the robotic horse in a physical simulation environment through predefined high-level parameterized instructions and task reward mechanism, so as to obtain a unified gait generation strategy for continuously adjusting gait type, gait frequency, speed and behavior pattern. The control module is used to control the quadrupedal robotic horse to perform corresponding actions through the unified gait generation strategy for continuously adjusting gait type, gait frequency, speed, and behavior pattern.

8. A computer device, characterized in that, It includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the quadrupedal robotic horse gait generation and control method as described in any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the quadrupedal robotic horse gait generation and control method as described in any one of claims 1 to 6.