A frequency-based deep reinforcement learning method for the motion control of a quadruped robot
By employing frequency-based deep reinforcement learning methods, quadruped robots can achieve efficient and robust motion control in complex environments, solving the problems of slow convergence speed and poor environmental adaptability in existing technologies, and achieving stable obstacle crossing in different tasks and terrains.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- UNIV OF SCI & TECH OF CHINA
- Filing Date
- 2026-02-25
- Publication Date
- 2026-05-26
AI Technical Summary
In existing technologies, the motion control of quadruped robots in complex environments suffers from slow convergence speed and low sample efficiency. Furthermore, model-based control methods require precise physical models and parameter adjustments, while model-free reinforcement learning methods are prone to instability in complex environments.
A frequency-based deep reinforcement learning method is adopted. The environmental and ontology perception information is projected from the time domain to the frequency domain through the spectral convolution module to extract frequency features. Then, a model-free trajectory tracker is trained using a constrained synchronous teacher-student reinforcement learning framework to achieve motion control of the quadruped robot.
It improves the motion control efficiency and robustness of quadruped robots in complex terrain, enabling them to dynamically adapt to the environment and reduce noise interference, and achieve diverse and robust obstacle-crossing skills.
Smart Images

Figure CN121721972B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of quadruped robot motion control technology, and in particular to a frequency-based deep reinforcement learning-based quadruped robot motion control method. Background Technology
[0002] Quadruped robots hold great potential in real-world scenarios requiring high mobility, such as disaster relief, factory inspection, and cargo handling. Despite significant progress in the motion control of legged robots, complex terrain, dynamic environments, and the gap between simulation and reality still pose serious challenges to their practical application.
[0003] Model-based control methods, utilizing known kinematic and dynamic models, are widely used for efficient motion planning. However, these methods typically require accurate physical models and fine-tuning of parameters. In contrast, model-free reinforcement learning methods learn the optimal policy through continuous trial and error, exhibiting more robust performance in real-world scenarios. However, model-free reinforcement learning methods require extensive trial and error during the learning process, and are prone to slow convergence and low sample efficiency in complex environments.
[0004] Frequency-based control methods can adapt to the environment and keep the system stable under external disturbances. Traditional methods improve modeling accuracy through frequency domain analysis and system identification, and are widely used in the fields of flexible robots, robotic arms, and mobile robots. Frequency-based methods enable controllers to be more efficient and less susceptible to noise interference by providing a high-level representation of system dynamics; however, their high-level representation is difficult to directly map to specific low-level motion commands for robot control.
[0005] In privileged learning, the teacher's strategy has full access to all simulation information and is trained through reinforcement learning, while the student's strategy is only supervised through ontology perception. This approach can improve the learning efficiency and real-world performance of the student's strategy. However, privileged learning depends on the quality of the teacher's strategy; if the teacher's strategy is biased in the simulation, it will affect the overall performance of the student's strategy.
[0006] In view of this, the present invention is hereby proposed. Summary of the Invention
[0007] The purpose of this invention is to provide a frequency-based deep reinforcement learning-based motion control method for quadruped robots, enabling quadruped robots to learn diverse and robust gait control skills, making quadruped robot motion control more efficient and free from noise interference, thereby solving the aforementioned technical problems existing in the prior art.
[0008] The objective of this invention is achieved through the following technical solution:
[0009] A frequency-based deep reinforcement learning-based motion control method for quadruped robots includes:
[0010] Step 1: Use a model-based reference trajectory generator to generate kinematically and dynamically continuous reference trajectories online according to different task objectives, and guide the trajectory tracking strategy of the model-free trajectory tracker to explore.
[0011] Step 2: Using the spectral convolution module, the environmental perception information and the body perception information are projected from the time domain to the frequency domain through discrete Fourier transform to obtain environmental frequency domain perception information and body frequency domain perception information, respectively. Environmental frequency features and body frequency features are extracted from the environmental frequency domain perception information and body frequency domain perception information, respectively. The motion features of the model-based trajectory planner are adjusted by the environmental frequency features and the control bandwidth of the model-free trajectory tracker is improved by the body frequency features.
[0012] Step 3: Train the model-free trajectory tracker using a constrained synchronous teacher-student reinforcement learning framework. The training process involves applying various physical constraints to enable the model-free trajectory tracker to explore within a safe domain, and synchronous teacher-student training is performed.
[0013] Step 4: Use a trained model-free trajectory tracker to control the motion of the quadruped robot.
[0014] Compared with existing technologies, the frequency-based deep reinforcement learning-based quadruped robot motion control method provided by this invention has the following advantages:
[0015] The proposed frequency-based deep reinforcement learning method for quadruped robot motion control captures frequency information about the environment and the robot's state from multimodal frequency-domain perception information using a spectral convolution module. Based on Fourier transform, environmental and body frequency features are extracted from the multimodal perception information, enabling the control output to dynamically adapt to the surrounding environment and improving the bandwidth of the network controller. Results show that the proposed method is applicable to different quadruped robots and different tasks, enabling the quadruped robot to perform diverse and robust obstacle-crossing skills on complex terrain. Attached Figure Description
[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 A flowchart of a frequency-based deep reinforcement learning-based quadruped robot motion control method provided in an embodiment of the present invention.
[0018] Figure 2 This is a schematic diagram of the architecture of a frequency-based deep reinforcement learning-based quadruped robot motion control method provided in an embodiment of the present invention. Detailed Implementation
[0019] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the specific content of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments, which do not constitute a limitation of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention.
[0020] First, the following explanations are provided for the terms that may be used in this article:
[0021] The term "and / or" means that either or both can be achieved simultaneously. For example, X and / or Y means that it includes both "X" or "Y" as well as the three cases of "X and Y".
[0022] The terms "comprising," "including," "containing," "having," or other similar semantic descriptions should be interpreted as non-exclusive inclusion. For example, including a technical feature element (such as raw material, component, ingredient, carrier, dosage form, material, size, part, component, mechanism, device, step, process, method, reaction conditions, processing conditions, parameter, algorithm, signal, data, product or article of manufacture, etc.) should be interpreted as including not only the expressly listed technical feature element, but also other technical feature elements that are not expressly listed and are well-known in the art.
[0023] The term "composed of" excludes any technical features not expressly listed. When used in a claim, it closes the claim to exclude all technical features other than those expressly listed, except for associated conventional impurities. If the term appears only in a clause of a claim, it limits the claim to the elements expressly listed in that clause; elements recited in other clauses are not excluded from the overall claim.
[0024] Unless otherwise explicitly specified or limited, the terms "installation," "connection," "linking," and "fixing," etc., should be interpreted broadly. For example, they can refer to fixed connections, detachable connections, or integral connections; they can refer to mechanical connections or electrical connections; they can refer to direct connections or indirect connections through an intermediate medium; and they can refer to the internal connection between two components. Those skilled in the art can understand the specific meaning of the above terms in this document according to the specific circumstances.
[0025] The terms “center,” “longitudinal,” “lateral,” “length,” “width,” “thickness,” “upper,” “lower,” “front,” “back,” “left,” “right,” “vertical,” “horizontal,” “top,” “bottom,” “inner,” “outer,” “clockwise,” and “counterclockwise” indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience and simplification of description and do not imply that the device or component referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this document.
[0026] The solution provided by this invention will be described in detail below. Contents not described in detail in the embodiments of this invention are prior art known to those skilled in the art. Where specific conditions are not specified in the embodiments of this invention, they shall be performed according to conventional conditions in the art or conditions recommended by the manufacturer. Reagents or instruments used in the embodiments of this invention whose manufacturers are not specified are all conventional products that can be purchased commercially.
[0027] like Figure 1 and Figure 2 As shown, this invention provides a frequency-based deep reinforcement learning-based motion control method for quadruped robots, comprising:
[0028] Step 1: Use a model-based reference trajectory generator to generate kinematically and dynamically continuous reference trajectories online according to different task objectives, and guide the trajectory tracking strategy of the model-free trajectory tracker to explore.
[0029] Step 2: Using the spectral convolution module, the environmental perception information and the body perception information are projected from the time domain to the frequency domain through discrete Fourier transform to obtain environmental frequency domain perception information and body frequency domain perception information, respectively. Environmental frequency features and body frequency features are extracted from the environmental frequency domain perception information and body frequency domain perception information, respectively. The motion features of the model-based trajectory planner are adjusted by the environmental frequency features and the control bandwidth of the model-free trajectory tracker is improved by the body frequency features.
[0030] Step 3: Train the model-free trajectory tracker using a constrained synchronous teacher-student reinforcement learning framework. The training process involves applying various physical constraints to enable the model-free trajectory tracker to explore within a safe domain, and synchronous teacher-student training is performed.
[0031] Step 4: Use a trained model-free trajectory tracker to control the motion of the quadruped robot.
[0032] Preferably, in step 1 of the above method, a kinematically and dynamically continuous reference trajectory is generated online according to different task objectives using a model-based reference trajectory generator, including:
[0033] Generated reference trajectory It includes: gait pattern reference trajectory, leg swing trajectory reference trajectory, foot landing reference trajectory, and trunk stability reference trajectory.
[0034] Preferably, in the above method, the gait pattern reference trajectory is the contact sequence of the quadruped robot's feet, and the gait pattern reference trajectory is a function of... for:
[0035] ;
[0036] in, It is the phase offset value of the i-th leg of the quadruped robot. This represents the time offset between the legs of a quadruped robot. This refers to the left hind leg of the quadruped robot. The phase offset value, This refers to the left front leg of the quadruped robot. The phase offset value, This refers to the right hind leg of the quadruped robot. The phase offset value, This refers to the right front leg of the quadruped robot. The phase offset value; t is the phase counter that determines the step frequency, which is updated at each control time step in the following manner:
[0037] ;
[0038] in, It is the desired gait frequency; It is the update increment of the phase counter for each control time step that determines the step frequency;
[0039] By modifying the target contact mode To achieve tripedal and bipedal walking gait in quadruped robots, and target contact modes. for:
[0040] ;
[0041] The reference trajectory for the leg swing is the target motion of the non-contact foot of the quadruped robot. The reference trajectory for the smooth leg swing motion of the non-contact foot is constructed using a third-order Bessel function as follows:
[0042] ;
[0043] in, It is the leg swing height; B() is the third-order Bessel function; It's the leg swing factor. It is a function of the gait pattern reference trajectory; , and These are the starting point height, highest point height, and ending point height of the control points for the leg swing trajectory reference trajectory, respectively. The third-order Bessel function form is:
[0044] ;
[0045] in, and These are the starting and ending points of the reference trajectory for the leg swing, respectively, and t is the phase counter that determines the step frequency;
[0046] The height of the highest point in the leg swing trajectory reference trajectory is calculated as follows:
[0047] ;
[0048] in, Based on online adjustment of the height of surrounding obstacles, the leg swing trajectory is only subject to reference constraints in the normal direction.
[0049] The foot landing reference trajectory uses a heuristic algorithm to select a suitable foot landing position to maintain stable motion. ,for:
[0050] ;
[0051] in, It is the projection of the quadruped robot's hip joints onto the ground; It is the duration of the supporting phase; and These are the target linear velocity and angular velocity, respectively. It is the distance from the hip joint of the quadruped robot to its base;
[0052] Expand the reference target point to the target area, and use the target landing point as the reference target point. A circular region centered at a given radius for:
[0053] ;
[0054] Where R represents the real number field; This represents the radius of the circular area of the landing point that can be adjusted online according to surrounding obstacles.
[0055] The trunk stability reference trajectory is approximated by a variable-height inverted pendulum model, the stability criterion of which is:
[0056] ;
[0057] in, It is the acceleration of the quadruped robot's center of mass; It is the location of the quadruped robot's center of mass; It is the set of contact points of the legs of a quadruped robot; It is the weight of the i-th foot contact point of the quadruped robot; It is the position of the i-th foot contact point of the quadruped robot; It is the acceleration of the quadruped robot's center of mass in the height direction; It is a gravity vector; It is the height of the quadruped robot's center of mass from the ground.
[0058] On flat ground, the target area degenerates into a point constraint, which can more effectively guide the robot's gait learning; on rugged terrain, the target landing area expands to provide greater obstacle-crossing flexibility.
[0059] Preferably, in step 2 of the above method, the environmental perception information is projected from the time domain to the frequency domain using a spectral convolution module via discrete Fourier transform to obtain environmental frequency domain perception information, and environmental frequency features are extracted from the environmental frequency domain perception information, including:
[0060] The discrete Fourier transform submodule of the spectral convolution module uses continuous environmental perception information. As input, environmental perception information Convert to the frequency domain to obtain the environmental spectrum distribution The low-pass filter submodule of the spectral convolution module is used to filter and eliminate environmental spectral distribution. High-frequency noise components and low-frequency components extracted As environmental frequency domain sensing information;
[0061] The feature extraction submodule of the spectral convolution module extracts the low-frequency components, which are part of the environmental frequency domain perception information. Calculate the average spectral amplitude And based on the average spectral amplitude Determine frequency components As extracted environmental frequency features.
[0062] The above-described extraction of frequency features from environmental perception information utilizes a spectral convolution method to extract key features. First, continuous frames of environmental perception information are input into a discrete Fourier transform module, which transforms them to the frequency domain to obtain a spectral distribution. In this distribution, low-frequency components reflect the overall trend of the signal, while high-frequency components are affected by sensor noise. A low-pass filter module is used to eliminate high-frequency noise and extract the main low-frequency components. Finally, a feature extraction module calculates the average spectral intensity of the low-frequency component distribution, representing the average frequency of the environment.
[0063] In step 2, the entity perception information is projected from the time domain to the frequency domain using a spectral convolution module and discrete Fourier transform to obtain entity frequency domain perception information, and entity frequency features are extracted from the entity frequency domain perception information, including:
[0064] Using the joint readings of the quadruped robot as the input of the discrete Fourier transform module of the spectrum convolution module, the energy density distribution is extracted from the quadruped robot's motion, and the original motion trajectory of the quadruped robot is transformed into the frequency domain to obtain the motion spectrum features as the body's frequency domain perception information.
[0065] The motion spectrum amplitude is obtained by calculating the motion spectrum features through the feature extraction submodule of the spectral convolution module. And based on the amplitude of the action spectrum Determine frequency characteristics As the extracted ontological frequency features.
[0066] The frequency features of the ontological perception information described above are extracted using the same spectral convolution method to calculate the frequency features of the control signal. First, a discrete Fourier transform module is used to transform the continuous control output trajectory to the frequency domain. Then, the average spectral intensity, representing the average frequency of the control output signal, is directly extracted from the spectral distribution of the control signal.
[0067] By extracting frequency features from environmental and ontological information, the control output can dynamically adapt to the surrounding environment and improve the bandwidth of the network controller.
[0068] Preferably, in the above method, in the process of projecting environmental perception information from the time domain to the frequency domain using discrete Fourier transform, the environmental perception information is obtained by measuring the elevation map of the ground relative to the height of the quadruped robot body.
[0069] Environmental sensing information as input Represented as: R represents the real number field. For each frame, the environmental information has the following dimensions: Dimensions, n frames in total;
[0070] Output frequency components It is shown as:
[0071] C represents the complex field, with dimension 1. dimension;
[0072] ; This refers to the frequency domain environment information obtained from the previous time step through the discrete Fourier transform. For each time step, the complex value is represented.
[0073] ;
[0074] ;
[0075] in, Represents a normal distribution; This represents the signal after frequency domain transformation, where k is the discrete sampling frequency component. , Represents an integer; Indicates the average spectral amplitude; This indicates an amplitude calculation operation; This is the environmental signal filtered by the low-pass filter submodule; This indicates a low-pass filter operation; m represents the number of frequency components. This refers to the frequency domain environment information obtained from the previous time step through the discrete Fourier transform submodule. For each time step, there is a complex value. This represents the positive frequency component of the complex value corresponding to each time step. The discrete Fourier transform operation is represented; C represents the complex field with dimension 1. dimension; Represents environmental frequency domain sensing information; This represents the low-frequency portion of environmental frequency domain sensing information.
[0076] Preferably, in step 2 of the above method, the motion features of the model-based trajectory planner are adjusted using the extracted environmental frequency features in the following manner:
[0077] By extracting frequency components as environmental frequency characteristics The motion features of the reference trajectory are linearly combined with the original motion features to adjust the motion features online. The adjusted motion features include: gait frequency. Leg swing trajectory height from the ground Radius of landing point and target base height ,in,
[0078] ;
[0079] ;
[0080] ;
[0081] ;
[0082] in, , , and These are the adaptive coefficients for the model-based reference trajectory generator, the adaptive coefficients for the leg swing trajectory planning, the adaptive coefficients for the foot landing point planning, and the adaptive coefficients for the centroid planning; the motion features of the original reference trajectory include gait frequency. Leg swing trajectory height from the ground Radius of landing point and target base height .
[0083] Preferably, in step 2 of the above method, the motion features of the model-based trajectory planner are adjusted using the extracted environmental frequency features in the following manner:
[0084] By extracting frequency components as environmental frequency characteristics The motion features of the reference trajectory are linearly combined with the original motion features to adjust the motion features online. The adjusted motion features include: gait frequency. Leg swing trajectory height from the ground Radius of landing point and target base height ,in,
[0085] ;
[0086] ;
[0087] ;
[0088] ;
[0089] in, , , and These are the adaptive coefficients for the model-based reference trajectory generator, the adaptive coefficients for the leg swing trajectory planning, the adaptive coefficients for the foot landing point planning, and the adaptive coefficients for the centroid planning; the motion features of the original reference trajectory include gait frequency. Leg swing trajectory height from the ground Radius of landing point and target base height .
[0090] Preferably, in step 2 of the above method, the joint readings of the quadruped robot are used. The given joint readings of the quadruped robot are used as inputs to the discrete Fourier transform module of the spectral convolution module. Represented as ;
[0091] The motion spectrum amplitude is calculated from the motion spectrum features using the feature extraction submodule of the spectral convolution module as follows: And based on the amplitude of the action spectrum Determine frequency characteristics The extracted ontology frequency features include:
[0092] The feature extraction submodule of the spectral convolution module calculates the action spectrum amplitude according to the following process. Fourier adaptive output of ontology-aware information as ontology frequency features Perform the calculation:
[0093] ;
[0094] ;
[0095] in, This represents the Discrete Fourier Transform operation; m represents the number of frequency components. Represents discrete sampling frequency components. , ;
[0096] Extracted frequency features The distribution of motion energy density is characterized by the following frequency-dependent reward function. During reinforcement learning, the encoding bandwidth is limited, and the frequency-related reward function... for:
[0097] ;
[0098] Where r is the original reward.
[0099] Preferably, in step 3 of the above method, the model-free trajectory tracker is trained using a constrained synchronous teacher-student reinforcement learning framework in the following manner:
[0100] This algorithm employs a normalized penalty policy optimization algorithm, which is based on the unconstrained proximal policy optimization algorithm and additionally penalizes actions that violate constraints. Its objective function is... for:
[0101] ;
[0102] Among them, the reward loss function ; Indicates the weighting coefficient of the constraint; The number indicates the constraint sequence. In this invention, there are a total of 5 constraints. The values are 1, 2, 3, 4, 5; E[] represents the expected value;
[0103] Constraint loss function ;
[0104] in, Indicates the probability ratio of the policy , This represents the updated policy distribution. Indicates an action, Indicates the complete state. () indicates the policy distribution before the update; It is a normalized reward advantage function; It is the first The dominant function of each constraint; It is a constraint violation value; It is a constraint threshold; The operation will crop the value to and Between, among The extent of control strategy updates.
[0105] Preferably, in step 3 of the above method, the training process involves applying various physical constraints, including:
[0106] (1) Joint constraint Joint position ,speed and torque Within the hardware limitations of the quadruped robot's motors:
[0107] ;
[0108] in, , and Hardware limitations that define the robot's joint angles, speeds, and torques; Let represent all the rotational joints of the quadruped robot, and j represent the j-th rotational joint of the quadruped robot;
[0109] (2) Link collision constraint The links of a quadruped robot should not collide with the ground or other links, and the contact force of the links should be minimized. It needs to be zero to avoid collisions:
[0110] ;
[0111] in, The symbol represents the link that the quadruped robot should avoid colliding with; l represents the l-th link of the quadruped robot;
[0112] (3) Contact sliding constraint The legs of a quadruped robot should maintain contact with the ground during movement; that is, the speed of the contact leg should be... Zero:
[0113] ;
[0114] in, i={1,2,3,4} represents the contact state of the four legs of the quadruped robot, where 1 indicates contact and 0 indicates no contact.
[0115] (4) Leg swing height constraint For each foot of the quadruped robot in the leg-swinging phase, the contact force of the foot. It needs to be zero to avoid obstacles:
[0116] ;
[0117] (5) Trunk stability constraints The ground reaction force must maintain the stability of the underactuated system during motion, i.e., the trunk stability reference trajectory in the reference trajectory of step 1. :
[0118] ;
[0119] in, It is the acceleration of the quadruped robot's center of mass; It is the location of the quadruped robot's center of mass; It is the set of contact points of the legs of a quadruped robot; s is the weight of the i-th foot contact point of the quadruped robot; i It is the position of the contact point of the i-th leg of the quadruped robot; This represents the acceleration of the quadruped robot's center of mass in the height direction; It is a gravity vector; It is the height of the quadruped robot's center of mass from the ground.
[0120] Preferably, in step 3 of the above method, the constrained synchronous teacher-student reinforcement learning framework optimizes the synchronous training of a teacher policy that has full access to privileged information through constrained policies, while allowing the student policy using ontology-aware information to minimize reconstruction loss through supervised learning. The overall loss function of this constrained synchronous teacher-student reinforcement learning framework is... for:
[0121] ;
[0122] in, For policy networks The objective function of the policy network The teacher and student agents, trained in parallel for both teacher and student strategies, share the same policy network. This policy network is a multilayer perceptron with an exponential linear unit activation function. objective function for:
[0123] ;
[0124] in, This represents the pruning loss based on the policy gradient; This represents the clipping loss for each constraint.
[0125] Policy Network Based on proprioceptive observation Motion estimator Estimated motion characteristics of the output and respectively by privileged encoders Or proprioceptive encoder Generated latent representations To calculate actions Rewarding critics and long-term cost commentators Process complete state And use generalized advantage estimation to evaluate rewards in parallel. and constraints ;
[0126] By minimizing the following loss function Training rewards critic and long-term cost commentators loss function for:
[0127] ;
[0128] Where E[] represents the expected value; R() represents the actual reward value of the strategy; Represents the estimated complete state Corresponding actions The value of C() represents the actual constraint violation value. () indicates the estimated constraint violation value;
[0129] Teacher group's intelligent physical fitness access complete status and utilize privileged encoders Extracting latent representations from the complete state The student-group intelligent agent only receives ontology perception observations. and utilizes a body-aware encoder observation sequence Encoding as a latent representation The student-group intelligent agent uses a motion estimator. From the perspective of proprioception Medium estimation features ;
[0130] Proprioceptive encoder The reconstruction loss function is updated by introducing a reconstruction loss to minimize the difference between privileged and ontology-aware latent representations. for:
[0131] ;
[0132] in, Indicates the complete state of using privileges. The calculated environmental characterization; This represents the environmental representation obtained by using a quadruped robot through proprioception prediction.
[0133] Motion estimator The loss function is trained by minimizing the mean squared error between the estimated value and the true value. for: .
[0134] In summary, the method of this invention, through the discrete Fourier transform module, can capture frequency information of the environment and the robot's own state. Based on both environment-aware and body-aware Fourier transforms, it extracts frequency features from multimodal frequency domain sensing information, enabling the control output to dynamically adapt to the surrounding environment and improving the bandwidth of the network controller. The results show that the quadruped robot motion control method of this invention is applicable to different robots and different tasks.
[0135] To more clearly demonstrate the technical solution and its effects provided by the present invention, the following detailed description of the solution provided by the embodiments of the present invention is provided with reference to specific examples.
[0136] Example 1
[0137] This embodiment provides a frequency-based deep reinforcement learning-based motion control method for quadruped robots, the architecture of which is as follows: Figure 2 As shown, by extracting frequency features from environmental perception information and airborne perception information, the output signal controlling the quadruped robot can dynamically adapt to the surrounding environment and improve the bandwidth of the network controller.
[0138] The frequency-based deep reinforcement learning-based quadruped robot motion control method includes the following steps:
[0139] Step 1: Use a model-based reference trajectory generator to generate kinematically and dynamically continuous reference trajectories online according to different task objectives, and guide the trajectory tracking strategy of the model-free trajectory tracker to explore.
[0140] Step 2: Use Discrete Fourier Transform to project the ontology perception and environment perception from the time domain to the frequency domain to obtain multimodal frequency domain perception information. Extract the main frequency features from the multimodal frequency domain perception information through the spectral convolution module to adjust the motion features of the model-based trajectory planner and improve the control bandwidth of the model-free trajectory tracker.
[0141] Step 3: Train the model-free trajectory tracker using a constrained synchronous teacher-student reinforcement learning framework. The training process explores the network within the safe domain by applying various physical constraints, avoiding tedious hyperparameter tuning. Synchronous teacher-student training can accelerate the distillation of expert knowledge from simulation to reality.
[0142] Step 1 of the above method is as follows:
[0143] The designed model-based reference trajectory generator, consisting of phase-based gait planning, Bezier-based swing leg trajectory planning, region-constrained foot placement planning, and an inverted pendulum-based center-of-mass planner, is used to generate kinematically and dynamically consistent reference trajectories. These reference trajectories, including gait pattern reference trajectory, leg swing trajectory reference trajectory, foot landing reference trajectory, and trunk stability reference trajectory, are key components of foot movement tasks.
[0144] The gait pattern reference trajectory is defined as the contact sequence of the robot's feet. To specify the desired contact sequence, a gait pattern function is used. Defined as:
[0145] ;
[0146] in, is the phase offset value of the i-th leg of the quadruped robot; t is the phase counter, which determines the step frequency and is updated at each control time step.
[0147] ;
[0148] here It is the desired gait frequency and phase offset value. This represents the time offset between the legs of a quadruped robot. These represent the phase offset values of the left hind foot, left forefoot, right hind foot, and right forefoot of the quadruped robot, respectively.
[0149] Triped and bipedal gait can be achieved by modifying the target contact mode. For example, setting the contact mode of the non-contact foot to 0 at all times will enable bipedal walking. for:
[0150] .
[0151] The aforementioned leg swing trajectory reference trajectory refers to the target motion of the non-contact foot. To achieve a smooth leg swing motion of the non-contact foot, the reference trajectory is constructed using a third-order Bessel function:
[0152] ;
[0153] in, It is the leg swing height; B() is the third-order Bessel function; It's the leg swing factor. , and These are the heights of the starting point, the highest point, and the ending point of the control point in the leg swing trajectory, expressed in the form of a third-order Bessel function:
[0154] ;
[0155] The height of the highest point of the leg swing trajectory is calculated as follows:
[0156] ;
[0157] in, The leg swing trajectory is adjusted online based on the height of surrounding obstacles. Reference constraints are applied only in the normal direction, ensuring a smooth leg swing while maintaining flexibility in the tangential foot landing point.
[0158] The aforementioned landing point reference trajectory maintains stable motion by selecting appropriate landing positions. To reduce computational load, this invention employs a heuristic algorithm:
[0159] ;
[0160] in, It is the distance from the hip joint to the base. It is the duration of the supporting phase. It is the projection of the hip joint onto the ground. and These are the target linear velocity and angular velocity, respectively. Its limitation is that it is only applicable to blind motion on flat terrain. This invention broadens the reference target point to a target area, a circular region with a given radius centered on the target landing point:
[0161] ;
[0162] On flat ground, the target area degenerates into a point constraint; while on rugged terrain, the target area expands to provide greater obstacle-crossing flexibility.
[0163] The aforementioned trunk stability reference trajectory maintains the center of pressure (CoP) within the supporting polygon by balancing the ground reaction force (GRF), thus ensuring dynamic stability during motion. To calculate the center of mass dynamics of the robot system with low complexity, a variable-height inverted pendulum model is used as an approximation. The stability criterion for the inverted pendulum model is defined as:
[0164] ;
[0165] in, It is the acceleration of the quadruped robot's center of mass; It is the location of the quadruped robot's center of mass; It is the set of contact points of the legs of a quadruped robot; It represents the contact point weight of the i-th leg of the quadruped robot; It is the position of the contact point of the i-th leg of the quadruped robot; It is the acceleration of the quadruped robot's center of mass in the height direction; It is a gravity vector; This refers to the height of the quadruped robot's center of mass above the ground. The variable-height inverted pendulum model plays a crucial role in generating a dynamically continuous motion reference trajectory because it effectively captures the flight phase dynamics caused by body height movement.
[0166] The specific processing of step 2 of the above method is as follows:
[0167] A frequency-based control strategy, namely Fourier Online Adaptive (FOA), is adopted. This strategy can adaptively adjust the model-based reference trajectory generator and the model-free trajectory tracker based on the frequency characteristics of real-time environmental perception information and ontology perception information.
[0168] Environmentally Aware Fourier Online Adaptation: First, the Discrete Fourier Transform module uses continuous environmental awareness information... As input, transform to the frequency domain to obtain the spectral distribution. Environmental perception information is acquired through elevation maps, which are obtained by measuring the ground height relative to the robot body. In the generated environmental spectrum, low-frequency components reflect the overall trend of the signal, while high-frequency components capture detailed information but are more susceptible to sensor noise. Therefore, a low-pass filter module is used to eliminate noise and extract the main frequency components, especially the low-frequency components. After acquiring these low-frequency components, key features are captured from the environmental spectrum using a feature extraction module. In this process, the present invention employs average spectral amplitude. It characterizes the energy density distribution at different frequencies. Given an environmental perception information input sequence. Output Represented as:
[0169] ;
[0170] ;
[0171] ;
[0172] ;
[0173] Where m represents the number of frequency components; The signal after frequency domain transformation; This is the filtered environmental signal; The average spectral amplitude; This represents the Discrete Fourier Transform operation; Indicates low-pass filtering operation; This indicates the amplitude calculation operation.
[0174] Extracting frequency components Then, the reference trajectory is adjusted based on the frequency density of the environment. Specifically, the gait frequency... Leg swing trajectory height from the ground Radius of landing point and target base height The system will be adjusted online to enable the quadruped robot to take more aggressive actions in unstructured terrain environments and behave more conservatively in stable environments. The environment-aware Fourier adaptive mechanism is defined as a linear combination of environmental energy density and the original motion characteristics.
[0175] ;
[0176] ;
[0177] ;
[0178] ;
[0179] in, , , and These are the adaptive coefficients for gait planning, leg swing trajectory planning, foot landing point planning, and center of mass planning. , , and As the original reference motion characteristics, , , and This refers to the motion characteristics after online adjustment.
[0180] Body-aware Fourier online adaptation: Similar to environment-aware adaptation, body-aware Fourier adaptation uses discrete Fourier transform to convert the original motion trajectory to the frequency domain. The difference lies in that body-aware adaptation uses the joint readings of the quadruped robot as input, directly extracting the energy density distribution from the quadruped robot's motion without filtering. Specifically, given a series of joint readings... Action spectrum amplitude Fourier adaptive output with proprioception The calculation process is as follows:
[0181] ;
[0182] ;
[0183] Extracted frequency features The distribution of motion frequency components can be used to limit encoding bandwidth during reinforcement learning. To enhance robustness to network noise output, this invention employs the following frequency-dependent reward function:
[0184] ;
[0185] By multiplying the original reward r by the exponential of the action spectrum amplitude, the training process encourages low-frequency actions that correspond to higher rewards, thereby achieving stable motion in real-world scenarios.
[0186] The specific processing of step 3 of the above method is as follows:
[0187] This invention employs a constrained policy optimization algorithm that enforces multiple constraints while students mimic the teacher's strategy. Specifically, it utilizes a normalized penalty policy optimization algorithm, based on an unconstrained proximal policy optimization algorithm, to additionally penalize actions that violate constraints. Its objective function... Defined as:
[0188] ;
[0189] in,
[0190] ;
[0191] ;
[0192] in, Indicates the probability ratio of the policy , This represents the updated policy distribution. Indicates an action, Indicates the complete state. () indicates the policy distribution before the update; It is a normalized reward advantage function; It is the first The dominant function of each constraint; It is a constraint violation value; It is a constraint threshold; The operation will crop the value to and between, Parameters for controlling the update magnitude of the strategy.
[0193] This invention introduces a total of 5 physical constraints:
[0194] (1) Joint restriction Joint position ,speed and torque It should be within the hardware limitations of the motor:
[0195] ;
[0196] (2) Linkage collision The connecting rod should not collide with the ground or other links to avoid damage; the contact force of the connecting rod... It needs to be zero to avoid collisions:
[0197] ;
[0198] (3) Contact sliding The legs of a quadruped robot should maintain contact with the ground during movement; that is, the speed of the contact leg should be... It is zero.
[0199] ;
[0200] in, i={1,2,3,4} represents the contact state of the four legs of the quadruped robot, where 1 indicates contact and 0 indicates no contact.
[0201] (4) Leg swing height For each foot during the leg swing phase, the contact force of the foot. It needs to be zero to avoid obstacles.
[0202] ;
[0203] (5) Trunk stability The ground reaction force must maintain the stability of the underactuated system during the motion, i.e., the trunk stability reference trajectory mentioned in step one.
[0204] ;
[0205] This invention further improves learning efficiency and ensures real-world performance by employing a synchronous teacher-student reinforcement learning framework. This framework optimizes the synchronous training of a teacher policy with full access to privileged information through policy constraints, while simultaneously allowing the student policy, which uses ontology awareness, to minimize reconstruction loss through supervised learning. The teacher and student policies are trained in parallel by dividing the agents into "teacher group agents" and "student group agents." Both groups of agents share the same policy network. This network is a multilayer perceptron with exponential linear unit activation functions. The policy network is based on ontology-aware observations. The estimated motion features output by the motion estimator and respectively by privileged encoders Or proprioceptive encoder Generated latent representations To calculate actions To ensure memory efficiency, critics are rewarded. and long-term cost commentators Process complete state And use generalized advantage estimation to evaluate rewards in parallel. and constraints The teacher group agents have access to the complete state. and utilize privileged encoders Extracting latent representations from the complete state The student-group intelligent agent only receives ontological observations. and utilizes a body-aware encoder observation sequence Encoding as a latent representation The student-group intelligent agent uses a motion estimator. From airborne sensors Medium estimation features .
[0206] Policy Network objective function Defined as:
[0207] ;
[0208] Rewarding critics and long-term cost commentators By minimizing the following loss Conduct training:
[0209] ;
[0210] By introducing reconstruction loss To update the ontology-aware encoder, minimizing the difference between the privileged and ontology-aware latent representations:
[0211] ;
[0212] The motion estimator uses the mean squared error between the estimated value and the true value as its loss function. To train and predict motion characteristics as accurately as possible:
[0213] ;
[0214] Overall loss function of a constrained synchronous teacher-student framework Defined as:
[0215] .
[0216] This invention provides a quadruped robot motion control method for training quadruped robots to perform diverse and robust obstacle-crossing skills on complex terrain. Through a spectral convolution module, it captures frequency information about the environment and the quadruped robot's own state. Based on both environment-aware Fourier transforms and body-aware Fourier transforms, it extracts frequency features from multimodal sensory information, enabling the output signal controlling the quadruped robot to dynamically adapt to the surrounding environment and improving the bandwidth of the network controller. Results show that this frequency-based deep reinforcement learning-based quadruped robot motion control method is applicable to different robots and different tasks.
[0217] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0218] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims. The information disclosed in the background section is intended only to enhance the understanding of the overall background technology of the present invention and should not be construed as an admission or implication in any way that such information constitutes prior art known to those skilled in the art.
Claims
1. A method for motion control of a quadruped robot based on frequency-based deep reinforcement learning, characterized in that, include: Step 1: Use a model-based reference trajectory generator to generate kinematically and dynamically continuous reference trajectories online according to different task objectives, and guide the trajectory tracking strategy of the model-free trajectory tracker to explore. Step 2: Using the spectral convolution module, the environmental perception information and the body perception information are projected from the time domain to the frequency domain through discrete Fourier transform to obtain environmental frequency domain perception information and body frequency domain perception information, respectively. Environmental frequency features and body frequency features are extracted from the environmental frequency domain perception information and body frequency domain perception information, respectively. The motion features of the model-based trajectory planner are adjusted by the environmental frequency features and the control bandwidth of the model-free trajectory tracker is improved by the body frequency features. In step 2, the motion characteristics of the model-based trajectory planner are adjusted using the extracted environmental frequency features in the following manner: By extracting frequency components as environmental frequency characteristics The motion features of the reference trajectory are linearly combined with the original motion features to adjust the motion features online. The adjusted motion features include: gait frequency. Leg swing trajectory height from the ground Radius of landing point and target base height ,in, ; ; ; ; in, , , and These are the adaptive coefficients for the model-based reference trajectory generator, the adaptive coefficients for the leg swing trajectory planning, the adaptive coefficients for the foot landing point planning, and the adaptive coefficients for the centroid planning; the motion features of the original reference trajectory include gait frequency. Leg swing trajectory height from the ground Radius of landing point and target base height ; Step 3: Train the model-free trajectory tracker using a constrained synchronous teacher-student reinforcement learning framework. The training process involves applying various physical constraints to enable the model-free trajectory tracker to explore within a safe domain, and synchronous teacher-student training is performed. Step 4: Use a trained model-free trajectory tracker to control the motion of the quadruped robot.
2. The frequency-based deep reinforcement learning-based quadruped robot motion control method according to claim 1, characterized in that, In step 1, a model-based reference trajectory generator is used online to generate kinematically and dynamically continuous reference trajectories according to different task objectives, including: Generated reference trajectory It includes: gait pattern reference trajectory, leg swing trajectory reference trajectory, foot landing reference trajectory, and trunk stability reference trajectory.
3. The frequency-based deep reinforcement learning-based quadruped robot motion control method according to claim 2, characterized in that, The gait pattern reference trajectory is the contact sequence of the quadruped robot's feet, and the function of the gait pattern reference trajectory is... for: ; in, It is the phase offset value of the i-th leg of the quadruped robot. This represents the time offset between the legs of a quadruped robot. This refers to the left hind leg of the quadruped robot. The phase offset value, This refers to the left front leg of the quadruped robot. The phase offset value, This refers to the right hind leg of the quadruped robot. The phase offset value, This refers to the right front leg of the quadruped robot. The phase offset value; t is the phase counter that determines the step frequency, which is updated at each control time step in the following manner: ; in, It is the desired gait frequency; It is the update increment of the phase counter for each control time step that determines the step frequency; By modifying the target contact mode To achieve tripedal and bipedal walking gait in quadruped robots, and target contact modes. for: ; The reference trajectory for the leg swinging motion is the target motion of the non-contact foot of the quadruped robot. The reference trajectory for the smooth leg swinging motion of the non-contact foot is constructed using a third-order Bessel function as follows: ; in, It is the leg swing height; B() is the third-order Bessel function; It's the leg swing factor. It is a function of the gait pattern reference trajectory; , and These are the starting point height, highest point height, and ending point height of the control points for the leg swing trajectory reference trajectory, respectively. The third-order Bessel function form is: ; in, and These are the starting and ending points of the reference trajectory for the leg swing, respectively, and t is the phase counter that determines the step frequency; The height of the highest point in the leg swing trajectory reference trajectory is calculated as follows: ; in, Based on online adjustment of the height of surrounding obstacles, the leg swing trajectory is only subject to reference constraints in the normal direction. The foot landing reference trajectory uses a heuristic algorithm to select a suitable foot landing position to maintain stable motion. ,for: ; in, It is the projection of the quadruped robot's hip joints onto the ground; It is the duration of the supporting phase; and These are the target linear velocity and angular velocity, respectively. It is the distance from the hip joint of the quadruped robot to its base; Expand the reference target point to the target area, and use the target landing point as the reference target point. A circular region with a given radius centered at a given radius for: ; Where R represents the real number field; This represents the radius of the circular area of the landing point that can be adjusted online according to surrounding obstacles. The trunk stability reference trajectory is approximated by a variable-height inverted pendulum model, the stability criterion of which is: ; in, It is the acceleration of the quadruped robot's center of mass; It is the location of the quadruped robot's center of mass; It is the set of contact points of the legs of a quadruped robot; It represents the contact point weight of the i-th leg of the quadruped robot; It is the position of the contact point of the i-th leg of the quadruped robot; It is the acceleration of the quadruped robot's center of mass in the height direction; It is a gravity vector; It is the height of the quadruped robot's center of mass from the ground.
4. The frequency-based deep reinforcement learning-based quadruped robot motion control method according to any one of claims 1-3, characterized in that, In step 2, the environmental perception information is projected from the time domain to the frequency domain using a spectral convolution module and discrete Fourier transform to obtain environmental frequency domain perception information, and environmental frequency features are extracted from the environmental frequency domain perception information, including: The discrete Fourier transform submodule of the spectral convolution module uses continuous environmental perception information. As input, environmental perception information Convert to the frequency domain to obtain the environmental spectrum distribution The low-pass filter submodule of the spectral convolution module is used to filter and eliminate environmental spectral distribution. High-frequency noise components and low-frequency components extracted As environmental frequency domain sensing information; The feature extraction submodule of the spectral convolution module extracts the low-frequency components, which are part of the environmental frequency domain perception information. Calculate the average spectral amplitude And based on the average spectral amplitude Determine frequency components As extracted environmental frequency features; In step 2, the entity perception information is projected from the time domain to the frequency domain using a spectral convolution module and discrete Fourier transform to obtain entity frequency domain perception information, and entity frequency features are extracted from the entity frequency domain perception information, including: Using the joint readings of the quadruped robot as the input of the discrete Fourier transform module of the spectrum convolution module, the energy density distribution is extracted from the quadruped robot's motion, and the original motion trajectory of the quadruped robot is transformed into the frequency domain to obtain the motion spectrum features as the body's frequency domain perception information. The motion spectrum amplitude is obtained by calculating the motion spectrum features through the feature extraction submodule of the spectral convolution module. And based on the amplitude of the action spectrum Determine frequency characteristics As the extracted ontological frequency features.
5. The frequency-based deep reinforcement learning-based quadruped robot motion control method according to claim 4, characterized in that, In the process of projecting environmental perception information from the time domain to the frequency domain using a spectral convolution module and discrete Fourier transform, the environmental perception information... This environmental perception information is obtained by measuring the elevation map of the ground relative to the height of the quadruped robot. Represented as: Where R represents the real number field, For each frame, the environmental information has the following dimensions: Dimensions, n frames in total; Output frequency components Represented as: ; ; ; ; in, Represents a normal distribution; The signal is after frequency domain transformation, where k is the discrete sampling frequency component. , , Represents an integer; Indicates the average spectral amplitude; This indicates an amplitude calculation operation; This is the environmental signal filtered by the low-pass filter submodule; This indicates a low-pass filter operation; m represents the number of frequency components. This refers to the frequency domain environment information obtained from the previous time step through the discrete Fourier transform submodule. For each time step, the complex value is represented. This represents the positive frequency component of the complex value corresponding to each time step. Represents environmental frequency domain sensing information; The discrete Fourier transform operation is represented; C represents the complex field with dimension 1. dimension; This represents the low-frequency portion of environmental frequency domain sensing information.
6. The frequency-based deep reinforcement learning-based quadruped robot motion control method according to claim 4, characterized in that, In step 2, the joint readings of the quadruped robot are used. The given joint readings of the quadruped robot are used as inputs to the discrete Fourier transform module of the spectral convolution module. Represented as ; The motion spectrum amplitude is calculated from the motion spectrum features using the feature extraction submodule of the spectral convolution module as follows: And based on the amplitude of the action spectrum Determine frequency characteristics The extracted ontology frequency features include: The feature extraction submodule of the spectral convolution module calculates the action spectrum amplitude according to the following process. Fourier adaptive output of ontology-aware information as ontology frequency features Perform the calculation: ; ; in, This represents the Discrete Fourier Transform operation; m represents the number of frequency components. Represents discrete sampling frequency components. , ; Extracted frequency features The distribution of motion energy density is characterized by the following frequency-dependent reward function. During reinforcement learning, the encoding bandwidth is limited, and the frequency-related reward function... for: ; Where r is the original reward.
7. The frequency-based deep reinforcement learning-based quadruped robot motion control method according to any one of claims 1-3, characterized in that, In step 3, the model-free trajectory tracker is trained using a constrained synchronous teacher-student reinforcement learning framework in the following manner: This algorithm employs a normalized penalty policy optimization algorithm, which is based on the unconstrained proximal policy optimization algorithm and additionally penalizes actions that violate constraints. Its objective function is... for: ; Among them, the reward loss function ; Indicates the weighting coefficient of the constraint; This indicates the constraint number; there are a total of 5 constraints. The values are 1, 2, 3, 4, 5; i={1,2,3,4} represents the index of the four legs of the quadruped robot; E[] represents the expected value; Constraint loss function ; in, Indicates the probability ratio of the policy , This represents the updated policy distribution. Indicates an action, Indicates the complete state. This represents the policy distribution before the update; It is a normalized reward advantage function; It is the first The dominant function of each constraint; It is a constraint violation value; It is a constraint threshold; The operation will crop the value to and Between, among The extent of control strategy updates.
8. The frequency-based deep reinforcement learning-based quadruped robot motion control method according to claim 7, characterized in that, In step 3, the training process involves applying various physical constraints, including: (1) Joint constraint Joint position ,speed and torque Within the hardware limitations of the quadruped robot's motors: ; in, , and Hardware limitations that define the joint angles, speeds, and torques of a quadruped robot; Let represent all the rotational joints of the quadruped robot, and j represent the j-th rotational joint of the quadruped robot; (2) Link collision constraint The links of a quadruped robot should not collide with the ground or other links, and the contact force of the links should be minimized. It needs to be zero to avoid collisions: ; in, The symbol represents the link that the quadruped robot should avoid colliding with; l represents the l-th link of the quadruped robot; (3) Contact sliding constraint The legs of a quadruped robot should maintain contact with the ground during movement; that is, the speed of the contact leg should be... Zero: ; in, This refers to the four legs of a quadruped robot. This represents the contact state of the i-th leg of the quadruped robot, where 1 indicates contact and 0 indicates no contact. (4) Leg swing height constraint Foot contact force of each leg of the quadruped robot during the leg-swinging phase. It needs to be zero to avoid obstacles: ; (5) Trunk stability constraints The ground reaction force must maintain the stability of the underactuated system during motion, i.e., the trunk stability reference trajectory in the reference trajectory of step 1. for: ; in, It is the acceleration of the quadruped robot's center of mass; It is the location of the quadruped robot's center of mass; s is the weight of the i-th foot contact point of the quadruped robot; i It is the position of the contact point of the i-th leg of the quadruped robot; It is the set of contact points of the legs of a quadruped robot; This represents the acceleration of the quadruped robot's center of mass in the height direction; It is a gravity vector; It is the height of the quadruped robot's center of mass from the ground.
9. The frequency-based deep reinforcement learning-based quadruped robot motion control method according to claim 7, characterized in that, In step 3, the constrained synchronous teacher-student reinforcement learning framework optimizes the synchronous training of a teacher policy that has full access to privileged information through constrained policies, while allowing the student policy using ontology-aware information to minimize reconstruction loss through supervised learning. The overall loss function of this constrained synchronous teacher-student reinforcement learning framework is... for: ; in, For policy networks The objective function of the policy network The teacher and student agents, trained in parallel for both teacher and student strategies, share the same policy network. This policy network is a multilayer perceptron with an exponential linear unit activation function. objective function for: ; in, This represents the pruning loss based on the policy gradient; This represents the clipping loss for each constraint. Policy Network Based on proprioceptive observation Motion estimator Estimated motion characteristics of the output and respectively by privileged encoders Or proprioceptive encoder Generated latent representations To calculate actions Rewarding critics and long-term cost commentators Process complete state And use generalized advantage estimation to evaluate rewards in parallel. and constraints ; By minimizing the following loss function Training rewards critic and long-term cost commentators , The weights and loss function of the critic network. for: ; in, This indicates the expectation value; This represents the actual reward value of the strategy; Represents the estimated complete state Corresponding actions value, This represents the actual constraint violation value. This represents the estimated constraint violation value; Teacher group's intelligent physical fitness access complete status and utilize privileged encoders Extracting latent representations from the complete state The student-group intelligent agent only receives ontology perception observations. and utilizes a body-aware encoder observation sequence Encoding as a latent representation The student-group intelligent agent uses a motion estimator. From the perspective of proprioception Medium estimation features ; Proprioceptive encoder The reconstruction loss function is updated by introducing a reconstruction loss to minimize the difference between privileged and ontology-aware latent representations. for: ; in, Indicates the complete state of using privileges. The calculated environmental characterization; This represents the environmental representation obtained by using a quadruped robot through proprioception prediction. Motion estimator The loss function is trained by minimizing the mean squared error between the estimated value and the true value. for: ; in, Represents the motion state parameters of a quadruped robot in a simulation environment; This represents the motion state parameters predicted by the quadruped robot through body perception.