Soft-hard collaborative navigation method of quadruped robot fusing dynamic scheduling of hardware resources

By constructing a comprehensive observation space and combining a deep reinforcement learning network with the safety constraint processing of an embedded controller, the contradiction between hardware resource scheduling and safety in quadruped robot navigation is resolved. This achieves coordination between dynamic power consumption reduction and stable operation, and improves the environmental adaptability and safety of autonomous navigation.

CN122632869APending Publication Date: 2026-08-25ANHUI SANHEYI INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611132328.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-29
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Existing quadruped robot navigation and hardware scheduling control schemes cannot simultaneously achieve dynamic power consumption reduction scheduling and robot physical operation safety. They suffer from hardware resource scheduling relying on upper-level intelligent strategies and lacking underlying security verification, resulting in high power consumption and unstable operation.

Method used

By acquiring multi-source heterogeneous state data in real time, a comprehensive observation space is constructed. The deep reinforcement learning upper-layer policy network outputs continuous motion instructions and candidate discrete hardware scheduling instructions. The embedded controller calculates the physical risk index for safety constraint processing, ensuring that hardware resources remain safe during dynamic scheduling.

Benefits of technology

It enables intelligent dynamic scheduling of hardware resources for quadruped robots in dynamic environments, reducing unnecessary consumption, improving battery life, and restoring perception and computing capabilities in the event of instability, thus ensuring operational safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122632869A_ABST
    Figure CN122632869A_ABST
Patent Text Reader

Abstract

The present application provides a kind of four-legged robot soft and hard collaborative navigation method of fusion hardware resource dynamic scheduling, storage medium and four-legged robot, it is related to four-legged robot autonomous navigation control field.The present application first real-time acquisition four-legged robot multi-source heterogeneous state data, fusion environment, body movement, foot end contact and hardware operation information are constructed comprehensive observation space, through the upper policy network of deep reinforcement learning single inference synchronous output motion instruction and hardware scheduling instruction, using environment and hardware working condition dynamic regulation and control perception, computing power equipment, reduce resource consumption, improve endurance;At the same time, through the real-time calculation of physical risk index by embedded controller to check scheduling instruction, when risk is over limit and candidate instruction is non full power gear, intercept original instruction and forced switching full power mode, restore complete sensing and computing ability.The method effectively balances hardware consumption and robot motion stability, improves the environmental adaptability and operation safety of four-legged robot autonomous navigation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of autonomous navigation and control technology for quadruped robots, and specifically to a software-hardware collaborative navigation method for quadruped robots that integrates dynamic scheduling of hardware resources. Background Technology

[0002] Quadruped robots, with their excellent adaptability to complex terrain, are widely used in autonomous operations in unstructured scenarios such as outdoor inspection, emergency search and rescue, and field exploration. As the intelligence level of robots continues to improve, the application of technologies such as multi-sensor fusion perception and deep learning strategy reasoning has significantly improved navigation accuracy. However, this has also brought about problems such as high-frequency operation of sensing devices and continuous full-power operation of computing platforms, resulting in high overall power consumption and limited battery life. This severely restricts the ability of quadruped robots to perform autonomous operations over long periods and large areas in the field. Therefore, it is urgent to achieve intelligent dynamic scheduling of hardware resources to balance navigation performance and power consumption.

[0003] In related technologies, most mainstream quadruped robot autonomous navigation solutions currently employ an architecture that combines upper-level intelligent decision-making with lower-level motion control, leveraging artificial intelligence algorithms such as reinforcement learning and deep learning to achieve environmentally adaptive navigation control. Conventional solutions typically construct observation dimensions by collecting robot environmental perception data and body motion state data, then inputting this data into an intelligent policy network after feature fusion, outputting corresponding body motion control commands to drive the robot to complete autonomous navigation. Meanwhile, some existing technologies attempt to improve overall power consumption by adjusting the sampling frequency of sensing devices and the operating frequency of the computing platform at fixed levels based on simple environmental conditions or operating states. By reducing the hardware load under unnecessary conditions, basic power consumption optimization control is achieved, thereby alleviating the problem of high power consumption during robot operation.

[0004] However, existing quadruped robot navigation and hardware scheduling control schemes have significant technical shortcomings, failing to simultaneously address the core requirements of dynamic power consumption reduction scheduling and robot physical operational safety. On one hand, traditional hardware power consumption scheduling methods mostly rely on fixed logic adjustments, achieving only simple, static hardware parameter adaptation. They cannot combine real-time environmental characteristics and the robot's operating state for dynamic, refined hardware resource scheduling, resulting in limited power consumption reduction and poor adaptability. On the other hand, existing solutions generally lack independent underlying security verification and fallback protection mechanisms. Hardware resource scheduling decisions depend entirely on upper-level intelligent strategy outputs, without real-time risk verification of scheduling instructions. When the robot is in dangerous conditions such as slipping, tilting, or instability, the low-power hardware scheduling instructions output by the upper-level strategy are prone to sensor failure and insufficient computing power, further exacerbating the risk of robot motion instability and loss of walking control. This fails to ensure the robot's underlying physical execution safety while achieving dynamic power consumption reduction, creating a technical contradiction between power consumption optimization and operational safety. Summary of the Invention

[0005] (a) Technical problems to be solved To address the shortcomings of existing technologies, this invention provides a software-hardware collaborative navigation method for quadruped robots that integrates dynamic scheduling of hardware resources, solving the technical problem of balancing dynamic energy-saving scheduling of hardware resources with underlying motion safety.

[0006] (II) Technical Solution To achieve the above objectives, the present invention provides the following technical solution: In a first aspect, the present invention provides a method for software-hardware cooperative navigation of a quadruped robot that integrates dynamic scheduling of hardware resources, comprising: Real-time acquisition of multi-source heterogeneous state data of quadruped robots, including environmental perception information, robot body motion information, foot contact information, and hardware operation status information; Spatiotemporal alignment and feature extraction are performed on multi-source heterogeneous state data to construct a comprehensive observation space that jointly represents the robot's motion state and hardware operation state; The current comprehensive observation space is input into the deep reinforcement learning upper-layer policy network, and continuous motion instructions and candidate discrete hardware scheduling instructions are output during the same policy inference process; among them, the candidate discrete hardware scheduling instructions are used to coordinate and adjust the working mode of the multimodal environment sensing device and the operating status of the host computer computing board. Candidate discrete hardware scheduling instructions are sent to the embedded controller. The embedded controller calculates the physical risk index, which characterizes the degree of instability risk of the quadruped robot, in real time based on the robot's motion information and foot contact information. Based on the physical risk index, the candidate discrete hardware scheduling instructions are subjected to safety constraint processing to determine the actual discrete hardware scheduling instructions. Specifically, when the physical risk index exceeds the preset safety threshold and the candidate discrete hardware scheduling instruction corresponds to a non-full power operation state, the candidate discrete hardware scheduling instruction is intercepted and the candidate discrete hardware scheduling instruction is overwritten with the actual discrete hardware scheduling instruction corresponding to the full power operation state. The working mode of the multimodal environment sensing device and the running status of the host computer computing board are adjusted based on actual discrete hardware scheduling instructions, and joint control instructions are generated based on continuous motion instructions to drive the joint drive motors of the quadruped robot.

[0007] Preferably, the integrated observation space is a multi-dimensional state vector that integrates environmental features, ontological motion features, and hardware state features; The fused environmental features include the relative coordinates of the target point in the robot coordinate system and the environmental point cloud features after dimensionality reduction. The body motion characteristics include robot body posture angle, body motion speed, joint rotation angle, and foot contact state information; The hardware status characteristics include the current battery remaining power percentage, the current core temperature of the host computer's computing board, and the hardware operating mode actually executed by the multimodal environment sensing device in the previous moment.

[0008] Preferably, the deep reinforcement learning upper-layer policy network adopts a dual-branch output structure, including a shared feature encoding layer, a continuous motion instruction output branch, and a candidate discrete hardware scheduling instruction output branch. The continuous motion instruction output branch and the candidate discrete hardware scheduling instruction output branch generate the continuous motion instruction and the candidate discrete hardware scheduling instruction respectively based on the same encoded feature output by the shared feature encoding layer in the same forward inference process.

[0009] Preferably, the continuous motion command is a continuous control quantity of the robot body, including the target linear velocity of the robot along the first horizontal direction, the target linear velocity along the second horizontal direction, and the target yaw rate in the robot body coordinate system.

[0010] Preferably, both the candidate discrete hardware scheduling instruction and the actual discrete hardware scheduling instruction are discrete control instructions with three-level binding settings. Each setting is associated with the working mode of the multimodal environment sensing device and the operating status of the host computer computing board. The three-level binding settings include a low-power setting, a normal setting, and a full-power setting. The low power consumption level corresponds to the low-frequency scanning mode or standby mode of the multimodal environment sensing device and the reduced frequency operation state of the host computer computing board. The normal gear corresponds to the intermediate frequency scanning mode of the multimodal environment sensing device and the normal operating state of the host computer computing board; The full power level corresponds to the high-frequency scanning mode of the multimodal environment sensing device and the full power operation state of the host computer computing board.

[0011] Preferably, the calculation process of the physical risk index includes: Based on the joint angles and angular velocities of each leg of the quadruped robot, forward kinematics calculations are performed to obtain the velocity of each leg's foot end, and the velocities of each leg's foot end and the actual translational velocity of the robot body are converted to the same coordinate system; Calculate the absolute value of the velocity difference between the velocity of each leg and the actual translational velocity of the fuselage, and sum the absolute values ​​of the velocity differences corresponding to the four legs to obtain the overall slip term; Extract the roll angle and pitch angle of the robot body, calculate the square root of the sum of their squares, and use it as the body tilt term to characterize the overall tilt of the body. The physical risk index is obtained by weighted summation of the overall slip term and the fuselage tilt term.

[0012] Preferably, the security constraint processing includes: When the candidate discrete hardware scheduling instruction corresponds to a low-power level or a normal level, the physical risk index is compared with the preset security threshold. If the physical risk index exceeds the preset security threshold, the candidate discrete hardware scheduling instruction is intercepted, the actual discrete hardware scheduling instruction is set to full power, and the hardware protection interruption flag is sent back to the deep reinforcement learning upper-layer policy network. If the physical risk index does not exceed the preset security threshold, then the candidate discrete hardware scheduling instruction is determined as the actual discrete hardware scheduling instruction.

[0013] Preferably, the deep reinforcement learning upper-layer policy network uses a composite reward function to complete iterative training, and the total reward value in a single time step includes a navigation guidance reward, a hardware resource consumption penalty, and a bottom-layer overwrite penalty. The navigation guidance reward items include positive rewards for approaching the target, positive rewards for reaching the target, negative penalties for obstacle collisions, and negative penalties for navigation timeouts; The hardware resource consumption penalty is calculated based on the estimated real-time power consumption of the multimodal environment sensing device and the host computer computing board corresponding to the actual discrete hardware scheduling instruction, as well as the weighted calculation of the degree of temperature exceeding the limit of the host computer computing board. The underlying overwrite penalty is a fixed negative reward. When the candidate discrete hardware scheduling instruction is overwritten by the embedded controller, the fixed negative reward is applied to the total reward value.

[0014] In a second aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein, when executed by a processor, the computer program causes the processor to perform the quadruped robot soft-hard cooperative navigation method as described above.

[0015] Thirdly, the present invention provides a quadruped robot, including a robot body, joint drive motors, multi-source heterogeneous sensing devices, and a heterogeneous hierarchical computing platform. The multi-source heterogeneous sensing devices include at least a multimodal environmental sensing device for acquiring environmental sensing information, and the heterogeneous hierarchical computing platform includes an upper-level computing board and a lower-level embedded MCU. The host computer computing board is used for: Real-time acquisition of multi-source heterogeneous state data of quadruped robots, including environmental perception information, robot body motion information, foot contact information, and hardware operation status information; Spatiotemporal alignment and feature extraction are performed on multi-source heterogeneous state data to construct a comprehensive observation space that jointly represents the robot's motion state and hardware operation state; The current comprehensive observation space is input into the deep reinforcement learning upper-layer policy network. During the same policy inference process, continuous motion instructions and candidate discrete hardware scheduling instructions are output and sent down to the underlying embedded MCU. Among them, the candidate discrete hardware scheduling instructions are used to coordinate and adjust the working mode of the multimodal environment sensing device and the operating status of the upper computer computing board. The underlying embedded MCU is used for: The physical risk index, which characterizes the degree of instability risk of the quadruped robot, is calculated in real time based on the robot's motion information and foot contact information. The candidate discrete hardware scheduling instructions are then subjected to safety constraints based on the physical risk index to determine the actual discrete hardware scheduling instructions. Specifically, when the physical risk index exceeds a preset safety threshold and the candidate discrete hardware scheduling instructions correspond to a non-full-power operating state, the candidate discrete hardware scheduling instructions are intercepted, and the actual discrete hardware scheduling instructions are set to the discrete hardware scheduling instructions corresponding to the full-power operating state. It also adjusts the working mode of the multimodal environment sensing device and the running status of the host computer computing board based on actual discrete hardware scheduling instructions, and generates joint control instructions based on continuous motion instructions to drive the joint drive motor.

[0016] (III) Beneficial Effects This invention provides a hardware-software cooperative navigation method for quadruped robots that integrates dynamic scheduling of hardware resources. Compared with existing technologies, it has the following advantages: This invention collects multi-source heterogeneous state data of a quadruped robot in real time, integrates environmental perception information, robot motion information, foot contact information, and hardware operating status information to construct a comprehensive observation space, and outputs continuous motion commands and candidate discrete hardware scheduling commands in the same policy inference process through a deep reinforcement learning upper-level policy network based on the current operating state. This dynamically adjusts the working mode of the multimodal environmental perception device and the operating state of the host computer computing board based on the robot's environment and its own hardware operating status, reducing unnecessary hardware resource consumption and improving overall battery life. Simultaneously, the embedded controller calculates the physical risk index in real time based on the robot's motion information and foot contact information, and performs safety constraint processing on the candidate discrete hardware scheduling commands. When the physical risk index exceeds a preset safety threshold and the candidate discrete hardware scheduling command corresponds to a non-full-power operating state, the candidate discrete hardware scheduling command is intercepted, and the actual discrete hardware scheduling command is set to a full-power operating state, enabling the multimodal environmental perception device and the host computer computing board to restore their corresponding perception and computing capabilities when the robot's instability risk increases. Therefore, this invention can coordinate the dynamic reduction of hardware resources with the stability of robot operation, thereby improving the environmental adaptability and operational safety of quadruped robots during autonomous navigation. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 A flowchart illustrating a hardware-software cooperative navigation method for a quadruped robot that integrates dynamic scheduling of hardware resources, provided in an embodiment of the present invention; Figure 2 This is a schematic diagram illustrating the control architecture and information interaction relationship between a heterogeneous computing platform for upper computers, an embedded real-time control layer for lower computers, and a quadruped robot body, as provided in an embodiment of the present invention. Figure 3 The training convergence curve of a policy network provided in an embodiment of the present invention under hardware power consumption penalty and underlying interception mechanism. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention are described clearly and completely. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] This application provides a quadruped robot software-hardware collaborative navigation method and a quadruped robot that integrates dynamic scheduling of hardware resources. It solves the problems of continuous high power consumption of multimodal environment perception devices and host computer computing boards during the autonomous navigation of existing quadruped robots, as well as the lack of underlying safety constraints in the upper-level intelligent decision-making model that are in line with the robot's real-time physical stability. It realizes continuous motion control and dynamic scheduling of hardware resources, and provides underlying constraints for hardware resource degradation actions when the risk of robot instability increases.

[0021] The technical solution in this application is to solve the above-mentioned technical problems, and the general idea is as follows: This application integrates the robot's physical state and hardware operating state into the comprehensive observation space of the deep reinforcement learning policy network, and combines continuous motion instructions and discrete hardware scheduling instructions into joint actions, enabling the policy network to simultaneously make navigation motion decisions and hardware resource scheduling decisions based on the current environment, robot motion state, and hardware operating state.

[0022] Furthermore, the configurable host computer computing board is responsible for the construction of the comprehensive observation space and the inference of the policy network. The embedded controller (such as the underlying embedded MCU) calculates the physical risk index at a control frequency higher than that of the upper-level policy network, and performs safety constraint processing on the candidate discrete hardware scheduling instructions output by the policy network, thereby determining the actual discrete hardware scheduling instructions that can be executed. When the physical risk index exceeds the preset safety threshold and the candidate instruction corresponds to a non-full-power operation state, the underlying embedded MCU intercepts the candidate instruction and sets the actual instruction to the full-power level.

[0023] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.

[0024] Please see Figure 1 , Figure 1 The overall flow of the software-hardware cooperative navigation method described in this embodiment is disclosed. The method of this embodiment includes steps S1 to S5, which correspond to multi-source perception and hardware state acquisition, comprehensive observation space construction, policy network decision-making and action output, bottom-level hard constraint fallback verification, and physical execution and hardware scheduling, respectively. Specifically, it includes: S1. Real-time acquisition of multi-source heterogeneous state data of the quadruped robot, including environmental perception information, robot body motion information, foot contact information, and hardware operation status information; S2. Perform spatiotemporal alignment and feature extraction on multi-source heterogeneous state data to construct a comprehensive observation space that jointly represents the robot's motion state and hardware operation state; S3. Input the current comprehensive observation space into the deep reinforcement learning upper-layer policy network, and output continuous motion instructions and candidate discrete hardware scheduling instructions in the same policy inference process; among them, the candidate discrete hardware scheduling instructions are used to coordinate and adjust the working mode of the multimodal environment sensing device and the running status of the host computer computing board. S4. Send candidate discrete hardware scheduling instructions to the embedded controller. The embedded controller calculates the physical risk index, which represents the degree of instability risk of the quadruped robot, in real time based on the robot's motion information and foot contact information. Based on the physical risk index, the candidate discrete hardware scheduling instructions are subjected to safety constraints to determine the actual discrete hardware scheduling instructions. Specifically, when the physical risk index exceeds the preset safety threshold and the candidate discrete hardware scheduling instruction corresponds to a non-full power operation state, the candidate discrete hardware scheduling instruction is intercepted and the candidate discrete hardware scheduling instruction is overwritten with the actual discrete hardware scheduling instruction corresponding to the full power operation state. S5. Adjust the working mode of the multimodal environment sensing device and the running status of the host computer computing board based on actual discrete hardware scheduling instructions, and generate joint control instructions based on continuous motion instructions to drive the joint drive motors of the quadruped robot.

[0025] This embodiment introduces hardware-dimensional features into the state space and action space, and embeds a safety fallback mechanism based on quadruped kinematic hard constraints in the underlying control, thereby enabling the quadruped robot to achieve efficient, low-power, and safe autonomous navigation.

[0026] Figure 2 An exemplary control architecture and information interaction relationship between a host computer heterogeneous computing platform, a slave computer embedded real-time control layer, and a quadruped robot body are disclosed.

[0027] It should be noted in advance that the scope of protection of this invention is not limited to... Figure 2 Limited to the module division, hardware carrier, and one-way data interaction form shown, any equivalent hardware layering architecture and module splitting method that can realize the fusion of multi-source sensing features to construct the observation space, the synchronous output of motion and hardware scheduling instructions by the dual-branch strategy network, and the real-time calculation of risk index by the bottom embedded layer and the execution of hardware instructions for bottom-line interception and verification of layered collaborative control logic, all fall within the protection scope of this invention.

[0028] The following combination Figure 2 The above technical solutions are described in detail below: Specifically, in step S1, multi-source heterogeneous state data of the quadruped robot are acquired in real time, including environmental perception information, robot body motion information, foot contact information, and hardware operating status information.

[0029] Combination Figure 2 The illustrated heterogeneous computing platform for the host computer includes at least multimodal environmental sensing devices for acquiring environmental perception information, such as 3D LiDAR, depth cameras, and RGB-D depth cameras. The robot body can be equipped with an Inertial Measurement Unit (IMU), joint encoders, and foot contact detection components. Hardware operating status can be obtained through the Battery Management System (BMS), the computing board system monitoring interface, and the multimodal environmental sensing device driver interface.

[0030] Environmental perception information can include external point cloud data, depth data, and obstacle distance data; robot body motion information can include body posture angle, body linear velocity, body angular velocity, joint angle, and joint angular velocity; foot contact information can indicate whether each leg of the quadruped robot is in a supporting state; hardware operating status can include remaining battery power, host computer computing board core temperature, hardware operating level of multimodal environmental perception devices, and computing board operating status; due to the different sampling frequencies of each perception device, a timestamp is added to each data and it is entered into a cache queue.

[0031] In step S2, spatiotemporal alignment and feature extraction are performed on the multi-source heterogeneous state data to construct a comprehensive observation space that jointly represents the robot's motion state and hardware operating state. This embodiment breaks through the limitation of traditional reinforcement learning that only relies on the environment and ontology state, and incorporates the hardware boundary into the observation. In one embodiment, the comprehensive observation space is a multi-dimensional state vector that integrates environmental features, ontology motion features and hardware state features. The fused environmental features include the relative coordinates of the target point in the robot coordinate system and the environmental point cloud features after dimensionality reduction. The body motion characteristics include robot body posture angle, body motion speed, joint rotation angle, and foot contact state information; The hardware status characteristics include the current battery remaining power percentage, the current core temperature of the host computer's computing board, and the hardware operating mode actually executed by the multimodal environment sensing device in the previous moment.

[0032] Specifically, time alignment can be achieved by recent timestamp matching, interpolation, caching, or historical frame stacking, and the point cloud, target point, and body state can be transformed to a unified coordinate system through coordinate transformation. The integrated observation space is defined as follows: In the formula, This represents the comprehensive observation space at the current time t; The current time t is the environmental feature vector, which includes the relative coordinates of the target point and the environmental perception latent features after dimensionality reduction by a feature extractor (such as PointNet or CNN). The body motion characteristics at the current time t include fuselage attitude, fuselage linear velocity, fuselage angular velocity, joint angles, joint angular velocities, and foot contact state; The hardware state feature vector is mathematically expressed as: =[ , , ] in, This represents the percentage of battery charge remaining at the current time t. The current core temperature of the host computer's computing board at time t; This refers to the hardware working step actually executed at the previous time, i.e., time t-1. The "previous time" here can be the previous strategy decision time step, which is not the same as the control cycle of the underlying MCU; the superscript T is the transpose symbol.

[0033] Furthermore, to avoid the gradient explosion problem caused by scale differences in multi-source heterogeneous data, the observations can be normalized to their maximum and minimum values. For example, the temperature of the computing board can be mapped to the [-1,1] interval. In addition, to handle the time delay caused by asynchronous sensors, a temporal stacking technique can be used to concatenate the historical observation vectors of the past H frames as input to provide temporal dynamic features.

[0034] In step S3, the current comprehensive observation space is input into the deep reinforcement learning upper-layer policy network, and continuous motion instructions and candidate discrete hardware scheduling instructions are output during the same policy inference process; wherein, the candidate discrete hardware scheduling instructions are used to coordinate and adjust the working mode of the multimodal environment sensing device and the operating status of the host computer computing board. Combination Figure 2 The heterogeneous computing platform shown in the figure adopts an actor-critic architecture for its deep reinforcement learning upper-layer policy network. The training algorithm can be a proximal policy optimization (PPO) algorithm, a soft actor-critic (SAC) algorithm, or other reinforcement learning algorithms suitable for continuous and discrete mixed action spaces.

[0035] More specifically, the deep reinforcement learning upper-layer policy network adopts a dual-branch output structure, including a shared feature encoding layer, a continuous motion instruction output branch, and a candidate discrete hardware scheduling instruction output branch. The continuous motion instruction output branch and the candidate discrete hardware scheduling instruction output branch generate the continuous motion instruction and the candidate discrete hardware scheduling instruction respectively based on the same encoded feature output by the shared feature encoding layer in the same forward inference process.

[0036] The continuous motion command is a continuous control quantity of the robot body, including the target linear velocity of the robot along the first horizontal direction, the target linear velocity along the second horizontal direction, and the target yaw rate in the robot body coordinate system.

[0037] Both the candidate discrete hardware scheduling instructions and the actual discrete hardware scheduling instructions are three-level bound discrete control instructions. Each level is associated with the working mode of the multimodal environment sensing device and the operating status of the host computer computing board. The three-level bound levels include a low-power level, a normal level, and a full-power level. The low power consumption level corresponds to the low-frequency scanning mode or standby mode of the multimodal environment sensing device and the reduced frequency operation state of the host computer computing board. The normal gear corresponds to the intermediate frequency scanning mode of the multimodal environment sensing device and the normal operating state of the host computer computing board; The full power level corresponds to the high-frequency scanning mode of the multimodal environment sensing device and the full power operation state of the host computer computing board.

[0038] Specifically, the upper-layer policy network adopts a multi-head output structure, and its action space is defined as a joint action vector: In the formula, Let be the joint action space at the current time t; This is the continuous motion command at the current time t. This continuous motion command is a continuous control quantity of the fuselage in the fuselage coordinate system, including the target linear velocity along the first horizontal direction of the fuselage. Target linear velocity along the second horizontal direction of the fuselage and target yaw rate The control quantities are output by the continuous motion head through the Tanh activation function and then mapped to the physical extreme range of the robot body motion. The discrete hardware scheduling instructions are defined as a set of categorical variables. ∈{0,1,2}: when When =0 (low power level), the multimodal environment sensing device is instructed to enter low-frequency scanning mode or standby mode, and the host computer computing board operates at reduced frequency; when When =1 (normal setting), the multimodal environment sensing device is instructed to enter the intermediate frequency scanning mode and the host computer computing board operates normally; when When =2 (full power setting), the multimodal environment sensing device is instructed to enter the full power high-frequency scanning mode and the host computer computing board is running at full power.

[0039] In step S4, candidate discrete hardware scheduling instructions are sent to the embedded controller. The embedded controller calculates a physical risk index in real time, which characterizes the degree of instability risk of the quadruped robot, based on the robot's motion information and foot contact information. Based on the physical risk index, the candidate discrete hardware scheduling instructions are subjected to safety constraints to determine the actual discrete hardware scheduling instructions. When the physical risk index exceeds a preset safety threshold and the candidate discrete hardware scheduling instructions correspond to a non-full-power operating state, the candidate discrete hardware scheduling instructions are intercepted and overwritten with the actual discrete hardware scheduling instructions corresponding to the full-power operating state.

[0040] This step is the core mechanism for isolating the high-level black-box model "illusion" from physical entities, combining... Figure 2 The lower-level embedded real-time control layer shown includes an embedded controller (such as a low-level embedded MCU) that includes a state estimator, a hardware intercept verifier, and a low-level PID (Proportional-Integral-Derivative) motion controller.

[0041] The state estimator can employ an extended Kalman filter (EKF) to fuse high-frequency measurement data from the fuselage IMU, joint motion data, and foot contact force feedback to obtain estimates of the actual translational velocity, attitude, and foot state.

[0042] Accordingly, the calculation process of the physical risk index specifically includes: S10. Based on the joint angles and angular velocities of each leg of the quadruped robot, perform forward kinematics calculations to obtain the velocity at the foot of each leg. And the speed of each leg end and the actual translational speed of the fuselage. Transform to the same coordinate system.

[0043] S20. Calculate the absolute value of the velocity difference between the velocity of each leg foot and the actual translational velocity of the fuselage. And sum the absolute values ​​of the speed differences corresponding to the four legs. Obtain the overall slip item; S30, Extracting robot body roll angle With pitch angle Calculate the square root of the sum of their squares. , which is the fuselage tilt term that characterizes the overall tilt of the fuselage; S40. The overall slip term and the fuselage tilt term are weighted and summed to obtain the physical risk index. The calculation formula is as follows: In the formula, α and β are the weighting coefficients corresponding to the overall slip term and the fuselage tilt term, respectively.

[0044] Furthermore, the hardware intercept verifier is used to perform security constraint processing, which in one embodiment includes: When the candidate discrete hardware scheduling instruction corresponds to a low-power level or a normal level, the physical risk index is compared with the preset security threshold. If the physical risk index exceeds the preset security threshold, the candidate discrete hardware scheduling instruction is intercepted, the actual discrete hardware scheduling instruction is set to full power, and the hardware protection interruption flag is sent back to the deep reinforcement learning upper-layer policy network. If the physical risk index does not exceed the preset safety threshold, then the candidate discrete hardware scheduling instruction is determined as the actual discrete hardware scheduling instruction. Specifically, the hardware intercept verifier will use the physical risk index. With preset safety threshold Comparison. When the candidate hardware level is low power or normal and When, intercept candidate instructions and set the actual discrete hardware scheduling instruction to full power level; when When an interception occurs, the candidate instruction is determined as the actual instruction. Simultaneously, the MCU returns a hardware protection interruption flag to the host computer's policy network.

[0045] More specifically, the forced rewrite and interception logic of the embedded controller is quantitatively expressed as follows: When the upper-layer policy network outputs When the instruction is less than 2 (requesting sleep or frequency reduction), the embedded controller performs the following judgment: In the formula, This refers to the actual hardware state that is ultimately allowed to be executed at the underlying level, i.e., the actual discrete hardware scheduling instructions. When When the system detects a situation where the quadruped robot is on rough terrain or at risk of slipping and becoming unstable, "safety priority" takes precedence over "power consumption priority." The embedded controller ignores the sleep request from the upper-layer policy network and forcibly overrides the system to full-power operation. =2), and at the same time, return the hardware protection interruption flag to the upper-layer policy network to trigger the downgrade protection policy.

[0046] To ensure that risk calculation and security constraint processing are executed before hardware scheduling instructions, in one example, the upper-layer policy network inference cycle is 20ms (corresponding to a 50Hz inference frequency), and the lower-layer embedded MCU hardware real-time control cycle is 1ms (corresponding to a 1000Hz control frequency). The lower-layer verification frequency is much higher than the upper-layer policy output frequency, effectively ensuring the real-time performance of the security fallback.

[0047] In step S5, the working mode of the multimodal environment sensing device and the running state of the host computer computing board are adjusted based on the actual discrete hardware scheduling instructions, and joint control instructions are generated based on continuous motion instructions to drive the joint drive motors of the quadruped robot.

[0048] This step involves hardware scheduling and motion control execution: on the one hand, based on actual discrete hardware scheduling instructions... The control interface of the multimodal environment sensing device adjusts the LiDAR rotation speed, emission status, sampling frequency, or depth camera frame rate and exposure mode, and the power management interface of the host computer computing board adjusts the CPU / GPU frequency, power consumption mode, or computing resource quota. On the other hand, such as Figure 2 As shown, continuous motion command The system enters the underlying PID motion controller in the underlying embedded real-time control layer. The underlying motion controller performs gait generation, joint trajectory tracking, and force / position / velocity control to generate the desired joint torque, desired joint position, or desired joint velocity, and drives the execution unit (such as driving the joint drive motor through the joint motor driver).

[0049] The above describes the complete hardware-software collaborative control process for the online navigation operation of the quadruped robot. To enable the deep reinforcement learning upper-layer policy network to autonomously learn hardware scheduling strategies that balance navigation stability and low power consumption, iterative training of the network needs to be completed during the offline simulation phase. In this embodiment, a composite reward function is designed to guide the training convergence for the policy network, as detailed below: Furthermore, the deep reinforcement learning upper-layer policy network described in this embodiment uses a composite reward function to complete iterative training. The total reward value at a single time step includes a navigation guidance reward, a hardware resource consumption penalty, and a bottom-layer overwrite penalty. The navigation guidance reward items include positive rewards for approaching the target, positive rewards for reaching the target, negative penalties for obstacle collisions, and negative penalties for navigation timeouts; The hardware resource consumption penalty is calculated based on the estimated real-time power consumption of the multimodal environment sensing device and the host computer computing board corresponding to the actual discrete hardware scheduling instruction, as well as the weighted calculation of the degree of temperature exceeding the limit of the host computer computing board. The underlying overwrite penalty is a fixed negative reward. When the candidate discrete hardware scheduling instruction is overwritten by the embedded controller, the fixed negative reward is applied to the total reward value.

[0050] Specifically, the total reward value for a single time step can be expressed as: In the formula, This represents the total reward value at the current time t.

[0051] The navigation guidance rewards include: target approach positive reward (a continuous reward that is positively correlated with the change in distance the quadruped robot moves from its current position to the final target point), target arrival positive reward (a positive leap reward given when the quadruped robot reaches the detection range of the final target point), obstacle collision negative penalty (a negative reward given when the sensing device detects that the distance between the quadruped robot and environmental obstacles is less than a preset safe radius), and navigation timeout negative penalty (a negative reward given when the number of steps executed in a single navigation round exceeds the maximum allowable value set by the system).

[0052] This is a penalty item for hardware resource consumption, and its calculation formula is as follows: In the formula, The total power of the sensing and computing system is estimated for the actual discrete hardware scheduling instructions executed at the current time t. For finding the maximum value operator; The temperature observation value of the host computer's computing board; λ1 and λ2 are the preset temperature alarm thresholds; λ1 and λ2 are the corresponding penalty weight factors.

[0053] To override the penalty term at the lower level, when the embedded MCU executes an instruction and intercepts it and sends the hardware protection interruption flag back to the upper-level policy network, the training phase will apply a fixed negative reward (e.g., -50 points) to the current moment. This will guide the policy network to learn the resource scheduling strategy of "sleeping to save power on flat terrain and actively activating full power sensing on complex terrain" to avoid low-power, regular-level scheduling actions that are prone to instability.

[0054] like Figure 3As shown, the initial success rate of the model during the initial training phase was only around 0.51, and the original success rate (dashed line) fluctuated significantly due to the influence of the random strategy. As the number of training steps continued to increase, under the dual constraints of hardware power consumption penalty and the underlying interception and overwriting of negative reward terms, the smoothed average success rate (solid line) showed a stable upward trend. When the total number of training steps reached 389,844, the smoothed navigation success rate of the model converged to 0.883. This curve verifies that the composite reward function designed in this invention can effectively guide the policy network to balance low-power hardware scheduling and aircraft motion safety. The network can autonomously learn to select low-power levels under stable operating conditions and actively avoid energy-saving scheduling behaviors under high-risk instability conditions, ultimately achieving a high navigation task completion rate while taking into account both overall battery life and operational safety.

[0055] Furthermore, the training process in this embodiment includes a virtual hardware state simulation mechanism within a reinforcement learning simulation environment (such as Isaac Gym), dynamically updating the virtual battery level and virtual temperature based on the actions output by the upper-layer policy network; if the action output in the current step... If the value is 0, then in the construction of the comprehensive observation space in the next time step, the data loss phenomenon caused by sensor dormancy in the real physical world is simulated by applying a mask to the point cloud data matrix to zero or adding extreme Gaussian noise.

[0056] Based on the same inventive concept, embodiments of the present invention also provide a computer-readable storage medium having a computer program stored thereon, wherein, when executed by a processor, the computer program causes the processor to perform the quadruped robot soft-hard cooperative navigation method as provided in any embodiment of the present invention.

[0057] In embodiments of the present invention, any combination of one or more storage media may be used. The storage medium may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable computer disk, a hard disk, RAM, ROM, erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in connection with an instruction execution system, apparatus, or device.

[0058] Based on the same inventive concept, this embodiment of the invention also provides a quadruped robot, which serves as the physical hardware carrier for executing the quadruped robot software-hardware cooperative navigation method provided in any embodiment of the invention. This embodiment does not limit the mechanical structure of the quadruped robot, such as its legs, body, and reducer, but only provides a detailed description of the perception and hierarchical heterogeneous computing platform that realizes software-hardware cooperative scheduling, underlying hardware real-time safety fallback, and other features.

[0059] Specifically, the quadruped robot mainly includes a robot body, joint drive motors, multi-source heterogeneous sensing devices, and a heterogeneous hierarchical computing platform. The multi-source heterogeneous sensing devices include at least a multimodal environmental sensing device for acquiring environmental perception information. The heterogeneous hierarchical computing platform includes two layers of hardware: an upper-level computing board and a lower-level embedded MCU. These two layers of hardware interact with each other through a communication link and respectively carry out reinforcement learning policy inference, hard real-time security verification, and low-level motion control functions.

[0060] In a practical application scenario, a commercial quadruped robot can be selected as the basic platform, such as a quadruped robot from a certain brand (e.g., Figure 2 The shown quadruped robot dog), B2 series model, and its original expansion peripherals complete the hardware setup. Figure 2 The illustrated quadruped robot dog's hierarchical control architecture serves as an exemplary carrier. More specifically, the host computer computing board employs an edge AI computing node (such as an embedded computing platform like Nvidia Jetson Orin NX or Nano) mounted on the back of the robot dog. It internally deploys ROS2 nodes and the PyTorch deep learning inference framework, providing an online inference environment for deep reinforcement learning models.

[0061] The deep reinforcement learning upper-layer policy network runs on this edge AI computing node, responsible for receiving and processing massive amounts of environmental point cloud and depth image perception data from the Livox Mid-360 3D LiDAR and Intel RealSense depth camera at high frequency, and based on the aforementioned comprehensive observation space A single inference synchronously calculates continuous motion commands. With candidate discrete hardware scheduling instructions When the policy network outputs scheduling instructions to reduce power consumption, such as... When the value is 0, the host computer computing board performs hardware power reduction control in two ways: First, it calls the official software development kit (SDK) of the multimodal environment perception device and sends a standby control command through Ethernet communication to control the internal motor of the lidar to stop rotating and stop laser emission, entering a low-power sleep state; Second, it calls NVIDIA's official computing power dynamic frequency adjustment tool nvpmodel to actively reduce the operating frequency of the CPU and GPU cores, thereby reducing the overall power consumption and chip heat generation from the hardware level and alleviating the power consumption and high temperature problems caused by long-term navigation.

[0062] To address the risk of instability caused by rugged and slippery terrain and ensure the safe movement of the robot, the heterogeneous hierarchical computing platform also includes an embedded MCU deployed at the bottom layer of the robot, i.e., the motion control board that comes pre-installed with the device. This embedded MCU interacts with the robot's IMU and the 12-channel leg joint motor drivers via two high-speed communication links: EtherCAT real-time industrial bus and CAN device bus. It should be noted that the IMU is an Inertial Measurement Unit used to collect the robot's roll, pitch, and velocity. The robot's attitude sensing hardware is not limited to an IMU; gyroscopes, combined inertial navigation systems, and other sensing components with attitude calculation capabilities can be used as equivalent replacements. This embodiment uses an IMU as a preferred example and does not constitute a limitation on the scope of the sensing hardware of this invention.

[0063] Furthermore, in the actual control link, the host computer computing board sends the hardware scheduling instructions to the underlying embedded MCU via User Datagram Protocol (UDP) or serial communication. The MCU then independently executes the hard real-time fallback mechanism described in step S4 using the written underlying control code. The MCU utilizes its control cycle of up to 1000Hz to calculate the physical risk index of the quadruped robot in real time. Once it is determined that the robot body is at risk of instability or falling, the MCU sends back the actual discrete hardware scheduling instruction to the host computer computing board to switch to the full power level. The host computer computing board then performs the operation of waking up the sensing device and adjusting the computing power to restore full power sensing, and sends back the hardware protection interruption flag bit to the policy network. Simultaneously, it calls the underlying motion control function library to generate anti-slip gait trajectory or control the robot to lower its center of gravity and lie flat, thus avoiding the risk of falling and equipment damage from both the sensing scheduling and body motion layers.

[0064] In summary, compared with existing technologies, it has the following beneficial effects: 1. This invention integrates environmental perception, body motion, foot contact, and hardware operating status to construct a comprehensive observation space, and outputs continuous motion instructions and candidate hardware scheduling instructions through the same policy network, which helps to reduce ineffective hardware power consumption and improve battery life.

[0065] 2. This invention calculates the physical risk index through the underlying embedded MCU and imposes safety constraints on candidate hardware instructions, so that non-full power resource degradation actions are not executed under high physical risk conditions, thereby improving the timeliness of sensing and computing resource recovery.

[0066] 3. This invention converts continuous motion commands into gait, joint trajectory and force / position / velocity control quantities through a bottom-level motion controller, thereby achieving the connection between upper-level navigation decision-making and lower-level joint drive control.

[0067] 4. This invention designs a reward function that includes penalties for hardware resource consumption and penalties for underlying overwriting. This function guides the policy network to actively learn resource scheduling strategies during the simulation training phase, such as "reducing frequency and sleeping to save power in open terrain and avoiding obstacles at full power in complex and rugged terrain". This achieves coordinated optimization of navigation efficiency, hardware lifespan and physical stability.

[0068] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0069] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for software-hardware cooperative navigation of a quadruped robot integrating dynamic scheduling of hardware resources, characterized in that, include: Real-time acquisition of multi-source heterogeneous state data of quadruped robots, including environmental perception information, robot body motion information, foot contact information, and hardware operation status information; Spatiotemporal alignment and feature extraction are performed on multi-source heterogeneous state data to construct a comprehensive observation space that jointly represents the robot's motion state and hardware operation state; The current comprehensive observation space is input into the deep reinforcement learning upper-layer policy network, and continuous motion instructions and candidate discrete hardware scheduling instructions are output during the same policy inference process; among them, the candidate discrete hardware scheduling instructions are used to coordinate and adjust the working mode of the multimodal environment sensing device and the operating status of the host computer computing board. Candidate discrete hardware scheduling instructions are sent to the embedded controller. The embedded controller calculates the physical risk index, which characterizes the degree of instability risk of the quadruped robot, in real time based on the robot's motion information and foot contact information. Based on the physical risk index, the candidate discrete hardware scheduling instructions are subjected to safety constraints to determine the actual discrete hardware scheduling instructions. Specifically, when the physical risk index exceeds the preset safety threshold and the candidate discrete hardware scheduling instruction corresponds to a non-full power operation state, the candidate discrete hardware scheduling instruction is intercepted and the candidate discrete hardware scheduling instruction is overwritten with the actual discrete hardware scheduling instruction corresponding to the full power operation state. The working mode of the multimodal environment sensing device and the running status of the host computer computing board are adjusted based on actual discrete hardware scheduling instructions, and joint control instructions are generated based on continuous motion instructions to drive the joint drive motors of the quadruped robot.

2. The quadruped robot soft-hard cooperative navigation method as described in claim 1, characterized in that, The integrated observation space is a multi-dimensional state vector that integrates environmental features, body motion features, and hardware state features. The fused environmental features include the relative coordinates of the target point in the robot coordinate system and the environmental point cloud features after dimensionality reduction. The body motion characteristics include robot body posture angle, body motion speed, joint rotation angle, and foot contact state information; The hardware status characteristics include the current battery remaining power percentage, the current core temperature of the host computer's computing board, and the hardware operating mode actually executed by the multimodal environment sensing device in the previous moment.

3. The quadruped robot soft-hard cooperative navigation method as described in claim 1, characterized in that, The deep reinforcement learning upper-layer policy network adopts a dual-branch output structure, including a shared feature encoding layer, a continuous motion instruction output branch, and a candidate discrete hardware scheduling instruction output branch. The continuous motion instruction output branch and the candidate discrete hardware scheduling instruction output branch generate the continuous motion instruction and the candidate discrete hardware scheduling instruction respectively based on the same encoded feature output by the shared feature encoding layer in the same forward inference process.

4. The quadruped robot soft-hard cooperative navigation method as described in claim 1, characterized in that, The continuous motion command is a continuous control quantity of the robot body, including the target linear velocity of the robot along the first horizontal direction, the target linear velocity along the second horizontal direction, and the target yaw rate in the robot body coordinate system.

5. The quadruped robot soft-hard cooperative navigation method as described in claim 1, characterized in that, Both the candidate discrete hardware scheduling instructions and the actual discrete hardware scheduling instructions are three-level bound discrete control instructions. Each level is associated with the working mode of the multimodal environment sensing device and the operating status of the host computer computing board. The three-level bound levels include a low-power level, a normal level, and a full-power level. The low power consumption level corresponds to the low-frequency scanning mode or standby mode of the multimodal environment sensing device and the reduced frequency operation state of the host computer computing board. The normal gear corresponds to the intermediate frequency scanning mode of the multimodal environment sensing device and the normal operating state of the host computer computing board; The full power level corresponds to the high-frequency scanning mode of the multimodal environment sensing device and the full power operation state of the host computer computing board.

6. The quadruped robot soft-hard cooperative navigation method as described in claim 1, characterized in that, The calculation process for the physical risk index includes: Based on the joint angles and angular velocities of each leg of the quadruped robot, forward kinematics calculations are performed to obtain the velocity of each leg's foot end, and the velocities of each leg's foot end and the actual translational velocity of the robot body are converted to the same coordinate system; Calculate the absolute value of the velocity difference between the velocity of each leg and the actual translational velocity of the fuselage, and sum the absolute values ​​of the velocity differences corresponding to the four legs to obtain the overall slip term; Extract the roll angle and pitch angle of the robot body, calculate the square root of the sum of their squares, and use it as the body tilt term to characterize the overall tilt of the body. The physical risk index is obtained by weighted summation of the overall slip term and the fuselage tilt term.

7. The quadruped robot soft-hard cooperative navigation method as described in claim 5, characterized in that, The security constraint processing includes: When the candidate discrete hardware scheduling instruction corresponds to a low-power level or a normal level, the physical risk index is compared with the preset security threshold. If the physical risk index exceeds the preset security threshold, the candidate discrete hardware scheduling instruction is intercepted, the actual discrete hardware scheduling instruction is set to full power, and the hardware protection interruption flag is sent back to the deep reinforcement learning upper-layer policy network. If the physical risk index does not exceed the preset security threshold, then the candidate discrete hardware scheduling instruction is determined as the actual discrete hardware scheduling instruction.

8. The quadruped robot soft-hard cooperative navigation method as described in claim 1, characterized in that, The deep reinforcement learning upper-layer policy network uses a composite reward function to complete iterative training. The total reward value at a single time step includes a navigation guidance reward, a hardware resource consumption penalty, and a bottom layer overwrite penalty. The navigation guidance reward items include positive rewards for approaching the target, positive rewards for reaching the target, negative penalties for obstacle collisions, and negative penalties for navigation timeouts; The hardware resource consumption penalty is calculated based on the estimated real-time power consumption of the multimodal environment sensing device and the host computer computing board corresponding to the actual discrete hardware scheduling instruction, as well as the weighted calculation of the degree of temperature exceeding the limit of the host computer computing board. The underlying overwrite penalty is a fixed negative reward. When the candidate discrete hardware scheduling instruction is overwritten by the embedded controller, the fixed negative reward is applied to the total reward value.

9. A computer-readable storage medium, characterized in that, It stores a computer program, wherein when executed by a processor, the computer program causes the processor to perform the quadruped robot soft-hard cooperative navigation method as described in any one of claims 1 to 8.

10. A quadruped robot, characterized in that, The system includes a robot body, joint drive motors, multi-source heterogeneous sensing devices, and a heterogeneous hierarchical computing platform. The multi-source heterogeneous sensing devices include at least a multimodal environmental sensing device for acquiring environmental sensing information, and the heterogeneous hierarchical computing platform includes a host computer computing board and a bottom-level embedded MCU. The host computer computing board is used for: Real-time acquisition of multi-source heterogeneous state data of quadruped robots, including environmental perception information, robot body motion information, foot contact information, and hardware operation status information; Spatiotemporal alignment and feature extraction are performed on multi-source heterogeneous state data to construct a comprehensive observation space that jointly represents the robot's motion state and hardware operation state; The current comprehensive observation space is input into the deep reinforcement learning upper-layer policy network. During the same policy inference process, continuous motion instructions and candidate discrete hardware scheduling instructions are output and sent down to the underlying embedded MCU. Among them, the candidate discrete hardware scheduling instructions are used to coordinate and adjust the working mode of the multimodal environment sensing device and the operating status of the upper computer computing board. The underlying embedded MCU is used for: The physical risk index, which characterizes the degree of instability risk of the quadruped robot, is calculated in real time based on the robot's motion information and foot contact information. The candidate discrete hardware scheduling instructions are then subjected to safety constraints based on the physical risk index to determine the actual discrete hardware scheduling instructions. Specifically, when the physical risk index exceeds a preset safety threshold and the candidate discrete hardware scheduling instructions correspond to a non-full-power operating state, the candidate discrete hardware scheduling instructions are intercepted, and the actual discrete hardware scheduling instructions are set to the discrete hardware scheduling instructions corresponding to the full-power operating state. It also adjusts the working mode of the multimodal environment sensing device and the running status of the host computer computing board based on actual discrete hardware scheduling instructions, and generates joint control instructions based on continuous motion instructions to drive the joint drive motor.