Hand-foot multiplexing under-actuated hexapod robot based on multi-agent reinforcement learning and control method

By employing multi-agent reinforcement learning and underactuated design, a hexapod robot was developed that addresses the shortcomings of hexapod robots with reusable hands and feet in terms of coordination, motion optimization, and terrain adaptability. This approach enables efficient integrated control of movement and manipulation, improving task success rate and the ability to adapt to complex scenarios.

CN121626320APending Publication Date: 2026-03-10NANJING TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-02
Publication Date
2026-03-10

Smart Images

  • Figure CN121626320A_ABST
    Figure CN121626320A_ABST
Patent Text Reader

Abstract

The invention discloses a hand-foot multiplexing under-actuated hexapod robot based on multi-agent reinforcement learning and a motion control method, and belongs to the technical field of intelligent robots, the hand-foot multiplexing under-actuated hexapod robot comprises a control platform main body, a power supply end cover is arranged at the top end of the control platform main body, and a multi-agent control unit is arranged in the control platform main body; the under-actuated leg joint is located outside the control platform main body; the micro spine acupuncture claw joints are located at the ends, away from the control platform body, of the under-actuated leg joints and used for being matched with the under-actuated leg joints for use; the multi-agent control unit is connected with the under-actuated leg joint and the micro-spine acupuncture claw joint and used for outputting a leg-claw coordination action instruction; and the power unit is arranged in the control platform main body. Through collaborative learning and task allocation among the intelligent agents, the problem of complex mobile operation integration can be better solved, and efficient action optimization and dynamic adjustment are realized under different terrain and task requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent robot technology, and more specifically, to a hexapod robot with hand and foot reuse based on multi-agent reinforcement learning and its control method. Background Technology

[0002] With the widespread application of special-purpose robots in emergency rescue, field exploration, and post-disaster reconstruction, the operational requirements in complex scenarios are becoming increasingly stringent. Robots must not only possess stable terrain adaptability but also achieve efficient integrated mobility and manipulation. In tasks such as retrieving survivors from rubble or collecting samples in jungles, robots must not only perform efficient terrain movement but also possess precise manipulation capabilities. Therefore, the ability to reuse limbs has become a crucial design goal for current special-purpose robots. Existing hexapod robots with reusable limbs face numerous bottlenecks in achieving this goal.

[0003] Insufficient leg-grip coordination is a significant issue. In traditional designs, leg movement and end effector manipulation are controlled separately. The legs rely primarily on a pre-defined gait library for support and swing, while the gripper operates on a single motor. This lack of dynamic coordination means that in rubble environments, when a robot needs to grasp suspended objects, if the legs fail to adapt to the gripper's operating angle, the robot may tilt, preventing the gripper from accurately contacting the target object and leading to mission failure. Therefore, enabling dynamic coordination between leg and gripper movements is crucial for improving the robot's overall performance.

[0004] Optimizing complex motion spaces is challenging, as hand-foot reuse involves a hybrid motion space of discrete and continuous actions. Specifically, discrete actions include leg support and swinging patterns, and claw grasping and releasing patterns, while continuous actions include leg center of gravity adjustment (COM adjustment), claw opening and closing angles, and operational force. Traditional control methods, such as rule-based control or single-agent reinforcement learning, struggle to balance the coupling and real-time performance of various action dimensions, often leading to action conflicts. For example, claw grasping actions may conflict with leg swinging actions, resulting in uncoordinated robot movements or low optimization efficiency. Therefore, ensuring smooth and coordinated movements while optimizing the motion space is crucial to solving this problem.

[0005] Poor terrain-task adaptability is a problem in existing hand-foot reusable robot designs, which often fail to adequately consider the close relationship between terrain features and task requirements. For example, on soft sand, the robot may be unable to effectively grip heavy objects due to insufficient leg support. In complex rubble or sloping terrain, the gripping force of the claws may also be insufficient, preventing the robot from accurately completing the task. Current designs fail to incorporate terrain features (rubble gaps, slope angles) and operational requirements (gripping force, operational accuracy) into a unified decision-making framework, significantly reducing the robot's performance in different terrain environments. Furthermore, existing designs lack the ability to dynamically adjust for movement-operation transitions, typically requiring response times exceeding 500ms, which is far from meeting the real-time requirements of complex scenarios.

[0006] No effective solutions have yet been proposed to address the problems in the relevant technologies. Summary of the Invention

[0007] To address the shortcomings of existing technologies, this invention proposes a hand-foot reuse underactuated hexapod robot and motion control method based on multi-agent reinforcement learning. Each leg and the distal micro-spine needle claw joint constitute an independent agent. The hybrid action space is optimized through a deep fusion algorithm of hybrid action space and multi-agent collaboration. The feasibility model of free gait transfer is extended to adapt to the feasibility judgment of hand-foot collaboration, realizing integrated intelligent control of movement and operation, and solving the problems of poor collaboration and difficulty in motion optimization in traditional designs.

[0008] According to one aspect of the present invention, a hand-foot reuse underactuated hexapod robot based on multi-agent reinforcement learning is provided, comprising:

[0009] The control platform consists of a main body with a power supply end cap at its top and a built-in multi-agent control unit. Underactuated leg joints are located symmetrically distributed off the main body. Micro-spine needle-claw joints are located at the end of the underactuated leg joints furthest from the main body and are used in conjunction with the underactuated leg joints. The multi-agent control unit connects to the underactuated leg joints and micro-spine needle-claw joints and outputs commands for coordinated leg-claw movements. A power unit, built into the main body, provides power management for the hand-foot reusable underactuated hexapod robot.

[0010] Furthermore, the underactuated leg joint includes a rotary motor connected to the main body of the control platform. A linear motor support base frame is located at the end of the rotary motor away from the main body of the control platform. A linear motor is installed within the linear motor support base frame. A first active link is located at the end of the linear motor away from the rotary motor. The end of the first active link away from the linear motor is connected to one end of a V-shaped transmission link. A second active link is connected to the other end of the V-shaped transmission link. A first passive link is connected to the middle of the V-shaped transmission link. A third active link is connected to the end of the second active link away from the V-shaped transmission link. A second passive link is connected to the end of the first passive link away from the V-shaped transmission link. The middle of the end-effector is connected to the end-effector. The end of the second passive link away from the first passive link is connected to one end of the end-effector. The other end of the end-effector is connected to the micro-spine needle claw joint.

[0011] Furthermore, the micro-spine needle claw joint includes a geared motor mounting housing connected to the underactuated leg joint. The geared motor is installed inside the geared motor mounting housing, and a micro-needle housing is installed on the side of the geared motor mounting housing. A claw housing connecting transmission component is installed between the micro-needle housing and the geared motor mounting housing. A needle claw elastic finger is installed in the middle of the micro-needle housing, and a micro-needle is installed at the end of the needle claw elastic finger.

[0012] Furthermore, the main body of the control platform includes an upper hollow frame connected to the power supply end cover, and a lower hollow frame is provided at the bottom of the upper hollow frame.

[0013] According to another aspect of the present invention, a control method for a hand-foot reuse underactuated hexapod robot based on multi-agent reinforcement learning is also provided, comprising:

[0014] Initialize a hexapod robot with hand-leg multiplexing and underactuated mechanism, and use the legs and claws of the hexapod robot to collect local observation signals respectively; based on the local observation signals, multiple agents independently sample discrete and continuous actions;

[0015] Based on the global state of a discrete action, continuous action, and hand-foot reuse underactuated hexapod robot, a transfer feasibility model for extended free gait is constructed. Using the transfer feasibility model for extended free gait and combined with the constraints of underactuated machinery, the legality of the state transition of leg movement and claw operation is obtained.

[0016] When the state transition is legal, multi-objective reward calculation and network training are performed to update the agent; the agent is then used to execute and dynamically optimize hand-foot reuse tasks.

[0017] Furthermore, multi-agent independent sampling of discrete and continuous actions includes:

[0018] Multi-agent independent sampling includes discrete actions of leg and paw patterns, as well as continuous actions including continuous parameters of the legs and paws.

[0019] Make discrete actions follow a category distribution, and make continuous actions follow a diagonal Gaussian distribution;

[0020] Among them, Q-networks with attention mechanisms are used to calculate the collaborative value of actions.

[0021] Furthermore, based on the global state of discrete actions, continuous actions, and hand-foot reuse underactuated hexapod robots, a feasibility model for extended free gait transfer is constructed, including:

[0022] Obtain the current global state of the underactuated hexapod robot with hand-leg reusability; execute the action at the next moment to obtain the global state at the next moment;

[0023] By utilizing the current and next global states of a hexapod robot with hand-foot reusability, and combining discrete and continuous actions, a transfer feasibility model for extended free gait is obtained.

[0024] Furthermore, the constraints for underactuated machines include:

[0025] Examine the leg-claw pattern combination constraints of an underactuated hexapod robot with hand-foot reusability;

[0026] Based on the physical limits of the underactuated leg joints and claw joints, the leg joint angles, claw opening and closing angles, and claw target distances are checked.

[0027] Based on a single rigid body dynamics model, with the main body of the control platform as the modeling object, a three-dimensional inertial coordinate system with the body center of mass of the underactuated hexapod robot with hand and foot reusable as the origin is adopted. The force balance is analyzed through the laws of rigid body dynamics to determine the dynamic feasibility of leg support and claw operation.

[0028] Furthermore, when the state transition is deemed valid, multi-objective reward calculation and network training are performed to update the agent, including:

[0029] A multi-objective reward function is constructed based on mobile efficiency rewards, operational accuracy rewards, leg-paw coordination rewards, energy consumption penalties, and terminal task rewards; the total multi-objective reward is calculated based on the multi-objective reward function.

[0030] When the state transition is valid, the total multi-objective reward is used as the reward for each agent. Input for network and policy network updates;

[0031] The Q-network is updated by minimizing the Bellman residual, and the policy network is updated by maximizing the cumulative reward with entropy regularization; discrete and continuous actions are automatically adjusted by temperature updates to balance exploration and exploitation.

[0032] For the diagonal Gaussian distribution characteristics of continuous motion, the objective function controls the entropy of the continuous parameters and their fluctuations. When the temperature parameter of continuous motion increases, the adjustment range of the continuous parameters expands. When the temperature parameter of continuous motion decreases, the continuous parameters tend to stabilize.

[0033] Furthermore, leveraging intelligent agents to perform and dynamically optimize hand-foot reuse tasks includes:

[0034] The parameters of the underactuated hexapod robot with reusable hands and feet are dynamically adjusted by using a multi-agent control unit to update the status in real time.

[0035] The multi-level drive unit receives action commands output by the multi-agent control unit to drive each agent to perform limb reuse tasks.

[0036] The beneficial effects of this invention are as follows:

[0037] (1) A leapfrog improvement in the synergy of hand and foot reuse, solving the pain points of traditional separation control:

[0038] While traditional leg-claw mechanisms possess reusability potential, the lack of collaborative algorithms leads to issues such as leg posture deviation during claw attachment and claw operation failure during leg movement. This invention utilizes the CTDE (Centralized Training-Distributed Execution) framework, treating each leg-claw group as an independent agent. A Q-network with attention mechanisms calculates action relevance in real time (prioritizing leg stability during claw attachment), resulting in more precise movement-operation collaborative decision-making. Experimental data shows that in rubble material retrieval tasks, the success rate is increased by over 40% compared to traditional rule-based control (96% in simulated environment vs. 65% traditional), the leg-claw action conflict rate is reduced by 90% (from 30% traditionally to below 3%), and the hand-foot mode switching latency is compressed to within 80ms (traditionally 520ms), resolving the core pain point of asynchronous movement and operation.

[0039] (2) Efficient optimization of complex mixed motion space to adapt to the needs of hand and foot reusable motion:

[0040] Hand and foot reuse involves a high-dimensional hybrid space of discrete leg / claw patterns and continuous COM / grasping force. Traditional rule-based control requires a large amount of pre-set logic, making it difficult to cover all scenarios. This invention employs a deep fusion algorithm of hybrid action space and multi-agent collaboration, modeling discrete actions as category distributions and continuous actions as diagonal Gaussian distributions. Through collaborative optimization of policy networks and Q-networks, end-to-end learning of the action space is achieved. Compared to single-agent algorithms, the training convergence speed is improved by 30% (stable performance can be achieved in 20,000 episodes), real-time inference latency is ≤80ms, and it can dynamically adapt to complex action combinations such as reducing leg thrust and increasing claw gripping force in soft terrain, and increasing leg extension and fine-tuning claw angle in hard terrain, avoiding the optimization stagnation problem caused by action dimension coupling in traditional algorithms.

[0041] (3) The high reliability of the mechanical configuration and algorithm synergy enhances the adaptability to complex terrain:

[0042] On the one hand, the underactuated diamond-shaped closed-loop mechanism can offset more than 80% of the terrain impact load (joint deformation ≤0.5mm when falling from a height of 200mm), and the passive adaptive characteristics reduce the pressure on algorithm compensation. On the other hand, the extended free gait transfer feasibility model can simultaneously determine the feasibility of leg movement and claw operation, that is, check whether the leg support force meets the requirements for grasping heavy objects and whether the claw angle is adapted to the target size, making the state transition safer. In complex terrains such as 30° slopes and rubble piles, the robot's grounding error is ≤3mm, while the traditional design is ≥8mm, and the claw grasping success rate reaches 98%, while the traditional rigid claw is only 60%. Even when faced with local rubble collapse, the algorithm can dynamically adjust the leg-claw movement, that is, switch two swinging legs for support and increase the claw gripping force to avoid body tilting or task interruption.

[0043] (4) Optimize the balance between energy consumption and cost, taking into account both practicality and economy:

[0044] In terms of energy consumption, this underactuated design requires only two drive units per leg: a linear motor and a rotary motor, reducing drive redundancy by 50% compared to fully actuated legs. Combined with the energy-optimized algorithm's reward system (penalizing overcurrent actions), the overall continuous operating energy consumption is reduced by 38%. On a full charge, the robot can operate for ≥4 hours in low-speed detection mode (0.3 m / s) and ≥2 hours in heavy-load transport mode (2 kg load, 0.15 m / s), far exceeding the endurance of traditional fully actuated reusable robots (≤1.5 hours). Regarding cost, the mechanical components utilize low-cost composite materials, namely ABS plastic and 6061-T6 aluminum alloy, eliminating the need for complex reducers and reducing mass production costs by 25% compared to traditional fully actuated robots. Furthermore, the algorithm is developed based on an open-source framework, requiring no additional hardware investment, thus balancing economic efficiency and practicality for industrial applications.

[0045] (5) Strong robustness and task adaptability to meet the diverse needs of emergency scenarios:

[0046] The multi-agent architecture of this invention possesses inherent redundancy. When a single claw fails (microneedle breakage) or a single leg malfunctions (motor overload), adjacent agents can quickly compensate (claw 2 / 6 increases gripping force when claw 1 fails, and leg 2 / 4 switches to support when leg 3 fails). The task success rate after a single component failure remains ≥82% (compared to ≤50% in traditional designs). Furthermore, the algorithm supports the entire process of movement-grabbing-transporting-releasing, adapting to different weights of materials (0.5-2kg) and terrain types (sand, rubble, slopes) without retraining. It demonstrates high adaptability in scenarios such as rubble rescue (grabbing survivors, transporting tools), jungle exploration (collecting samples), and mountain transportation (transferring supplies), providing a reliable technical solution for the emergency applications of special robots. Attached Figure Description

[0047] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0048] Figure 1 This is a three-dimensional structural schematic diagram of a hexapod robot with hand and foot reuse based on multi-agent reinforcement learning according to an embodiment of the present invention.

[0049] Figure 2 This is a complete three-dimensional structural view of an underactuated arm of an underactuated hexapod robot based on multi-agent reinforcement learning according to an embodiment of the present invention.

[0050] Figure 3 This is a complete structural view of the underactuated joint of an underactuated hexapod robot based on multi-agent reinforcement learning according to an embodiment of the present invention.

[0051] Figure 4 This is a schematic diagram of the micro-spine needle claw at the end of the underactuated arm in a hexapod robot based on multi-agent reinforcement learning according to an embodiment of the present invention.

[0052] Figure 5 This is a schematic diagram of the main body of the control platform in a hexapod robot with hand-foot reuse based on multi-agent reinforcement learning according to an embodiment of the present invention.

[0053] Figure 6 This is a schematic diagram of a power supply end cap in a hand-foot reuse underactuated hexapod robot based on multi-agent reinforcement learning according to an embodiment of the present invention.

[0054] Figure 7 This is a block diagram of the control platform body, multi-agent control unit, and power unit in a hand-foot reuse underactuated hexapod robot based on multi-agent reinforcement learning according to an embodiment of the present invention.

[0055] Figure 8 This is one of the overall motion control flowcharts of a control method for a hand-foot reuse underactuated hexapod robot based on multi-agent reinforcement learning according to an embodiment of the present invention;

[0056] Figure 9 This is the second overall motion control flowchart of a control method for a hand-foot reuse underactuated hexapod robot based on multi-agent reinforcement learning according to an embodiment of the present invention;

[0057] Figure 10 This is a flowchart of a control method for an underactuated hexapod robot based on multi-agent reinforcement learning according to an embodiment of the present invention.

[0058] In the picture:

[0059] 1. Power supply end cap; 2. Control platform main body; 201. Upper hollow frame; 202. Lower hollow frame; 3. Underactuated leg joint; 301. Rotary motor; 302. Linear motor; 303. Linear motor support base frame; 304. First active link; 305. V-shaped transmission link; 306. Second active link; 307. First passive link; 308. Third active link; 309. Second passive link; 310. End effector link; 4. Micro-spine needle claw joint; 401. Gear motor; 402. Gear motor mounting housing; 403. Claw housing connecting transmission component; 404. Microneedle housing; 405. Needle claw elastic finger; 406. Microneedle; 5. Multi-agent control unit; 6. Power unit. Detailed Implementation

[0060] To further illustrate the various embodiments, the present invention provides accompanying drawings, which are part of the disclosure of the present invention. These drawings are mainly used to illustrate the embodiments and can be used in conjunction with the relevant descriptions in the specification to explain the operating principles of the embodiments. With reference to these drawings, those skilled in the art should be able to understand other possible implementation methods and the advantages of the present invention.

[0061] According to embodiments of the present invention, an underactuated hexapod robot with reusable limbs and a motion control method are provided. A composite solution integrating multi-agent reinforcement learning (MARL) is employed, breaking through the limitations of traditional limb-reusable designs through deep coupling of mechanical configuration and intelligent algorithms. Through collaborative learning and task allocation among agents, complex mobility-operation integration problems can be better handled, and efficient motion optimization and dynamic adjustment can be achieved under different terrains and task requirements.

[0062] The present invention will now be further described in conjunction with the accompanying drawings and specific embodiments, such as... Figures 1-7 As shown, according to one embodiment of the present invention, an underactuated hexapod robot with reusable hands and feet is provided, comprising:

[0063] The control platform body 2 has a power supply end cap 1 at its top and a multi-agent control unit 5 built into it. Underactuated leg joints 3 are located symmetrically distributed outside the control platform body 2. Micro-spine needle-claw joints 4 are located at the end of the underactuated leg joints 3 furthest from the control platform body 2 and are used in conjunction with the underactuated leg joints 3. The multi-agent control unit 5 is connected to the underactuated leg joints 3 and the micro-spine needle-claw joints 4 and is used to output commands for coordinated leg-claw movements. A power unit 6 is built into the control platform body 2 and is used to provide power management for the hand-foot reusable underactuated hexapod robot.

[0064] In one embodiment, the underactuated leg joint 3 includes a rotary motor 301 connected to the control platform body 2. A linear motor support frame 303 is provided at one end of the rotary motor 301 away from the control platform body 2. A linear motor 302 is disposed within the linear motor support frame 303. A first active link 304 is provided at one end of the linear motor 302 away from the rotary motor 301. The end of the first active link 304 away from the linear motor 302 is connected to one end of a V-shaped transmission link 305. The other end of the V-shaped transmission link 305 is connected to a second active link 306. The middle of the connecting rod 305 is connected to the first passive connecting rod 307; the end of the second active connecting rod 306 away from the V-shaped transmission connecting rod 305 is connected to the third active connecting rod 308; the end of the first passive connecting rod 307 away from the V-shaped transmission connecting rod 305 is connected to the second passive connecting rod 309; the end of the third active connecting rod 308 away from the second active connecting rod 306 is connected to the middle of the end-acting connecting rod 310; the end of the second passive connecting rod 309 away from the first passive connecting rod 307 is connected to one end of the end-acting connecting rod 310; and the other end of the end-acting connecting rod 310 is connected to the micro-spine needle claw joint 4.

[0065] In one embodiment, the micro-spine needle claw joint 4 includes a geared motor mounting housing 402 connected to the underactuated leg joint 3. A geared motor 401 is disposed inside the geared motor mounting housing 402. A micro-needle housing 404 is disposed on the side of the geared motor mounting housing 402. A claw housing connecting transmission member 403 is disposed between the micro-needle housing 404 and the geared motor mounting housing 402. A needle claw elastic finger 405 is disposed in the middle of the micro-needle housing 404, and a micro-needle 406 is disposed at the end of the needle claw elastic finger 405.

[0066] In one embodiment, the control platform body 2 includes an upper hollow frame 201 connected to the power supply end cover 1, and a lower hollow frame 202 is provided at the bottom of the upper hollow frame 201.

[0067] like Figure 10 As shown, according to another embodiment of the present invention, a control method for a hexapod robot with hand-foot reuse based on multi-agent reinforcement learning is also provided, comprising:

[0068] S1. Initialize the hand-leg multiplexing underactuated hexapod robot, and use the legs and claws of the hand-leg multiplexing underactuated hexapod robot to collect local observation signals respectively; based on the local observation signals, multiple agents independently sample discrete actions and continuous actions.

[0069] S2. Based on the global state of the underactuated hexapod robot with discrete actions, continuous actions, and hand-foot reuse, construct a transfer feasibility model for extended free gait. Using the transfer feasibility model for extended free gait and combined with the constraints of underactuated machinery, obtain the legality of the state transition of leg movement and claw operation.

[0070] S3. When the state transition is legal, multi-objective reward calculation and network training are performed to update the agent; the agent is then used to execute and dynamically optimize the hand-foot reuse task.

[0071] In one embodiment, multi-agent independent sampling of discrete and continuous actions includes:

[0072] Multi-agent independent sampling includes discrete actions in leg and paw patterns, as well as continuous actions including continuous parameters of the legs and paws; discrete actions are made to follow a class distribution, and continuous actions are made to follow a diagonal Gaussian distribution; among them, a Q-network with attention mechanism is used to calculate the collaborative value of actions.

[0073] In one embodiment, based on the global state of a discrete action, continuous action, and a hexapod robot with hand-foot reuse, a feasibility model for extending free gait is constructed, including:

[0074] Obtain the global state of the underactuated hexapod robot with hand-foot reusability at the current moment; execute an action at the next moment to obtain the global state at the next moment; use the current and next global states of the underactuated hexapod robot with hand-foot reusability, and combine discrete and continuous actions to obtain a transfer feasibility model for extended free gait.

[0075] In one embodiment, the constraints of the underactuated machine include:

[0076] The constraints of the leg-claw mode combination of the underactuated hexapod robot with hand-foot reusability are checked; based on the physical limits of the underactuated leg joints and claw joints, the leg joint angles, claw opening and closing angles, and claw target distances are checked; based on the single rigid body dynamics model, taking the main body 2 of the control platform as the modeling object, a three-dimensional inertial coordinate system with the body center of mass of the underactuated hexapod robot with hand-foot reusability as the origin is adopted, and the force balance is analyzed through the laws of rigid body dynamics to determine the dynamic feasibility of leg support and claw operation actions.

[0077] In one embodiment, when the state transition is legal, multi-objective reward calculation and network training are performed to update the agent, including:

[0078] A multi-objective reward function is constructed based on mobility efficiency rewards, operational accuracy rewards, leg-paw coordination rewards, energy consumption penalties, and terminal task rewards. The total multi-objective reward is calculated according to this function. When the state transition is valid, the total multi-objective reward is used as the reward for each agent. The inputs for updating the network and policy network are: updating the Q-network by minimizing the Bellman residual and updating the policy network by maximizing the cumulative reward with entropy regularization; automatically adjusting discrete and continuous actions through temperature updates to balance exploration and exploitation; for the diagonal Gaussian distribution of continuous actions, the objective function controls the entropy of the continuous parameters and their volatility; when the temperature parameter of the continuous action increases, the adjustment range of the continuous parameters expands; when the temperature parameter of the continuous action decreases, the continuous parameters tend to stabilize.

[0079] In one embodiment, using an intelligent agent to perform and dynamically optimize hand-foot reuse tasks includes:

[0080] The multi-agent control unit (5) is used to update the status in real time and dynamically adjust the parameters of the underactuated hexapod robot with hand and foot reuse; the multi-level drive unit receives the action instructions output by the multi-agent control unit (5) to drive each agent to perform the hand and foot reuse task.

[0081] To facilitate understanding of the above technical solutions of the present invention, the working principle of the present invention in actual process will be described in detail below.

[0082] A hand-foot reusable underactuated hexapod robot and motion control method are proposed for complex unstructured environments (ruin rescue, jungle exploration, mountain material transportation and emergency material grasping). In particular, it addresses the core pain points of poor coordination between robot leg movement and end-effector operation in extreme scenarios and the difficulty in optimizing complex motion space. By integrating multi-agent reinforcement learning (MARL) with underactuated mechanical configuration, it achieves intelligent coordination of leg-claw movements, thereby improving the robot's adaptability and reliability in complex tasks involving movement and manipulation.

[0083] The underactuated hexapod robot with reusable limbs has overall dimensions of 600mm × 600mm × 240mm (length × width × height) and a weight of ≤5kg. It adopts a modular assembly design, with components connected via quick-release interfaces (M3 high-strength bolts) for easy maintenance and component replacement. It includes a control platform body 2, 6 sets of leg-claw intelligent agents, a multi-level drive unit, an intelligent control unit (i.e., a multi-agent control unit 5), and a power unit 6. The 6 sets of leg-claw intelligent agents are symmetrically distributed around the control platform body in a 360° configuration, with adjacent agents forming a 60° angle. Each leg-claw intelligent agent consists of an underactuated leg joint and a micro-spine needle-like claw joint connected in series.

[0084] The main body of the control platform adopts a two-layer nested hexagonal hollow frame structure, with the frame material being 6061-T6 aluminum alloy. The upper hollow frame 201 has an outer circle diameter of 176mm, an inner circle diameter of 152mm, and a frame wall thickness of 5mm. It has 12 evenly distributed leg mounting holes (3mm in diameter) and 6 evenly distributed power supply end cap connection holes (5mm in diameter). The lower hollow frame 202 has the same dimensions as the upper layer and is fixed to the upper hollow frame 201 by 6 pentagonal central support columns. The two ends of the columns are connected to the frame by threads, and the verticality error is ≤0.1mm. The hollowed-out area houses an intelligent control unit. A bowl-shaped top cover, made of ABS engineering plastic with a wall thickness of 4mm, an inner diameter of 136mm, and a height of 47mm, is detachably connected to the top of the upper frame. The top cover and platform form a sealed cavity containing a built-in power unit 6. The intelligent control unit integrates a multi-agent decision-making module, which includes a deep fusion algorithm network of hybrid motion space and multi-agent collaboration, and an extended free gait transfer feasibility model, used to output leg-claw coordinated movement commands. The underactuated leg joint 3 and the distal micro-spine needle-like claw joint 4 are equipped with torque sensors, gripping force sensors, and position encoders, respectively, for collecting local observation signals from the intelligent agent. Combined with… Figures 1-7 The following explanation is provided.

[0085] Function of the hollowed-out area: The hollowed-out area of ​​the lower hollowed-out frame 202 is used to fix the development board of the intelligent control unit through an insulating bracket. For example, for the NVIDIA Jetson Orin NX Super, the bracket and the column are connected by a snap-fit, leaving 10mm of space for cable routing; the top of the upper frame is detachably connected to the bowl-shaped top cover by 6 M5 bolts. Figure 6 The material is ABS engineering plastic, with a wall thickness of 4mm, an inner diameter of 136mm, and a height of 47mm. A waterproof sealing ring is installed between the power supply end cover 1 and the frame to form an IP65 protection-rated cavity, which houses the power unit 6.

[0086] The six legs of the underactuated leg joint 3 are symmetrically distributed around the control platform body 2 in a 360° pattern, with an included angle of 60° between adjacent legs. Each leg is connected to the flange of the rotary motor 301, and the coaxiality error of the flange is ≤0.05mm. The rotary motor 301 is fixed to the control platform body 2 through the M3 mounting holes of the upper and lower frame bodies.

[0087] The underactuated leg joint 3 is composed of a diamond-shaped closed-loop mechanism [7 composite material connecting rods hinged together: first active connecting rod 304 (50mm), second active connecting rod 306 (75mm), third active connecting rod 308 (80mm), V-type transmission connecting rod 305 (55mm), first passive connecting rod 307 (70mm), second passive connecting rod 309 (75mm), end actuation connecting rod 310 (30mm)], linear motor 302 (model LASF30-024D, rated thrust 80N, stroke 30mm, rated voltage 24V), and rotary motor 30 The system consists of a (model FEETECH-SM60CL, rated torque 2 N·m, rated speed 35 RPM, power supply voltage 12V) and a joint torque sensor (range 0-200 N·m, accuracy ±0.5%). The joint torque sensor is installed at the tail hinge of the 303 linear motor support base frame and the 3 under-driven leg joint thigh covering structure. It is used to collect torque feedback during leg support / swing, such as the ground hardness and support stability, as a local observation signal for the intelligent agent; one end of the diamond closed-loop mechanism is fixed to the frame (7075-T6 aluminum alloy material). The leg (110mm×34mm×32mm in size) is fixed to the thigh covering structure (ABS plastic, semi-closed sleeve, 130mm in length, 42mm in width, and 4mm in thickness). The output shaft of the rotary motor 301 is coaxially connected to the thigh covering structure, driving the leg to rotate horizontally around the platform by ±60°. The micro-spine needle claw joint 4 at the end consists of an outer frame (ABS engineering plastic, 38mm×30mm in size), zigzag needle claw elastic fingers 405 (Shore A hardness 60A silicone, free end tilted up 30°), and a high-strength microneedle array (304 stainless steel). The system consists of a hardness HV≥200, dimensions Φ1mm×10mm, a geared motor 401 (model GA12-N20, reduction ratio 1:100, rated torque 0.8N・m, voltage 6V), a gripping force sensor (range 0-50N, accuracy ±0.2N), and an operating position encoder (resolution 0.01mm). The awl claw's elastic fingers 405 are fixed to the inner mounting slot of the outer frame using M2 bolts. The output shaft of the geared motor 401 passes through the frame and drives the awl claw's elastic fingers 405, allowing adjustment of the opening angle of the awl claw's elastic fingers 405 to ±60°. The gripping force sensor is installed at the connection end between the awl claw's elastic fingers 405 and the outer frame. The operating position encoder is integrated into the output shaft of the geared motor 401, collecting the real-time gripping force of the claw and the relative position of the target, respectively, as local observation signals for the intelligent agent.

[0088] The power unit 6 houses two sets of ternary lithium batteries, each with a capacity of 2500mAh, connected in series to a voltage of 24V, for a total capacity of 20000mAh. It supports fast charging for 2 hours and slow charging for 8 hours. The battery packs are attached to the inner wall of the top cover via insulating mounting brackets (made of EVA foam, 5mm thick). The power management module is fixed to the side of the battery pack and electrically connected to the battery pack and intelligent control unit via wires. Model LT8705, with a conversion efficiency ≥92%. The power management module features 24V to 5V (3A output current) and 24V to 12V (2A output current) voltage conversion functions, as well as overload protection (threshold current 10A) and low battery warning (audible and visual alert triggered when remaining battery ≤20%).

[0089] Combination Figure 2 (Complete structure of underactuated arm) Figure 3 (Underacted joint structure) Figure 4 (End-effector micro-spine needle claw structure) The specific details of the underactuated leg joint 3, the micro-spine needle claw joint 4, the multi-level drive unit, and the intelligent control unit are explained.

[0090] The underactuated leg joint 3 is the core of the robot's terrain adaptation. It adopts a closed-loop underactuated mechanism + composite drive design. The length of a single leg is 170mm in the retracted state and 290mm in the extended state. It is composed of 7 composite material links hinged together to form a diamond-shaped closed-loop structure. The length of each link is as follows: (active link) Figure 3 First active link 304, second active link 306, third active link 308: lengths 50mm, 75mm, 80mm, respectively, hinged together at both ends by butt bolts; passive link ( Figure 3 The first passive link 307 and the second passive link 309 are 70mm and 75mm long, respectively, and are hinged to other links at both ends by butt bolts; the end actuator link 310 is 30mm long, with one end connected to the third active link 308 and the second passive link 309, and the other end docked with the geared motor mounting housing of the end micro-spine needle claw joint 4; the fixed frame, i.e. Figure 3 The linear motor support base frame 303 is made of 7075-T6 aluminum alloy, with dimensions of 110mm×34mm×32mm and an L-shaped base. The linear motor 302 (model LASF30-024D, rated thrust 80N, stroke 30mm, rated voltage 24V) is fixed by bolts. The output shaft of the linear motor 302 is hinged to the first active connecting rod 304 through the intermediate connecting slider. When the drive mechanism extends or retracts, the closed-loop structure can convert the linear motion into a composite motion of bending and stretching of the underactuated structure, thus offsetting the impact of the terrain.

[0091] The thigh support structure is made of ABS plastic and is in the shape of a semi-closed sleeve (130mm in length, 42mm in width, and 4mm in thickness). The linear motor support base frame 303 is fixed inside the sleeve by bolts. The other end is coaxially connected to the output shaft of the rotary motor 301 (model such as FEETECH-SM60CL, rated torque 2N·m, rated speed 35RPM, power supply voltage 12V). The rotary motor 301 is fixed to the leg mounting holes of the frame of the control platform body 2 through the tail threaded hole. The output shaft of the rotary motor 301 coincides with the central axis of the thigh support structure, which can drive the leg to rotate horizontally around the platform by ±60°.

[0092] The micro-spine needle claw joint 4 at the end is designed for gripping complex terrain, employing passive elastic deformation and active angle fine-tuning. The outer frame, i.e., the micro-needle shell 404, is a hollow rectangular shell made of ABS engineering plastic, measuring 38mm × 30mm. Its top is detachably connected to the end actuator 310 of the underactuated leg joint 3 via a butt bolt. Two mounting slots are symmetrically formed on the inner wall. The needle claw elastic fingers 405 are made of silicone with a Shore A hardness of 60A, featuring a zigzag bending structure with the free end curving outwards at a 30° angle. One end is embedded in the mounting slot and secured with an M2 bolt. High strength... The microneedle 406 is made of 304 stainless steel (hardness HV≥200), with a specification of Φ1mm×10mm (diameter×length), and is fixed to the end of the elastic finger 405 of the needle claw. The geared motor 401 (model GA12-N20, reduction ratio 1:100, rated torque 0.8N・m, voltage 6V) is connected to the end actuator 310 through the geared motor mounting housing 402. The output shaft of the geared motor 401 passes through the frame and is connected to the microneedle housing 404 through the claw housing connecting transmission component 403. The finger opening and closing angle can be adjusted to ±60°, and it can only be activated in the scene of slope >30° or vertical rock wall.

[0093] Specifically, the passive elastic deformation + active angle fine-tuning mechanism eliminates the need for motor drive in the passive elastic deformation. Utilizing a zigzag-shaped elastic finger (30° angled upwards at the free end) made of 60A Shore A silicone, it conforms to irregular targets (rubble, irregularly shaped materials) through the material's inherent elasticity and special structure, increasing contact area, buffering impact loads, and compensating for minor leg position deviations. This is further enhanced by a 304 stainless steel microneedle array (hardness HV≥200, specifications Φ1mm×10mm) to improve gripping stability. The active angle fine-tuning mechanism relies on a GA12-N20 geared motor 401 with a reduction ratio of 1:100 and a rated torque of 0.8N·m. Combined with a 0.01mm resolution operating position encoder, it precisely controls the finger's ±60° opening and closing angle, adapting to targets of different sizes (from 5mm plant samples to 100mm material boxes) and task types (gripping / operation). The two mechanisms work together to actively establish a basic angle framework and passively fill in the gaps in detail adaptation, ensuring gripping accuracy while improving fault tolerance, resulting in a 98% gripping success rate.

[0094] The intelligent control unit is equipped with hardware such as an 8-core CPU, a 1024-core GPU, 16GB of LPDDR5 memory, 256GB of NVMe storage, and a PCB board measuring 100mm × 80mm. It is fixed to the lower frame of the platform via an insulating bracket, with a 15mm gap between the PCB and the support column. The software includes a deep fusion algorithm based on PyTorch 2.1 for hybrid action space and multi-agent collaboration, integrated with JetPack SDK 6.1.1 and CUDA 11.8 acceleration. The experience replay buffer uses circular storage with a batch size of 1024. The algorithm parameters include a learning rate of 10. -4 Discount factor Target network update coefficients The training cycle consists of 20,000 episodes, with a maximum step size of 100 per episode. It employs an integrated PCB design combining a main control module, a multi-agent decision-making module, a driver module, and a communication module. Core chips include an 8-core Arm Cortex-A78AE CPU, a 1024-core GPU, an AI computing power of 157 TOPS in Super Mode, a memory bandwidth of 102 GB / s, and 16 GB of LPDDR5 memory.

[0095] In the multi-agent decision-making module of the intelligent control unit, the deep fusion algorithm network for hybrid action space and multi-agent collaboration includes a discrete-continuous hybrid policy network, a Q-network with attention mechanism, and a target network; the discrete-continuous hybrid policy network outputs hybrid actions. ,in It consists of 6 discrete combinations of leg mode (support=1 / swing=0) and claw mode (grasp=1 / release=0 / operation=2). The parameters are 6-dimensional continuous parameters, including the leg COM adjustment amount / joint angle increment, and the claw opening and closing angle increment / grip force increment.

[0096] The multi-agent decision-making module integrates a deep fusion algorithm network of hybrid action space and multi-agent collaboration, including a discrete-continuous hybrid strategy network (outputting leg-claw action commands), a Q-network with attention mechanism (evaluating the value of action collaboration), and a free gait transition feasibility extension module (judging the feasibility of the state transition between leg movement and claw operation); the drive module includes 6 leg drive channels (adapted to linear motor 302 and rotary motor 301) and 6 claw operation drive channels (adapted to geared motor 401), supporting position mode (joint angle / claw opening and closing angle control + force control mode (leg thrust / claw gripping force control) switching, with a maximum output current of 5A per channel; the communication module has 1 CAN bus (connecting the motor and sensor), 1 USB 3.2, 1 GbE Ethernet, and 1 Wi-Fi 6 / Bluetooth 5.0, supporting industrial protocols and edge-cloud low-latency communication, and a reserved M.2 slot for connecting an external 5G / LTE module.

[0097] like Figures 8-9 A control method for an underactuated hexapod robot with hand and foot reuse based on multi-agent reinforcement learning, comprising:

[0098] I. System Initialization: System initialization begins with observation and data acquisition. After starting the robot, the intelligent control unit completes a hardware self-test, including motors, sensors, and algorithm modules. Initialization is performed on the parameters of the deep fusion algorithm for hybrid motion space and multi-agent collaboration: policy network φ, Q-network θ, and target network. (Initial value) ), (Initial value) ), Experience replay buffer (Capacity is 10) Discrete action temperature parameters Continuous operating temperature parameters .

[0099] Local observations were collected from 6 groups of legs and claws. ( The observation signal relies entirely on mechanical sensors; the joint angles observed in the leg... (Accuracy ±0.1°), Joint torque (Reverse ground hardness), leg COM coordinates (Relative to the fuselage center); Claw observation: gripping force (Determine the target contact state), opening and closing angle (Current operating posture), relative position to the target ( The real-time coordinates of claw i The target coordinates sent to the remote terminal) are displayed in the leg mode (support=1 / swing=0) and claw mode (grab=1 / release=0 / operation=2) of the first 3 time steps of the historical observation. This is used to avoid action conflicts. The historical observation is not collected independently, but is acquired and stored synchronously with the real-time local observation, without the need for additional hardware.

[0100] II. Multi-Agent Hybrid Action Decision-Making: Multi-agent hybrid action decision-making based on a deep fusion algorithm of hybrid action space and multi-agent collaboration, where each agent performs hybrid actions. Discrete actions Divided into leg mode (support) / swing ) and claw mode (grabbing) / release / operate The combination of ) The discrete options follow a class distribution. ( Discrete policy probabilities, depending only on observations Discrete part Indicates the probability of leg pattern Claw mode probability. Continuous actions. Leg continuous parameters (COM adjustment amount) Joint angle increment ; It is a 3-dimensional real space, meaning that the COM adjustment includes independent adjustment components in the x, y, and z axes. (in 1-dimensional real space) + continuous parameters of the claw (increment of opening and closing angle) Seize the incremental force It is a 6-dimensional continuous space that follows a diagonal Gaussian distribution. ,in, For continuous policy probabilities, depending on observation With selected discrete actions ; continuous part This represents the probability product of four consecutive parameters. The final joint probability can efficiently calculate the rationality of each mixed action and provide accurate gradients for policy updates.

[0101] In distributed execution during policy reasoning and collaborative optimization, each agent bases its actions on local observations. Independent sampling discrete action Continuous actions During centralized evaluation, the Q-network with attention mechanism calculates the collaborative value of actions, using the following formula:

[0102] ;

[0103] in, For the first The collaborative value (Q value) of the group of intelligent agents and This is a global observation set, containing local observations from 6 groups of agents. , ..., ; It is a global discrete action set, containing discrete actions of 6 agents. , ..., ; A global set of continuous actions, containing the continuous actions of 6 agents. , ..., ; For the first The value fusion function of the group of intelligent agents; No. State-action feature extraction network for group agents; To exclude the first Index of other agents outside the group ; For the first Group agents on the first Attention weights of group agents; For the first State-action feature extraction network for group agents; It is a weighted sum of the features of other agents.

[0104] III. Feasibility Assessment of State Transition: The feasibility model for transitioning to extended free gait is as follows: ,in, Let be the set of global states of the robot at time t; The set of global actions of the robot at time t, i.e., the combined actions of the six agents. ; Execute action Then, the global state at candidate time t+1, combined with the kinematic and dynamic constraints of the underactuated machine, is used to determine the legality of the state transition between leg movement and claw operation. This includes checking the mode combination of leg-claw actions, kinematics, dynamics, and the feasibility of the claw operation. Step three includes:

[0105] 1. Leg-claw pattern combination constraint Filter out illegal pattern combinations, including leg swing ( When the claw performs the gripping ( ) or operation ( ), meaning that without support, the aircraft is prone to imbalance, and leg support ( The claw continues to apply force even when it is not being operated. and This refers to wasting energy.

[0106] 2. Kinematic constraints Based on the physical limits of the underactuated leg joints and claw joints, the leg joint angles were examined. Claw opening and closing angle (Adapting to the physical limits of flexible fingers), the distance between the claw and the target is ||pri||≤290mm, which is the maximum extension length of the leg, ensuring that the claw can reach the target.

[0107] 3. Dynamic constraints Based on a single rigid body dynamics model, this study takes the rigid body of the robot's control platform as the modeling object and employs a three-dimensional inertial coordinate system with the body's center of mass as the origin. Ignoring elastic deformation of components, the study analyzes the force balance using rigid body dynamics laws to determine the dynamic feasibility of the leg-support + claw-operation action. The core constraints include that the sum of the supporting forces of all supporting legs in the vertical direction must be greater than or equal to the sum of the robot's weight and the total gripping force of the claws. Combined with the weight constraints of the robot body and the target, the leg supporting force... ( For fuselage quality , (Total gripping force of all claws, ensuring fuselage stability); claw gripping force ( (For target quality, 1.2 is a safety factor to ensure the gripper does not detach); If so, then the state transition is feasible. ;otherwise And record the conflict type (for reward and punishment).

[0108] IV. Reward Function Design and Network Training: A multi-objective reward function is designed and updated through network training, considering factors such as mobility efficiency, operational accuracy, collaboration, and energy consumption. This includes:

[0109] (1) The formula for the multi-objective reward function is as follows:

[0110] ;

[0111] in, The total reward for multiple objectives at a single time step is also The core input for network and strategy network updates is that the higher the total reward, the better the current leg-claw action combination mode and parameters are adapted to the integrated movement-operation task.

[0112] As a reward for mobile efficiency, among which For the leg COM (center of mass) coordinates (relative to the fuselage center, observation item, acquired by the position encoder); The target coordinates (coordinates of materials in the ruins, coordinates of jungle samples) sent to the remote terminal; 300 represents the straight-line distance between the leg's COM and the target, in mm; 300 represents the maximum stride length (matching the maximum leg extension length of 290mm, with redundancy); 0.5 represents the bonus coefficient (to ensure...). (To avoid a single reward dominating); its physical meaning is that the closer the distance, the higher the reward.

[0113] For the operational precision bonus (assessing the accuracy and reliability of gripper attachment / operation), values ​​are assigned in three scenarios:

[0114] 1. The claw performs the grasping task (claw mode) =1) and the real-time gripping force of the claw. Safety gripping force threshold ( ), that is, when the gripping force meets the standard, its value is assigned to 5.

[0115] 2. The claw performs fine operations (claw mode) =2) and the relative positional deviation between the claw and the target mm ( (This refers to the relative position of the claw and the target), meaning that when the operational deviation is small, it is assigned a value of 2. Wherein, the relative position... For the first Current position of each paw (or foot) Subtract target position That is, used to represent the first The relative positional relationship between the current positions of each claw (or foot).

[0116] 3. If the claw attempts to grasp the target but fails, and the claw mode is met but the grasping force does not reach the minimum threshold, the grasping fails. And the claw mode is for gripping. ), that is, when the capture fails, its value is assigned to -3.

[0117] in, The target mass is 0.5~2kg. (Gravity acceleration), 1.2 is a safety factor to prevent the gripper from falling off.

[0118] Rewards leg-paw coordination, penalties for conflicting actions, and rewards for non-conflictual actions. Non-conflictual actions (leg support) +claw attachment Leg swing +Claw release (Conforming to the constraints of the transfer feasibility model) its value is assigned to 1; there is a conflict (leg swing). +claw attachment Leg support +Claw without operation applying force (If the constraint of the transfer feasibility model is violated, the value is -2; the conflict judgment directly depends on the constraint result of the transfer feasibility model.)

[0119] To penalize energy consumption, suppress high-energy-consuming actions, and extend the overall battery life, the segmented formula is as follows:

[0120] ;

[0121] in, The real-time current of the motors, including the leg linear motor LASF30-024D, the rotary motor FEETECH-SM60CL, and the claw geared motor GA12-N20, is collected by the drive module. This refers to the motor's rated current; when the current exceeds the limit, the higher the percentage of over-limit, the heavier the penalty.

[0122] As a terminal task reward, it evaluates the completion effect of the entire movement + operation task and is only triggered when the task ends; if the task is successfully completed, the target is grabbed and moved to the designated area, it is assigned a value of 10; if the task fails (timeout > 100 steps, motor overload, claw attachment failure), it is assigned a value of -5.

[0123] (2) The Q-network and policy network are updated through experience replay to automatically adjust the temperature parameters. Policy evaluation: The Q-network is updated by minimizing the Bellman residual, as shown in the following formula:

[0124] ;

[0125] in, For the first The loss function of the group agent Q network (which needs to be minimized by gradient descent); For the first Trainable parameters (weights, biases, etc.) of a group agent Q-network; For experience replay buffer The samples in the dataset are taken from the expected value (to reduce data correlation and avoid training oscillations). This is the set of global observations at time t / t+1 (including 6 sets of local observations of the leg-claw). ); The set of global actions at time t / t+1 (including the mixed actions of 6 agents) ); The instantaneous reward at time t, i.e. Multi-objective total reward; Discount factor ( ); The current state-action value (current Q value) output by the Q network for the i-th agent group; The parameters of the target Q-network for the j-th agent group (compared to the current Q-network) (Structure consistent) The next state-action value (target Q value) output by the target Q network; For all target Q networks ( The output of ) takes the minimum value.

[0126] Policy Improvement: Update the policy network by maximizing the cumulative reward with entropy regularization, as shown in the following formula:

[0127] ;

[0128] in, For the first The objective function of the group agent policy network (the objective that needs to be maximized through gradient ascent); For the first Trainable parameters (weights, biases, etc.) of the group agent policy network; For action Obedience Policy Network The probability distribution of actions sampled from the policy. The state-action value output by the Q-network for the i-th group of agents; Temperature parameters for discrete / continuous actions (initializing discrete action temperature parameters) Continuous operating temperature parameters ); Discrete Actions The logarithmic probability (the logarithm of the discrete action probability output by the policy network). For continuous action Logarithmic probability density (given discrete action) and observation ).

[0129] Temperature update: Automatic adjustment and To balance exploration and utilization; considering the category distribution characteristics of discrete actions, the objective function ensures that the discrete policy entropy is stable at a preset threshold, avoiding excessive rigidity or randomness in mode selection; corresponding to the six discrete combinations of leg mode (support=1 / swing=0) and claw mode (grasp=1 / release=0 / operation=2), the randomness of discrete actions is controlled. Enlarging the intelligent agent makes it easier to attempt combinations of leg swinging + claw pre-grasping with multi-leg support + claw fine-tuning; To minimize the impact, prioritize 3-leg support + claw attachment:

[0130] ;

[0131] in, For the first The discrete action-temperature objective function of each agent needs to be minimized through gradient descent. Automatic updates; For the first The discrete action temperature parameters of an agent, and the entropy of the discrete policy; For the first Each agent is based on local observation The output discrete policy distribution follows a categorical distribution, indicated by the parentheses " "Represents all possible discrete actions; Indicates the first An intelligent agent in Local observation at time ; For the first The log probability of a discrete action of an agent reflects the known value of the discrete action; the lower the probability, the smaller the log probability. The target entropy of the discrete policy is a fixed value used to constrain the stable range of the discrete action entropy; No. Local observation signals of an intelligent agent.

[0132] For the diagonal Gaussian distribution characteristics of continuous motion, the objective function controls the entropy of continuous parameters to avoid excessive parameter fluctuations leading to instability or excessively small parameters leading to insufficient adaptability; corresponding to the leg COM adjustment amount. Increment of claw opening and closing angle Seize the incremental force Six-dimensional continuous parameters are used to control the volatility of continuous parameters. When the value is increased, the adjustment range of the continuous parameter expands; As the parameter decreases, the continuous parameter tends to stabilize.

[0133] ;

[0134] in, For the first The objective function for the continuous actions of an agent, namely temperature, needs to be minimized through gradient descent. Automatic updates; For the first The temperature parameters of the continuous actions of an agent, and the entropy of the continuous control policy; For continuous actions Expected value after sampling; For from the first Continuous policy distribution of an agent The continuous action samples in the middle are continuously distributed according to a diagonal Gaussian distribution; For the first An agent observes locally. Discrete Actions The probability density of continuous policies under the given conditions follows a diagonal Gaussian distribution; For the first Discrete actions of an agent; For the first The log probability density of a series of actions of an agent reflects the optimality of continuous parameters; the lower the probability density, the smaller the log probability. The target entropy of the continuous policy is a fixed value used to constrain the stable range of the entropy of continuous actions; For the first Local observation signals of an intelligent agent.

[0135] V. Hand-Foot Reuse Task Execution and Dynamic Optimization: A multi-level drive unit drives the leg-claw agent to perform actions. Linear motor 302 adjusts the thrust according to the terrain (50N for soft soil / 80N for hard soil), and geared motor 401 adjusts the claw opening and closing angle. The intelligent control unit updates the status in real time, dynamically adjusting parameters to address operational deviations, insufficient gripping force, and excessive energy consumption. The multi-level drive unit receives action commands from a deep fusion algorithm that integrates hybrid motion space and multi-agent collaboration, driving each agent to execute the commands. Specifically, this includes:

[0136] 1. The rotating motor 301 of the leg adjusts the leg posture (horizontal rotation ±60°), and the linear motor 302 drives the underactuated mechanism to extend and retract (speed 6mm / s) to achieve support / swing; the thrust is set to 50N for soft terrain and 80N for hard terrain; the reduction motor 401 of the claw adjusts the opening and closing angle (position mode), and maintains the gripping force (deviation ≤±1N) in force control mode to complete gripping / release / operation.

[0137] 2. The intelligent control unit updates the status and observations every 100ms and makes real-time adjustments for the following issues: if the operation deviation is >5mm, the position adjustment amount of the claw is increased. At the same time, fine-tune the leg COM Compensation for deviation; if the gripping force is less than the safety value, increase the gripping force increment of the claw. If it is still insufficient, switch the legs to enhanced support mode (increase the number of support legs to 4); when the energy consumption exceeds the limit (current > 120% of the rated value), reduce the force of non-critical movements (reduce the claw's operating force by 20%) and extend the gait cycle to 1.5s.

[0138] 3. When receiving instructions from a remote terminal (switching from moving to grasping), the multi-agent decision-making module completes the action mode switch within 50ms without needing to retrain.

[0139] VI. If a fault is detected (motor overload, claw gripping failure, low battery), immediately trigger the shutdown of unnecessary drive units (legs / claws not involved in support / operation); send fault information (fault location, cause, current status) to the remote terminal; and activate redundant control (if claw 1 fails, the adjacent claw 2 / claw 6 will compensate; if the support legs are insufficient, switch the swing leg to support mode) to ensure machine stability. Figure 8 Line segments ① and ② in the diagram are respectively with... Figure 9 Connect line segment ① and line segment ② in the diagram.

[0140] This invention addresses the core pain points of traditional hand-foot reusable robots, namely, the disconnect between leg and claw coordination, the difficulty in optimizing complex hybrid motion spaces, and poor adaptability to complex terrain. The robot's main body adopts a two-layer nested hexagonal hollow frame structure, with six groups of leg-claw agents symmetrically distributed around the main body at 360° (adjacent angles are 60°). Each group of agents consists of an underactuated leg joint (a rhomboid closed-loop mechanism + linear / rotary motor) connected in series with an end-effector micro-spine needle-prick claw joint (a zigzag elastic finger + a high-strength micro-needle array + a geared motor). It also integrates a multi-level drive unit (adapted to leg-claw motor drive), an intelligent control unit (containing a deep fusion algorithm network for hybrid motion spaces and multi-agent coordination, and an extended free gait transfer feasibility model to achieve leg-claw collaborative decision-making), and a power unit 6 (two sets of series-connected ternary lithium batteries + a power management module). The motion control method achieves integrated intelligent control of movement and operation through a process of system initialization and observation acquisition → multi-agent hybrid action decision-making (discrete leg / paw mode + continuous parameter optimization) → transfer feasibility model state transition feasibility judgment (filtering illegal actions and constraining kinematic / dynamic parameters) → multi-objective reward calculation and network training (considering mobility efficiency, operational accuracy, coordination, and energy consumption) → hand-foot reuse task execution and dynamic optimization → emergency protection. This invention is applicable to scenarios with strict requirements for mobility, energy consumption, and cost, such as rubble rescue, jungle exploration, and material transportation in complex terrain. It can reduce the leg-paw action conflict rate by more than 90%, compress the hand-foot mode switching delay to within 80ms, improve the overall battery life by 38% compared to traditional designs, and maintain a mission success rate of ≥82% even after a single component failure, while also considering coordination, adaptability, economy, and robustness.

[0141] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A multi-agent reinforcement learning based hand-foot reuse underactuated hexapod robot, characterized in that, The utility model relates to a kind of hand-foot multipurpose underactuated hexapod robot, including: Control platform body (2), the top of the control platform body (2) is provided with power end cover (1), the control platform body (2) is built-in multi-agent control unit (5); Underactuated leg joint (3), located outside the control platform body (2), symmetrically distributed; Microspinal acupuncture claw joint (4), located at the end of the underactuated leg joint (3) away from the control platform body (2), for cooperating with the underactuated leg joint (3); The multi-agent control unit (5) is connected with the underactuated leg joint (3) and the microspinal acupuncture claw joint (4), for outputting leg claw cooperative action instruction; Power unit (6), built-in in the control platform body (2), for providing power management for hand-foot multipurpose underactuated hexapod robot.

2. The multi-agent reinforcement learning based hand-foot cooperation underactuated hexapod robot according to claim 1, wherein, The underactuated leg joint (3) includes a rotary motor (301) connected with the control platform body (2), a linear motor bearing base frame (303) is provided at the end of the rotary motor (301) away from the control platform body (2), a linear motor (302) is provided in the linear motor bearing base frame (303), a first driving link (304) is provided at the end of the linear motor (302) away from the rotary motor (301), one end of the first driving link (304) away from the linear motor (302) is connected with one end of a V-shaped transmission link (305), the other end of the V-shaped transmission link (305) is connected with a second driving link (306), and a first passive link (307) is connected with the middle part of the V-shaped transmission link (305). The second driving link (306) is connected with a third driving link (308) at the end away from the V-shaped transmission link (305), the first passive link (307) is connected with a second passive link (309) at the end away from the V-shaped transmission link (305), the third driving link (308) is connected with the middle part of an end execution link (310) at the end away from the second driving link (306), and the second passive link (309) is connected with one end of the end execution link (310) at the end away from the first passive link (307), and the other end of the end execution link (310) is connected with the microspinal acupuncture claw joint (4).

3. The multi-agent reinforcement learning based hand-foot cooperation underactuated hexapod robot according to claim 1, wherein, The microspinal acupuncture claw joint (4) includes a reduction motor mounting shell (402) connected with the underactuated leg joint (3), a reduction motor (401) is provided in the reduction motor mounting shell (402), a micro-needle shell (404) is provided at the side of the reduction motor mounting shell (402), and a claw shell connecting transmission member (403) is provided between the micro-needle shell (404) and the reduction motor mounting shell (402). A needle acupuncture claw elastic finger (405) is provided in the middle part of the micro-needle shell (404), and a micro-needle (406) is provided at the end of the needle acupuncture claw elastic finger (405).

4. The multi-agent reinforcement learning based hand-foot cooperation underactuated hexapod robot according to claim 1, wherein, The control platform body (2) comprises an upper hollow frame (201) connected with the power end cover (1), and the bottom end of the upper hollow frame (201) is provided with a lower hollow frame (202).

5. A control method of a multi-agent reinforcement learning-based leg-multiplex underactuated hexapod robot, for controlling the multi-agent reinforcement learning-based leg-multiplex underactuated hexapod robot according to claims 1-4, characterized in that, The method comprises the steps of: Initializing a hand-foot-reuse underactuated hexapod robot, and collecting local observation signals by using the legs and claws of the hand-foot-reuse underactuated hexapod robot; Based on the local observation signals, multiple agents independently sample discrete actions and continuous actions; Based on the discrete actions, the continuous actions and the global state of the hand-foot-reuse underactuated hexapod robot, a transition feasibility model of the extended free gait is constructed; the transition feasibility model of the extended free gait is used, and the state transition legality of leg movement and claw operation is obtained in combination with the constraint conditions of the underactuated mechanism; When the state transition legality is legal, multi-objective reward calculation and network training are performed to update the agent; the agent is used to perform and dynamically optimize the hand-foot-reuse task.

6. The control method of the multi-agent reinforcement learning-based hand-foot multiplexed underactuated hexapod robot according to claim 5, characterized in that, The multiple agents independently sample discrete actions and continuous actions, which include: The multiple agents independently sample discrete actions including leg modes and claw modes, and continuous actions including leg continuous parameters and claw continuous parameters; The discrete actions are subjected to a category distribution, and the continuous actions are subjected to a diagonal Gaussian distribution; The action coordination value is calculated by using a Q network with an attention mechanism.

7. The control method of the multi-agent reinforcement learning-based hand-foot multiplexed underactuated hexapod robot according to claim 5, wherein, The transition feasibility model of the extended free gait is constructed based on the discrete actions, the continuous actions and the global state of the hand-foot-reuse underactuated hexapod robot, which comprises the steps of: The global state of the hand-foot-reuse underactuated hexapod robot at the current time is obtained; the global state at the next time is obtained by performing actions at the next time; The global states of the hand-foot-reuse underactuated hexapod robot at the current time and the next time are used, and the transition feasibility model of the extended free gait is fused based on the discrete actions and the continuous actions.

8. The control method of the multi-agent reinforcement learning-based hand-foot multiplex underactuated hexapod robot according to claim 5, characterized in that, The constraint conditions of the underactuated mechanism include: The leg-claw mode combination constraint of the hand-foot-reuse underactuated hexapod robot is checked; Based on the physical limits of the underactuated leg joints and claw joints, the leg joint angle, the claw opening angle and the claw target distance are checked; Based on the single rigid body dynamics model, the control platform body (2) is taken as the modeling object, a three-dimensional inertial coordinate system with the center of mass of the body of the hand-foot-reuse underactuated hexapod robot as the origin is adopted, and the dynamics feasibility of the leg support and claw operation actions is judged by analyzing the force balance based on the rigid body dynamics law.

9. The control method of the multi-agent reinforcement learning-based hand-foot multiplex underactuated hexapod robot according to claim 5, wherein, When the state transition legality is legal, multi-objective reward calculation and network training are performed to update the agent, which comprises the steps of: A multi-objective reward function is constructed based on the movement efficiency reward, the operation precision reward, the leg-claw coordination reward, the energy consumption penalty and the terminal task reward; the multi-objective total reward is calculated according to the multi-objective reward function; When the state transition legality is legal, the multi-objective total reward is taken as the input of the policy network update of each agent the input of the network and the policy network update The Q network is updated by minimizing the Bellman residual, and the policy network is updated by maximizing the cumulative reward with entropy regularization; the discrete actions and the continuous actions are automatically adjusted by temperature update to balance exploration and utilization. For the diagonal Gaussian distribution characteristics of continuous action, the target function controls the entropy of continuous parameters and controls the volatility of continuous parameters; when the continuous action temperature parameter increases, the continuous parameter adjustment range expands; when the continuous action temperature parameter decreases, the continuous parameter tends to be stable.

10. The control method of the multi-agent reinforcement learning-based hand-foot multiplex underactuated hexapod robot according to claim 5, wherein, The execution and dynamic optimization of the task of hand-foot multiplexing by the agent includes: The multi-agent control unit (5) is used to update the state in real time and dynamically adjust the parameters of the hand-foot multiplexing underactuated hexapod robot; The multi-level driving unit receives the action instruction output by the multi-agent control unit (5) to drive each agent to execute the hand-foot multiplexing task.