Samander-imitated robot path tracking method and system based on guided diffusion model

Through a hierarchical control scheme combining guided diffusion model with traditional planning, the problems of control complexity and training instability in the path tracking of imitation salamander robots are solved, and stable path tracking and autonomy are achieved in the real world.

CN120295291APending Publication Date: 2025-07-11NANKAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510231200.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The prior art has problems such as complex control, high environmental dependence, unstable training and non-convergence in the path tracking of newt imitation robots, especially when applied in the real world.

Method used

A hierarchical control scheme based on a guided diffusion model is adopted, combining upper-level strategies and underlying controllers, and path tracking is used to train upper-level strategies through offline data to generate action instructions, and actions are performed through traditionally planned underlying controllers.

Benefits of technology

It realizes stable path tracking without environmental interaction in the real world, improves the autonomy and security of the newt imitation robot, and avoids the dependence of traditional methods on in-depth domain knowledge and training instability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295291A_ABST
    Figure CN120295291A_ABST
Patent Text Reader

Abstract

The invention discloses a salamander-imitating robot path tracking method and system based on a guiding type diffusion model, and belongs to the field of electromechanical system control, and the method comprises the steps: obtaining path state information; inputting the path state information into a trained upper-layer strategy based on guide diffusion to obtain an action instruction; the action instructions are input into a bottom layer controller based on traditional planning, and position instructions of all steering engines of the salamander-imitating robot are obtained. On one hand, an upper strategy based on guide diffusion is used as an action decision maker, and long-term planning and flexible behavior sampling capability of a diffusion model are fully utilized; and on the other hand, the bottom-layer controller analyzes the upper-layer action command by adopting a traditional gait planning method and executes a specific action. The method has the beneficial effects that the layering scheme provided by the invention can make full use of offline trajectory data without interaction with the environment, and the diffusion model coupled with the bottom layer controller is applied to the robot executing tasks in the real world.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of mechatronic system control, and particularly relates to a path tracking method and system for a salamander-like robot based on a guided diffusion model. Background Technique

[0002] The statements in this part only mention the background techniques related to this application and do not necessarily constitute prior art.

[0003] Legged locomotion can solve complex and harsh environments that cannot be solved by tracked or wheeled vehicles, which greatly expands the scope of robot systems. As a typical legged robot, the salamander-like robot combines the advantages of snake-like robots and is more flexible while maintaining stability. With such capabilities, salamander-like robots have the potential to be applied in a wide range of scenarios, such as disaster relief, environmental monitoring, and surrounding exploration. Developing intelligent controllers on salamander-like robots to complete tasks such as path tracking can not only help us better understand animal behavior but also enable us to achieve the automation and autonomy of robot systems.

[0004] Although traditional controllers have demonstrated impressive results in the movement of legged robots in the real world, traditional methods often require in-depth domain knowledge and a large amount of manual adjustment, and may not be well transferred to new scenarios due to ignoring some real-world factors. Unfortunately, these drawbacks will make it worse for salamander-like robots to perform path tracking tasks because the robot needs to execute upper-layer tasks while maintaining its own coordination, which poses higher requirements for its control. Recently, the progress of reinforcement learning (RL) has greatly improved the control of legged robots. Models based on the learning paradigm eliminate complex modeling processes, reduce prior knowledge in controller design, and meet the requirements of optimal control and real-time control. However, due to various sensor interferences and actuator differences in the real world, transferring learning-based policies to the real world is a persistent and challenging problem in robotics. There are some methods that focus on directly applying deep reinforcement learning to train legged robots in the real world. However, it is not always possible to interact between the agent and the environment in the real world. Training robots in environments such as laboratories heavily relies on manual reset between episodes, and this process cannot guarantee the safety of training. More seriously, in applications such as healthcare and autonomous driving, the exploration of untrained policies in the real world may have fatal consequences.

[0005] To address this issue, a large number of literature studies have focused on recovering effective policies entirely from previously collected offline data without interacting with the offline decision-making environment. Using offline data (e.g., human demonstrations) is much cheaper than online interaction. However, although methods based on offline reinforcement learning eliminate the need for real-world exploration, they are not without flaws. Since they still follow the spirit of traditional temporal difference (TD) learning, they face challenges from the "fatal triad" in reinforcement learning, leading to unstable training and non-convergence. On the other hand, generative sequence modeling methods view reinforcement learning as a general sequence generation problem and model the joint distribution of state, action, and reward sequences, thus avoiding the challenges of the "fatal triad" and opening up a new and compelling alternative for the offline reinforcement learning problem. Taking diffusion models as an example, they not only exhibit remarkable performance in image generation but also possess some attractive characteristics in sequence modeling problems. First, diffusion models have a powerful ability to capture multimodal distributions; second, diffusion models focus on the accuracy of the entire generation trajectory, so they are not prone to single-step errors and can handle long-term planning gracefully; third, diffusion models model the environment dynamics and behavior during training while allowing for flexible guidance according to different task constraints during the sampling process. Summary of the Invention

[0006] To address the deficiencies of the prior art, the present application provides a salamander robot path tracking method and system based on a guided diffusion model; for learning the motion control of a salamander robot using offline data. The objective of the present invention is to propose a salamander robot path tracking technology based on a guided diffusion model, which couples the diffusion model with the underlying controller for hierarchical control of the salamander robot to perform path tracking tasks.

[0007] In a first aspect, the present application provides a salamander robot path tracking method based on a guided diffusion model;

[0008] A salamander robot path tracking method based on a guided diffusion model, comprising:

[0009] Obtain path state information;

[0010] Input the path state information into the trained upper-layer policy based on guided diffusion to obtain an action instruction;

[0011] Input the action instruction into the underlying controller based on traditional planning to obtain the position instructions for all the servos of the salamander robot.

[0012] In a second aspect, the present application provides a salamander robot path tracking system based on a guided diffusion model;

[0013] A path tracking system for a salamander-like robot based on a guided diffusion model, comprising:

[0014] An acquisition module, which is configured to: acquire path status information;

[0015] An upper-layer policy module, which is configured to: input the path status information into the trained upper-layer policy based on guided diffusion to obtain an action instruction;

[0016] A lower-layer controller module, which is configured to: input the action instruction into a lower-layer controller based on traditional planning to obtain position instructions for all the servos of the salamander-like robot.

[0017] In a third aspect, the present application further provides an electronic device, comprising: one or more processors, one or more memories, and one or more computer programs; wherein, the processor is connected to the memory, and the above-mentioned one or more computer programs are stored in the memory. When the electronic device runs, the processor executes the one or more computer programs stored in the memory so that the electronic device executes the method described in the first aspect above.

[0018] In a fourth aspect, the present application further provides a computer-readable storage medium for storing computer instructions, and when the computer instructions are executed by a processor, the method described in the first aspect is completed.

[0019] In a fifth aspect, the present application further provides a computer program (product), comprising a computer program, and when the computer program runs on one or more processors, it is used to implement the method of any item in the foregoing first aspect.

[0020] Compared with the prior art, the beneficial effects of the present application are:

[0021] A path tracking technology for a salamander-like robot based on a guided diffusion model, which consists of an upper-layer policy based on guided diffusion and a lower-layer controller based on traditional planning. Among them, the upper-layer policy based on guided diffusion inputs and selects path status information and outputs an action instruction; the lower-layer controller based on traditional planning inputs and selects the action instruction and outputs position instructions for all the servos of the salamander-like robot. On the one hand, the upper-layer policy based on guided diffusion in the present invention is used as an action decision maker, making full use of the long-term planning and flexible behavior sampling capabilities of the diffusion model; on the other hand, the lower-layer controller uses a traditional gait planning method to parse the upper-layer action command and execute specific actions. Compared with a completely learning-based method, the hierarchical scheme proposed in the present invention can make full use of offline trajectory data without interacting with the environment. In addition, different from the diffusion model used for sequence planning in simulation, the present invention applies the diffusion model coupled with the lower-layer controller to a robot performing tasks in the real world.

[0022] Advantages of additional aspects of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The accompanying drawings forming a part of this application are used to provide a further understanding of this application. The schematic embodiments and descriptions thereof of this application are used to explain this application and do not constitute an improper limitation to this application.

[0024] Figure 1 It is a schematic diagram of a new path tracking method for a salamander-like robot based on a guided diffusion model.

[0025] Figure 2 It is a network structure diagram of the diffusion model for a new path tracking method for a salamander-like robot based on a guided diffusion model.

[0026] Figure 3 It is a network structure diagram of the encoding module of the diffusion model for a new path tracking method for a salamander-like robot based on a guided diffusion model.

[0027] Figure 4 It is an internal network structure diagram of the first basic block of the diffusion model for a new path tracking method for a salamander-like robot based on a guided diffusion model.

[0028] Figure 5 It is an internal network structure diagram of the third basic block of the diffusion model for a new path tracking method for a salamander-like robot based on a guided diffusion model. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0029] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used in the present invention have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs.

[0030] It should be noted that the terms used here are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to this application. As used here, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these process, method, product, or device.

[0031] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0032] Embodiment 1

[0033] This embodiment provides a path tracking method for a salamander-like robot based on a guided diffusion model;

[0034] As shown in Figure 1 a path tracking method for a salamander-like robot based on a guided diffusion model includes:

[0035] S101: Obtain path status information;

[0036] S102: Input the path status information into the trained upper-layer policy based on guided diffusion to obtain an action instruction;

[0037] As one or more embodiments, as shown in Figure 1 the upper-layer policy based on guided diffusion includes:

[0038] Trajectory initialization and guided sampling of the diffusion model neural network.

[0039] The trajectory initialization, where as an alternative implementation, the trajectory representation form is a two-dimensional array x k (τ) concatenated by an action sequence and a state sequence on a trajectory horizon H at a diffusion time step of k, and the actions and states are determined by the action space and the state space.

[0040] As one or more embodiments, the action space is an abstract motion command (i.e., forward, backward, left turn, right turn). Specifically, for a salamander-like robot, the motion command is achieved by adjusting the step length of the legs and the offset of the spine. Thus, the action space a can be defined as

[0041]

[0042] where l left and l right respectively represent the step lengths of the left and right legs of the salamander-like robot, is the offset term of the spine of the salamander-like robot. When l left = l right , the robot tends to go straight; when l left ≠ l right , the robot tends to turn, and the spine offset term assists in turning.

[0043] As one or more embodiments, the state space is the coordinates of all scattered points inside the tracking window on the target path in the robot coordinate system. Specifically, the target path is directional, starting from the end close to the robot and ending at the end far from the robot. The entire target path is a series of discrete points, i.e.,

[0044] (R P 1,R P2, …, R P N ),

[0045] where R P i = ( R x i , R y i ), i = 1, 2, …, N, the subscript R at the lower left indicates the vector in the robot coordinate system, and N represents the coordinates of the i-th scatter point in the robot coordinate system. The window is defined as a set of W consecutive points, denoted as ( R P m , P m+1 , …, R P m+W-1 ). The criterion for selecting the tracked window is to always track the nearest window and always update the window along the target path towards the path end point. Specifically, at each moment t, the point closest to the robot is selected as the starting point of the target window After determining the target window, all points in front of the target window on the target path are deleted When the distance between the robot and the target window meets the preset minimum distance, it is considered that the window has been fully tracked, and the points inside the window are deleted from the path. The entire path tracking process is considered terminated when the number of remaining points on the target path is less than the window size.

[0046] For the initialization noise process of the trajectory, Gaussian noise is used for initialization, and its expression is

[0047]

[0048] The initialized noise is input into the trained diffusion model neural network for K times of guided sampling to obtain a noise-free trajectory;

[0049] As one or more embodiments, as Figure 2 shown, the diffusion model network, the network structure includes:

[0050] An encoding module structure, a first module structure, a second module structure, a third module structure, an intermediate module structure, a fourth module structure, a fifth module structure, and a sixth module structure connected in sequence;

[0051] As Figure 3 shown, the encoding module structure includes:

[0052] A position encoding layer, a first fully connected layer, a first Mish function layer, and a second fully connected layer connected in sequence.

[0053] The output ends of the second fully-connected layer of the encoding module are respectively connected to the input ends of the first module structure, the second module structure, the third module structure, the intermediate module structure, the fourth module structure, the fifth module structure, and the sixth module structure;

[0054] The first module structure is three first basic blocks, a second basic block, and a third basic block connected in sequence;

[0055] The second module structure is three fourth basic blocks, a fifth basic block, and a sixth basic block connected in sequence;

[0056] The third module structure is three seventh basic blocks, an eighth basic block, and a ninth basic block connected in sequence;

[0057] The intermediate module structure is three tenth basic blocks, an eleventh basic block, and a twelfth basic block connected in sequence;

[0058] The fourth module structure is three thirteenth basic blocks, a fourteenth basic block, and a fifteenth basic block connected in sequence;

[0059] The fifth module structure is three sixteenth basic blocks, a seventeenth basic block, and an eighteenth basic block connected in sequence;

[0060] The sixth module structure is three nineteenth basic blocks, a twentieth basic block, and a twenty-first basic block connected in sequence;

[0061] Among them, the input end of the first basic block is the input trajectory, the input end of the encoding module structure is the diffusion time step, and the output end of the twenty-first basic block is the output trajectory;

[0062] Among them, the third basic block is connected to the fourth basic block, the sixth basic block is connected to the seventh basic block, the ninth basic block is connected to the tenth basic block, the outputs of the ninth basic block and the twelfth basic block are spliced and connected to the thirteenth basic block, the outputs of the sixth basic block and the fifteenth basic block are spliced and connected to the sixteenth basic block, and the outputs of the third basic block and the eighteenth basic block are spliced and connected to the nineteenth basic block;

[0063] Among them, the output ends of the encoding module structure are respectively connected to the input ends of the first basic block, the second basic block, the fourth basic block, the fifth basic block, the seventh basic block, the eighth basic block, the tenth basic block, the twelfth basic block, the thirteenth basic block, the fourteenth basic block, the sixteenth basic block, the seventeenth basic block, the nineteenth basic block, and the twentieth basic block;

[0064] Among them, the internal structures of the first basic block, the second basic block, the fourth basic block, the fifth basic block, the seventh basic block, the eighth basic block, the tenth basic block, the twelfth basic block, the thirteenth basic block, the fourteenth basic block, the sixteenth basic block, the seventeenth basic block, the nineteenth basic block and the twentieth basic block are the same;

[0065] Among them, the internal structures of the third basic block, the sixth basic block, the ninth basic block, the eleventh basic block, the fifteenth basic block, the eighteenth basic block and the twenty-first basic block are the same;

[0066] As Figure 4 shown, the internal structure of the first basic block includes:

[0067] Parallel first and second branches;

[0068] The first branch includes a second Mish function layer and a third fully connected layer connected in sequence;

[0069] The second branch includes a first convolutional layer, a first group normalization layer and a third Mish function layer connected in sequence;

[0070] After the data at the output ends of the third fully connected layer and the third Mish function layer are added, they are sequentially connected to a second convolutional layer, a second group normalization layer and a fourth Mish function layer. After the input end of the first convolutional layer is added to the output end of the fourth Mish function layer, it is output from the output end of the first basic block.

[0071] The output ends of the encoding module structure are respectively connected to the input ends of the second Mish function layers of the first basic block, the second basic block, the fourth basic block, the fifth basic block, the seventh basic block, the eighth basic block, the tenth basic block, the twelfth basic block, the thirteenth basic block, the fourteenth basic block, the sixteenth basic block, the seventeenth basic block, the nineteenth basic block and the twentieth basic block;

[0072] The input end of the first convolutional layer of the first basic block is the input trajectory. The output end of the first basic block is connected to the input end of the first convolutional layer of the second basic block. The output end of the fourth basic block is connected to the input end of the first convolutional layer of the fifth basic block. The output end of the seventh basic block is connected to the input end of the first convolutional layer of the eighth basic block. The output ends of the ninth basic block and the twelfth basic block are concatenated and connected to the input end of the first convolutional layer of the thirteenth basic block. The output end of the thirteenth basic block is connected to the input end of the first convolutional layer of the fourteenth basic block. The output end of the sixteenth basic block is connected to the input end of the first convolutional layer of the seventeenth basic block. The output end of the nineteenth basic block is connected to the input end of the first convolutional layer of the twentieth basic block.

[0073] As Figure 5As shown, the internal structure of the third basic block includes:

[0074] A third convolutional layer and an attention layer connected in sequence;

[0075] The data obtained by adding the input end of the third convolutional layer and the output end of the attention layer is output from the output end of the third basic block.

[0076] The output end of the second basic block is connected to the input end of the third convolutional layer of the third basic block. The output end of the attention layer of the third basic block is connected to the input end of the fourth basic block. The output end of the fifth basic block is connected to the input end of the third convolutional layer of the sixth basic block. The output end of the attention layer of the sixth basic block is connected to the input end of the seventh basic block. The output end of the eighth basic block is connected to the input end of the third convolutional layer of the ninth basic block. The output end of the attention layer of the ninth basic block is connected to the input end of the tenth basic block. The output end of the tenth basic block is connected to the input end of the third convolutional layer of the eleventh basic block. The output end of the attention layer of the eleventh basic block is connected to the input end of the twelfth block. The output end of the fourteenth basic block is connected to the input end of the third convolutional layer of the fifteenth basic block. The data at the output ends of the attention layers of the sixth basic block and the fifteenth basic block are concatenated and connected to the input end of the sixteenth basic block. The output end of the seventeenth basic block is connected to the input end of the third convolutional layer of the eighteenth basic block. The data at the output ends of the attention layers of the third basic block and the eighteenth basic block are concatenated and connected to the input end of the nineteenth basic block. The output end of the twentieth is connected to the input end of the third convolutional layer of the twenty-first basic block. The output end of the attention layer of the twenty-first basic block and the output end of the nineteenth basic block are the output trajectories.

[0077] As one or more embodiments, the training steps of the diffusion model network include:

[0078] Construct a training set, which is an offline data set that tracks multiple trajectories with known storage paths;

[0079] Construct a diffusion model network;

[0080] Input the training set into the constructed diffusion model network, and train the diffusion model network. When the loss function of the diffusion model reaches the minimum value, stop training to obtain a trained diffusion model network. The loss function is a reconstruction loss function, denoted as

[0081]

[0082] where \(x_0(\tau)\) is the noise-free trajectory sampled from the offline data set, and \(x\) θ is the diffusion model network with \(\theta\) as the parameter.

[0083] Furthermore, as one or more embodiments, the guided sampling process includes:

[0084] Sampling under the gradient guidance of noise-free trajectory estimation, mean estimation, return estimation, and initial state limitation.

[0085] For the noise-free trajectory estimation, the trajectory at time step k and the time step k are input into the trained diffusion model network to obtain the estimated noise-free trajectory, that is

[0086] x0(τ) = x θ (x k (τ), k);

[0087] For the mean estimation, the mean is estimated according to the noise-free trajectory estimation and the trajectory at time step k Its expression is

[0088]

[0089] where α k = 1 - β k , β k is a constant between 0 and 1;

[0090] For the sampling under the gradient guidance of the return estimation, the expression of its gradient g is

[0091]

[0092] where γ is the reward discount factor, s t , a t , are the state, action, and reward function of the robot at time t respectively, is the total cumulative discounted reward. The reward function is defined in the form of the potential energy difference between the current state and the previous state, and its expression is

[0093]

[0094] where is the potential energy function describing the distance between the robot and the tracked window, and its expression is

[0095]

[0096] where (( R x i , R y i ) represents the coordinates of the points within the target window at the current moment, and w i represents the weights assigned to different points within the tracked window. The expression for guided sampling is

[0097]

[0098] The initial state constraint is to reset the initial state of the trajectory using the current state;

[0099] Repeat the processes of noise-free trajectory estimation, mean estimation, sampling under the gradient guidance of return estimation, and initial state constraint K times to obtain a noise-free trajectory x0, and select the first action in x0 as the action instruction output by the upper-level policy.

[0100] S103: Input the action instruction into the low-level controller based on traditional planning to obtain the position instructions of all servos of the salamander robot.

[0101] As one or more embodiments, the low-level controller based on traditional planning includes:

[0102] A spine controller and a leg controller.

[0103] The spine controller periodically assigns an angular position to each spine joint. Setting the spine controller in the form of a sine signal can be defined as

[0104]

[0105] where i = 1, 2, 3 is the spine joint index, q s,i (t) represents the control angle of the spine joint at time t, a i and respectively represent the amplitude of the sine signal and the offset term used to control the turning radius of the robot, f is the oscillation frequency of the spine, which needs to be coordinated with the leg movement, and φ i represents the initial angle and is related to the configuration of the spine joint.

[0106] The leg controller can usually solve for the joint positions based on the inverse kinematics given the end-effector motion trajectory of the leg, so as to achieve the purpose of tracking. However, due to the redundancy of the degrees of freedom of the salamander robot's legs, the Jacobian matrix cannot be directly inverted, and the pseudo-inverse method is unstable near the singular points. Here, solving the inverse kinematics is transformed into finding the optimal Δq to minimize the following formula:

[0107] ||Δp - JΔq|| 2 + λ||Δq|| 2

[0108] where Δq represents the difference between the target and the current joint position, Δp represents the difference between the target and the current end-effector position in the Cartesian space, J represents the Jacobian matrix, and λ is a damping term to prevent the joint from reaching the singular point. The above formula can be solved by setting the partial derivative of Δq to zero to obtain

[0109]

[0110] The salamander - like robot mimics the movement of a salamander. Optionally, it adopts a periodic crawling gait with a duty cycle of 0.75. The salamander - like robot combines the above - mentioned leg controller with a spine controller that generates sine signals, and always keeps the projection of the center of gravity within the support triangle area formed by three legs during the standing phase, achieving synchronization between the spine and the legs while maintaining static stability.

[0111] Embodiment Two

[0112] This embodiment provides a path - tracking system for a salamander - like robot based on a guided diffusion model;

[0113] A path - tracking system for a salamander - like robot based on a guided diffusion model includes:

[0114] An acquisition module, which is configured to: acquire path status information;

[0115] An upper - layer policy module, which is configured to: input the path status information into a trained upper - layer policy based on guided diffusion to obtain an action instruction;

[0116] A lower - layer controller module, which is configured to: input the action instruction into a lower - layer controller based on traditional planning to obtain position instructions for all the servos of the salamander - like robot.

[0117] It should be noted here that the above - mentioned acquisition module and saliency target detection module correspond to steps S101 to S103 in Embodiment One. The examples and application scenarios implemented by the above - mentioned modules and the corresponding steps are the same, but are not limited to the content disclosed in Embodiment One above. It should be noted that the above - mentioned modules, as part of the system, can be executed in a computer system such as a set of computer - executable instructions.

[0118] In the above - mentioned embodiments, the descriptions of each embodiment have their own focuses. For parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0119] The proposed system can be implemented in other ways. For example, the above - described system embodiments are merely illustrative. For example, the above - mentioned module division is only a logical function division. In actual implementation, there can be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.

[0120] Embodiment Three

[0121] This embodiment also provides an electronic device, including: one or more processors, one or more memories, and one or more computer programs; wherein, the processor is connected to the memory, and the one or more computer programs are stored in the memory. When the electronic device runs, the processor executes the one or more computer programs stored in the memory, so that the electronic device executes the method described in Embodiment 1 above.

[0122] It should be understood that in this embodiment, the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0123] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.

[0124] In the implementation process, each step of the above method may be completed by the integrated logic circuit in the hardware of the processor or the instructions in the form of software.

[0125] The method in Embodiment 1 may be directly embodied as being executed and completed by the hardware processor, or executed and completed by a combination of the hardware and software modules in the processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.

[0126] Those of ordinary skill in the art can realize that, in combination with the units and algorithm steps of the examples described in this embodiment, they can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0127] Embodiment 4

[0128] This embodiment also provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the method described in Embodiment 1 is completed.

[0129] The above are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A path tracking method for a salamander-like robot based on a guided diffusion model, characterized in that, Including: Obtain path status information; Input the path status information into the trained upper-level policy based on guided diffusion to obtain action instructions; Input the action instructions into the lower-level controller based on traditional planning to obtain the position instructions of all servos of the salamander-like robot.

2. The path tracking method of the salamander-like robot based on the guided diffusion model according to claim 1, characterized in that The upper-level policy based on guided diffusion includes: Trajectory initialization and guided sampling of the diffusion model neural network; In the trajectory initialization, the trajectory expression form is a two-dimensional array concatenated by an action sequence and a state sequence, and the actions and states are determined by the action space and the state space.

3. A path tracking method for a salamander-like robot based on a guided diffusion model according to claim 2, characterized in that, The action space is an abstract motion command. For the salamander-like robot, the motion command is realized by adjusting the step length of the legs and the offset of the spine; The state space is the coordinates of all scattered points inside the tracking window on the target path in the robot coordinate system; The criteria for selecting the tracking window are to always track the nearest window and always update the window along the target path towards the path end point; after determining the target window, all points in front of the target window on the target path are deleted; when the distance between the robot and the target window meets the preset minimum distance, it is considered that the window has been fully tracked, and the points inside the window are deleted from the path; the entire path tracking process is regarded as terminated when the number of remaining points on the target path is less than the window size.

4. The path tracking method of the salamander-like robot based on the guided diffusion model according to claim 2, characterized in that, The guided sampling process includes: Sampling under the gradient guidance of noise-free trajectory estimation, mean estimation, and return estimation, and initial state restriction.

5. A path tracking method for a salamander-like robot based on a guided diffusion model according to claim 1, characterized in that, The lower-level controller based on traditional planning includes: A spine controller and a leg controller, The spine controller periodically assigns an angular position to each spine joint and sets the spine controller in the form of a sine signal; The leg controller transforms the solution of inverse kinematics into finding the optimal difference between the target and the current position of the joint, and solves it by setting the partial derivative of the difference between the target and the current position of the joint to zero; The salamander-like robot imitates the movement of a salamander and adopts a periodic crawling gait with a duty cycle of 0.

75. The salamander-like robot uses the above-mentioned leg controller in cooperation with a spine controller that generates a sine signal, and always keeps the projection of the center of gravity within the support triangle area formed by three legs during the standing phase, realizing the synchronization of the spine and the legs while maintaining static stability.

6. The path tracking method of the salamander-like robot based on the guided diffusion model according to claim 2, characterized in that The training steps of the diffusion model neural network include: Construct a training set, which is an offline data set that stores multiple trajectories of known paths; Construct a diffusion model network; Input the training set into the constructed diffusion model network, train the diffusion model network, and stop training when the loss function of the diffusion model reaches the minimum value to obtain a trained diffusion model network.

7. A method for path tracking of a salamander-like robot based on a guided diffusion model according to claim 4, characterized in that For the sampling under the gradient guidance of the return estimation, the gradient g expression is where γ is the reward discount factor, s t , a t , are the state, action, and reward function of the robot at time t, is the total cumulative discounted reward; where the reward function is defined in the form of the potential energy difference between the current state and the previous state, and the expression is Among them is the potential energy function for describing the distance between the robot and the tracked window, and its expression is Among them ( R x i , R y i ) represents the coordinates of the points within the target window at the current moment, and w i represents the weights assigned to different points within the tracked window.

8. A path tracking system for a salamander-like robot based on a guided diffusion model, which is used to implement the method described in any one of claims 1-7, characterized in that, Including: An acquisition module, which is configured to: obtain path status information; An upper-level policy module, which is configured to: input the path status information into the trained upper-level policy based on guided diffusion to obtain action instructions; The underlying controller module is configured to: input an action instruction into the underlying controller based on traditional planning to obtain the position instructions of all the servos of the salamander-like robot.

9. An electronic device, characterized in that, Comprising: One or more processors, one or more memories, and one or more computer programs; wherein, the processor is connected to the memory, and the one or more computer programs are stored in the memory. When the electronic device runs, the processor executes the one or more computer programs stored in the memory so that the electronic device executes the method described in any one of claims 1-7 above.

10. A computer-readable storage medium for storing computer instructions, characterized in that: When the computer instructions are executed by the processor, the method described in any one of claims 1-7 is completed.