Robot swing arm trajectory optimization method based on dynamic planning
By building a central mode generator model and dynamic planning decision-making mechanism, the swing trajectory of the humanoid robot is optimized in real time, and the problems of insufficient adaptability and trajectory optimization in the existing technology are solved, and high-precision and stable swing trajectory control are achieved.
Patent Information
- Application Number
- CN202510835950.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-08-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing humanoid robot arm swing control method is difficult to adapt to multi-dimensional perturbation in complex environments, lacks adaptive adjustment capabilities, feedback lags and trajectory optimization lacks globality, resulting in insufficient posture stability and trajectory accuracy.
By building a central mode generator model, combining attitude disturbance modeling and dynamic programming decision-making mechanisms, we collect attitude state data in real time to build disturbance vectors, evaluate track stability and generate trajectory correction parameters, and realize dynamic planning to search for the optimal swing arm trajectory.
It improves the posture stability and trajectory accuracy of the robot in dynamic running scenarios, has good real-time responsiveness and engineering applicability, and significantly improves the adaptive adjustment and coordination of swing arm movement.
Smart Images

Figure CN120439306A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of robot control and intelligent motion planning, and in particular to a robot swing arm trajectory optimization method based on dynamic programming. Background Art
[0002] With the continuous expansion of humanoid robot application scenarios, their movement ability and posture stability in complex environments have become the focus of research and engineering practice. As a key link for humanoid robots to maintain balance and coordinate running, arm swinging motion is of great significance for improving the dynamic stability and movement flexibility of robots. At present, common humanoid robot arm swinging control methods mostly use rhythmic generation technology based on central pattern generators (CPGs) or fixed trajectories to achieve basic arm swinging motion drive. Although these methods can ensure the basic coordination of the robot's gait in standard experimental environments, under dynamic working conditions such as actual running, the robot faces complex influences such as various posture disturbances, trajectory deviations, and changes in the center of gravity, resulting in obvious limitations of existing control technologies.
[0003] In existing technologies, optimization of the robot's swing arm trajectory suffers from the following main deficiencies: Poor adaptability: Traditional control methods often rely on static parameters or single sensor feedback, making it difficult to timely perceive and respond to multi-dimensional disturbances and anomalies that occur during the robot's complex motion, and lack the ability to adaptively adjust to sudden dynamic changes. Feedback lag and difficulty in prediction: Most existing methods use real-time feedback correction, lacking the ability to predict and optimize future motion trends. This results in delayed adjustment of the robot's swing arm trajectory and reduced stability when the track becomes unstable or there is external interference. Trajectory optimization lacks globality and systematicity: Some control strategies rely primarily on empirical rules or single-step adjustments, failing to systematically plan the optimal trajectory correction path over multiple cycles based on the evolution trend of the disturbance, resulting in unsatisfactory optimization results. Ineffective coupling of disturbances and trajectory parameters: Traditional solutions fail to quantify posture disturbances as dynamic adjustment inputs for specific trajectory parameters, ignoring the inherent relationship between disturbances and control parameters, which limits control accuracy and coordination.
[0004] Therefore, how to provide a robot swing arm trajectory optimization method based on dynamic programming is a problem that technicians in this field urgently need to solve. Summary of the Invention
[0005] One purpose of the present invention is to propose a humanoid robot arm swing trajectory optimization method based on dynamic programming. The present invention integrates central pattern generator modeling, posture perturbation modeling and dynamic programming decision-making mechanism, and describes in detail the construction of perturbation vectors by real-time collection of posture state data during the running of the humanoid robot, and the use of track stability indicators to trigger the generation of trajectory correction parameters. Finally, the control strategy of searching for the optimal arm swing trajectory through dynamic programming has the advantages of timely trajectory adjustment response, strong adaptability and high posture control accuracy.
[0006] A robot arm trajectory optimization method based on dynamic programming according to an embodiment of the present invention includes the following steps:
[0007] S1. Construct a central pattern generator model to generate rhythmic arm swing trajectories including arm swing frequency parameters, arm swing amplitude parameters, and arm swing phase parameters;
[0008] S2. Acquire posture state data of the humanoid robot during running;
[0009] S3, converting the rhythmic arm swing trajectory and posture state data into a posture disturbance input vector;
[0010] S4. Calculate a periodic orbit stability index based on the attitude disturbance input vector and determine whether it exceeds a preset deviation threshold;
[0011] S5. When the periodic orbit stability index exceeds a preset deviation threshold, the controllable periodic orbit adjustment algorithm is called to generate a set of candidate trajectory adjustment parameters;
[0012] S6. Construct a dynamic programming model, use the attitude disturbance input vector as the state variable, and the trajectory adjustment candidate parameter set as the strategy space, perform trajectory optimization strategy search, and output the trajectory correction parameter set;
[0013] S7. Input the trajectory correction parameter set into the central pattern generator model, and output a corrected rhythmic arm swing trajectory for driving the humanoid robot's next cycle of arm swing motion.
[0014] The present invention generates rhythmic arm swing trajectories by constructing a central pattern generator model, and combines multi-source attitude state data such as attitude angle offset, arm swing trajectory deviation, and center of gravity change rate during the robot's running process to construct an attitude disturbance input vector and evaluate the periodic orbit stability index in real time. When the stability index exceeds the threshold, the system calls the controllable periodic orbit adjustment algorithm to generate multiple correction candidate parameters, and performs trajectory optimization search on them using a dynamic programming model, outputting the optimal correction parameter set to update the arm swing trajectory for the next cycle. This method realizes adaptive adjustment and closed-loop control of arm swing motion, significantly improving the attitude stability, trajectory accuracy, and movement coordination of humanoid robots in dynamic running scenarios, and has good real-time responsiveness and engineering applicability.
[0015] Optionally, the S1 specifically includes:
[0016] S11. Constructing a neural oscillatory network structure of a central pattern generator model. The central pattern generator model includes a disturbance mapping layer, a trajectory decoupling control layer, an asymmetric coupling layer, a trajectory correction input layer, and an output generation layer from the input layer to the output layer.
[0017] S12, the disturbance mapping layer includes three input nodes, which are used to receive the attitude angle offset value, the swing arm trajectory deviation value and the center of gravity change rate value respectively, and map the received data into the trajectory adjustment factor vector;
[0018] S13. The trajectory decoupling control layer includes a plurality of three-parameter oscillator units, each of which is used to output a rhythmic arm swing trajectory signal. The rhythmic arm swing trajectory signal is:
[0019] θ i (t) = A i (t)·sin(2πf i (t)t+φ i (t));
[0020] Among them, A i (t) is the swing arm amplitude parameter, f i (t) is the swing frequency parameter, φ i (t) is the swing arm phase parameter;
[0021] S14, the asymmetric coupling layer includes multiple phase connection units, which construct a coupling graph structure between the oscillators and set coupling weights based on a dynamic adjustment mechanism of the phase difference between the left and right channels. The coupling weights are updated in real time driven by a trajectory adjustment factor vector;
[0022] S15. The trajectory correction input layer includes three trajectory parameter receiving channels, which are used to receive frequency correction value, amplitude correction value and phase correction value respectively, and superimpose the trajectory correction parameter set with the original trajectory parameters in the trajectory decoupling control layer to generate updated trajectory parameters;
[0023] S16. The output generation layer includes multiple trajectory output channels for outputting updated rhythmic arm swing trajectories.
[0024] The present invention realizes the structured design of the central pattern generator model by constructing a neural oscillator network including a disturbance mapping layer, a trajectory decoupling control layer, an asymmetric coupling layer, a trajectory correction input layer and an output generation layer. The system receives multi-source disturbance information such as attitude angle offset, swing arm trajectory deviation and center of gravity change rate, maps it into trajectory adjustment factors, and drives the oscillator to generate a rhythmic swing arm trajectory signal containing frequency, amplitude and phase. The coupling layer realizes dynamic coordination of multiple oscillators through a phase difference adjustment mechanism, and the correction input layer further integrates the adjustment parameters to complete the trajectory update, and finally outputs a high-precision control signal. This method effectively improves the adaptability and coordination of the swing arm trajectory generation.
[0025] Optionally, the S2 specifically includes: obtaining the posture state data of the humanoid robot during running, including: collecting the posture angle offset value through the inertial measurement unit installed on the torso, collecting the swing arm trajectory deviation value through the angle encoder set at the swing arm joint, and collecting and calculating the center of gravity change rate value through the pressure sensors and ankle accelerometers distributed under the feet. The three types of data are aligned by time stamps to form a posture state data vector.
[0026] This method installs an inertial measurement unit on the robot's trunk to collect attitude angle offsets, places angle encoders at the arm joints to obtain arm trajectory deviations, and deploys pressure sensors and ankle accelerometers under both feet to calculate the rate of change of center of gravity. These three types of data are timestamped and then constructed into an attitude state data vector. This enables high-precision, multi-dimensional perception of the robot's posture changes during running. This method effectively improves the timeliness and integrity of attitude data.
[0027] Optionally, the S3 specifically includes:
[0028] S31, extracting the rhythmic arm swing trajectory generated by the central pattern generator model as the ideal arm swing trajectory reference value;
[0029] S32, obtaining the actual trajectory value of the swing arm measured by the angle encoder, and calculating the difference between the ideal swing arm trajectory reference value and the actual trajectory value at each time point to obtain the swing arm trajectory error component;
[0030] S33, reading the data of the three axial angular velocity sensors in the inertial measurement unit, calculating the current attitude angle offset value by integration, performing a difference operation on the attitude angle offset value and the set reference attitude angle to obtain the attitude angle disturbance component;
[0031] S34, extracting the plantar contact force data from the pressure sensor and the vertical acceleration data from the ankle accelerometer, calculating the center of gravity change rate by the rate of change within the continuous time window, and subtracting the average rate value in the initial equilibrium state of the movement to obtain the center of gravity disturbance component;
[0032] S35. Scale the above three disturbance components to a uniform numerical range according to the normalization factor, align them according to the timestamp, and combine them to construct a posture disturbance input vector. The posture disturbance input vector is used to describe the degree of deviation of the current motion state from the ideal posture.
[0033] The present invention extracts the rhythmic arm swing trajectory output by the central pattern generator as the ideal trajectory reference, and combines it with the actual arm swing trajectory obtained by real-time measurement of the angle encoder to calculate the difference between the two to generate a trajectory error component; at the same time, the inertial measurement unit is used to obtain angular velocity data and integrate it to obtain the attitude angle, which is then compared with the reference attitude angle to obtain the attitude angle disturbance component; the center of gravity change rate is further calculated from the plantar pressure and ankle acceleration, and the center of gravity disturbance component is obtained after deducting the initial equilibrium state rate. The three disturbance components are normalized and time-aligned and then combined into a posture disturbance input vector, which is used to quantify the degree of deviation of the current motion state. This method realizes the precise modeling of posture disturbances, improves the sensitivity of the disturbance response mechanism and the pertinence of the control strategy.
[0034] Optionally, the S4 specifically includes:
[0035] S41, decomposing the attitude disturbance input vector generated in the current cycle into three disturbance components, namely, attitude angle disturbance component, swing arm trajectory error component and center of gravity disturbance component;
[0036] S42, calculating the root mean square value of each disturbance component within the sliding time window to obtain a stability assessment value of each disturbance component;
[0037] S43. Assign a preset weight factor to each stability evaluation value, and calculate a periodic orbit stability index according to a weighted summation method. The orbit stability index is in the following form: the orbit stability index is equal to the root mean square value of the attitude angle disturbance multiplied by a first weight, the root mean square value of the swing arm trajectory error multiplied by a second weight, and the root mean square value of the center of gravity disturbance multiplied by a third weight.
[0038] S44, comparing the periodic orbit stability index with the orbit stability deviation threshold set in the current motion mode;
[0039] S45. When the track stability index is less than or equal to the deviation threshold, the current rhythmic trajectory output is maintained unchanged; when the track stability index is greater than the deviation threshold, it is marked as a track instability state and the trajectory correction parameter generation process is entered.
[0040] The present invention decomposes the attitude disturbance input vector into three components: attitude angle disturbance, swing arm trajectory error, and center of gravity disturbance. The root mean square value of each component is calculated within a sliding time window to evaluate the disturbance intensity. Different weights are then assigned to each component, and a weighted summation method is used to construct a periodic orbit stability index. This index can comprehensively reflect the stability of the current motion state and compare and judge it with the preset threshold. When the index does not exceed the threshold, the system maintains the original trajectory output. If the threshold is exceeded, it is considered that the orbit is unstable, triggering the subsequent trajectory correction process. This method effectively improves the dynamic perception capability of orbit stability, realizes automatic stability monitoring based on quantitative evaluation of multi-source disturbances, provides a precise trigger mechanism for trajectory optimization, and enhances the response intelligence and steady-state recognition capability of the control system.
[0041] Optionally, the S5 specifically includes:
[0042] S51, constructing a historical sequence of attitude disturbance input vectors, extracting attitude disturbance input vectors of the current cycle and the previous multiple consecutive cycles, calculating the first-order derivative and the second-order derivative respectively, and forming a disturbance velocity sequence and a disturbance acceleration sequence;
[0043] S52, based on the disturbance velocity sequence and the disturbance acceleration sequence, the attitude disturbance input vector of the next T cycles is predicted by the sliding regression model to obtain the disturbance prediction sequence
[0044] S53, inputting the disturbance prediction sequence into the orbit stability index function for prediction and judgment, and activating the controllable periodic orbit adjustment algorithm when the predicted stability index of any period exceeds the deviation threshold;
[0045] S54, the controllable periodic orbit adjustment algorithm includes the following processing flow:
[0046] S541, set the trajectory control parameter boundary and the maximum adjustment increment, including the swing arm frequency parameter, swing arm amplitude parameter and swing arm phase parameter. The current cycle values are f t 、A t 、φ t , the adjustable range is set to [f min ,f max ]、[A min ,A max ]、[φ min ,φ max ], the corresponding maximum change is
[0047] S542. Based on the three disturbance components in the disturbance prediction sequence, a disturbance response direction is constructed according to a fixed mapping rule, namely, the attitude angle disturbance prediction component is mapped to the swing arm phase parameter adjustment direction, the swing arm trajectory error prediction component is mapped to the swing arm amplitude parameter adjustment direction, and the center of gravity disturbance prediction component is mapped to the swing arm frequency parameter adjustment direction.
[0048] S543. Construct multiple candidate trajectory parameter combinations within the parameter adjustment constraints, each candidate combination including a frequency correction value, an amplitude correction value, and a phase correction value;
[0049] S544. Each candidate trajectory parameter combination and the corresponding disturbance prediction vector combination are input into a stability index function for scoring calculation. The stability score calculation method is as follows: the difference between the attitude angle disturbance prediction value and the phase correction value is multiplied by a first weighting coefficient, the difference between the swing arm trajectory error prediction value and the amplitude correction value is multiplied by a second weighting coefficient, and the difference between the center of gravity disturbance prediction value and the frequency correction value is multiplied by a third weighting coefficient. The three results are summed to obtain the stability score of the parameter combination.
[0050] S545 , screening out several trajectory parameter combinations with the best scores and satisfying parameter boundary restrictions to form a trajectory adjustment candidate parameter set, wherein the trajectory adjustment candidate parameter set includes multiple frequency correction value, amplitude correction value and phase correction value triplets for optimizing the input.
[0051] The present invention constructs a historical sequence of attitude disturbance input vectors, extracts disturbance data for the current and previous consecutive cycles, calculates its first-order and second-order derivatives to form disturbance velocity and acceleration sequences, and uses a sliding regression model to predict disturbance trends for the next T cycles. When the predicted stability index exceeds the deviation threshold, the system activates the controllable periodic orbit adjustment algorithm and generates multiple sets of trajectory correction parameter combinations based on the disturbance response mapping relationship within the set frequency, amplitude and phase parameter boundaries. Each set of candidate parameters is scored by combining the predicted disturbance with the input stability function, and its stability index is calculated by weighted difference. Several correction parameters with the best scores and meeting the constraints are selected to form a candidate set for trajectory adjustment. This method realizes a forward-looking trajectory adjustment design for disturbance prediction, effectively improving the initiative and stability response capability of trajectory control.
[0052] Optionally, the S6 specifically includes:
[0053] S61, inputting the current periodic attitude disturbance input vector into the dynamic programming model to generate an initial state node;
[0054] S62, inputting the trajectory adjustment candidate parameter set into the dynamic programming model to generate an optional strategy node;
[0055] S63, performing forward recursion on each policy node in turn according to the state transition function and cost function defined in the dynamic programming model to form a state transition path that is expanded cycle by cycle;
[0056] S64. Record the correspondence between the state nodes and the policy nodes in each state transition path, and construct a state policy path tree;
[0057] S65. Based on the state strategy path tree, traverse backward from the end cycle state node to the current cycle initial state node to determine the strategy path with the lowest global cost;
[0058] S66 , extracting the trajectory adjustment candidate parameter triples corresponding to the current cycle in the global optimal path, and generating a trajectory correction parameter set.
[0059] The present invention inputs the posture disturbance input vector of the current cycle as the initial state into the dynamic programming model, and uses the generated trajectory adjustment candidate parameter set as the strategy space. Relying on predefined state transition functions and cost functions, forward recursion is performed on each strategy node to form a state transition path that includes multi-cycle state evolution. The system records the correspondence between each state and strategy in the path, constructs a complete state strategy path tree, and searches for the global optimal strategy path with the minimum total cost by traversing backward from the terminal state node. The optimal parameter triple corresponding to the current cycle is extracted to generate a trajectory correction parameter set. This method takes global optimization as the goal, effectively avoids the problem of local adjustment error accumulation, and significantly improves the accuracy of trajectory correction and the stability of gait control.
[0060] Optionally, the dynamic programming model specifically includes:
[0061] Receive the attitude disturbance input vector, which contains three components: attitude angle disturbance component, swing arm trajectory error component and center of gravity disturbance component, as the current state variable S in the state space t ;
[0062] Read the candidate parameter set for trajectory adjustment. Each parameter set in the candidate parameter set for trajectory adjustment consists of frequency correction value, amplitude correction value and phase correction value, which constitute the action variable u in the strategy space. t ;
[0063] Construct a state transfer function, input the state variables and action variables into the disturbance mapping matrix, and predict the state variables of the next cycle. The state transfer function is defined as:
[0064] S t+1 =S t +M·u t ;
[0065] in, represents the disturbance response matrix, ut is the current strategy vector, S t is the current state vector;
[0066] The cost function is constructed based on the state variables. The cost function consists of the state deviation cost term and the policy disturbance penalty term, which is defined as follows:
[0067]
[0068] Among them, S ref is the target reference state, λ is the regularization coefficient;
[0069] For each feasible strategy combination u t , perform state transfer and cost accumulation within the set time range T to form a strategy sequence and state path mapping;
[0070] Traverse all policy paths, select the policy sequence with the minimum total cost function value, and extract the optimal policy vector for the current cycle from the sequence The output is a set of trajectory correction parameters.
[0071] The dynamic programming model described in the present invention constructs a state space and a policy space by taking the attitude disturbance input vector as the current state variable and the trajectory adjustment candidate parameters composed of frequency, amplitude, and phase correction values as action variables. The system constructs a state transfer function based on the disturbance response matrix to predict the state of the next cycle under the action of the strategy. At the same time, a cost function containing state deviation terms and policy disturbance terms is designed to measure the impact of each step of the strategy on stability. Within a set time range, state transfer and cost accumulation are performed on all strategy combinations to form a policy path mapping, and the path with the minimum total cost is selected from it. The optimal policy vector corresponding to the current cycle is extracted as the trajectory correction parameter output. This method realizes the global optimization matching from disturbance input to correction parameters, improving the stability of attitude control and the precision of action adjustment.
[0072] Optionally, the S7 specifically includes: inputting the trajectory correction parameter set into the central pattern generator model to output a corrected rhythmic arm swing trajectory, including: weighted superposition of the frequency correction value, amplitude correction value and phase correction value in the trajectory correction parameter set with the current frequency parameter, amplitude parameter and phase parameter of the oscillator corresponding to the output layer of the central pattern generator model, and the updated parameters are propagated through the internal coupling of the model to generate rhythmic arm swing control signals for the left arm and the right arm, which are used to drive the synchronous output of the left and right arm swings of the humanoid robot in the next cycle.
[0073] The present invention inputs the trajectory correction parameter set output by dynamic programming into a central pattern generator model. The resulting frequency, amplitude, and phase correction values are then weighted and superimposed with the current frequency, amplitude, and phase parameters of each oscillator in the model's output layer to form updated trajectory control parameters. After the updated parameters propagate through the coupling mechanism within the model, rhythmic swing control signals are generated for the left and right arms, respectively, achieving synchronous drive for the next swing cycle. This method achieves smooth transition and rapid response of the control signal by deeply integrating the correction results with the rhythmic control model.
[0074] The beneficial effects of the present invention are:
[0075] Unlike existing arm swing control methods based on fixed trajectories or simple feedback regulation, this invention builds a central pattern generator model that integrates a disturbance mapping layer, a trajectory decoupling control layer, and an asymmetric coupling mechanism to enable dynamic response in rhythmic arm swing trajectories. By introducing a posture perturbation input vector, the posture angle offset, arm swing trajectory error, and center of gravity change rate are unified in the model, enabling comprehensive perception and state representation of multi-source disturbances. This enhances the robot's perceptual robustness in complex running environments and the real-time adaptability of trajectory generation.
[0076] This paper, for the first time, decouples orbit stability evaluation from the generation of trajectory correction parameters, proposing a controllable periodic orbit adjustment algorithm based on a disturbance prediction sequence. This algorithm can predict potential instability risks at the earliest stages of disturbance trends and generate candidate parameter combinations. By explicitly mapping the adjustment directions between disturbance components and swing arm parameters (frequency, amplitude, and phase), a correction strategy space is systematically constructed, providing structured support for subsequent dynamic programming optimization.
[0077] During the optimization decision phase, the present invention utilizes a dynamic programming model with perturbation inputs as state variables and trajectory adjustment parameters as the action space. Through forward expansion and backtracking along the path with the minimum total cost, the method accurately searches for the globally optimal trajectory correction strategy. This mechanism significantly differs from existing approaches based on heuristic local adjustments, possessing multi-cycle planning capabilities and optimality guarantees, significantly improving the robot's arm swing control quality and posture stability during complex motion tasks.
[0078] Through the coupled control of trajectory correction parameters and the central pattern generator, the present invention achieves real-time synchronous updating and coordinated driving of the robot's left and right arm swing movements, ensuring the continuity, symmetry and rhythmicity of the trajectory during movement, providing a solid motion control foundation for high-intensity dynamic tasks, and has the practical application advantages of high stability, strong robustness and high control accuracy. It is particularly suitable for humanoid robot application scenarios such as high-speed running, human-machine collaboration and strict dynamic balance requirements. BRIEF DESCRIPTION OF THE DRAWINGS
[0079] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0080] Figure 1 This is a flow chart of a robot arm swing trajectory optimization method based on dynamic programming proposed by the present invention;
[0081] Figure 2 Schematic diagram of the neural oscillation structure of the central pattern generator model proposed in this invention;
[0082] Figure 3 This is a processing flow chart of the controllable periodic orbit adjustment algorithm proposed in the present invention. DETAILED DESCRIPTION
[0083] The present invention will now be described in further detail with reference to the accompanying drawings, which are simplified schematic diagrams that illustrate the basic structure of the present invention in a schematic manner.
[0084] refer to Figure 1-3 , a robot arm swing trajectory optimization method based on dynamic programming, comprising the following steps:
[0085] S1. Construct a central pattern generator model to generate rhythmic arm swing trajectories including arm swing frequency parameters, arm swing amplitude parameters, and arm swing phase parameters;
[0086] S2. Acquire posture state data of the humanoid robot during running;
[0087] S3, converting the rhythmic arm swing trajectory and posture state data into a posture disturbance input vector;
[0088] S4. Calculate a periodic orbit stability index based on the attitude disturbance input vector and determine whether it exceeds a preset deviation threshold;
[0089] S5. When the periodic orbit stability index exceeds a preset deviation threshold, the controllable periodic orbit adjustment algorithm is called to generate a set of candidate trajectory adjustment parameters;
[0090] S6. Construct a dynamic programming model, use the attitude disturbance input vector as the state variable, and the trajectory adjustment candidate parameter set as the strategy space, perform trajectory optimization strategy search, and output the trajectory correction parameter set;
[0091] S7. Input the trajectory correction parameter set into the central pattern generator model, and output a corrected rhythmic arm swing trajectory for driving the humanoid robot's next cycle of arm swing motion.
[0092] In this embodiment, S1 specifically includes:
[0093] S11. Constructing a neural oscillatory network structure of a central pattern generator model. The central pattern generator model includes a disturbance mapping layer, a trajectory decoupling control layer, an asymmetric coupling layer, a trajectory correction input layer, and an output generation layer from the input layer to the output layer.
[0094] S12, the disturbance mapping layer includes three input nodes, which are used to receive the attitude angle offset value, the swing arm trajectory deviation value and the center of gravity change rate value respectively, and map the received data into the trajectory adjustment factor vector;
[0095] S13. The trajectory decoupling control layer includes a plurality of three-parameter oscillator units, each of which is used to output a rhythmic arm swing trajectory signal. The rhythmic arm swing trajectory signal is:
[0096] θ i (t) = A i (t)·sin(2πf i (t)t+φ i (t));
[0097] Among them, A i (t) is the swing arm amplitude parameter, f i (t) is the swing frequency parameter, φ i (t) is the swing arm phase parameter;
[0098] S14, the asymmetric coupling layer includes multiple phase connection units, which construct a coupling graph structure between the oscillators and set coupling weights based on a dynamic adjustment mechanism of the phase difference between the left and right channels. The coupling weights are updated in real time driven by a trajectory adjustment factor vector;
[0099] S15. The trajectory correction input layer includes three trajectory parameter receiving channels, which are used to receive frequency correction value, amplitude correction value and phase correction value respectively, and superimpose the trajectory correction parameter set with the original trajectory parameters in the trajectory decoupling control layer to generate updated trajectory parameters;
[0100] S16. The output generation layer includes multiple trajectory output channels for outputting updated rhythmic arm swing trajectories.
[0101] In this embodiment, S2 specifically includes: obtaining the posture state data of the humanoid robot during running, including: collecting the posture angle offset value through the inertial measurement unit installed on the torso, collecting the swing arm trajectory deviation value through the angle encoder set at the swing arm joint, and collecting and calculating the center of gravity change rate value through the pressure sensors and ankle accelerometers distributed under the feet. The three types of data are aligned by time stamps to form a posture state data vector.
[0102] In this embodiment, S3 specifically includes:
[0103] S31, extracting the rhythmic arm swing trajectory generated by the central pattern generator model as the ideal arm swing trajectory reference value;
[0104] S32, obtaining the actual trajectory value of the swing arm measured by the angle encoder, and calculating the difference between the ideal swing arm trajectory reference value and the actual trajectory value at each time point to obtain the swing arm trajectory error component;
[0105] S33, reading the data of the three axial angular velocity sensors in the inertial measurement unit, calculating the current attitude angle offset value by integration, performing a difference operation on the attitude angle offset value and the set reference attitude angle to obtain the attitude angle disturbance component;
[0106] S34, extracting the plantar contact force data from the pressure sensor and the vertical acceleration data from the ankle accelerometer, calculating the center of gravity change rate by the rate of change within the continuous time window, and subtracting the average rate value in the initial equilibrium state of the movement to obtain the center of gravity disturbance component;
[0107] S35. Scale the above three disturbance components to a uniform numerical range according to the normalization factor, align them according to the timestamp, and combine them to construct a posture disturbance input vector. The posture disturbance input vector is used to describe the degree of deviation of the current motion state from the ideal posture.
[0108] In this embodiment, the S4 specifically includes:
[0109] S41, decomposing the attitude disturbance input vector generated in the current cycle into three disturbance components, namely, attitude angle disturbance component, swing arm trajectory error component and center of gravity disturbance component;
[0110] S42, calculating the root mean square value of each disturbance component within the sliding time window to obtain a stability assessment value of each disturbance component;
[0111] S43. Assign a preset weight factor to each stability evaluation value, and calculate a periodic orbit stability index according to a weighted summation method. The orbit stability index is in the following form: the orbit stability index is equal to the root mean square value of the attitude angle disturbance multiplied by a first weight, the root mean square value of the swing arm trajectory error multiplied by a second weight, and the root mean square value of the center of gravity disturbance multiplied by a third weight.
[0112] S44, comparing the periodic orbit stability index with the orbit stability deviation threshold set in the current motion mode;
[0113] S45. When the track stability index is less than or equal to the deviation threshold, the current rhythmic trajectory output is maintained unchanged; when the track stability index is greater than the deviation threshold, it is marked as a track instability state and the trajectory correction parameter generation process is entered.
[0114] In this embodiment, the S5 specifically includes:
[0115] S51, constructing a historical sequence of attitude disturbance input vectors, extracting attitude disturbance input vectors of the current cycle and the previous multiple consecutive cycles, calculating the first-order derivative and the second-order derivative respectively, and forming a disturbance velocity sequence and a disturbance acceleration sequence;
[0116] S52, based on the disturbance velocity sequence and the disturbance acceleration sequence, the attitude disturbance input vector of the next T cycles is predicted by the sliding regression model to obtain the disturbance prediction sequence
[0117] S53, inputting the disturbance prediction sequence into the orbit stability index function for prediction and judgment, and activating the controllable periodic orbit adjustment algorithm when the predicted stability index of any period exceeds the deviation threshold;
[0118] S54, the controllable periodic orbit adjustment algorithm includes the following processing flow:
[0119] S541, set the trajectory control parameter boundary and the maximum adjustment increment, including the swing arm frequency parameter, swing arm amplitude parameter and swing arm phase parameter. The current cycle values are f t 、A t 、φ t , the adjustable range is set to [f min ,f max ]、[A min ,A max ]、[φ min ,φ max ], the corresponding maximum change is
[0120] S542. Based on the three disturbance components in the disturbance prediction sequence, a disturbance response direction is constructed according to a fixed mapping rule, namely, the attitude angle disturbance prediction component is mapped to the swing arm phase parameter adjustment direction, the swing arm trajectory error prediction component is mapped to the swing arm amplitude parameter adjustment direction, and the center of gravity disturbance prediction component is mapped to the swing arm frequency parameter adjustment direction.
[0121] S543. Construct multiple candidate trajectory parameter combinations within the parameter adjustment constraints, each candidate combination including a frequency correction value, an amplitude correction value, and a phase correction value;
[0122] S544. Each candidate trajectory parameter combination and the corresponding disturbance prediction vector combination are input into a stability index function for scoring calculation. The stability score calculation method is as follows: the difference between the attitude angle disturbance prediction value and the phase correction value is multiplied by a first weighting coefficient, the difference between the swing arm trajectory error prediction value and the amplitude correction value is multiplied by a second weighting coefficient, and the difference between the center of gravity disturbance prediction value and the frequency correction value is multiplied by a third weighting coefficient. The three results are summed to obtain the stability score of the parameter combination.
[0123] S545 , screening out several trajectory parameter combinations with the best scores and satisfying parameter boundary restrictions to form a trajectory adjustment candidate parameter set, wherein the trajectory adjustment candidate parameter set includes multiple frequency correction value, amplitude correction value and phase correction value triplets for optimizing the input.
[0124] In this embodiment, S6 specifically includes:
[0125] S61, inputting the current periodic attitude disturbance input vector into the dynamic programming model to generate an initial state node;
[0126] S62, inputting the trajectory adjustment candidate parameter set into the dynamic programming model to generate an optional strategy node;
[0127] S63, performing forward recursion on each policy node in turn according to the state transition function and cost function defined in the dynamic programming model to form a state transition path that is expanded cycle by cycle;
[0128] S64. Record the correspondence between the state nodes and the policy nodes in each state transition path, and construct a state policy path tree;
[0129] S65. Based on the state strategy path tree, traverse backward from the end cycle state node to the current cycle initial state node to determine the strategy path with the lowest global cost;
[0130] S66 , extracting the trajectory adjustment candidate parameter triples corresponding to the current cycle in the global optimal path, and generating a trajectory correction parameter set.
[0131] In this embodiment, the dynamic programming model specifically includes:
[0132] Receive the attitude disturbance input vector, which contains three components: attitude angle disturbance component, swing arm trajectory error component and center of gravity disturbance component, as the current state variable S in the state space t ;
[0133] Read the candidate parameter set for trajectory adjustment. Each parameter set in the candidate parameter set for trajectory adjustment consists of frequency correction value, amplitude correction value and phase correction value, which constitute the action variable u in the strategy space. t ;
[0134] Construct a state transfer function, input the state variables and action variables into the disturbance mapping matrix, and predict the state variables of the next cycle. The state transfer function is defined as:
[0135] S t+1 =S t +M·u t ;
[0136] Where M∈R 3×3 represents the disturbance response matrix, u t is the current strategy vector, S t is the current state vector;
[0137] The cost function is constructed based on the state variables. The cost function consists of the state deviation cost term and the policy disturbance penalty term, which is defined as follows:
[0138]
[0139] Among them, S ref is the target reference state, λ is the regularization coefficient;
[0140] For each feasible strategy combination u t , perform state transfer and cost accumulation within the set time range T to form a strategy sequence and state path mapping;
[0141] Traverse all policy paths, select the policy sequence with the minimum total cost function value, and extract the optimal policy vector for the current cycle from the sequence The output is a set of trajectory correction parameters.
[0142] In this embodiment, S7 specifically includes: inputting the trajectory correction parameter set into the central pattern generator model to output a corrected rhythmic swing trajectory, including: weighted superposition of the frequency correction value, amplitude correction value and phase correction value in the trajectory correction parameter set with the current frequency parameter, amplitude parameter and phase parameter of the oscillator corresponding to the output layer of the central pattern generator model, and the updated parameters are propagated through the internal coupling of the model to generate rhythmic swing control signals for the left arm and the right arm, which are used to drive the synchronous output of the left and right swing arms of the humanoid robot in the next cycle.
[0143] Example 1:
[0144] To verify the feasibility of this invention, it was applied to a humanoid robot running posture control platform at the Shanghai Joint Experimental Center for Artificial Intelligence and Robotics. This platform primarily studies the high-speed locomotion capabilities of humanoid robots in complex terrain and under disturbed conditions. Existing platforms use a fixed-parameter central pattern generator (CPG) for arm swing trajectory control. However, when subjected to speed changes or disturbances, the robot often experiences posture imbalance, manifesting as torso tilt, gait instability, or even falls. This problem is particularly prominent during acceleration.
[0145] In this scenario, the researchers integrated the proposed dynamic programming-based humanoid robot swing trajectory optimization method into the control system of the experimental humanoid robot "RT-X2." The system integration specifically includes the following key technical features:
[0146] The central pattern generator model consists of a five-layer neural oscillator architecture, including a perturbation mapping layer, a trajectory decoupling control layer, an asymmetric coupling layer, a trajectory correction input layer, and an output generation layer. The perturbation mapping layer receives input data from three sensor types: attitude angle offset values obtained by a gyroscope, arm trajectory deviation values obtained by an angle encoder, and the center of gravity change rate calculated by a plantar pressure sensor and ankle accelerometer. The data from these three channels are aligned with a unified timestamp and normalized before being combined into an attitude perturbation input vector, which is then fed into the model as a state variable.
[0147] At the end of each motion cycle, the system evaluates the orbital stability metric. If the metric exceeds a set deviation threshold, the controllable periodic orbit adjustment algorithm is invoked. This algorithm constructs a prediction sequence by extracting historical disturbance data. Based on the disturbance trend, it derives the optimal correction direction and constraint boundaries for the three parameters of frequency, amplitude, and phase, generating multiple candidate trajectory correction triplets. Next, the state variables and strategy variables are input into a dynamic programming model. Using the state transition function and cost function, multi-cycle path expansion and minimum total cost search are performed to extract the optimal correction strategy for the current cycle.
[0148] This strategy outputs frequency, amplitude, and phase corrections, which are directly fed into the trajectory correction input layer of the central pattern generator. These are then superimposed with the original trajectory parameters, driving the output layer to generate a new rhythmic control signal. The robot's arms ultimately achieve synchronized swing adjustments based on this rhythmic signal, significantly improving their dynamic response to disturbances.
[0149] The experiment set up three typical running scenarios: steady running at a constant speed (2.5m / s), accelerated running (2.5m / s→3.5m / s), and sudden disturbance (crosswind + 2° ground slope). The experiment lasted two weeks, and the robot ran back and forth 60 times in each scenario, using both traditional fixed trajectory control and the proposed method. The data collected from the experiment is as follows:
[0150] Table 1 Comparison experimental data between the method of the present invention and the traditional method
[0151] Test indicators Traditional methods Method of the present invention Average swing arm trajectory deviation (°) 12.7 4.1 RMS value of attitude angle disturbance (rad) 0.113 0.037 <![CDATA[Rate of change of center of gravity fluctuation (m / s 2 )]]> 0.92 0.35 Fall rate 5.2% 0.3% Action adjustment response time (ms) 250 80 Left and right arm synchronization phase error (%) 9.4% 2.1% Average running time (seconds) 6.87 6.44 Increased energy consumption per unit cycle of the control system / +3.8%
[0152] The experimental results show that the proposed method outperforms traditional fixed-trajectory control methods across all performance indicators. Regarding arm swing trajectory deviation, the average angular deviation was reduced from 12.7° to 4.1°, demonstrating that the proposed method can more accurately generate rhythmic arm swing trajectories during motion, significantly reducing motion deviation. Regarding attitude angle disturbance control, the root mean square value was reduced from 0.113 to 0.037, indicating a significant improvement in robot posture stability and a significant reduction in trunk sway.
[0153] The center of gravity change rate fluctuates from 0.92m / s 2Down to 0.35m / s 2 This demonstrates that the present invention, by intervening in arm swinging motion, effectively suppresses center-of-gravity drift during running and improves dynamic balance. The fall rate under conventional control is 5.2%, while the present invention is only 0.3%, demonstrating that this method significantly reduces the risk of instability and falls, improving the robot's safety in sudden disturbance scenarios.
[0154] The motion adjustment response time was shortened from 250ms to 80ms, indicating faster control feedback, enabling high-frequency and high-sensitivity posture adjustments. Furthermore, the synchronous phase error between the left and right arms was reduced from 9.4% to 2.1%, further demonstrating that the swing trajectory output by the present invention is superior in coordination, effectively ensuring the symmetry and rhythmicity of the robot's movements.
[0155] In terms of overall task completion efficiency, the average time for the robot to complete the running task was reduced from 6.87 seconds to 6.44 seconds, demonstrating the effectiveness of trajectory optimization in promoting gait smoothness and efficiency. Although this method introduces more complex model calculations into the control process, resulting in a 3.8% increase in energy consumption per cycle, the trade-off is an overall improvement in system stability, responsiveness, and safety.
[0156] From the above analysis, it can be seen that the present invention realizes dynamic optimization control of the swing arm trajectory of the humanoid robot through the organic combination of the central pattern generator model and the dynamic programming optimization strategy, successfully solving the problems of large trajectory deviation, poor stability, reaction delay and frequent falls existing in the traditional method, and verifies its excellent application effect and engineering feasibility in complex dynamic environments.
[0157] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.
Claims
1. A robot arm trajectory optimization method based on dynamic programming, characterized in that: The following steps are included: S1. Construct a central pattern generator model to generate rhythmic arm swing trajectories including arm swing frequency parameters, arm swing amplitude parameters, and arm swing phase parameters; S2. Acquire posture state data of the humanoid robot during running; S3, converting the rhythmic arm swing trajectory and posture state data into a posture disturbance input vector; S4. Calculate a periodic orbit stability index based on the attitude disturbance input vector and determine whether it exceeds a preset deviation threshold; S5. When the periodic orbit stability index exceeds a preset deviation threshold, the controllable periodic orbit adjustment algorithm is called to generate a set of candidate trajectory adjustment parameters; S6. Construct a dynamic programming model, use the attitude disturbance input vector as the state variable, and the trajectory adjustment candidate parameter set as the strategy space, perform trajectory optimization strategy search, and output the trajectory correction parameter set; S7. Input the trajectory correction parameter set into the central pattern generator model, and output a corrected rhythmic arm swing trajectory for driving the humanoid robot's next cycle of arm swing motion.
2. The robot arm trajectory optimization method based on dynamic programming according to claim 1, characterized in that: Said S1 specifically includes: S11. Constructing a neural oscillatory network structure of a central pattern generator model. The central pattern generator model includes a disturbance mapping layer, a trajectory decoupling control layer, an asymmetric coupling layer, a trajectory correction input layer, and an output generation layer from the input layer to the output layer. S12, the disturbance mapping layer includes three input nodes, which are used to receive the attitude angle offset value, the swing arm trajectory deviation value and the center of gravity change rate value respectively, and map the received data into the trajectory adjustment factor vector; S13. The trajectory decoupling control layer includes a plurality of three-parameter oscillator units, each of which is used to output a rhythmic arm swing trajectory signal. The rhythmic arm swing trajectory signal is: i i (t)=A i (t)·sin(2πf i (t)t+φ i (t)); Among them, A i (t) is the swing arm amplitude parameter, f i (t) is the swing frequency parameter, φ i (t) is the swing arm phase parameter; S14, the asymmetric coupling layer includes multiple phase connection units, which construct a coupling graph structure between the oscillators and set coupling weights based on a dynamic adjustment mechanism of the phase difference between the left and right channels. The coupling weights are updated in real time driven by a trajectory adjustment factor vector; S15. The trajectory correction input layer includes three trajectory parameter receiving channels, which are used to receive frequency correction value, amplitude correction value and phase correction value respectively, and superimpose the trajectory correction parameter set with the original trajectory parameters in the trajectory decoupling control layer to generate updated trajectory parameters; S16. The output generation layer includes multiple trajectory output channels for outputting updated rhythmic arm swing trajectories.
3. The robot arm trajectory optimization method based on dynamic programming according to claim 1, characterized in that: The S2 specifically includes: obtaining the posture state data of the humanoid robot during running, including: collecting the posture angle offset value through the inertial measurement unit installed on the torso, collecting the swing arm trajectory deviation value through the angle encoder set at the swing arm joint, and collecting and calculating the center of gravity change rate value through the pressure sensors and ankle accelerometers distributed under the feet. The three types of data are aligned by time stamps to form a posture state data vector.
4. The robot arm swing trajectory optimization method based on dynamic programming according to claim 1, characterized in that: The S3 specifically includes: S31, extracting the rhythmic arm swing trajectory generated by the central pattern generator model as the ideal arm swing trajectory reference value; S32, obtaining the actual trajectory value of the swing arm measured by the angle encoder, and calculating the difference between the ideal swing arm trajectory reference value and the actual trajectory value at each time point to obtain the swing arm trajectory error component; S33, reading the data of the three axial angular velocity sensors in the inertial measurement unit, calculating the current attitude angle offset value by integration, performing a difference operation on the attitude angle offset value and the set reference attitude angle to obtain the attitude angle disturbance component; S34, extracting the plantar contact force data from the pressure sensor and the vertical acceleration data from the ankle accelerometer, calculating the center of gravity change rate by the rate of change within the continuous time window, and subtracting the average rate value in the initial equilibrium state of the movement to obtain the center of gravity disturbance component; S35. Scale the above three disturbance components to a uniform numerical range according to the normalization factor, align them according to the timestamp, and combine them to construct a posture disturbance input vector. The posture disturbance input vector is used to describe the degree of deviation of the current motion state from the ideal posture.
5. The robot arm swing trajectory optimization method based on dynamic programming according to claim 1, characterized in that: The S4 specifically includes: S41, decomposing the attitude disturbance input vector generated in the current cycle into three disturbance components, namely, attitude angle disturbance component, swing arm trajectory error component and center of gravity disturbance component; S42, calculating the root mean square value of each disturbance component within the sliding time window to obtain a stability assessment value of each disturbance component; S43. Assign a preset weight factor to each stability evaluation value, and calculate a periodic orbit stability index according to a weighted summation method. The orbit stability index is in the following form: the orbit stability index is equal to the root mean square value of the attitude angle disturbance multiplied by a first weight, the root mean square value of the swing arm trajectory error multiplied by a second weight, and the root mean square value of the center of gravity disturbance multiplied by a third weight. S44, comparing the periodic orbit stability index with the orbit stability deviation threshold set in the current motion mode; S45. When the track stability index is less than or equal to the deviation threshold, the current rhythmic trajectory output is maintained unchanged; when the track stability index is greater than the deviation threshold, it is marked as a track instability state and the trajectory correction parameter generation process is entered.
6. The robot arm swing trajectory optimization method based on dynamic programming according to claim 1, characterized in that: The S5 specifically includes: S51, constructing a historical sequence of attitude disturbance input vectors, extracting attitude disturbance input vectors of the current cycle and the previous multiple consecutive cycles, calculating the first-order derivative and the second-order derivative respectively, and forming a disturbance velocity sequence and a disturbance acceleration sequence; S52, based on the disturbance velocity sequence and the disturbance acceleration sequence, the attitude disturbance input vector of the next T cycles is predicted by the sliding regression model to obtain the disturbance prediction sequence S53, inputting the disturbance prediction sequence into the orbit stability index function for prediction and judgment, and activating the controllable periodic orbit adjustment algorithm when the predicted stability index of any period exceeds the deviation threshold; S54, the controllable periodic orbit adjustment algorithm includes the following processing flow: S541, set the trajectory control parameter boundary and the maximum adjustment increment, including the swing arm frequency parameter, swing arm amplitude parameter and swing arm phase parameter. The current cycle values are f t 、A t 、φ t , the adjustable range is set to [f min ,f max ]、[A min ,A max ]、[φ min ,φ max ], the corresponding maximum change is S542. Based on the three disturbance components in the disturbance prediction sequence, a disturbance response direction is constructed according to a fixed mapping rule, namely, the attitude angle disturbance prediction component is mapped to the swing arm phase parameter adjustment direction, the swing arm trajectory error prediction component is mapped to the swing arm amplitude parameter adjustment direction, and the center of gravity disturbance prediction component is mapped to the swing arm frequency parameter adjustment direction. S543. Construct multiple candidate trajectory parameter combinations within the parameter adjustment constraints, each candidate combination including a frequency correction value, an amplitude correction value, and a phase correction value; S544. Each candidate trajectory parameter combination and the corresponding disturbance prediction vector combination are input into a stability index function for scoring calculation. The stability score calculation method is as follows: the difference between the attitude angle disturbance prediction value and the phase correction value is multiplied by a first weighting coefficient, the difference between the swing arm trajectory error prediction value and the amplitude correction value is multiplied by a second weighting coefficient, and the difference between the center of gravity disturbance prediction value and the frequency correction value is multiplied by a third weighting coefficient. The three results are summed to obtain the stability score of the parameter combination. S545 , screening out several trajectory parameter combinations with the best scores and satisfying parameter boundary restrictions to form a trajectory adjustment candidate parameter set, wherein the trajectory adjustment candidate parameter set includes multiple frequency correction value, amplitude correction value and phase correction value triplets for optimizing the input.
7. The robot arm swing trajectory optimization method based on dynamic programming according to claim 1, characterized in that: The S6 specifically includes: S61, inputting the current periodic attitude disturbance input vector into the dynamic programming model to generate an initial state node; S62, inputting the trajectory adjustment candidate parameter set into the dynamic programming model to generate an optional strategy node; S63, performing forward recursion on each policy node in turn according to the state transition function and cost function defined in the dynamic programming model to form a state transition path that is expanded cycle by cycle; S64. Record the correspondence between the state nodes and the policy nodes in each state transition path, and construct a state policy path tree; S65. Based on the state strategy path tree, traverse backward from the end cycle state node to the current cycle initial state node to determine the strategy path with the lowest global cost; S66 , extracting the trajectory adjustment candidate parameter triples corresponding to the current cycle in the global optimal path, and generating a trajectory correction parameter set.
8. The robot arm swing trajectory optimization method based on dynamic programming according to claim 1, characterized in that: The dynamic programming model is characterized by comprising the following steps: Receive the attitude disturbance input vector, which contains three components: attitude angle disturbance component, swing arm trajectory error component and center of gravity disturbance component, as the current state variable S in the state space t ; Read the candidate parameter set for trajectory adjustment. Each parameter set in the candidate parameter set for trajectory adjustment consists of frequency correction value, amplitude correction value and phase correction value, which constitute the action variable u in the strategy space. t ; Construct a state transfer function, input the state variables and action variables into the disturbance mapping matrix, and predict the state variables of the next cycle. The state transfer function is defined as: S t+1 =S t +M·u t ; in, represents the disturbance response matrix, u t is the current strategy vector, S t is the current state vector; The cost function is constructed based on the state variables. The cost function consists of the state deviation cost term and the policy disturbance penalty term, which is defined as follows: Among them, S ref is the target reference state, λ is the regularization coefficient; For each feasible strategy combination u t , perform state transfer and cost accumulation within the set time range T to form a strategy sequence and state path mapping; Traverse all policy paths, select the policy sequence with the minimum total cost function value, and extract the optimal policy vector for the current cycle from the sequence The output is a set of trajectory correction parameters.
9. The robot arm trajectory optimization method based on dynamic programming according to claim 1, characterized in that: The S7 specifically includes: inputting the trajectory correction parameter set into the central pattern generator model to output a corrected rhythmic arm swing trajectory, including: weighted superposition of the frequency correction value, amplitude correction value and phase correction value in the trajectory correction parameter set with the current frequency parameter, amplitude parameter and phase parameter of the oscillator corresponding to the output layer of the central pattern generator model, and the updated parameters are propagated through the internal coupling of the central pattern generator model to generate rhythmic arm swing control signals for the left arm and the right arm.