A method and system for controlling a two-wheeled robot to walk in a straight line
The method employs a pulse neural network with Liapunov exponent detection and meta-cognitive mechanisms to improve robot control precision and adaptability in dynamic environments by addressing non-linear disturbances and chaotic features.
Patent Information
- Application Number
- CN202510614189.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-05-13
AI Technical Summary
When facing nonlinear perturbation and chaotic characteristics, existing robot dynamic control systems have insufficient control accuracy and are unable to effectively deal with complex dynamic environments.
The pulsed neural network (SNN) is used to combine the Lyapnov index for dynamic state estimation and control. By initializing the pre-trained weight matrix of the pulsed neural network, calibrating the sensor, generating the status code and initial pulse sequence, dynamic state estimation is performed, and the Lyapnov index detects chaotic characteristics, generating nonlinear perturbation compensation, fusing linear velocity and nonlinear perturbation compensation, generating mixed control instructions and performing PWM signal output, building a performance evaluation function, and triggering the metacognitive arbitration mechanism to optimize the pulsed neural network weight.
It realizes precise control of the robot in complex dynamic environments, improves adaptability and control accuracy, can respond to nonlinear perturbations in real time, optimizes control strategies, and enhances the autonomy and flexibility of the robot.
Smart Images

Figure CN120134326B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent control technology, and particularly to a method and system for controlling a two-wheeled robot to walk straight. Background Art
[0002] With the continuous development of robot technology, robot autonomous control systems have been widely applied in multiple fields. Especially in dynamic environments, robots require precise attitude and motion control to ensure efficient operation. In recent years, spiking neural networks (SNNs), as a new computational model that mimics the biological nervous system, have gradually been applied to robot control, especially showing strong advantages when dealing with nonlinear complex systems. However, the current robot dynamic control systems still face some challenges. Although existing spiking neural networks have been applied to a certain extent in state estimation and control instruction generation, for the nonlinear perturbation compensation of chaotic systems, existing technologies still lack effective real-time response capabilities and cannot fully adapt to the changing complex situations in dynamic environments.
[0003] Currently, most technologies in practical applications fail to effectively combine chaotic feature detection with dynamic weight adjustment, which limits the adaptability and control accuracy of robots in complex environments. Traditional linear control methods often ignore the nonlinear perturbations of the system and show large control errors when facing various external environmental factors. To address such problems, a dynamic state estimation and control method based on spiking neural networks is proposed. By introducing Lyapunov exponents for chaotic feature detection, the control accuracy and stability of robots in nonlinear environments are significantly improved. Summary of the Invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides a method for controlling a two-wheeled robot to walk straight to solve the problem of insufficient control accuracy in effectively dealing with nonlinear perturbations and chaotic features in existing robot dynamic control systems.
[0006] To solve the above technical problems, the present invention provides the following technical solutions:
[0007] In a first aspect, the present invention provides a method for controlling a two-wheeled robot to walk straight, which includes initializing the pre-trained weight matrix of a spiking neural network, calibrating sensors, generating a status code and an initial pulse sequence through hardware self-check, triggering synchronous acquisition by sensors, and outputting a preprocessed data packet;
[0008] Inputting into the spiking neural network for dynamic state estimation, and outputting the robot's attitude angle, linear velocity, and attitude angle change rate;
[0009] Calculate the Lyapunov exponent based on the change rate of the attitude angle, generate a non-linear disturbance compensation amount when chaotic characteristics are detected, and fuse the linear velocity and the non-linear disturbance compensation amount;
[0010] Generate a hybrid control command through dynamic weight allocation, convert it into a PWM signal, and output the actual torque feedback;
[0011] Construct a performance evaluation function based on the actual torque feedback and the linear velocity, trigger the metacognitive arbitration mechanism, and optimize the weights of the pulsed neural network;
[0012] Incrementally learn and update the initial weights by storing the optimized weights of the pulsed neural network.
[0013] As a preferred scheme of the method for controlling the linear walking of a two-wheeled robot according to the present invention, wherein: initialize the pre-trained weight matrix of the pulsed neural network, calibrate the sensors, generate a status code and an initial pulse sequence through hardware self-check, trigger the synchronous acquisition of sensors, and output a preprocessed data packet. The specific steps are as follows.
[0014] Set the input sensor data and the output robot state data, and pre-train the weight matrix of the pulsed neural network using the STDP rule;
[0015] Use static and dynamic calibration methods to adjust the bias and scale factor of the sensors, and correct the accelerometer and gyroscope data;
[0016] Perform hardware self-check, detect the working status of sensors, communication and computing units, and generate a status code;
[0017] Generate a corresponding initial pulse sequence according to the status code. If the sensor is normal, output a high-frequency pulse. If the sensor is abnormal, the pulse frequency is reduced;
[0018] Use a hardware timer to trigger the synchronous acquisition of sensors, perform denoising and filtering processing on the sensor data, and output a preprocessed data packet.
[0019] As a preferred scheme of the method for controlling the linear walking of a two-wheeled robot according to the present invention, wherein: input the pulsed neural network for dynamic state estimation, and output the robot attitude angle, linear velocity and attitude angle change rate. The specific steps are as follows.
[0020] Convert the sensor data into a pulse sequence through time encoding and input it into the pulsed neural network;
[0021] Update the neuron state using the membrane potential equation to estimate the attitude angle of the robot;
[0022] By fusing the accelerometer and gyroscope data, combining with the estimated attitude angle, use the integration method and the numerical differentiation method to calculate the robot linear velocity and attitude angle change rate;
[0023] Output the attitude angle, linear velocity, and attitude angle change rate of the robot.
[0024] As a preferred solution of the method for controlling the straight-line walking of a two-wheeled robot according to the present invention, wherein: calculating the Lyapunov exponent according to the attitude angle change rate, generating a non-linear perturbation compensation amount when detecting chaotic characteristics, and fusing the linear velocity with the non-linear perturbation compensation amount. The specific steps are as follows.
[0025] Calculate the maximum Lyapunov exponent based on the attitude angle change rate, set a Lyapunov exponent threshold, and detect whether chaotic characteristics appear.
[0026] When it is detected that the Lyapunov exponent is less than the threshold, no chaotic characteristics appear, and there is no need to generate a non-linear perturbation compensation amount. The robot moves at the current linear velocity.
[0027] When it is detected that the Lyapunov exponent is greater than the threshold, chaotic characteristics appear. Generate a non-linear perturbation compensation amount by combining the Lyapunov exponent and the sigmoid function to suppress the chaotic characteristics and obtain a corrected non-linear perturbation compensation amount.
[0028] Fuse the corrected non-linear perturbation compensation amount with the current linear velocity, calculate the corrected linear velocity, and perform speed dynamic adjustment.
[0029] As a preferred solution of the method for controlling the straight-line walking of a two-wheeled robot according to the present invention, wherein: converting the generation of a hybrid control command through dynamic weight allocation into a PWM signal and outputting an actual torque feedback. The specific steps are as follows.
[0030] Calculate the dynamic weight coefficient according to the real-time attitude angle, linear velocity, and non-linear perturbation compensation amount of the robot, and determine the priority of the parameters in the control command.
[0031] Use the dynamic weight coefficient to perform weighted fusion on the attitude angle, linear velocity, and non-linear perturbation compensation amount to form a hybrid control command.
[0032] Convert the hybrid control command into a PWM signal to drive the motor.
[0033] Collect the actual torque output by the motor through a sensor, compare it with the preset torque, calculate the error, and feedback to adjust the control command.
[0034] As a preferred solution of the method for controlling the straight-line walking of a two-wheeled robot according to the present invention, wherein: constructing a performance evaluation function according to the actual torque feedback and the linear velocity, triggering a metacognitive arbitration mechanism, and optimizing the weights of the pulsed neural network. The specific steps are as follows.
[0035] Based on the actual torque feedback and the linear velocity, construct a performance evaluation function, preset a performance evaluation function threshold, and quantify the operation efficiency.
[0036] Adjust the behavior of the spiking neural network through the metacognitive arbitration mechanism according to the evaluation value of the performance evaluation function;
[0037] When the evaluation value exceeds the preset threshold of the performance evaluation function, trigger the metacognitive arbitration mechanism to update the weights of the spiking neural network;
[0038] Adjust the weights of the updated spiking neural network using the gradient descent method to optimize the weights of the spiking neural network.
[0039] As a preferred solution of the method for controlling the straight-line walking of a two-wheeled robot according to the present invention, wherein: storing the optimized weights of the spiking neural network and incrementally learning to update the initial weights, the specific steps are as follows.
[0040] Store the optimized weights of the spiking neural network in an external storage unit;
[0041] Extract the optimized initial weights of the spiking neural network from the external storage unit as the starting point for incremental learning;
[0042] Based on the current input sensor data and robot state data, calculate the error, and update the weights of the spiking neural network through the backpropagation algorithm to incrementally learn and update the initial weights.
[0043] In a second aspect, the present invention provides a system for controlling the straight-line walking of a two-wheeled robot, including an initialization module, a state estimation module, a disturbance compensation module, a control instruction module, a performance evaluation module, and a weight optimization module;
[0044] The initialization module is used to initialize the pre-trained weight matrix of the spiking neural network, calibrate the sensors, generate a status code and an initial pulse sequence through hardware self-check, trigger synchronous acquisition of the sensors, and output a preprocessed data packet;
[0045] The state estimation module is used to input the spiking neural network for dynamic state estimation and output the robot's attitude angle, linear velocity, and attitude angle change rate;
[0046] The disturbance compensation module is used to calculate the Lyapunov exponent according to the attitude angle change rate, generate a non-linear disturbance compensation amount when detecting chaotic characteristics, and fuse the linear velocity with the non-linear disturbance compensation amount;
[0047] The control instruction module is used to generate a hybrid control instruction through dynamic weight allocation, convert it into a PWM signal, and output an actual torque feedback;
[0048] The performance evaluation module is used to construct a performance evaluation function based on the actual torque feedback and the linear velocity, trigger the metacognitive arbitration mechanism, and optimize the weights of the spiking neural network;
[0049] The weight optimization module is used to update the initial weight by incremental learning through storing the optimized weights of the spiking neural network.
[0050] In a third aspect, the present invention provides a computer device, including a memory and a processor, where the memory stores a computer program, and: when the computer program is executed by the processor, any step of the method for controlling a two-wheeled robot to walk straight as described in the first aspect of the present invention is implemented.
[0051] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and: when the computer program is executed by the processor, any step of the method for controlling a two-wheeled robot to walk straight as described in the first aspect of the present invention is implemented.
[0052] The beneficial effects of the present invention are as follows: By combining the spiking neural network (SNN) with the chaotic feature detection of Lyapunov exponents, the present invention realizes the precise control of the robot in a complex dynamic environment. First, through real-time state estimation by the spiking neural network, the robot can accurately perceive its own attitude angle, linear velocity, and attitude angle change rate, providing reliable data support for the generation of subsequent control instructions. Second, the Lyapunov exponent is introduced as a means of chaos detection. When a non-linear disturbance is detected, a compensation amount can be quickly generated, effectively overcoming the problem of non-linear disturbances that cannot be handled by traditional control methods. Most importantly, by optimizing the weights of the spiking neural network through the metacognitive mechanism, the system has the ability of incremental learning, can continuously optimize the control strategy according to the changes of the environment in practical applications, and improves the adaptability and control accuracy of the robot in a dynamic and complex environment. Description of the Drawings
[0053] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0054] Figure 1 It is a flowchart of the method for controlling a two-wheeled robot to walk straight in Embodiment 1.
[0055] Figure 2 It is a module diagram of the system for controlling a two-wheeled robot to walk straight in Embodiment 1. Detailed Embodiments
[0056] To make the above objects, features, and advantages of the present invention more obvious and understandable, the detailed embodiments of the present invention will be described in detail below with reference to the drawings in the specification.
[0057] In the following description, numerous specific details are set forth to provide a thorough understanding of the present invention. However, the present invention may be practiced in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the spirit of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0058] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation manner of the present invention. The appearances of "in one embodiment" in different places in this specification do not all refer to the same embodiment, nor are they separate or alternative embodiments that are mutually exclusive with other embodiments.
[0059] Example 1, referring to Figure 1 and Figure 2 , is the first embodiment of the present invention. This embodiment provides a method for controlling a two-wheeled robot to walk in a straight line, including the following steps:
[0060] S1. Initialize the pre-trained weight matrix of the spiking neural network, calibrate the sensors, generate a status code and an initial pulse sequence through hardware self-check, trigger synchronous acquisition of the sensors, and output a preprocessed data packet.
[0061] Furthermore, set the input sensor data and the output robot state data, and pre-train the weight matrix of the spiking neural network using the STDP rule;
[0062] Specifically, utilize the spike timing correlation of STDP to extract spatio-temporal features (such as the phase difference between the accelerometer and gyroscope signals) from the sensor data through unsupervised learning, enabling the network to have the basic ability of motion pattern recognition before deployment. In contrast, traditional robot state recognition networks use supervised learning for weight initialization, requiring a large amount of labeled data and being difficult to adapt to dynamic environments. Through comparison, it is shown that the network pre-trained by STDP has a 60% improvement in convergence speed (compared with supervised learning) in sudden motion recognition tasks, and the false alarm rate decreases by 22%. This feature is particularly suitable for robot sudden posture adjustment scenarios (such as fall detection) because it relies on the spike firing timing rather than the amplitude information, and the ability to resist signal amplitude perturbation is significantly enhanced;
[0063] It should be noted that the STDP rule, as a unique unsupervised learning mechanism of the spiking neural network, its weight adjustment follows the biological plasticity principle of "enhancing the connection when the pre-pulse triggers the post-pulse, and weakening it otherwise" (in line with the Hebbian theory). Compared with the traditional backpropagation algorithm, STDP only needs to record the spike timing difference between adjacent neurons in hardware implementation, reducing the memory occupancy of the computing unit (reducing the storage overhead by about 40%). Selecting the STDP rule for pre-training has the advantage of hardware adaptability.
[0064] Adjust the bias and scale factor of the sensor using static and dynamic calibration methods to correct the accelerometer and gyroscope data;
[0065] Specifically, in the initialization stage, static compensation (eliminating the initial bias) and dynamic compensation (fitting the scale factor curve through a preset standard motion trajectory of the robot) are synchronously executed. The two-parameter error calibration method is used. Most existing calibration schemes adopt offline calibration (such as static compensation before leaving the factory), which cannot correct the parameter drift of the sensor during long-term operation. The measured data shows that the composite calibration reduces the gyroscope angular velocity integration error to 0.05° / min (the traditional single calibration is 0.15° / min). During 10 minutes of continuous operation, the deviation of the robot's heading angle estimation is reduced from 3.2° to 0.8°. It is directly related to the accuracy of subsequent data packets, avoiding the amplification of errors in the filtering link;
[0066] It should be noted that static calibration refers to the bias compensation of the sensor in the zero-input state (such as eliminating the zero drift when the gyroscope is stationary), and dynamic calibration calculates the scale factor nonlinear error through a preset motion trajectory (such as a three-axis rotation sequence). The combination of static calibration and dynamic calibration can cover the full working condition range of the sensor. Traditional methods only adopt a single calibration mode and cannot cope with the parameter time-varying problems caused by temperature drift (such as the temperature drift of MEMS accelerometers can reach 0.2mg / ℃) and mechanical stress;
[0067] Perform a hardware self-check to detect the working status of the sensor, communication, and computing units and generate a status code;
[0068] Specifically, convert the status code into a pulse frequency signal (high frequency 20 - 50kHz indicates normal, low frequency 1 - 5kHz indicates abnormal), so that the status information can be directly embedded in the sensor data stream (through pulse interval coding) to achieve in-band monitoring. Existing hardware self-checks report abnormalities through interrupt signals, resulting in the coupling of the control flow and the data flow. Under the condition of limited communication bandwidth (such as CAN bus), it saves the resource occupancy of the independent status channel, and the pulse frequency and the sensor data are sampled by the same hardware timer to ensure the strict synchronization of the status and the data (jitter less than 0.1μs). After testing, when the sensor has a sudden failure, the delay from abnormal recognition to the execution of the security protocol is reduced from 15ms to 8ms;
[0069] It should be noted that the status code adopts 8-bit binary coding (such as 00000001 indicates accelerometer abnormality, 00000010 indicates gyroscope abnormality), and the working status of the sensor, communication, and computing units is mapped in real time through hardware registers. Compared with the traditional status word, its coding density is increased by 3 times, and it can be directly converted into a pulse frequency modulation signal (such as high-frequency pulses corresponding to the normal status code 0xFF), avoiding the delay introduced by additional encoding and decoding circuits.
[0070] Generate the corresponding initial pulse sequence according to the status code. If the sensor is normal, output high-frequency pulses; if the sensor is abnormal, the pulse frequency decreases.
[0071] It should be noted that the status code is encoded into a pulse frequency signal, and the system health is directly characterized by the time-domain characteristics (frequency, duty cycle) of the pulse sequence. The status information is embedded in the sensor data pulse stream, saving communication resources. The pulse frequency signal and the sensor data share the same hardware timer clock source (such as the APB2 bus clock), eliminating the clock offset in traditional multi-channel transmission.
[0072] High-frequency pulses refer to square wave signals of 20 - 50 kHz (such as 30 kHz pulses are output when the robot is running normally), which match the input bandwidth of mainstream spiking neural network (SNN) hardware (such as the pulse event queue of the Intel Loihi chip supports up to 50 kHz at most). The high-frequency characteristics can improve the integration speed of the membrane potential of SNN neurons and accelerate network convergence.
[0073] Low-frequency pulses refer to square wave signals of 1 - 5 kHz (such as 2 kHz pulses are output when the sensor is abnormal). The duty cycle of the low-frequency pulses decreases (such as the duty cycle of 2 kHz pulses is 10%) to reduce the power consumption in abnormal states. By the significant frequency difference (high-frequency / low-frequency ratio ≥ 10:1), it is ensured that the downstream SNN layer can quickly identify abnormal events (the response time of the SNN synaptic filter to events with a frequency ratio > 10:1 is < 0.5 ms).
[0074] Use the hardware timer to trigger the synchronous acquisition of sensors, and perform denoising and filtering processing on the sensor data, and output the preprocessed data packet.
[0075] It should be noted that the hardware comparison and matching signal of the timer (such as TIM_TRGO of STM32) is used to trigger the start of conversion of all sensor ADCs simultaneously to achieve physical layer synchronization. Asynchronous sampling of multiple sensors will cause data phase deviation (such as the sampling moments of the accelerometer and the gyroscope are misaligned). Software timestamp compensation is adopted, but the hardware interrupt delay cannot be eliminated.
[0076] Sensor data (such as accelerometer and gyroscope data) introduces various noise sources during the acquisition process, including high-frequency noise: generated by the coupling of circuit thermal noise (Johnson-Nyquist noise) and mechanical vibration, with a frequency band usually above 1 kHz (for example, the typical noise density of a MEMS gyroscope is 0.015° / s / √Hz), low-frequency drift: caused by sensor bias temperature drift (such as the zero-bias temperature drift of an accelerometer can reach ±1 mg / ℃) and integration cumulative error (the angular velocity integration error of a gyroscope grows linearly with time), and impulse interference: caused by electromagnetic compatibility (EMC) problems or mechanical shocks, manifested as instantaneous amplitude mutations (such as the output peak of an accelerometer during a collision can reach ±20 g);
[0077] Specifically, the time alignment error of multi-sensor data is reduced from ±50 μs of software synchronization to within ±5 μs, improving the attitude fusion accuracy of the subsequent Kalman filter by 18%. When the robot rotates at a high speed (>300° / s), it can avoid the accumulation of angular velocity integration errors caused by data asynchronization;
[0078] High-frequency noise is preferentially eliminated through wavelet denoising to avoid false pulse triggering. The discrete wavelet transform (DWT) decomposes the signal into multi-scale coefficients, processes the high-frequency detail coefficients, compensates for low-frequency drift through Kalman filtering to ensure the long-term stability of state estimation. The Kalman filter performs recursive estimation through sensor data, removes the noise of the sensor, and smooths the prediction of the operating state. The moving average filter suppresses impulse interference and prevents abnormal update of the weight matrix. The moving average filter refers to a simple and effective signal processing method that removes noise in the signal, smooths the data sequence, and suppresses high-frequency noise in the data by averaging the current data point and its neighboring data points, while retaining the main trend of the signal.
[0079] S2. Input into the pulsed neural network for dynamic state estimation, and output the robot's attitude angle, linear velocity, and attitude angle change rate.
[0080] Furthermore, the sensor data is converted into a pulse sequence through time encoding and input into the pulsed neural network;
[0081] It should be noted that sensor data (such as accelerometer and gyroscope data) needs to be converted into a pulse sequence and input into the pulsed neural network (SNN). Time encoding, as an effective neural network input method, converts continuous sensor data into pulsed time-series information. Through the time encoding method, the changes of sensor data in time can be accurately simulated, improving the processing ability of the pulsed neural network for time-series data. Time encoding itself can maintain the timeliness and dynamic characteristics of sensor data, while the pulsed neural network can efficiently process this time-series information for more accurate dynamic state estimation;
[0082] Specifically, the continuous sensor signal is converted into a sparse pulse stream through time encoding while preserving the signal dynamic characteristics (such as the dense area of pulse timing corresponding to the acceleration mutation). When the robot makes a rapid turn (angular velocity > 180° / s), the pulse interval generated by the encoder is compressed from 20 ms to 2 ms, reducing the inter-layer propagation delay of the SNN from 15 ms to 3 ms, meeting the real-time control requirements. The number of pulses output by the encoder is reduced to 1 / 3 of the traditional scheme (128 pulses / second vs. 384 pulses / second), and the energy consumption of SNN synaptic operations is reduced by 65% (from 9.6 nJ / OP to 3.3 nJ / OP).
[0083] Preferably, traditional rate encoding has insufficient pulse density in the low dynamic range (such as no pulses when the accelerometer is stationary), while phase delay encoding maps the full-range data through the time dimension (even when a(t)=0, reference pulses are still generated). The time-encoded pulses can be directly input into the SNN hardware (such as Intel Loihi) through GPIO pins without an additional ADC module.
[0084] Update the neuron state using the membrane potential equation to estimate the pose angle of the robot.
[0085] It should be noted that the spiking neural network (SNN) updates its state based on the membrane potential equation of neurons. Through the membrane potential equation, the state change of neurons is combined with external input signals (i.e., sensor data) to dynamically calculate the activation state of neurons, reflecting the changes of the robot in real time and accurately estimating the pose angle of the robot. The membrane potential equation can simulate the activities of neurons in the neural network, enabling the network to gradually deduce the motion state of the robot in time. The membrane potential equation can dynamically respond to input signals (such as sensor data) and calculate the pose angle in real time.
[0086] Specifically, using the pulse timing nonlinear dynamics characteristics of the SNN, directly calculate the pose angle through membrane potential integration. When the robot makes a large-angle movement (pitch angle > 45°), the traditional EKF has an error of up to 8.2° due to ignoring high-order terms, while the SNN adaptively adjusts the integration gain through the pulse firing frequency, and the error is controlled within 2.3°. EKF needs to calculate the Jacobian matrix in real time (inverting a 6×6 matrix), while LIF only requires the iteration of a first-order differential equation (the computational complexity is reduced by 89%).
[0087] By fusing the data of the accelerometer and gyroscope, combining the estimated pose angle, use the integration method and numerical differentiation method to calculate the linear velocity and the change rate of the pose angle of the robot.
[0088] It should be noted that by fusing the accelerometer and gyroscope data to estimate the attitude angle, the integration method and numerical differentiation method are used to calculate the change rates of the linear velocity and the attitude angle. The accelerometer provides the measurement of the robot's acceleration, while the gyroscope measures the angular velocity of the robot. By combining the attitude angle, linear velocity, and change rate data of the attitude angle, the motion state of the robot can be accurately described;
[0089] Specifically, taking the attitude angle estimated by the SNN as the benchmark, the integration process is corrected backward. The pure gyroscope integration will increase the error. The integration initial conditions are reset every 1 ms through the angle output by the SNN, making the cumulative error approach the SNN estimation error. The direct differentiation of the accelerometer will amplify the high-frequency noise. The differential result is constrained by the first-order smoothness of the angle estimated by the SNN. The maximum error of the linear velocity estimation in the ramp road test is <0.05 m / s. The trapezoidal integration coefficient is used to compensate the influence of the attitude inclination angle in real time. The frequency response bandwidth of the attitude angle change rate is extended to 50 Hz, and attitude mutation events at the millisecond level can be captured.
[0090] Output the attitude angle, linear velocity, and attitude angle change rate of the robot.
[0091] It should be noted that the pulse neural network outputs the attitude angle, linear velocity, and attitude angle change rate of the robot, which are the core parameters of robot control and are used for further dynamic adjustment and path planning. The attitude angle reflects the current orientation of the robot, while the linear velocity and attitude angle change rate are the speed and turning rate of the robot's motion. The output values provide real-time and accurate data support for subsequent control decisions.
[0092] S3. Calculate the Lyapunov exponent according to the attitude angle change rate, generate a non-linear perturbation compensation amount when chaotic characteristics are detected, and fuse the linear velocity and the non-linear perturbation compensation amount.
[0093] Furthermore, calculate the maximum Lyapunov exponent based on the attitude angle change rate, set the Lyapunov exponent threshold, and detect whether chaotic characteristics appear;
[0094] It should be noted that the Lyapunov exponent (LE) is a key indicator for measuring dynamic stability, and its positive and negative values directly reflect whether chaotic characteristics are shown. According to the non-linear dynamics theory, when the maximum Lyapunov exponent is greater than the Lyapunov exponent threshold, the system presents chaotic characteristics. The threshold is set as a positive value close to zero, which can not only avoid misjudgment caused by noise interference (such as sensor noise may cause a small positive exponent), but also capture the chaotic trend in time. The sigmoid function has continuous differentiability and non-linear saturation characteristics, which can limit the mapping relationship between the Lyapunov exponent and the compensation amount within a limited range, avoiding oscillation of control commands caused by sudden changes in the compensation amount.
[0095] When the Lyapunov exponent is detected to be less than the threshold, no chaotic characteristics appear, and there is no need to generate a non-linear perturbation compensation amount. The robot moves at the current linear velocity.
[0096] It should be noted that in a stable situation, the robot does not need to perform complex non-linear perturbation compensation. When the Lyapunov exponent is less than the Lyapunov exponent threshold, it remains in a stable state. The robot can continue to walk straight at the current linear velocity without additional adjustment, avoiding unnecessary computational overhead and optimizing the control efficiency.
[0097] There is no need to perform complex compensation calculations, which can save computing resources and improve the processing speed. Especially in a real-time control environment, it helps to reduce the system burden and improve the response ability. In the absence of chaotic characteristics, maintaining the existing control strategy can ensure the stable operation of the system, while avoiding over-adjustment and ensuring the simplicity and stability of the control.
[0098] When the Lyapunov exponent is detected to be greater than the threshold, chaotic characteristics appear. By combining the Lyapunov exponent and the sigmoid function, a non-linear perturbation compensation amount is generated to suppress the chaotic characteristics and obtain the corrected non-linear perturbation compensation amount.
[0099] It should be noted that by combining the Lyapunov exponent and the sigmoid function to generate a non-linear perturbation compensation amount, the Lyapunov exponent reflects the intensity of the chaotic characteristics, while the sigmoid function has the characteristics of smoothness and gradual change, and can adjust the perturbation compensation amount according to the degree of chaos. When chaotic behavior is detected, precise non-linear perturbation compensation is performed to avoid overreaction or compensation. Combining the Lyapunov exponent and the sigmoid function for non-linear perturbation compensation makes the adjustment of the compensation amount smoother and more flexible, and can be adaptively adjusted according to the intensity of the chaotic characteristics, thereby effectively suppressing chaotic behavior.
[0100] Specifically, combining the Lyapunov exponent and the sigmoid function to generate a non-linear perturbation compensation amount, the formula is:
[0101] ;
[0102] Where represents the non-linear perturbation compensation amount, represents the maximum Lyapunov exponent calculated in real time, represents the Lyapunov exponent threshold, = , represents the amount exceeding the threshold of the chaos intensity, which directly determines the basic amplitude of the compensation amount, represents the adaptive adjustment factor, represents the error function, Denoted as the slope factor, controlling the steepness of the function transition region, Denoted as the dynamically adjusted weight of the differential term, Denoted as the hyperbolic tangent function, Denoted as the change rate of the Lyapunov exponent, Denoted as a tiny change amount, Denoted as a tiny change in time, Denoted as the differential time constant, Denoted as the base of the natural logarithm, Denoted as the attitude-related decay coefficient, Denoted as within the time window for the integral of, Denoted as the time window length, Denoted as the current moment, Denoted as the current moment minus the time window length , Denoted as the historical moment of the chaotic over-threshold quantity, Denoted as for each moment within the traversed time window , , Denoted as the infinitesimal increment with respect to time .
[0103] Fuse the corrected non-linear perturbation compensation quantity with the current linear velocity, calculate the corrected linear velocity, and perform dynamic velocity adjustment.
[0104] It should be noted that by fusing the corrected non-linear perturbation compensation quantity with the current linear velocity, a corrected linear velocity can be obtained. Through the dynamic adjustment of the velocity, it is adjusted according to the change of the chaotic characteristics, so as to better control the motion state of the robot;
[0105] Preferably, by fusing the non-linear perturbation compensation quantity and the linear velocity, the motion velocity of the robot can be dynamically adjusted, enabling the robot to adapt to the changing environment and dynamic behaviors. Compared with the traditional static control method, this dynamic adjustment improves the adaptability of the robot to complex environmental changes. Through real-time velocity adjustment, it can quickly respond when chaotic characteristics appear, correct the motion state of the robot, and thus effectively suppress unstable factors and maintain the running stability of the robot.
[0106] S4. Generate a hybrid control command through dynamic weight allocation, convert it into a PWM signal, and output the actual torque feedback.
[0107] Furthermore, based on the real-time attitude angle, linear velocity, and non-linear disturbance compensation amount of the robot, calculate the dynamic weight coefficient to determine the priority of the parameters in the control instruction;
[0108] It should be noted that the robot controls straight walking and calculates the dynamic weight coefficient based on the real-time feedback attitude angle, linear velocity, and non-linear disturbance compensation amount. The dynamic weight coefficient reflects the relative importance of each control parameter in the current environment or task execution state. For example, when the robot performs a quick turn, the attitude angle may require a higher weight, while in the case of a large change in linear velocity, the linear velocity may have a higher priority. By calculating the weight coefficient in real time, the priority of the parameters in the control instruction can be dynamically adjusted, so as to more accurately respond to the changes in the current environment;
[0109] Specifically, according to the real-time state of the robot (attitude angle error , linear velocity deviation , non-linear disturbance compensation amount ), the priority parameter dynamically adjusted, the expression for calculating the dynamic weight coefficient is:
[0110] ;
[0111] Among them, represents the dynamic weight coefficient, with a range of [0, 1], represents the high membership function, represents the real-time attitude angle error, represents the medium membership function, represents the linear velocity deviation, represents the low membership function, represents the non-linear disturbance compensation amount, represents the quantified severity of the attitude angle error , represents that when the linear velocity deviation is close to the nominal value, the weight of speed control is moderately increased, represents that when the disturbance compensation amount is small, the weight of the disturbance term is reduced.
[0112] Use the dynamic weight coefficient to perform weighted fusion on the attitude angle, linear velocity, and non-linear disturbance compensation amount to form a hybrid control instruction;
[0113] It should be noted that the hybrid control instruction is formed by integrating the attitude angle, linear velocity, and non-linear disturbance compensation input signal according to the actual needs of the robot, and forms the final control instruction to control the straight-line walking behavior of the robot. The dynamic weight coefficient ensures that the contribution of each parameter can be adjusted according to the real-time state of the robot at any time, thus optimizing the output of the control instruction, making the control instruction more comprehensively reflect the comprehensive state of the robot, avoiding the limitations of relying only on a single parameter, and significantly improving the adaptability of the system to complex dynamic environments.
[0114] Convert the hybrid control instruction into a PWM signal to drive the motor;
[0115] It should be noted that the hybrid control instruction is converted into a pulse width modulation (PWM) signal, which then drives the motor. The PWM signal refers to a common motor control method that precisely controls the speed and torque of the motor by adjusting the duty cycle of the signal. By converting the control instruction into a PWM signal, precise control of the motor drive can be achieved, thereby affecting the movement of the robot.
[0116] Collect the actual torque output by the motor through the sensor, compare it with the preset torque, calculate the error, and feedback to adjust the control instruction.
[0117] It should be noted that the sensor collects the output torque of the motor in real time, compares it with the preset torque, and by calculating the error, the control instruction can be dynamically adjusted to optimize the actual output of the motor, ensuring that the robot can make necessary adjustments according to the actual execution results to correct the deviation. Through real-time error calculation and feedback adjustment, it can ensure that the motor output is always consistent with the expected target, improving the control accuracy and robustness, and endowing the robot with higher adaptability.
[0118] S5. Construct a performance evaluation function based on the actual torque feedback and linear velocity, trigger the metacognitive arbitration mechanism, and optimize the weights of the pulse neural network.
[0119] Furthermore, based on the actual torque feedback and linear velocity, construct a performance evaluation function, preset the threshold of the performance evaluation function, and quantify the operating efficiency;
[0120] It should be noted that the performance evaluation function quantifies the operating efficiency of the robot system based on the actual torque feedback and linear velocity. This function combines two key indicators: the actual output torque and the desired movement speed (linear velocity). The torque feedback can reflect the mechanical performance of the robot when performing tasks, while the linear velocity reflects the accuracy and stability of the robot's movement. By presetting the threshold of the performance evaluation function, the operating efficiency of the robot can be monitored and quantified in real time, providing data support for subsequent optimization and adjustment;
[0121] Specifically, the performance evaluation function is used to quantify the operating efficiency of the robot control system and evaluate the system state based on torque feedback and linear velocity. The expression is:
[0122]
[0123] Wherein, represents the operating efficiency and performance of the current control system of the robot, represents the actual torque feedback of the robot motor, represents the current linear velocity of the robot. q represents the weight coefficient of torque feedback, represents the weight coefficient of linear velocity, represents the decay factor. p represents the adjustment factor, represents the disturbance decay factor, represents the current moment, represents the past moment within the time interval.
[0124] According to the evaluation value of the performance evaluation function, the behavior of the spiking neural network is adjusted through the metacognitive arbitration mechanism;
[0125] It should be noted that the metacognitive arbitration mechanism refers to using the feedback value of the evaluation function to adjust the behavior of the spiking neural network (SNN) to achieve better control effects. The metacognitive arbitration mechanism uses the output of the evaluation function to dynamically adjust the learning and decision-making processes of the neural network, enabling the robot to optimize its behavior according to the requirements of its current task. The metacognitive arbitration mechanism can automatically learn and adapt to task requirements, improving the flexibility and autonomy of the robot;
[0126] Preferably, combining the metacognitive arbitration mechanism with the performance evaluation function can adaptively adjust the decision-making process of the neural network according to the real-time performance of the robot. The metacognitive arbitration mechanism can dynamically adjust the strategy according to changes during task execution, ensuring a more efficient control and decision-making process, especially in complex environments or tasks. The robot no longer solely relies on preset control strategies but can continuously optimize its behavior based on real-time feedback, enhancing the robot's autonomous learning ability and enabling it to flexibly respond in different scenarios.
[0127] When the evaluation value exceeds the preset performance evaluation function threshold, the metacognitive arbitration mechanism is triggered to update the weights of the spiking neural network;
[0128] It should be noted that when the performance evaluation value exceeds the preset performance evaluation function threshold, the system triggers the metacognitive arbitration mechanism to update the weights. Here, "weight update" refers to optimizing the performance of the neural network by adjusting the connection strength in the spiking neural network, realizing self-regulation and optimization based on real-time performance feedback, and improving the control accuracy and task execution efficiency of the robot.
[0129] The gradient descent method is used to adjust the weights of the updated spiking neural network to optimize the weights of the spiking neural network.
[0130] It should be noted that the gradient descent method is an optimization algorithm widely used to adjust the weights of neural networks. By calculating the gradient of the loss function and adjusting the weights of the neural network along the gradient direction, the gradient descent method can effectively reduce errors and optimize the performance of the network. Using the gradient descent method to adjust the weights of the spiking neural network can continuously optimize the control strategy according to the task requirements after each update of the neural network;
[0131] The gradient descent method is an optimization algorithm widely used to adjust the weights of neural networks. By calculating the gradient of the loss function and adjusting the weights of the neural network along the gradient direction, the gradient descent method can effectively reduce errors and optimize the performance of the network. Using the gradient descent method to adjust the weights of the spiking neural network can continuously optimize the control strategy according to the task requirements after each update of the neural network.
[0132] S6. Incremental learning updates the initial weights by storing the optimized weights of the spiking neural network.
[0133] Furthermore, the optimized weights of the spiking neural network are stored in an external storage unit;
[0134] It should be noted that the training and optimization of the spiking neural network is a long process involving a large amount of computation and data processing. The optimized weights after one training represent the optimal state of the robot's straight-line walking control at a certain stage. If these optimized weights are not saved, the system will not be able to utilize the previous optimization results after restarting and must start training from scratch, wasting a large amount of computational resources and time. Storing the optimized weights for a long time ensures that the system can continue to execute the existing optimization strategies at different time nodes without losing past knowledge;
[0135] The external storage unit refers to a stable and large-capacity storage medium used to save the training results of the spiking neural network to ensure that the optimization information during the training process is not lost. Common storage units include hard disks, solid-state drives, or cloud storage. Since the training process of the spiking neural network is computationally intensive and time-consuming, saving the weights is crucial for subsequent use.
[0136] Extract the optimized initial weights of the spiking neural network from the external storage unit as the starting point for incremental learning;
[0137] It should be noted that incremental learning refers to a method of gradually learning by adding new sensor data, rather than training the weights of the spiking neural network from scratch. Incremental learning can effectively handle the situation where the training data is constantly changing and can learn new knowledge without losing the knowledge learned previously;
[0138] The key role of incremental learning is that during continuous operation, according to new sensor data and the robot's state, it can gradually adjust the weights of the spiking neural network, make full use of the existing network weights, and the robot can continuously self-optimize and adapt to new environmental conditions through incremental learning.
[0139] Based on the current input sensor data and the robot's state data, calculate the error, and update the weights of the spiking neural network through the backpropagation algorithm. Incremental learning updates the initial weights.
[0140] It should be noted that the backpropagation algorithm plays a crucial role in the training of neural networks. Combined with error calculation, the backpropagation algorithm accurately updates the weights of the spiking neural network. Each time an error is calculated based on the data collected by the sensor, the robot can adjust the network weights in real time according to the error, so as to perform control and prediction more accurately.
[0141] This embodiment also provides a system for controlling a two-wheeled robot to walk in a straight line, including: an initialization module, a state estimation module, a disturbance compensation module, a control instruction module, a performance evaluation module, and a weight optimization module;
[0142] The initialization module is used to initialize the pre-trained weight matrix of the spiking neural network, calibrate the sensor, generate a status code and an initial pulse sequence through hardware self-check, trigger synchronous acquisition of the sensor, and output a preprocessed data packet; the state estimation module is used to input the spiking neural network for dynamic state estimation and output the robot's attitude angle, linear velocity, and attitude angle change rate; the disturbance compensation module is used to calculate the Lyapunov exponent according to the attitude angle change rate, generate a non-linear disturbance compensation amount when detecting chaotic characteristics, and fuse the linear velocity with the non-linear disturbance compensation amount; the control instruction module is used to generate a hybrid control instruction through dynamic weight allocation, convert it into a PWM signal, and output actual torque feedback; the performance evaluation module is used to construct a performance evaluation function based on the actual torque feedback and the linear velocity, trigger the metacognitive arbitration mechanism, and optimize the weights of the spiking neural network; the weight optimization module is used to store the optimized weights of the spiking neural network and update the initial weights through incremental learning.
[0143] This embodiment also provides a computer device applicable to the situation of the method for controlling a two-wheeled robot to walk in a straight line, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the method for controlling a two-wheeled robot to walk in a straight line as proposed in the above embodiment.
[0144] The computer device may be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covering the display screen, or may be a button, a trackball, or a touchpad provided on the housing of the computer device, or may also be an external keyboard, touchpad, or mouse, etc.
[0145] This embodiment also provides a storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the method for controlling the linear walking of a two-wheeled robot as proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (abbreviated as SRAM), electrically erasable programmable read-only memory (abbreviated as EEPROM), erasable programmable read-only memory (abbreviated as EPROM), programmable read-only memory (abbreviated as PROM), read-only memory (abbreviated as ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc.
[0146] In summary, the present invention realizes precise control of a robot in a complex dynamic environment by combining the chaotic feature detection of a spiking neural network (SNN) and Lyapunov exponents. First, through real-time state estimation by the spiking neural network, the robot can accurately sense its own attitude angle, linear velocity, and the change rate of the attitude angle, providing reliable data support for the generation of subsequent control instructions. Second, the Lyapunov exponent is introduced as a means of chaos detection. When a nonlinear perturbation is detected, a compensation amount can be quickly generated, effectively overcoming the problem of nonlinear perturbations that cannot be addressed by traditional control methods. Most importantly, by optimizing the weights of the spiking neural network through a metacognitive mechanism, the system has the ability of incremental learning, can continuously optimize the control strategy according to environmental changes in practical applications, and improves the adaptability and control accuracy of the robot in a dynamic and complex environment.
[0147] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A method for controlling a two-wheeled robot to walk straight, characterized in that: Include, Initialize the pre-trained weight matrix of the spiking neural network, calibrate the sensors, generate a status code and an initial pulse sequence through hardware self-check, trigger synchronous acquisition of the sensors, and output a preprocessed data packet; Input into the spiking neural network for dynamic state estimation, and output the robot's attitude angle, linear velocity, and attitude angle change rate; Calculate the Lyapunov exponent based on the attitude angle change rate, generate a non-linear perturbation compensation amount when chaotic characteristics are detected, and fuse the linear velocity and the non-linear perturbation compensation amount; Generate a hybrid control instruction through dynamic weight allocation, convert it into a PWM signal, and output the actual torque feedback; Construct a performance evaluation function based on the actual torque feedback and the linear velocity, trigger the metacognitive arbitration mechanism, and optimize the weights of the spiking neural network; Incrementally learn and update the initial weights by storing the optimized weights of the spiking neural network.
2. The method for controlling a two-wheeled robot to walk in a straight line according to claim 1, wherein: The steps of initializing the pre-trained weight matrix of the spiking neural network, calibrating the sensors, generating a status code and an initial pulse sequence through hardware self-check, triggering synchronous acquisition of the sensors, and outputting a preprocessed data packet are as follows. Set the input sensor data and the output robot state data, and pre-train the weight matrix of the spiking neural network using the STDP rule; Adjust the bias and scale factor of the sensors using static and dynamic calibration methods, and correct the accelerometer and gyroscope data; Perform hardware self-check, detect the working status of the sensors, communication, and computing units, and generate a status code; Generate a corresponding initial pulse sequence according to the status code. If the sensor is normal, output high-frequency pulses. If the sensor is abnormal, the pulse frequency decreases; Use the hardware timer to trigger synchronous acquisition of the sensors, denoise and filter the sensor data, and output a preprocessed data packet.
3. The method for controlling a two-wheeled robot to walk in a straight line according to claim 2, characterized in that: The steps of inputting into the spiking neural network for dynamic state estimation and outputting the robot's attitude angle, linear velocity, and attitude angle change rate are as follows. Convert the sensor data into a pulse sequence through time encoding and input it into the spiking neural network; Update the neuron state using the membrane potential equation to estimate the robot's attitude angle; Computer the robot's linear velocity and attitude angle change rate by fusing the accelerometer and gyroscope data, combining with the estimated attitude angle, and using the integration method and numerical differentiation method; Output the robot's attitude angle, linear velocity, and attitude angle change rate.
4. The method for controlling a two-wheeled robot to walk in a straight line according to claim 3, wherein: The steps of calculating the Lyapunov exponent based on the attitude angle change rate, generating a non-linear perturbation compensation amount when chaotic characteristics are detected, and fusing the linear velocity and the non-linear perturbation compensation amount are as follows. Calculate the maximum Lyapunov exponent based on the attitude angle change rate, set the Lyapunov exponent threshold, and detect whether chaotic characteristics appear; When it is detected that the Lyapunov exponent is less than the threshold, no chaotic characteristics appear, and there is no need to generate a non-linear perturbation compensation amount. The robot moves at the current linear velocity; When it is detected that the Lyapunov exponent is greater than the threshold, chaotic characteristics appear. Generate a non-linear perturbation compensation amount by combining the Lyapunov exponent and the sigmoid function to suppress the chaotic characteristics and obtain the corrected non-linear perturbation compensation amount; Fuse the corrected non-linear perturbation compensation amount with the current linear velocity, calculate the corrected linear velocity, and perform speed dynamic adjustment.
5. The method for controlling a two-wheeled robot to walk straight according to claim 4, characterized in that: The conversion of the hybrid control instruction generated by dynamic weight allocation into a PWM signal outputs the actual torque feedback, and the specific steps are as follows: Calculate the dynamic weight coefficient according to the real-time attitude angle, linear velocity and nonlinear disturbance compensation amount of the robot, and determine the priority of the parameters in the control instruction; Use the dynamic weight coefficient to perform weighted fusion on the attitude angle, linear velocity and nonlinear disturbance compensation amount to form a hybrid control instruction; Convert the hybrid control instruction into a PWM signal to drive the motor; Collect the actual torque output by the motor through the sensor, compare it with the preset torque, calculate the error, and feedback to adjust the control instruction.
6. The method for controlling a two-wheeled robot to walk straight according to claim 5, wherein: The construction of the performance evaluation function according to the actual torque feedback and the linear velocity triggers the metacognitive arbitration mechanism to optimize the weights of the pulse neural network, and the specific steps are as follows: Based on the actual torque feedback and the linear velocity, construct a performance evaluation function, preset the threshold of the performance evaluation function, and quantify the operation efficiency; According to the evaluation value of the performance evaluation function, adjust the behavior of the pulse neural network through the metacognitive arbitration mechanism; When the evaluation value exceeds the preset threshold of the performance evaluation function, trigger the metacognitive arbitration mechanism to update the weights of the pulse neural network; Use the gradient descent method to adjust the weights of the updated pulse neural network to optimize the weights of the pulse neural network.
7. The method for controlling a two-wheeled robot to walk in a straight line according to claim 6, characterized in that: The incremental learning updates the initial weights by storing the optimized weights of the pulse neural network, and the specific steps are as follows: Store the optimized weights of the pulse neural network in the external storage unit; Extract the optimized initial weights of the pulse neural network from the external storage unit as the starting point of incremental learning; Based on the current input sensor data and the robot state data, calculate the error, and update the weights of the pulse neural network through the backpropagation algorithm to incrementally learn and update the initial weights.
8. A system for controlling a two-wheeled robot to walk in a straight line, based on the method for controlling a two-wheeled robot to walk in a straight line according to any one of claims 1 to 7, characterized in that: It includes an initialization module, a state estimation module, a disturbance compensation module, a control instruction module, a performance evaluation module and a weight optimization module; The initialization module is used to initialize the pre-trained weight matrix of the pulse neural network, calibrate the sensor, generate a status code and an initial pulse sequence through hardware self-check, trigger the synchronous acquisition of the sensor, and output a preprocessed data packet; The state estimation module is used to perform dynamic state estimation by inputting the pulse neural network, and output the robot attitude angle, linear velocity and attitude angle change rate; The disturbance compensation module is used to calculate the Lyapunov exponent according to the attitude angle change rate, generate a nonlinear disturbance compensation amount when detecting chaotic characteristics, and fuse the linear velocity and the nonlinear disturbance compensation amount; The control instruction module is used to convert the hybrid control instruction generated by dynamic weight allocation into a PWM signal and output the actual torque feedback; The performance evaluation module is used to construct a performance evaluation function according to the actual torque feedback and the linear velocity, trigger the metacognitive arbitration mechanism, and optimize the weights of the pulse neural network; The weight optimization module is used to incrementally learn and update the initial weights by storing the optimized weights of the pulse neural network.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the method for controlling the two-wheeled robot to walk straight as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the steps of the method for controlling a two-wheeled robot to walk straight according to any one of claims 1 to 7.
Citation Information
Patent Citations
Online stably controlled humanoid robot based on bionic reinforcement learning type cerebellum model
CN112060082A
Model predictive control techniques for autonomous systems
CN114746872A