A Three-Degree-of-Freedom Pectoral Fin Cooperative Motion Control and Its Gait Optimization Method
By establishing a three-degree of freedom pectoral fin kinematic model and multi-layer order neural network optimization control signal, combined with deep reinforcement learning and sliding mode control, the problems of three-degree of freedom pectoral fin collaborative motion control and gait optimization are solved, and efficient propulsion and posture stability of robot fish in complex waters are achieved.
Patent Information
- Application Number
- CN202510618662.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-05-14
AI Technical Summary
The prior art is difficult to achieve coordinated motion control and gait optimization of three-degree-of-freedom rigid pectoral fins, resulting in limited propulsion efficiency and posture stability of robotic fish in complex waters, and it is difficult to meet the needs of high degree-of-freedom attitude adjustment and multi-task control.
A three-degree of freedom pectoral fin kinematic model is established, combined with a multimodal perception system, a deep reinforcement learning algorithm and a sliding mode control method, and through a multi-layer order neural network and a central mode generator network, the control signal is optimized to achieve coordinated movement and gait optimization of pectoral fins, and an extended state observer is used for perturbation compensation to ensure that the pectoral fins maintain efficient propulsion and posture stability in complex environments.
It significantly improves the propulsion efficiency and attitude control capabilities of the robot fish, realizes accurate adjustment and stable control of pectoral fin movement, adapts to complex underwater environments, and meets the needs of high-performance bionic propulsion.
Smart Images

Figure CN120122464B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of underwater robots, and relates to a three-degree-of-freedom pectoral fin coordinated motion control and its gait optimization method, in particular to the multi-degree-of-freedom motion trajectory control and gait optimization of rigid pectoral fins. Background Art
[0002] Bionic robotic fish are a new type of underwater robot with advantages such as high maneuverability, high concealment, and low noise, and are widely used in fields such as underwater resource exploration, underwater equipment inspection, underwater biological observation, and underwater archaeology. Compared with traditional propulsion methods, bionic fish can achieve fine control and efficient propulsion in complex waters through flexible fin movements, and the pectoral fins play an important role in attitude adjustment, steering control, and stability maintenance. The pectoral fins are an important part of the central fin propulsion system, especially showing highly flexible control capabilities in fish such as manta rays and kitefish; research shows that the flapping motion of the pectoral fins can not only provide partial propulsion force, but also assist the fish body to complete attitude control such as pitching, rolling, and yawing; although some research has made certain progress in the control of rigid pectoral fins with single-degree-of-freedom or two-dimensional plane motion, their degrees of freedom of motion are limited and it is difficult to meet the requirements of high-degree-of-freedom attitude adjustment and multi-task control.
[0003] Currently, most pectoral fin control methods mainly focus on single-axis symmetric flapping or two-dimensional trajectory generation, and a complete three-degree-of-freedom rigid pectoral fin flapping model and control method have not been established; especially for the control method of three-degree-of-freedom coordinated flapping of pectoral fins such as forward and backward flapping, up and down flapping, and wing rocking through three servo drive units in a bionic structure, relevant research and literature are rarely publicly available; in existing patented technologies, there is also no systematic solution for the coordinated control mechanism and gait optimization strategy for three-degree-of-freedom rigid pectoral fins; in addition, the coordinated control of multi-degree-of-freedom pectoral fin motion not only faces the problem of kinematic coupling, but also involves complex rhythm generation, parameter synchronization, and motion efficiency optimization; existing methods are difficult to achieve rhythm matching, phase coordination, and gait optimization between multi-degree-of-freedom flapping, thus limiting their application potential in high-performance bionic propulsion.
[0004] In view of this, there is an urgent need for a coordinated motion control and its gait optimization method applicable to a three-degree-of-freedom flapping system of rigid pectoral fins. Through systematic modeling, rhythm control, and optimization solution of the three-degree-of-freedom drive model, the unity of pectoral fin propulsion efficiency and attitude stability is achieved, filling the gap in the existing research on three-degree-of-freedom pectoral fin flapping control and rhythm optimization. Summary of the Invention
[0005] In view of the problems in the technical background, the present invention provides a three - degree - of - freedom pectoral fin coordinated motion control and its gait optimization method, which can achieve multi - degree - of - freedom motion trajectory control and gait optimization of rigid pectoral fins, and is mainly used to improve the propulsion efficiency, attitude stability and disturbance robustness of robotic fish in complex water environments.
[0006] The present invention provides a three - degree - of - freedom pectoral fin coordinated motion control and its gait optimization method, and the method includes the following steps: (1) Referring to the typical motion behaviors of real fish pectoral fins during underwater swimming, establish a basic motion gait model of three - degree - of - freedom rigid pectoral fins, including forward - backward flapping motion, up - down flapping motion and wing - rocking motion, and use periodic functions to model the kinematic model equation of three - degree - of - freedom pectoral fins. With the generalized coordinate vector q(t)=[θ ro (t), θ fl (t), θ s (t)] T describing the pectoral fin attitude, θ ro (t) is the angular displacement of forward - backward flapping motion, θ fl (t) is the angular displacement of up - down flapping motion, θ s (t) is the angular displacement of wing - rocking motion, t is time, and define the control signal vector c(t) as the complete control parameter set of the pectoral fin; (2) Use a multi - modal perception system to collect the hydrodynamic feedback and attitude information of the pectoral fin under different combinations of control input parameters in real time. The perception system includes a distributed pressure sensor and an inertial measurement unit, which collect the thrust F T (t), lift F L (t), lateral force F A (t) and attitude angle θ imu (t); (3) Based on the real - time perception feedback data, construct a pectoral fin - environment interaction model, and use a deep reinforcement learning algorithm to train a high - level policy network to optimize the basic control signal c base (t), with propulsion efficiency and attitude stability as multi - objective reward functions, and the attitude stability is determined by lateral force, lift and attitude angle; (4) Construct a multi - layer hierarchical neural control network to decode the basic control signal c base (t) and output an incremental control signal Δc(t). The first layer is used to extract environmental dynamic features and policy patterns, and the second layer outputs the amplitude adjustment amount, frequency adjustment amount, phase - difference adjustment amount and bias adjustment amount of each degree of freedom; (5) Introduce a Central Pattern Generator (CPG) network, adopt a coupled non - linear oscillator structure, generate the rhythmic angular displacement reference trajectories of each degree of freedom according to the control signal c(t)=c base (t)+Δc(t), and construct a coordinated rhythmic relationship between degrees of freedom through the phase - difference , is the current oscillator phase of the central pattern generator network, is the selected reference oscillator phase of the central pattern generator network, and t is time; (6) The sliding mode control method based on the Extended State Observer (ESO) is adopted, and the sliding mode controller is used to construct the disturbance compensation control law u i (t), and the dynamic compensation of external disturbances and modeling errors is realized by estimating the disturbance term z i3 (t), is the difference between the desired trajectory and the actual trajectory of the pectoral fin, is the derivative of the difference between the desired trajectory and the actual trajectory of the pectoral fin, is a positive real adjustment parameter, and t is time; (7) During the continuous operation process, the current gait output effect is dynamically evaluated according to the performance evaluation index function J(t). If it deviates from the threshold ε, the incremental control signal Δc(t) is adjusted online based on the deviation direction to correct the control parameters, so as to keep the pectoral fin movement always in the optimal propulsion and stable state to achieve closed-loop optimization, and improve the propulsion efficiency and attitude control ability of the robotic fish.
[0007] As a further technical solution, taking the center of the three-degree-of-freedom movement of the pectoral fin as the origin, a local coordinate system {P} of the pectoral fin movement is defined: the x-axis points to the forward direction of the robotic fish, the y-axis is the left-right direction of the robotic fish, and the z-axis is perpendicular to the ventral-dorsal direction of the fish body to form a right-handed coordinate system, and a generalized coordinate system of the pectoral fin pose is established; the coordination of the three-degree-of-freedom movement determines the overall movement effect of the pectoral fin, and can realize the basic movement gaits of the pectoral fin, including figure-eight and elliptical curves; the kinematic model equation of the pectoral fin is:
[0008] ;
[0009] where θ ro (t) is the angular displacement of the forward and backward flapping motion, θ fl (t) is the angular displacement of the up and down flapping motion, θ s (t) is the angular displacement of the rocking motion, A ro is the amplitude of the forward and backward flapping motion, A fl is the amplitude of the up and down flapping motion, A s is the amplitude of the rocking motion, b ro is the offset of the forward and backward flapping motion, b fl is the offset of the up and down flapping motion, b s is the offset of the rocking motion, ω ro is the frequency of the forward and backward flapping motion, ω fl is the frequency of the up and down flapping motion, ω s is the frequency of the rocking motion, δ ro is the phase difference of the forward and backward flapping motion, δ flis the phase difference of the up-and-down flapping motion, δ s is the phase difference of the wing-rocking motion, t is time; when 2ω ro = ω fl , δ ro = π / 2, δ fl = 0, the basic motion gait of the pectoral fin is an "8"-shaped trajectory. When ω ro = ω fl , δ ro = π / 2, δ fl = 0, the basic motion gait of the pectoral fin is a "0"-shaped trajectory;
[0010] The control signal vector c(t) of the three-degree-of-freedom model of the pectoral fin contains the following parameters:
[0011] c(t) = [A ro (t), A fl (t), A s (t), b ro (t), b fl (t), b s (t), ω ro (t), ω fl (t), ω s (t), δ ro (t), δ fl (t), δ s (t)] T , where A ro (t) is the amplitude of the forward and backward flapping motion of the pectoral fin at time t, A fl (t) is the amplitude of the up-and-down flapping motion of the pectoral fin at time t, A s (t) is the amplitude of the wing-rocking motion of the pectoral fin at time t, b ro (t) is the bias of the forward and backward flapping motion of the pectoral fin at time t, b fl (t) is the bias of the up-and-down flapping motion of the pectoral fin at time t, b s (t) is the bias of the wing-rocking motion of the pectoral fin at time t, ω ro (t) is the frequency of the forward and backward flapping motion of the pectoral fin at time t, ω fl (t) is the frequency of the up-and-down flapping motion of the pectoral fin at time t, ω s (t) is the frequency of the wing-rocking motion of the pectoral fin at time t, δ ro (t) is the phase difference of the forward and backward flapping motion of the pectoral fin at time t, δ fl (t) is the phase difference of the up-and-down flapping motion of the pectoral fin at time t, δ s (t) is the phase difference of the wing-rocking motion of the pectoral fin at time t.
[0012] As a further technical solution, the data sampling frequency of the multi-modal perception system is as follows: the sampling frequency of the pressure sensor ≥ 100 Hz, the sampling frequency of the inertial measurement unit ≥ 200 Hz, and the attitude angle θ imu (t), thrust F T (t), lift F L (t), lateral force F A (t) are recorded synchronously to construct a high-precision feedback closed-loop and achieve a rapid response to complex dynamic environments.
[0013] As a further technical solution, during the training process of the high-level policy network, a sliding window strategy is adopted to evaluate the convergence and stability, and the policy network parameters are updated in combination with historical feedback; the form of the reward function R(t) is:
[0014] ;
[0015] In the formula, F T (t) is the thrust, P(t) is the power consumption, F L (t) is the lift, F A (t) is the lateral force, Var(F L (t)) and Var(F A (t)) respectively represent the fluctuation ranges of the lift and the lateral force, θ imu (t) is the real-time attitude angle, θ ref (t) is the target attitude angle, and w1, w2, w3, and w4 are the weighting coefficients of the reward function.
[0016] As a further technical solution, the first layer of the multi-layer hierarchical neural network is a structure based on Graph Neural Networks (GNN) or Long Short-Term Memory (LSTM), which is used to identify the trend of environmental state transition; the second layer is a feedforward neural network or a convolutional neural network, which is used to output the incremental control signal Δc(t) for adjusting the three-degree-of-freedom motion.
[0017] As a further technical solution, the neural oscillator of the central pattern generator network adopts an amplitude-phase coupling model, and its expression is:
[0018] ;
[0019] In the formula, the superscript "·" represents the first derivative of the corresponding variable, is the phase of the i-th oscillator, is the phase of the j-th oscillator, ω i is the local frequency, κ ij is the coupling strength, is the oscillator phase target value, A i is the current amplitude of the oscillator, A target is the target value amplitude of the oscillator, α i is the amplitude adjustment rate, b i is the bias output, b target is the bias target value, β i is the bias adjustment rate, ω i is the frequency output, ω target is the frequency target value, γ i is the bias adjustment rate, δ i is the phase difference, δ target is the phase difference target value, is the phase difference adjustment rate; dynamically adjust the oscillator amplitude A through the state equation i , frequency ω i , bias b i and the oscillator phase , ensuring the coupling fluctuation and rhythm continuity between degrees of freedom.
[0020] As a further technical solution, the disturbance compensation control law u i (t) is converted into the angular acceleration α i through the equivalent moment of inertia J i (t), and then integrated to obtain the target angular velocity ω ref,i (t), which is used for servo drive control:
[0021] ;
[0022] where ω0 is the initial angular velocity (which can be set as the velocity at the end of the previous cycle), and t is the time.
[0023] As a further technical solution, the performance evaluation index function J(t) is defined as:
[0024] ;
[0025] where F T (t) is the thrust, is the target thrust, F L (t) is the lift force, F A (t) is the lateral force, Var(F L (t)) and Var(F A (t)) respectively represent the fluctuation ranges of the lift force and the lateral force, θ imu (t) is the real-time attitude angle, θ ref(t) is the target attitude angle, w1, w2, w3, and w4 are the weighting coefficients of the reward function, and t is the time; when the performance evaluation index function J(t) > ε, it indicates that the current gait performance deviates greatly, and the feedback adjustment of each control component in the incremental control signal Δc(t) needs to be triggered, including:
[0026] ;
[0027] where i ∈ {ro, fl, s} respectively represent the three degrees of freedom of the pectoral fin's forward and backward flapping, up and down flapping, and rocking, t is the time, and η A , η ω , η b , η δ are the step adjustment factors of each parameter. The adjusted incremental control signal Δc(t) will be used to update the control signal c(t) = c base (t) + Δc(t), driving the central pattern generator network to generate a new displacement reference trajectory .
[0028] As a further technical solution, in the "8"-shaped trajectory movement, the parameter ranges of the control signal include: the amplitude A ro ∈ [-90°, 90°] of the forward and backward flapping movement, the frequency ω ro ∈ [1 Hz, 4 Hz] of the forward and backward flapping movement, the bias b ro ∈ [-45°, 45°] of the forward and backward flapping movement, the phase difference δ ro ∈ [-180°, 180°] of the forward and backward flapping movement, the amplitude A fl ∈ [-65°, 65°] of the up and down flapping movement, the frequency ω fl ∈ [1 Hz, 8 Hz] of the up and down flapping movement, the bias b fl ∈ [-30°, 30°] of the up and down flapping movement, the phase difference δ fl ∈ [-180°, 180°] of the up and down flapping movement; the amplitude A s ∈ [-180°, 180°] of the rocking movement, the frequency ω s ∈ [1 Hz, 4 Hz] of the rocking movement, the bias b s ∈ [-180°, 180°] of the rocking movement, the phase difference δ s ∈ [-180°, 180°] of the rocking movement; in the "0"-shaped trajectory movement, the amplitude A ro ∈ [-90°, 90°] of the forward and backward flapping movement, the frequency ω ro ∈ [1 Hz, 4 Hz] of the forward and backward flapping movement, the bias bro ∈ [-45°, 45°], the phase difference δ of the front and rear flapping motion ro ∈ [-180°, 180°], the amplitude A of the up and down flapping motion fl ∈ [-65°, 65°], the frequency ω of the up and down flapping motion fl ∈ [1 Hz, 4 Hz], the bias b of the up and down flapping motion fl ∈ [-30°, 30°], the phase difference δ of the up and down flapping motion fl ∈ [-180°, 180°]; the amplitude A of the wing-rocking motion s ∈ [-180°, 180°], the frequency ω of the wing-rocking motion s ∈ [1 Hz, 4 Hz], the bias b of the wing-rocking motion s ∈ [-180°, 180°], the phase difference δ of the wing-rocking motion s ∈ [-180°, 180°]; when the pectoral fin moves in two trajectories of "8" shape and "0" shape, it needs to be adaptively optimized according to the task requirements and hydrodynamic feedback to achieve the joint optimization of the propulsion efficiency and attitude control ability of the robotic fish.
[0029] As a further technical solution, the weights w1, w2, w3, and w4 of each performance index in the reward function and the performance evaluation index function can be dynamically adjusted according to the environmental complexity and task objectives to adapt to the propulsion efficiency priority or attitude stability priority strategy mode.
[0030] The additional aspects and advantages of the present invention will be partially given in the following description, partially will become obvious from the following description, or be understood through the practice of the present invention.
[0031] The beneficial effects of the present invention are as follows:
[0032] (1) The optimized three-degree-of-freedom pectoral fin mechanism has biomimicry in three dimensions of morphology, structure, and function.
[0033] (2) The adopted multi-modal perception system can collect the hydrodynamic feedback and attitude information of the pectoral fin under different combinations of control input parameters in real time.
[0034] (3) Based on the real-time perception feedback data, a pectoral fin-environment interaction model is constructed, and a deep reinforcement learning algorithm is used to train the high-level policy network to optimize the control signal with the propulsion efficiency and attitude stability as the multi-objective reward function; at the same time, a multi-layer hierarchical neural control network is constructed to decode the control signal and output the incremental control signal Δc(t).
[0035] (4) The present invention adopts a sliding mode control method based on an Extended State Observer (ESO), and dynamically evaluates the current gait output effect according to the performance evaluation index function J(t) during continuous operation. If it deviates from the threshold ε, the incremental control signal Δc(t) is adjusted online based on the deviation direction to correct the control parameters, realizing precise adjustment and stable control of the pectoral fin motion parameters, and significantly improving the actual application performance and environmental adaptability of the bionic pectoral fin robot.
[0036] (5) According to the method of the present invention, the pectoral fin motion is always in the optimal propulsion and stable state to achieve closed-loop optimization, improving the propulsion efficiency and attitude control ability of the robotic fish. Description of the Drawings
[0037] Figure 1 It is a flowchart of the three-degree-of-freedom pectoral fin collaborative motion control and its gait optimization method according to the embodiment of the present invention.
[0038] Figure 2 It is a motion displacement diagram of the "8"-shaped basic motion gait of the three-degree-of-freedom pectoral fin in the example of the present invention.
[0039] Figure 3 It is a motion displacement diagram of the "0"-shaped basic motion gait of the three-degree-of-freedom pectoral fin in the example of the present invention.
[0040] Figure 4 It is a schematic diagram of the description of the generalized coordinate system of the three-degree-of-freedom pectoral fin in the example of the present invention.
[0041] Figure 5 It is a schematic diagram of the actual motion of the "8"-shaped and "0"-shaped trajectories of the three-degree-of-freedom pectoral fin in the example of the present invention.
[0042] Figure 6 It is a schematic diagram of the layout and acquisition of the sensing system of the three-degree-of-freedom pectoral fin in the example of the present invention.
[0043] Figure 7 It is a flowchart of the optimization of the deep reinforcement learning algorithm during the motion of the three-degree-of-freedom pectoral fin in the example of the present invention.
[0044] Figure 8 It is a structure diagram of the multi-layer hierarchical neural control network during the motion of the three-degree-of-freedom pectoral fin in the example of the present invention.
[0045] Figure 9 It is a flowchart of the CPG network generating motion control signals for the three-degree-of-freedom pectoral fin in the example of the present invention.
[0046] Figure 10 It is a schematic diagram of the smooth switching of the "8"-shaped and "0"-shaped trajectories of the three-degree-of-freedom pectoral fin in the example of the present invention.
[0047] Figure 11 It is the block diagram of sliding mode control based on an extended state observer in the three - degree - of - freedom pectoral fin coordinated motion in the embodiments of the present invention.
[0048] Figure 12 It is the overall control block diagram of the three - degree - of - freedom pectoral fin coordinated motion control and its gait optimization method in the embodiments of the present invention. Specific embodiments
[0049] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.
[0050] Please refer to the attached Figures 1 to 12 , the embodiments of the present invention provide a three - degree - of - freedom pectoral fin coordinated motion control and its gait optimization method, and the method includes the following steps: (1) Referring to the typical motion behaviors of real fish pectoral fins during underwater swimming, establish a basic motion gait model of three - degree - of - freedom rigid pectoral fins, including forward - backward flapping motion, up - down flapping motion and yawing motion, and use periodic functions to model the kinematic model equation of the three - degree - of - freedom pectoral fins. With the generalized coordinate vector q(t)=[θ ro (t), θ fl (t), θ s (t)] T to describe the pectoral fin attitude, where θ ro (t) is the angular displacement of the forward - backward flapping motion, θ fl (t) is the angular displacement of the up - down flapping motion, θ s (t) is the angular displacement of the yawing motion, t is time, and define the control signal vector c(t) as the complete control parameter set of the pectoral fin; (2) Use a multi - modal perception system to collect the hydrodynamic feedback and attitude information of the pectoral fin under different combinations of control input parameters in real time. The perception system includes a distributed pressure sensor and an inertial measurement unit, and collects the thrust F T (t), lift F L (t), lateral force F A (t) and attitude angle θ imu (t); (3) Based on the real - time perception feedback data, construct a pectoral fin - environment interaction model, and use a deep reinforcement learning algorithm to train a high - level policy network to optimize the basic control signal c base (t), with the propulsion efficiency and attitude stability as multi - objective reward functions, and the attitude stability is determined by the lateral force, lift and attitude angle; (4) Construct a multi - layer hierarchical neural control network for the basic control signal c base(t) is decoded to output an incremental control signal Δc(t). The first layer is used to extract environmental dynamic features and policy patterns, and the second layer outputs the amplitude adjustment amount, frequency adjustment amount, phase difference adjustment amount, and bias adjustment amount for each degree of freedom; (5) A Central Pattern Generator (CPG) network is introduced, which adopts a coupled nonlinear oscillator structure. According to the control signal c(t) = c base (t) + Δc(t), the rhythmic angular displacement reference trajectory for each degree of freedom is generated , and the coordinated rhythmic relationship between degrees of freedom is constructed through the phase difference . is the current oscillator phase of the central pattern generator network, is the selected reference oscillator phase of the central pattern generator network, and t is time; (6) An extended state observer (ESO)-based sliding mode control method is adopted. Using the sliding mode controller , a disturbance compensation control law u i (t) is constructed. By estimating the disturbance term z i3 (t), dynamic compensation for external disturbances and modeling errors is achieved, is the difference between the desired trajectory and the actual trajectory of the pectoral fin, is the derivative of the difference between the desired trajectory and the actual trajectory of the pectoral fin, is a positive real adjustment parameter, and t is time; (7) During continuous operation, the current gait output effect is dynamically evaluated according to the performance evaluation index function J(t). If it deviates from the threshold ε, the incremental control signal Δc(t) is adjusted online based on the deviation direction to correct the control parameters, keeping the pectoral fin movement always in the optimal propulsion and stable state to achieve closed-loop optimization, and improving the propulsion efficiency and attitude control ability of the robotic fish.
[0051] The following further describes specific embodiments of the present invention.
[0052] Refer to Figure 1 . This method is based on the motion modeling of the flapping behavior of biological pectoral fins, and combines modules such as a multi-modal perception system, a deep reinforcement learning strategy, a hierarchical neural network controller, a CPG network, an ESO, and sliding mode control to achieve fine, rhythmic, and robust closed-loop control and adaptive optimization of the bionic pectoral fin.
[0053] S1. Referring to the research results on the pectoral fin movement gait of real fish during swimming, the basic movement gaits of the rigid pectoral fin include the "8"-shaped and "0"-shaped trajectories. At the same time, the pectoral fin movement behavior of biological fish is systematically analyzed, and referring to its typical flapping-wing movement mechanism, the three-degree-of-freedom basic movement modes of the rigid pectoral fin are extracted: forward and backward flapping-wing movement, up and down flapping-wing movement, and rocking-wing movement. The movement of each degree of freedom can be modeled by a periodic function and form effective propulsion through specific phase relationships. Based on this, the kinematic model of the three degrees of freedom of the pectoral fin is as follows:
[0054] ;
[0055] In the formula, θ ro (t) is the angular displacement of the forward and backward flapping-wing movement, θ fl (t) is the angular displacement of the up and down flapping-wing movement, θ s (t) is the angular displacement of the rocking-wing movement, A ro is the amplitude of the forward and backward flapping-wing movement, A fl is the amplitude of the up and down flapping-wing movement, A s is the amplitude of the rocking-wing movement, b ro is the offset of the forward and backward flapping-wing movement, b fl is the offset of the up and down flapping-wing movement, b s is the offset of the rocking-wing movement, ω ro is the frequency of the forward and backward flapping-wing movement, ω fl is the frequency of the up and down flapping-wing movement, ω s is the frequency of the rocking-wing movement, δ ro is the phase difference of the forward and backward flapping-wing movement, δ fl is the phase difference of the up and down flapping-wing movement, δ s is the phase difference of the rocking-wing movement, and t is time; referring to Figure 2 , when 2ω ro = ω fl , δ ro = π / 2, δ fl = 0, the basic movement gait formed by the forward and backward flapping-wing and up and down flapping-wing of the pectoral fin is the "8"-shaped trajectory; referring to Figure 3 , when ω ro = ω fl , δ ro = π / 2, δ fl = 0, the basic movement gait formed by the forward and backward flapping-wing and up and down flapping-wing of the said pectoral fin is the "0"-shaped trajectory.
[0056] Furthermore, referring to Figure 4, To uniformly represent the three - degree - of - freedom motion state of the pectoral fin in space, a generalized coordinate system is introduced to describe it. As a rigid planar structure, the actual motion of the pectoral fin occurs within the local fish - body coordinate system. First, taking the center of the three - degree - of - freedom motion of the pectoral fin as the origin, a local coordinate system {P} for the pectoral fin motion is defined: the x - axis points in the forward direction of the robotic fish, the y - axis is in the left - right direction of the robotic fish, and the z - axis is perpendicular to the ventral - dorsal direction of the fish body, forming a right - hand coordinate system. The motion of the pectoral fin is relative to this local coordinate system.
[0057] Specifically, define the generalized coordinate vector q(t) as a set of state variables describing the current posture of the pectoral fin, including the angular displacement θ ro (t) of the forward - backward flapping motion, the angular displacement θ fl (t) of the up - down flapping motion, and the angular displacement θ s (t) of the wing - rocking motion. It is expressed as: q(t)=[θ ro (t), θ fl (t), θ s (t)] T 。
[0058] Specifically, this generalized coordinate vector can serve as the core input of the controller and optimizer, connecting the perception, control, and execution modules. By tracking the changes of q(t) in real - time, the system can achieve dynamic modeling and control planning of the complex motion behavior of the pectoral fin, realizing the efficient propulsion and attitude adjustment of the robotic fish.
[0059] Furthermore, to establish an input - output mapping relationship for subsequent optimization and control, define the control - signal vector c(t) as a set of control parameters for the pectoral - fin motion, used to drive the three - degree - of - freedom motion of the pectoral fin. The specific form of c(t) is: c(t)=[A ro (t), A fl (t), A s (t), b ro (t), b fl (t), b s (t), ω ro (t), ω fl (t), ω s (t), δ ro (t),δ fl (t), δ s (t)] T ,where A ro (t) is the amplitude of the forward - backward flapping motion of the pectoral fin at time t, A fl (t) is the amplitude of the up - down flapping motion of the pectoral fin at time t, A s (t) is the amplitude of the wing - rocking motion of the pectoral fin at time t, b ro (t) is the offset of the forward - backward flapping motion of the pectoral fin at time t, b fl(t) is the offset of the pectoral fin flapping motion up and down at time t, b s (t) is the offset of the pectoral fin rowing motion at time t, ω ro (t) is the frequency of the pectoral fin flapping motion back and forth at time t, ω fl (t) is the frequency of the pectoral fin flapping motion up and down at time t, ω s (t) is the frequency of the pectoral fin rowing motion at time t, δ ro (t) is the phase difference of the pectoral fin flapping motion back and forth at time t, δ fl (t) is the phase difference of the pectoral fin flapping motion up and down at time t, δ s (t) is the phase difference of the pectoral fin rowing motion at time t; the control signal c(t) serves as the interface signal between the neural control network and the output of the reinforcement learning strategy, is used to describe the expected pectoral fin motion pattern, and is further generated by the controller to generate an execution instruction to drive the pectoral fin to achieve the target motion behavior.
[0060] Specifically, the control signal c(t) can be further divided into the combined relationship between the basic control signal c base (t) output by the strategy and the incremental control signal Δc(t) adjusted by the neural network. Its mathematical expression is: c(t) = c base (t) + Δc(t), where c base (t) represents the basic control signal directly generated by the reinforcement learning strategy, and Δc(t) is the dynamic incremental control signal output by the multi-layer neural control network according to the environmental feedback; the two work together to form the final control signal that drives the pectoral fin to execute, realizing the efficient decoupling and coordination between the strategy layer and the execution layer.
[0061] Specifically, refer to Figure 5 , based on the established kinematic model, a simulation tool is used to visually simulate the pectoral fin motion trajectory under different fluctuation parameters. The simulation results show that the model can accurately describe the three-dimensional coordinated motion and gait structure coordination of the fish pectoral fin, and can provide an accurate reference for subsequent rhythm control and optimization learning.
[0062] S2. To implement the closed-loop adjustment mechanism for the three-degree-of-freedom motion modeling and control of the pectoral fin, a multi-modal perception system is introduced into the pectoral fin system of the robotic fish to collect the hydrodynamic feedback data and motion posture information of the pectoral fin under different control input parameters (amplitude A, frequency ω, offset b, phase difference δ, etc.) in real time; refer to Figure 6, the perception system is mainly composed of a pressure sensor and an inertial measurement unit. Adopting a right-angled isosceles triangle layout strategy, three pressure sensors are respectively installed at the near root, the central area of the leading edge, and the end edge of the pectoral fin to record the normal pressure distribution generated during the interaction between the pectoral fin and the fluid, and then calculate hydrodynamic characteristics such as thrust, lateral force, and lift. This layout can cover the main working surface of the pectoral fin to the greatest extent, improve the information content and identification ability of sensor data, and the sensor layout positions on the front and back sides of the pectoral fin are mirror-symmetric about the fin surface; the inertial measurement unit module is installed at the base of the pectoral fin or the connection structure to continuously obtain the three-axis angular velocity, linear acceleration, and spatial attitude angle change data of the pectoral fin.
[0063] Furthermore, during the data acquisition process, by setting gait input parameters (such as amplitude A, frequency ω, bias b, phase difference δ, etc.) to drive the pectoral fin to perform periodic motion, and using the perception system to record the hydrodynamic and attitude data of the pectoral fin under each input condition in real time; for each control cycle T, the sampling frequency is set to f c = 500 Hz to obtain high-resolution data; the above data includes the hydrodynamic forces on the pectoral fin, the attitude angles and angular velocities in the three-axis directions.
[0064] Specifically, the calculation of hydrodynamic forces is based on the pressure integration method. Let the normal pressure distribution function on the surface of the pectoral fin be p n (x, y, t), and the time is t. Then the thrust F T 、lift F L and lateral force F A expressions can be respectively expressed as:
[0065] ;
[0066] In the formula, α(x, y) represents the angle between the normal of the surface element of the pectoral fin and the propulsion direction, β(x, y) represents the angle between the normal and the lateral direction, S is the force-bearing surface area of the pectoral fin, and x and y are the force-bearing directions of the pectoral fin respectively.
[0067] Specifically, the attitude data is obtained by fusing the gyroscope and accelerometer in the inertial measurement unit to obtain the actual attitude angle θ imu (t); all sensor data are timestamped and synchronized with the controller signal, so that the pectoral fin attitude q(t), the control signal c(t), and the thrust F T (t), lift F L (t), lateral force F A (t) in the feedback force value maintain a one-to-one correspondence, providing continuous and high-precision input for subsequent gait optimization and control strategy learning.
[0068] Specifically, the data collected by the perception system can not only be used to evaluate the execution effect of the three-degree-of-freedom motion of the pectoral fin, but also provide a basis for the parameter adjustment of the subsequent control strategy and the construction of the reward function of the optimization algorithm; by comparing the perception data with the generalized coordinate output in real time, the control error can be effectively monitored, supporting the update of the state-action mapping of reinforcement learning, so as to realize the dynamic adaptive control of the pectoral fin motion mode; and it is used to optimize the controller parameters and construct a hydrodynamic feedback mapping model, providing an input basis for the reward and punishment mechanism in reinforcement learning, and used for error correction between the actual motion and the ideal gait in the subsequent closed-loop control link.
[0069] S3. Refer to Figure 7 , on the basis of completing the three-degree-of-freedom motion modeling of the pectoral fin and the acquisition of perception data, a deep reinforcement learning algorithm is further introduced to realize the online optimization training of the basic motion gait of the pectoral fin; the optimization goal not only focuses on the improvement of the propulsion efficiency, but also takes into account the attitude stability and lift control, ensuring that the bionic robot fish has good motion performance in a complex underwater environment.
[0070] Furthermore, the reinforcement learning framework takes the real-time feedback data provided by the perception system as the input of the environmental state, and combines the action output generated by the neural dynamics model to form a complete state-action space; the pectoral fin attitude angle vector q(t) and attitude angle change q̇(t) represented by the generalized coordinates are used as state features, combined with the basic control signal c base (t) and the feedback hydrodynamic values of thrust F T (t), lift F L (t), and lateral force F A (t) to construct an interaction model between the pectoral fin and the environment.
[0071] Specifically, the Deep Deterministic Policy Gradient (DDPG) algorithm is used for policy training. Its policy network outputs high-level control parameters for three degrees of freedom, including continuous variables such as amplitude A, frequency ω, bias b, and phase difference δ, and maps them to actual control commands through the action network to guide the pectoral fin to complete a complete gait cycle.
[0072] Specifically, to achieve a balance among propulsion efficiency, attitude stability, and lift constraint, the following reward function R(t) is constructed:
[0073] ;
[0074] In the formula, F T (t) is the thrust, P(t) is the power consumption, F L (t) is the lift, F A (t) is the lateral force, Var(F L(t)) and Var(F A (t)) represent the fluctuation ranges of the lift and the lateral force respectively, and θ imu (t) is the real-time attitude angle, θ ref (t) is the target attitude angle, w1, w2, w3, and w4 are the weighting coefficients of the reward function, and t is the time; while encouraging the propulsion efficiency, this reward function suppresses the attitude fluctuations and the lift instability.
[0075] In each training iteration, the agent samples trajectories by continuously interacting with the environment, performs policy evaluation and policy update, and gradually forms an efficient and stable motion pattern; the finally output high-level control policy is defined as the basic control signal c base (t). This control policy contains the basic control parameters of the three degrees of freedom of the pectoral fin, and its structural form is: c base (t) = [A rob (t), A flb (t), A sb (t), b rob (t), b flb (t), b sb (t), ω rob (t), ω flb (t), ω sb (t), δ rob (t), δ flb (t), δ sb (t)] T , where A rob (t) is the basic amplitude of the pectoral fin flapping back and forth at time t, A flb (t) is the basic amplitude of the pectoral fin flapping up and down at time t, A sb (t) is the basic amplitude of the pectoral fin rocking at time t, b rob (t) is the basic offset of the pectoral fin flapping back and forth at time t, b flb (t) is the basic offset of the pectoral fin flapping up and down at time t, b sb (t) is the basic offset of the pectoral fin rocking at time t, ω rob (t) is the basic frequency of the pectoral fin flapping back and forth at time t, ω flb (t) is the basic frequency of the pectoral fin flapping up and down at time t, ω sb (t) is the basic frequency of the pectoral fin rocking at time t, δ rob (t) is the basic phase difference of the pectoral fin flapping back and forth at time t, δ flb (t) is the basic phase difference of the pectoral fin flapping up and down at time t, δ sb (t) is the basic phase difference of the pectoral fin rocking at time t; the above parameters serve as the input benchmarks for the decoding and adjustment of the incremental control signal Δc(t) in the subsequent neural network module.
[0076] S4. Refer to Figure 8 , based on the basic control signal c base (t) output by reinforcement learning, a multi-layer hierarchical neural control network structure is designed to generate an incremental control signal Δc(t) for executable pectoral fin motion; this neural network control system consists of a policy analysis layer and a degree-of-freedom decoding layer, and by decoupling the policy parameters layer by layer, while maintaining the flexibility of high-degree-of-freedom control, the coordinated control between the three degrees of freedom of the pectoral fin is achieved.
[0077] Furthermore, the first-layer policy analysis network takes the current basic control signal c base (t), the generalized coordinate state q(t), and the feedback attitude error θ imu (t) - θ ref (t) as inputs, extracts the current environmental state features through a feedforward neural network, and dynamically selects an appropriate basic gait pattern type (such as high-thrust type, stable type, or energy-saving type). The intermediate variable output by this layer is the policy pattern vector s(t), which is used to guide the refinement process of the lower-layer control.
[0078] Furthermore, the second-layer degree-of-freedom parameter decoding network takes the policy pattern vector s(t) and the basic control signal c base (t) as inputs, extracts multi-scale information using a convolutional network, embeds an LSTM network structure to model the temporal dynamic characteristics of the pectoral fin, and outputs a refined adjustment incremental control signal Δc(t).
[0079] Specifically, the expression of the incremental control signal Δc(t) is: Δc(t) = [ΔA ro (t), ΔA fl (t), ΔA s (t), Δb ro (t), Δb fl (t), Δb s (t), Δω ro (t), Δω fl (t), Δω s (t), Δδ ro (t), Δδ fl (t), Δδ s (t)] T , where ΔA ro (t) is the amplitude adjustment amount of the pectoral fin's front-back flapping motion at time t, ΔA fl (t) is the amplitude adjustment amount of the pectoral fin's up-down flapping motion at time t, ΔA s (t) is the amplitude adjustment amount of the pectoral fin's rocking motion at time t, Δb ro (t) is the offset adjustment amount of the pectoral fin's front-back flapping motion at time t, Δb fl$(t)$ is the offset adjustment amount of the pectoral fin's up-and-down flapping motion at time $t$, $\Delta b$ s $(t)$ is the offset adjustment amount of the pectoral fin's pitching motion at time $t$, $\Delta\omega$ ro $(t)$ is the frequency adjustment amount of the pectoral fin's front-to-back flapping motion at time $t$, $\Delta\omega$ fl $(t)$ is the frequency adjustment amount of the pectoral fin's up-and-down flapping motion at time $t$, $\Delta\omega$ s $(t)$ is the frequency adjustment amount of the pectoral fin's pitching motion at time $t$, $\Delta\delta$ ro $(t)$ is the phase difference adjustment amount of the pectoral fin's front-to-back flapping motion at time $t$, $\Delta\delta$ fl $(t)$ is the phase difference adjustment amount of the pectoral fin's up-and-down flapping motion at time $t$, $\Delta\delta$ s $(t)$ is the phase difference adjustment amount of the pectoral fin's pitching motion at time $t$, $t$ is time; the above control command is the dynamic adjustment amount output of the basic control signal $c$ in the high-level strategy base $(t)$, representing the incremental adjustment of the original policy parameters by the neural network according to the environmental feedback; this command will be added to the basic control signal $c$ base $(t)$ in the original policy to obtain the final control signal $c(t)$, which is used as the input of the subsequent CPG to achieve fine dynamic control of the gait.
[0080] Specifically, the multi-layer neural network is completed through an end-to-end training mechanism, learning from historical perception data and real control effects to optimize network weights; at the same time, it supports online dynamic adjustment, correcting the output command through error feedback to ensure high responsiveness and adaptability; the finally output control signal will be transmitted to the CPG module as a rhythm reference signal to drive the pectoral fin to perform fine movements.
[0081] Specifically, the multi-layer neural control architecture realizes a complete mapping process from policy generation to parameter decoding and then to execution control, with high scalability and environmental adaptability, providing flexible and efficient motion control capabilities for the bionic robotic fish in complex environments.
[0082] S5. Refer to Figure 9 , to ensure the rhythmic, coordinated and continuous composite motion control of the pectoral fin in a three-degree-of-freedom space, the system introduces a CPG network downstream of the neural network control structure as a periodic motion signal generator; this network realizes the real-time collaborative output of multi-degree-of-freedom rhythm signals by coupling multiple non-linear neural oscillators.
[0083] Furthermore, each degree of freedom of the pectoral fin (front-to-back flapping motion, up-and-down flapping motion, pitching motion) corresponds to an independent CPG neural oscillation unit, and the basic form of the oscillator is the Matsuoka neuron model or the Hopf oscillator model. To adapt to the complete structure of the pectoral fin's fluctuation control signal, the CPG module includes the dynamic generation and adjustment of four types of parameters: amplitude, frequency, phase difference and offset, and their state equations can be respectively expressed as:
[0084] ;
[0085] In the formula, the superscript "·" represents the first derivative of the corresponding variable. is the phase of the i-th oscillator, is the phase of the j-th oscillator, ω i is the local frequency, κ ij is the coupling strength, is the target value of the oscillator phase, A i is the current amplitude of the oscillator, A target is the target amplitude of the oscillator, α i is the amplitude adjustment rate, b i is the bias output, b target is the target value of the bias, β i is the bias adjustment rate, ω i is the frequency output, ω target is the target value of the frequency, γ i is the bias adjustment rate, δ i is the phase difference, δ target is the target value of the phase difference, is the phase difference adjustment rate.
[0086] Furthermore, to achieve the coordinated control between multiple degrees of freedom, the system uses the parameters in the control signal c(t) as the reference input of the CPG network to drive each CPG unit to stably generate the required periodic output. The output expression of the CPG network is: , where the phase difference δ i (t) is not generated independently, but is dynamically constructed through the difference between the current oscillator phase of the CPG network and the phase of the selected reference oscillator of the CPG network, that is , t is time; this trajectory model reflects the coordinated adjustment ability of the movement rhythm ( ), amplitude (A i ), bias (b i ) and phase difference (δ i ), ensuring the rhythm, stability and flexibility of the pectoral fin movement, and the output is used as the target movement trajectory of the i-th degree of freedom of the pectoral fin. i can take the values ro, fl, s, representing the forward and backward flapping movement, up and down flapping movement and wing rocking movement respectively.
[0087] At the same time, the coupling mechanism ensures that a fixed phase synchronization relationship is maintained between the three-channel outputs, thereby generating continuous coordinated movement gaits such as "0" shape and "8" shape, and realizing the natural switching of movement modes such as propulsion and turning; see Figure 10, the CPG network output has both flexibility and robustness while ensuring the fluctuation rhythm.
[0088] S6. Refer to Figure 11 , to further improve the robustness and control stability of the bionic fish in a dynamically complex underwater environment, after the output of the rhythm control layer, this system introduces a disturbance suppression strategy combining an Extended State Observer (ESO) and Sliding Mode Control (SMC) to online estimate and real-time compensate for external disturbances, hydrodynamic changes, and model uncertainties.
[0089] Furthermore, take the target trajectory output by the CPG as the reference input of the control system, define the reference pose as , and construct an error signal by combining the actual pose response q(t) of the pectoral fin:
[0090] ;
[0091] In the formula, is the derivative of the actual pose of the pectoral fin, is the derivative of the reference pose of the pectoral fin, is the difference between the desired trajectory and the actual trajectory of the pectoral fin, is the derivative of the difference between the desired trajectory and the actual trajectory of the pectoral fin, and t is time.
[0092] Furthermore, based on the above error, construct a sliding mode controller:
[0093] ;
[0094] In the formula, λ i is a positive real number adjustment parameter, and t is time.
[0095] Furthermore, construct an extended state observer to estimate the system disturbance term d i (t) and the unmodeled dynamics. The structure of the ESO is as follows:
[0096] ;
[0097] In the formula, y i is the actual measured output of the system (the current pose q i (t) of the pectoral fin), u i is the control input, z i1 is the estimated output, z i2 is the estimated speed, z i3To estimate the disturbance term (including uncertainty, modeling error, environmental interference, etc.), β1, β2, and β3 are the observation gains of the ESO, which can affect the estimation speed and stability.
[0098] Furthermore, the controller design adopts a typical sliding mode control law:
[0099] ;
[0100] In the formula, the first term is the disturbance suppression sliding mode term, and the second term is the compensation term estimated by the ESO in real time; is the sliding mode output, is the sign function of the sliding mode surface. When , , when , , when , , is the gain, ,, is the system estimated disturbance term.
[0101] Furthermore, the disturbance compensation control quantity u i (t) output by the sliding mode controller is converted into the angular velocity ω i that can be executed by the driver through the equivalent moment of inertia J ref,i (t). This process includes mapping the control compensation signal to the angular acceleration α i (t), and its expression is: ; Then, the target angular velocity trajectory ω ref,i (t) required by the pectoral fin motor is generated through integration, and its expression is ; In the formula, J i is the equivalent moment of inertia of the drive mechanism system, ω0 is the initial angular velocity (which can be set as the velocity at the end of the previous cycle), t is the time, and ω ref,i (t) is the desired angular velocity input executed by the drive motor, which is directly used to control the servo motor or actuator to drive the pectoral fin to achieve the target pose response.
[0102] Furthermore, this mapping process ensures the dynamic response consistency between the control compensation signal and the actuator, improves the tracking accuracy and attitude robustness in a disturbed environment, and realizes the dynamic compensation response to actual disturbances.
[0103] S6. Refer to Figure 12 , based on the hydrodynamic feedback information collected in real time by the multi-modal perception system, dynamically optimize the pectoral fin gait parameters, and output the basic control signal c base(t) and the incremental control signal Δc(t) to construct a closed-loop control mechanism, so that the pectoral fin movement is always in the optimal gait interval of high efficiency and stability, ensuring the closed-loop adaptability and long-term stable operation of the pectoral fin control system.
[0104] Furthermore, in the control execution stage, the perception system continuously collects the hydrodynamic response thrust F T (t), lift F L (t), lateral force F A (t) and attitude angle θ imu (t) as well as the target pose and the actual pose q(t), thereby constructing a gait performance evaluation index function J(t), which is used to measure the difference between the actual motion effect and the desired target, and its structure can be expressed as:
[0105] ;
[0106] In the formula, is the thrust, is the target thrust, F L (t) is the lift, F A (t) is the lateral force, Var(F L (t)) represents the fluctuation degree of the lift, Var(F A (t)) represents the fluctuation degree of the lateral force, θ imu (t) is the real-time attitude angle, θ ref (t) is the target attitude angle, w1, w2, w3, w4 are weight coefficients, and t is time; the system calculates the performance evaluation index function J(t) in real time and compares it with the set threshold ε. If J(t) > ε, it means that the current gait performance deviates greatly and the parameter fine-tuning mechanism needs to be triggered.
[0107] Furthermore, during the gait parameter fine-tuning process, the system retains the basic control signal c base (t) generated by the reinforcement learning strategy, and at the same time adjusts all four types of key control quantities in the incremental control signal Δc(t), including the amplitude adjustment amount ΔA i (t), the frequency adjustment amount Δω i (t), the bias adjustment amount Δb i (t) and the phase difference adjustment amount Δδ i (t); based on the thrust error, attitude error and lateral force fluctuation as the adjustment basis, the system can construct the following adaptive adjustment law:
[0108] ;
[0109] In the formula, i ∈{ro, fl, s} respectively represent the three degrees of freedom of the pectoral fin's front and back flapping, up and down flapping, and rocking, t is time, ΔAi (t + 1) is the amplitude adjustment amount at time t + 1, Δb i (t + 1) is the bias adjustment amount at time t + 1, Δω i (t + 1) is the frequency adjustment amount at time t + 1, Δδ i (t + 1) is the phase difference adjustment amount at time t + 1, η A and η ω and η b and η δ are the step adjustment factors for each parameter. The adjusted incremental control signal Δc(t) will be used to update the control signal c(t) = c base (t) + Δc(t), driving the CPG network to generate a new displacement reference trajectory .
[0110] Furthermore, this feedback fusion process realizes fast local correction without affecting the stability of the main strategy, improves the adaptive ability of the robotic fish in dynamic disturbances and task switching, and finally realizes a closed-loop fusion control framework of policy learning, neuromodulation, rhythm generation and feedback correction, so that the pectoral fin movement is always in a state of coordinated optimization with high propulsion efficiency and stable posture.
[0111] According to the method of the present invention, it is possible to balance the stability of the robotic fish movement while ensuring the pectoral fin propulsion speed, improve the swimming efficiency and performance of the robotic fish, and enable it to have the ability to operate stably.
[0112] The background part of the present invention may include background information about the problems or environments of the present invention, rather than necessarily describing the prior art. Therefore, the content included in the background art section is not an admission by the applicant of the prior art.
[0113] The above content is a further detailed description of the present invention in combination with specific / preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A three - degree - of - freedom pectoral fin coordinated motion control and its gait optimization method, characterized in that, It includes the following steps: S1. Refer to the typical motion behaviors of the real fish pectoral fin during underwater swimming, establish a basic motion gait model of a three-degree-of-freedom rigid pectoral fin, including forward and backward flapping motions, up and down flapping motions, and wing rocking motions, and use periodic functions to model the kinematic model equation of the three-degree-of-freedom pectoral fin, and define the control signal c(t) as the complete control parameter set of the pectoral fin; S2. Use a multi-modal perception system to collect the hydrodynamic feedback and attitude information of the pectoral fin in real time under different combinations of control input parameters. The perception system includes a distributed pressure sensor and an inertial measurement unit, which respectively collect the thrust F T (t), lift F L (t), lateral force F A (t) and attitude angle θ imu (t); S3. Based on the real-time perception feedback data, construct a pectoral fin-environment interaction model, and use a deep reinforcement learning algorithm to train a high-level policy network to optimize the basic control signal c base (t), with propulsion efficiency and attitude stability as multi-objective reward functions, and the attitude stability is determined by lateral force, lift force, and attitude angle; S4. Construct a multi-layer hierarchical neural control network to decode the basic control signal c base (t) and output the incremental control signal Δc(t). The first layer is used to extract environmental dynamic features and policy patterns, and the second layer outputs the amplitude adjustment amount, frequency adjustment amount, phase difference adjustment amount, and bias adjustment amount for each degree of freedom; S5. Introduce a central pattern generator network, adopt a coupled non-linear oscillator structure, and generate the rhythmic angular displacement reference trajectories of each degree of freedom according to the control signal c(t) = c base (t) + Δc(t) , and construct the cooperative rhythm relationship between degrees of freedom through the phase difference , where is the current oscillator phase of the central pattern generator network, is the selected reference oscillator phase of the central pattern generator network, and t is time; S6. Adopt a sliding mode control method based on an extended state observer, and use a sliding mode controller to construct a disturbance compensation control law u i (t), and realize the dynamic compensation for external disturbances and modeling errors by estimating the disturbance term z i3 (t). is the difference between the desired trajectory and the actual trajectory of the pectoral fin, is the derivative of the difference between the desired trajectory and the actual trajectory of the pectoral fin, is a positive real number adjustment parameter, and t is time; S7. Dynamically evaluate the current gait output effect according to the performance evaluation index function J(t) during the continuous operation. If it deviates from the threshold ε, then online adjust the incremental control signal Δc(t) based on the deviation direction to correct the control parameters, and keep the pectoral fin motion always in the optimal propulsion and stable state to achieve closed-loop optimization, and improve the propulsion efficiency and attitude control ability of the robotic fish.
2. The method for controlling the coordinated movement of three-degree-of-freedom pectoral fins and optimizing its gait according to claim 1, characterized in that, Taking the center of the three-degree-of-freedom motion of the pectoral fin as the origin, define the local coordinate system {P} of the pectoral fin motion: the x-axis points to the forward direction of the robotic fish, the y-axis is the left and right direction of the robotic fish, and the z-axis is perpendicular to the ventral and dorsal directions of the fish body to form a right-handed coordinate system, and establish the generalized coordinate system of the pectoral fin pose; the coordination of the three-degree-of-freedom motion determines the overall motion effect of the pectoral fin and can realize the basic motion gaits of the pectoral fin, including figure-eight and zero-shaped; the kinematic model equation of the pectoral fin is: ; where θ ro (t) is the angular displacement of the forward and backward flapping motion, θ fl (t) is the angular displacement of the up and down flapping motion, θ s (t) is the angular displacement of the wing-rocking motion, A ro is the amplitude of the forward and backward flapping motion, A fl is the amplitude of the up and down flapping motion, A s is the amplitude of the wing-rocking motion, b ro is the offset of the forward and backward flapping motion, b fl is the offset of the up and down flapping motion, b s is the offset of the wing-rocking motion, ω ro is the frequency of the forward and backward flapping motion, ω fl is the frequency of the up and down flapping motion, ω s is the frequency of the wing-rocking motion, δ ro is the phase difference of the forward and backward flapping motion, δ fl is the phase difference of the up and down flapping motion, δ s is the phase difference of the wing-rocking motion, t is time, and the pectoral fin attitude is described by the generalized coordinate vector q(t) = [θ ro (t), θ fl (t), θ s (t)] T When 2ω ro = ω fl , δ ro = π / 2, δ fl = 0, the basic motion gait of the pectoral fin is an "8"-shaped trajectory. When ω ro = ω fl , δ ro = π / 2, δ fl = 0, the basic motion gait of the pectoral fin is a "0"-shaped trajectory; the control signal c(t) of the three-degree-of-freedom model of the pectoral fin contains the following parameters: c(t) = [A ro (t), A fl (t), A s (t), b ro (t), b fl (t), b s (t), ω ro (t), ω fl (t), ω s (t), δ ro (t), δ fl (t), δ s (t)] T , where A ro (t) is the amplitude of the pectoral fin's forward and backward flapping motion at time t, A fl (t) is the amplitude of the pectoral fin's up and down flapping motion at time t, A s (t) is the amplitude of the pectoral fin's rocking motion at time t, b ro (t) is the offset of the pectoral fin's forward and backward flapping motion at time t, b fl (t) is the offset of the pectoral fin's up and down flapping motion at time t, b s (t) is the offset of the pectoral fin's rocking motion at time t, ω ro (t) is the frequency of the pectoral fin's forward and backward flapping motion at time t, ω fl (t) is the frequency of the pectoral fin's up and down flapping motion at time t, ω s (t) is the frequency of the pectoral fin's rocking motion at time t, δ ro (t) is the phase difference of the pectoral fin's forward and backward flapping motion at time t, δ fl (t) is the phase difference of the pectoral fin's up and down flapping motion at time t, δ s (t) is the phase difference of the pectoral fin's rocking motion at time t.
3. The three-degree-of-freedom pectoral fin coordinated motion control and gait optimization method according to claim 1, characterized in that The data sampling frequencies of the multi-modal perception system are as follows: the sampling frequency of the pressure sensor ≥ 100 Hz, the sampling frequency of the inertial measurement unit ≥ 200 Hz, and the attitude angle θ imu (t), thrust F T (t), lift F L (t), lateral force F A (t) are recorded synchronously to construct a high-precision feedback closed loop and achieve a fast response to complex dynamic environments.
4. A three-degree-of-freedom pectoral fin coordinated motion control and gait optimization method according to claim 1, characterized in that During the training process of the high-level policy network, the sliding window strategy is adopted to evaluate the convergence and stability, and the policy network parameters are updated in combination with historical feedback; the form of the reward function R(t) is: ; Where, F T (t) is the thrust, P(t) is the power consumption, F L (t) is the lift force, F A (t) is the lateral force, Var(F L (t)) and Var(F A (t)) respectively represent the fluctuation ranges of the lift force and the lateral force, θ imu (t) is the real-time attitude angle, θ ref (t) is the target attitude angle, and w1, w2, w3, and w4 are the weighting coefficients of the reward function.
5. A three - degree - of - freedom pectoral fin coordinated motion control and its gait optimization method according to claim 1, characterized in that The first layer of the multi-layer hierarchical neural network is a structure based on a graph neural network or a long short-term memory network, which is used to identify the trend of environmental state transition; the second layer is a feedforward neural network or a convolutional neural network, which is used to output an incremental control signal Δc(t) for adjusting the three-degree-of-freedom motion; Δc(t) = [ΔA ro (t), ΔA fl (t), ΔA s (t), Δb ro (t), Δb fl (t), Δb s (t), Δω ro (t), Δω fl (t), Δω s (t), Δδ ro (t), Δδ fl (t), Δδ s (t)] T , where ΔA ro (t) is the regulation amount of the flapping motion amplitude of the pectoral fin at time t, ΔA fl (t) is the regulation amount of the up and down flapping motion amplitude of the pectoral fin at time t, ΔA s (t) is the regulation amount of the rocking motion amplitude of the pectoral fin at time t, Δb ro (t) is the regulation amount of the bias of the flapping motion of the pectoral fin at time t, Δb fl (t) is the regulation amount of the bias of the up and down flapping motion of the pectoral fin at time t, Δb s (t) is the regulation amount of the bias of the rocking motion of the pectoral fin at time t, Δω ro (t) is the regulation amount of the flapping frequency of the pectoral fin at time t, Δω fl (t) is the regulation amount of the up and down flapping frequency of the pectoral fin at time t, Δω s (t) is the regulation amount of the rocking frequency of the pectoral fin at time t, Δδ ro (t) is the regulation amount of the phase difference of the flapping motion of the pectoral fin at time t, Δδ fl (t) is the regulation amount of the phase difference of the up and down flapping motion of the pectoral fin at time t, Δδ s (t) is the regulation amount of the phase difference of the rocking motion of the pectoral fin at time t, and t is time.
6. The three-degree-of-freedom pectoral fin coordinated motion control and gait optimization method according to claim 1, wherein The neural oscillator of the central pattern generator network adopts an amplitude-phase coupling model, and its expression is: ; In the formula, the superscript "·" represents the first derivative of the corresponding variable. is the phase of the i-th oscillator. is the phase of the j-th oscillator, ω i is the local frequency, κ ij is the coupling strength. is the target value of the oscillator phase, A i is the current amplitude of the oscillator, A target is the target value amplitude of the oscillator, α i is the amplitude adjustment rate, b i is the bias output, b target is the target value of the bias, β i is the bias adjustment rate, ω i is the frequency output, ω target is the target value of the frequency, γ i is the bias adjustment rate, δ i is the phase difference, δ target is the target value of the phase difference. is the phase difference adjustment rate; the oscillator amplitude A is dynamically adjusted through the state equation. i frequency ω i bias b i and the oscillator phase to ensure the coupling fluctuation and rhythm continuity between degrees of freedom.
7. A three - degree - of - freedom pectoral fin coordinated motion control and gait optimization method according to claim 1, characterized in that, The disturbance compensation control law u i (t) is converted into the angular acceleration α i (t) through the equivalent moment of inertia J i , and then integrated to obtain the target angular velocity ω ref,i (t) for servo drive control: ; In the formula, ω0 is the initial angular velocity, and t is the time.
8. A three-degree-of-freedom pectoral fin coordinated motion control and gait optimization method according to claim 1, characterized in that The performance evaluation index function J(t) is defined as: ; Where, F T (t) is the thrust force, is the target thrust force, F L (t) is the lift force, F A (t) is the lateral force, Var(F L (t)) and Var(F A (t)) respectively represent the fluctuation ranges of the lift force and the lateral force, θ imu (t) is the real-time attitude angle, θ ref (t) is the target attitude angle, w1, w2, w3, and w4 are the weighting coefficients of the reward function, and t is the time; When the performance evaluation index function J(t) > ε, it indicates that the current gait performance deviates greatly, and it is necessary to trigger the feedback adjustment of each control component in the incremental control signal Δc(t), including: ; where \(i\in\{ro, fl, s\}\) represents the three degrees of freedom of the pectoral fin, namely the forward and backward flapping, the up and down flapping, and the rocking, \(t\) is the time, and \(\Delta A\) i (t + 1) is the amplitude adjustment amount at time \(t + 1\), \(\Delta b\) i (t + 1) is the bias adjustment amount at time \(t + 1\), \(\Delta\omega\) i (t + 1) is the frequency adjustment amount at time \(t + 1\), \(\Delta\delta\) i (t + 1) is the phase difference adjustment amount at time \(t + 1\), \(\eta\) A 、\(\eta\) ω 、\(\eta\) b 、\(\eta\) δ are the parameter step adjustment factors, and the adjusted incremental control signal \(\Delta c(t)\) will be used to update the control signal \(c(t)=c\) base (t)+\(\Delta c(t)\) to drive the central pattern generator network to generate a new displacement reference trajectory .
9. A three-degree-of-freedom pectoral fin coordinated motion control and gait optimization method according to claim 2, characterized in that In the "8"-shaped trajectory motion, the control signal parameter ranges include: the amplitude A of the front and rear flapping wing motion ro ∈ [-90°, 90°], the frequency ω of the front and rear flapping wing motion ro ∈ [1 Hz, 4 Hz], the offset b of the front and rear flapping wing motion ro ∈ [-45°, 45°], the phase difference δ of the front and rear flapping wing motion ro ∈ [-180°, 180°], the amplitude A of the up and down flapping wing motion fl ∈ [-65°, 65°], the frequency ω of the up and down flapping wing motion fl ∈ [1 Hz, 8 Hz], the offset b of the up and down flapping wing motion fl ∈ [-30°, 30°], the phase difference δ of the up and down flapping wing motion fl ∈ [-180°, 180°]; the amplitude A of the wing rocking motion s ∈ [-180°,180°], the frequency ω of the wing rocking motion s ∈ [1 Hz, 4 Hz], the offset b of the wing rocking motion s ∈ [-180°, 180°], the phase difference δ of the wing rocking motion s ∈ [-180°, 180°]; in the "0"-shaped trajectory motion, the amplitude A of the front and rear flapping wing motion ro ∈[-90°, 90°], the frequency ω of the front and rear flapping wing motion ro ∈ [1 Hz, 4 Hz], the offset b of the front and rear flapping wing motion ro ∈ [-45°, 45°], the phase difference δ of the front and rear flapping wing motion ro ∈ [-180°, 180°], the amplitude A of the up and down flapping wing motion fl ∈ [-65°, 65°], the frequency ω of the up and down flapping wing motion fl ∈ [1 Hz, 4 Hz], the offset b of the up and down flapping wing motion fl ∈ [-30°, 30°], the phase difference δ of the up and down flapping wing motion fl ∈ [-180°, 180°]; the amplitude A of the wing rocking motion s ∈ [-180°,180°], the frequency ω of the wing rocking motion s ∈ [1 Hz, 4 Hz], the offset b of the wing rocking motion s ∈ [-180°, 180°], the phase difference δ of the wing rocking motion s ∈ [-180°, 180°]; When the pectoral fins move in the "8"-shaped and "0"-shaped trajectories, they need to be adaptively optimized according to the task requirements and hydrodynamic feedback to achieve the joint optimization of the propulsion efficiency and attitude control ability of the robotic fish.
10. A three - degree - of - freedom pectoral fin coordinated motion control and gait optimization method according to claim 1, characterized in that The performance index weights w1, w2, w3, and w4 in the reward function and the performance evaluation index function are dynamically adjusted according to the environmental complexity and task objectives to adapt to the policy modes of giving priority to propulsion efficiency or attitude stability.
Citation Information
Patent Citations
Implementation method for rapid great pitch angle change motion of pectoral fin paddling type robotic fish
CN104477357A
Bionic robotic fish synergistically propelled by pectoral fins and tail fins and control method of bionic robotic fish
CN117864352A