A biomimetic tuna dorsal fin control method based on policy search

By employing a policy search-based biomimetic tuna dorsal fin control method, utilizing a four-bar linkage and CPG model, combined with a policy gradient algorithm and reinforcement learning network, the movement of the dorsal and caudal fins is coordinated, solving the problems of insufficient maneuverability and stability in existing technologies, and achieving efficient steering and energy utilization.

CN118502243BActive Publication Date: 2025-10-31XIAMEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410570127.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-09
Publication Date
2025-10-31
Estimated Expiration
2044-05-09

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively coordinate the coordinated movements of the dorsal and caudal fins of biomimetic tuna, resulting in insufficient maneuverability and stability in complex environments. Furthermore, traditional policy gradient algorithms face difficulties in convergence or require excessively long training times in high-dimensional action spaces.

Method used

A biomimetic tuna dorsal fin control method based on policy search is adopted. It utilizes a four-bar linkage and a CPG model, combined with a policy gradient algorithm and a reinforcement learning network, to adjust the parameter space of the CPG model in real time, coordinating the movement of the dorsal and caudal fins. The initial parameter space is determined by the Monte Carlo method to reduce variance and improve convergence quality and speed.

Benefits of technology

It achieves optimized coordinated movement of the dorsal and caudal fins in complex environments, improving the maneuverability and stability of the biomimetic tuna, simplifying the control of the dorsal fin opening and closing angle, and enhancing turning efficiency and energy utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118502243B_ABST
    Figure CN118502243B_ABST
Patent Text Reader

Abstract

This invention discloses a biomimetic tuna dorsal fin control method based on policy search. The method includes: controlling the dorsal and caudal fins of the biomimetic tuna based on a policy gradient algorithm; real-time acquisition of power parameters and motion posture parameters of the dorsal and caudal fins during the biomimetic fish's movement; inputting the power parameters and motion posture parameters into a pre-trained reinforcement learning network model and outputting predicted motion parameter spaces for each CPG model; the reinforcement learning network model is constructed based on a policy gradient algorithm; and adjusting the parameter space of the CPG model using the predicted real-time parameter space. This invention can help achieve a more optimized coordinated motion parameter space for the dorsal and caudal fins of the biomimetic tuna.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biomimetic technology, and in particular to a biomimetic tuna dorsal fin control method based on strategy search. Background Technology

[0002] The coordination between fish fins plays a crucial role in determining their propulsion performance, stability, and maneuverability. Therefore, the study of the interaction mechanisms between fish fins and their application in guiding the design and enhancement of robotic fish to improve swimming performance has aroused great interest among scholars in this field.

[0003] Certain body structures in fish, such as tuna, play a crucial role in their swimming performance. For example, the foldable dorsal fin of the tuna is key to improving its turning efficiency. However, existing biological inspiration and biomimetic technologies face challenges in mimicking these complex biological structures. Most current Central Pattern Generator (CPG) models focus on the parameter space of multi-jointed tail movements or the parameter space of coordinated pectoral and caudal fin movements. CPG models targeting the median fins (dorsal and anal fins) mostly focus on the rotational movement of the median fin along the vertical axis, which differs from the movement of real fish. Therefore, existing solutions are insufficient for addressing the problem of improving the turning efficiency of a biomimetic tuna's foldable dorsal fin.

[0004] Furthermore, numerous scholars have applied reinforcement learning methods to the field of biomimetic underwater robots. In classic policy gradient learning algorithms, the policy parameter vector θ is directly learned and explored within the action space. However, this often results in high variance of historical samples, causing gradient estimation to be affected by noise. Moreover, the action space of biomimetic underwater robotic fish can be very large, especially in multi-degree-of-freedom scenarios. Traditional policy gradient algorithms may face challenges in handling high-dimensional action spaces, leading to difficulties in convergence or excessively long training times.

[0005] In summary, existing technologies are insufficient in their ability to enable biomimetic tuna to deploy their biomimetic structures at the right time and angle, which limits the effectiveness and flexibility of biomimetic devices in practical applications. Summary of the Invention

[0006] In view of this, the purpose of this invention is to propose a biomimetic tuna dorsal fin control method based on strategy search, which can help to achieve a more optimized motion parameter space for the coordinated motion of the biomimetic tuna dorsal and caudal fins.

[0007] According to one aspect of the present invention, a biomimetic tuna dorsal fin control method based on strategy search is provided, wherein the dorsal fin is composed of a four-bar linkage consisting of several fin bones and connecting rods, and the outer layer of the fin bones is covered with silicone material; the method includes the following steps

[0008] Bionic tuna dorsal and caudal fins controlled based on CPG model;

[0009] Real-time acquisition of power parameters and motion posture parameters of the dorsal and caudal fins of the bionic fish during movement;

[0010] The power parameters and motion posture parameters are input into a pre-trained reinforcement learning network model, and the predicted motion parameter space of each CPG model is output; the reinforcement learning network model is constructed based on the policy gradient algorithm.

[0011] The parameter space of the CPG model is adjusted in real time using the predicted CPG model parameter space.

[0012] In the aforementioned technical solution, firstly, this application focuses on defining the dorsal fin mechanism of the biomimetic tuna to improve the synergy between hardware and software. The advantage of this configuration is that the opening and closing angle of the dorsal fin can be easily and precisely controlled via a servo motor. While some existing configurations primarily use hydraulic pumps or shape memory alloys to control the opening and closing of linkage mechanisms, these solutions are relatively difficult to achieve precise control of the opening and closing angle, lacking sufficient biomimetic dexterity, and thus present significant challenges in applying them to the solution of this application.

[0013] Secondly, CPG models are set up in various parts, especially the CPG model of the dorsal fin, where each CPG unit corresponds to the control servo of the caudal and dorsal fins. A neural network learning method based on the policy gradient algorithm is proposed. This method aims to coordinate the cooperative movement of the dorsal and caudal fins by generating various parameters in the CPG network offline, thereby achieving optimal maneuverability and stability.

[0014] Finally, experiments were conducted to verify the actual performance under the above parameters, thus validating the correctness of the algorithm. The core objective of this application is to address the limitations of current CPG models in terms of simplicity and the challenge of dorsal fin control in complex flow fields. A biomimetic tuna dorsal fin control method based on policy search is proposed to coordinate the cooperative movement of the dorsal fin and other body parts, aiming to achieve optimal maneuverability and stability in complex environments.

[0015] In some embodiments, controlling the biomimetic tuna dorsal and caudal fins based on a CPG model includes:

[0016] The initial parameter space of the CPG model is determined based on the Monte Carlo method.

[0017] In the above technical solution, the policy gradient algorithm proposed in this invention learns the distribution of policy parameters, rather than the policy parameters themselves. Therefore, by using the Monte Carlo method, the initial parameter space of the CPG model is determined, shifting the exploration from the action space to the parameter space, reducing variance, improving convergence quality and convergence speed, and significantly improving the reliability of the algorithm.

[0018] In some embodiments, the CPG model specifically:

[0019] A biomimetic tuna central pattern generator (CPG) model is established, with each CPG unit corresponding to the tail fin control servo and the dorsal fin control servo, respectively. The specific control network equations are as follows:

[0020]

[0021]

[0022]

[0023]

[0024]

[0025]

[0026]

[0027] Where: b represents the offset state, B represents the advanced control command for the offset amount, and k b A positive constant determines the speed at which b converges to B, m represents the amplitude state, M represents the high-level control command for the amplitude, and k m The positive constant represents the velocity at which m converges to M, φ represents the phase state, ω represents the high-level control command of angular velocity (frequency), α is the rotation angle of the servo motor, β is an intermediate variable, and R represents the time ratio, which represents the time ratio between the two stages that form an oscillation cycle; that is, the CPG model has four control parameters: (M, ω, B, R).

[0028] In the aforementioned technical solution, a biomimetic tuna CPG model is introduced, where each CPG unit corresponds to a control servo for each part. This model is based on the Ijspreet model (used to control the movement of amphibious snake robots). Most CPG models applied to robotic fish are controlled by multiple motors. This biomimetic robotic fish has a compliant tail, so one advantage of this model is that only one motor is needed to generate the waveform. In addition, the time ratio R can also effectively control the robotic fish to turn efficiently. It should be noted that tuna belong to the Thunniform swimmers, and the midline body wave wavelength of this type of swimmer is relatively large, so the midline curve of the body does not form a complete waveform. Therefore, a single servo can fit its body movement midline. Therefore, in order to make the tail curve more similar to that of a real tuna tail, a compliant structure is formed by spring steel plates to better simulate the movement of a real tuna; that is, the tail is connected to the body by several spring steel plates.

[0029] In some embodiments, the parameter space of the CPG model is adjusted using the predicted real-time parameter space of the CPG model, and then the method further includes:

[0030] The biomimetic fish motion is controlled by the parameter space of the adjusted CPG model, and power parameters and motion posture parameters are collected in real time.

[0031] The real-time acquired power parameters and motion posture parameters are input into a pre-established reward function, which outputs the motion evaluation results.

[0032] In the above technical solution, a reward function is designed to evaluate the quality of specific actions in each action set. The purpose is to measure the energy consumed per unit angle of turn during the bionic tuna's turn, thereby characterizing its turn efficiency.

[0033] In some embodiments, the reward function is specifically as follows:

[0034]

[0035] in: Let m represent the average input power, m represent the mass of the biomimetic tuna, and g represent the acceleration due to gravity, taken as 9.81 m / s². 2 ω represents the angular velocity of the tuna during its movement; this formula is used to characterize the energy consumed per unit angle of the bionic tuna's turn, representing the turning efficiency of the bionic tuna.

[0036] In the aforementioned technical solutions, a common formula for evaluating the linear motion efficiency of biomimetic robotic fish is: COT = P / (mgv), which measures the energy consumed per unit distance of linear motion. This application primarily focuses on improving the turning efficiency of a biomimetic tuna. Previous studies on improving turning efficiency only addressed the turning radius or turning angular velocity. This application aims to introduce the concept of energy into the turning process, suggesting that the improvement in turning efficiency should be a result of the coupling of energy and angular velocity. The reward function used in this application is a modification of the aforementioned COT formula, aiming to incorporate energy efficiency.

[0037] According to another aspect of the present invention, a biomimetic tuna dorsal fin control device based on strategy search is provided, which is based on the above-described method and includes a CPG module, a data acquisition module, a prediction module and a control module connected in sequence.

[0038] The CPG module is used to control the dorsal and caudal fins of a biomimetic tuna based on the CPG model.

[0039] The acquisition module is used to collect the power parameters and motion posture parameters of the dorsal and caudal fins of the bionic fish in real time during movement.

[0040] The prediction module is used to input power parameters and motion posture parameters into a pre-trained reinforcement learning network model and output the predicted motion parameter space of each CPG model.

[0041] The control module is used to adjust the parameter space of the CPG model in real time using the predicted CPG model parameter space.

[0042] In order to better utilize the above methods, this application proposes a biomimetic tuna dorsal fin control method based on strategy search. Each module corresponds to a step of the above method, and its specific principle has been described above and will not be repeated here.

[0043] According to another aspect of the present invention, a biomimetic tuna dorsal fin control method based on strategy search is provided, comprising:

[0044] At least one processor and a memory communicatively connected to said at least one processor;

[0045] The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the method described above.

[0046] In the above technical solution, to better operate and process the method, the method is stored in memory, and the processor executes the stored method. It should be noted that the principle and effect of each step have been described above and will not be elaborated upon here.

[0047] According to another aspect of the present invention, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the above-described method.

[0048] In the above technical solution, to better operate and use the method, the method is stored in a computer-readable storage medium and implemented using a processor. It should be noted that the principle and effect of each step have been described above and will not be elaborated upon here. Attached Figure Description

[0049] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0050] Figure 1 This is a flowchart illustrating an embodiment of a biomimetic tuna dorsal fin control method based on strategy search according to the present invention.

[0051] Figure 2 This is a schematic diagram of the dorsal fin structure of an embodiment of a biomimetic tuna dorsal fin control method based on strategy search according to the present invention.

[0052] Figure 3 This is a schematic diagram of the biomimetic tuna structure and CPG model mapping of an embodiment of a biomimetic tuna dorsal fin control method based on strategy search according to the present invention.

[0053] Figure 4 This is a motion control schematic diagram of an embodiment of a biomimetic tuna dorsal fin control method based on strategy search according to the present invention.

[0054] Figure 5 This is a schematic diagram of an embodiment of a biomimetic tuna dorsal fin control method based on strategy search according to the present invention. Detailed Implementation

[0055] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be particularly noted that the following embodiments are for illustrative purposes only and do not limit the scope of the invention. Similarly, the following embodiments are only some, not all, embodiments of the present invention, and all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0056] This invention provides a biomimetic tuna dorsal fin control method based on strategy search, which can help achieve a more optimized motion parameter space for the coordinated movement of the biomimetic tuna dorsal and caudal fins.

[0057] Example 1

[0058] Please see Figure 1 and Figure 2 A biomimetic tuna dorsal fin control method based on strategy search, wherein the dorsal fin is composed of several fin bones A and connecting rods B forming a four-bar linkage mechanism, and the outer layer of the fin bones is covered with silicone material (not shown in the figure); the method includes the following steps:

[0059] S101: Bionic tuna dorsal and caudal fins controlled based on CPG model; please refer to Figure 3 In the figure, the biomimetic tuna 1 represents the dorsal fin control model (CPG1), the caudal fin control model (CPG2), the IMU (Inertial Measurement Unit) 2, the PC 3, the Raspberry Pi main control unit 4, and the power acquisition module 5. In this embodiment, the motion mechanism of the biomimetic tuna 1 includes the dorsal and caudal fins. CPG1 represents the control model for the dorsal fin, and CPG2 represents the control model for the caudal fin. The biomimetic tuna motion used in this embodiment follows the slender body theory proposed by Lighthill. More specifically, its body motion midline can be described as:

[0060] y(x,t)=(c1x+c2x 2 sin(kx+ωt)

[0061] Where: y(x,t) represents the lateral deflection of the midline of the biomimetic tuna body, c1 is the linear amplitude envelope, and c2 is the quadratic amplitude envelope. The body wavenumber is λ (where λ is the body wavelength). The biomimetic tuna dorsal fin consists of several fin bones and connecting rods. The shape of the fin bones is designed to mimic the radial spines of a real tuna dorsal fin. One of the fin bones is controlled by a servo motor to rotate and open / close. The connecting rods further drive the remaining fin bones to open and close synchronously. The connecting rods and fin bones meet the requirements of a four-bar linkage mechanism. The outer layer of the fin bones is covered with silicone material.

[0062] In this embodiment, S101 includes determining the initial parameter space of the CPG model based on the Monte Carlo method. The parameter space in the initial experiment of the bionic tuna is obtained through the Monte Carlo method, randomly generating the required model parameters within the range of values, and inputting the generated random parameter model into the CPG model, so that the bionic tuna begins to perform the expected task and learn.

[0063] In this embodiment, the CPG model specifically includes:

[0064] A biomimetic tuna central pattern generator (CPG) model is established, with each CPG unit corresponding to the tail fin control servo and the dorsal fin control servo, respectively. The specific control network equations are as follows:

[0065]

[0066]

[0067]

[0068]

[0069]

[0070]

[0071]

[0072] Where: b represents the offset state, B represents the advanced control command for the offset amount, and k b A positive constant determines the speed at which b converges to B, m represents the amplitude state, M represents the high-level control command for the amplitude, and k mThe positive constant represents the velocity at which m converges to M, φ represents the phase state, ω represents the high-level control command of angular velocity (frequency), α is the rotation angle of the servo motor, β is an intermediate variable, and R represents the time ratio, which represents the time ratio between the two stages that form an oscillation cycle; that is, the CPG model has four control parameters: (M, ω, B, R).

[0073] In this embodiment, a biomimetic tuna CPG model is introduced, where each CPG unit corresponds to a control servo for each part. This model is based on the Ijspreet model (used to control the movement of amphibious snake robots). Most CPG models applied to robotic fish are controlled by multiple motors. This biomimetic robotic fish has a compliant tail, so one advantage of this model is that only one motor is needed to form the waveform. In addition, the time ratio R can also effectively control the robotic fish to turn efficiently. It should be noted that tuna belong to the Thunniform swimmers, and the midline body wave wavelength of this type of swimmer is relatively large, so the midline curve of the body does not form a complete waveform. Therefore, a single servo can fit its body movement midline. Therefore, in order to make the tail curve closer to that of a real tuna tail, a compliant structure is formed by spring steel plates to better simulate the movement of a real tuna, that is, the tail is connected to the fish body by several spring steel plates. It should be noted that other compliant tail settings are acceptable as long as they can fit the fish body wave curve of the tuna movement well; the spring steel plate connection is a relatively convenient solution.

[0074] S102: Real-time acquisition of power parameters and motion posture parameters of various parts of the bionic fish during movement;

[0075] In this embodiment, power parameters are obtained through power acquisition module 5, which is connected to the propulsion waterproof servo circuit and has both current and voltage acquisition functions. The turning angular velocity during the movement of the bionic tuna is obtained through IMU inertial measurement unit 2. The data acquired by IMU inertial measurement unit 2 and power acquisition module 5 are uploaded to Raspberry Pi main control unit 4, which is connected to PC 3 via WIFI. It should be noted that in order to obtain parameters, the above-mentioned module is embedded in the bionic tuna 1 in this embodiment; other acquisition methods are also possible.

[0076] S103: Input power parameters and motion posture parameters into a pre-trained reinforcement learning network model, and output the predicted motion parameter space of each CPG model; the reinforcement learning network model is constructed based on the policy gradient algorithm; the reinforcement learning network is constructed using the policy gradient algorithm, which can better realize interaction and optimization with external water areas, thereby optimizing the parameter space of the CPG model and achieving better turning efficiency. The policy gradient algorithm maximizes the expected reward by using policy gradient ascent and updates the policy:

[0077]

[0078] Where: γ is the learning rate. θ is the policy gradient, and θ is the policy parameter vector.

[0079] In this embodiment, the reinforcement learning network model is constructed based on the policy gradient algorithm, as follows:

[0080] Using environmental state parameters as input and action selection distribution as output, the expected return is maximized through policy gradient ascent, and the policy is updated with control parameters (M, ω, B, R) of the CPG model.

[0081]

[0082] Where: γ is the learning rate. θ is the policy gradient, and θ is the policy parameter vector;

[0083] One reason for the large variance in the policy gradient is that an empirical average is taken at each time step, which is caused by the randomness of the policy. In this reinforcement learning network model, a linear deterministic policy is:

[0084] π(a|s,θ)=θs

[0085] In addition, we introduce a hyperparameter ρ: p(θ|ρ) is the prior distribution of the policy parameter θ, and the expected reward of the hyperparameter ρ is expressed as:

[0086] J(ρ)=∫∫p(h|θ)p(θ|ρ)R(h)dhdθ

[0087] Therefore, the policy gradient for:

[0088]

[0089] Since the entire history h is determined by only a single sample of the parameter θ in this formula, the variance of the gradient estimate is expected to be reduced.

[0090] The cumulative reward function maximized through policy search is set as the turning efficiency; the power parameters and motion attitude parameters collected by the power acquisition module and the IMU inertial measurement unit 2 are input into a pre-trained policy gradient reinforcement learning network model, and the predicted CPG model parameter space is output; the parameter space of the CPG model is adjusted in real time using the predicted CPG model parameter space.

[0091] This policy gradient reinforcement learning network model involves parameterized policies acting directly on the parameter space. By updating the policy, trajectories with higher rewards are made more likely to occur, while significant policy changes are avoided or entry into undesirable states is prevented. Furthermore, the policy gradient method avoids exploration in the action space, instead shifting the exploration to the parameter space, thereby reducing variance and improving the algorithm's reliability and convergence speed.

[0092] S104: Adjust the parameter space of the CPG model in real time using the predicted CPG model parameter space.

[0093] In this embodiment, S104, followed by:

[0094] The biomimetic fish motion is controlled by the parameter space of the adjusted CPG model, and power parameters and motion posture parameters are collected in real time.

[0095] The real-time acquired power parameters and motion posture parameters are input into a pre-established reward function, which outputs the motion evaluation results.

[0096] In this embodiment, a reward function is designed to evaluate the quality of specific actions in each action set. The purpose is to measure the energy consumed per unit angle of turn during the bionic tuna's turn, thereby characterizing its turn efficiency.

[0097] In this embodiment, the reward function is as follows:

[0098]

[0099] in: Let m represent the average input power, m represent the mass of the biomimetic tuna, and g represent the acceleration due to gravity, taken as 9.81 m / s². 2ω represents the turning angular velocity during the tuna's movement; this formula characterizes the energy consumed per unit turning angle during the bionic tuna's turning process, representing the bionic tuna's turning efficiency. A common formula for evaluating the linear motion efficiency of bionic robotic fish is: COT = P / (mgv), used to measure the energy consumed per unit distance of linear motion. This application mainly focuses on improving the turning efficiency of bionic tuna. Previous studies on improving turning efficiency only addressed the turning radius or turning angular velocity. This application aims to introduce the concept of energy into the turning process; the improvement in turning efficiency should be a result of the coupling of energy and angular velocity. The reward function used in this application is a modification of the aforementioned COT formula, aiming to incorporate energy efficiency. The reward function is used to evaluate the impact of a single concentrated dorsal fin fold on the turning efficiency of the bionic tuna, specifically... The average input power is obtained through a power acquisition module connected to the propulsion waterproof servo circuit, and it also has current and voltage acquisition functions; W represents the turning angular velocity during the tuna's movement, obtained through an inertial measurement unit (IMU). The aforementioned IMU has a three-axis gyroscope and a three-axis accelerometer, which can calculate the attitude angle of the bionic tuna's movement through an algorithm. The mounting point is located at the center of mass of the bionic tuna, which can most realistically reflect the tuna's movement; the specific location of the center of mass is obtained through relevant computer software.

[0100] In this embodiment, please refer to Figure 4 , Figure 4 This is a schematic diagram of motion control in this embodiment. First, this application defines the dorsal fin mechanism of the biomimetic tuna to improve the coordination between hardware and software. The advantage of this configuration is that the opening and closing angle of the dorsal fin can be easily and precisely controlled via a servo motor. Some existing configurations mainly control the opening and closing of the linkage mechanism through hydraulic pumps or shape memory alloys, but these solutions are relatively difficult to achieve precise control of the opening and closing angle, lacking sufficient biomimetic dexterity, and thus present significant difficulties in applying them to the solution of this application.

[0101] Secondly, CGP models are set up for various parts, especially the CPG model of the dorsal fin, where each CPG unit corresponds to the control servo of the caudal and pectoral fins. A neural network learning method based on the policy gradient algorithm is proposed. This method aims to coordinate the cooperative movement of the dorsal fin and other parts by generating the parameters in the CPG network offline, thereby achieving optimal maneuverability and stability.

[0102] Finally, experiments were conducted to verify the actual performance under the above parameters, thus validating the correctness of the algorithm. The core objective of this application is to address the limitations of current CPG models in terms of simplicity and the challenge of dorsal fin control in complex flow fields. A biomimetic tuna dorsal fin control method based on policy search is proposed to coordinate the cooperative movement of the dorsal fin and other body parts, aiming to achieve optimal maneuverability and stability in complex environments.

[0103] Example 2

[0104] Please see Figure 5 A biomimetic tuna dorsal fin control device based on strategy search, characterized in that it is based on the method described in one of the embodiments; including a CPG module, a data acquisition module, a prediction module and a control module connected in sequence;

[0105] The CPG module is used to control the movement of at least one part of a biomimetic tuna based on the CPG model, which includes the biomimetic tuna dorsal fin.

[0106] The data acquisition module is used to collect power parameters and motion posture parameters of the bionic fish in real time.

[0107] The prediction module is used to input power parameters and motion posture parameters into a pre-trained reinforcement learning network model and output the predicted motion parameter space of each CPG model.

[0108] The control module is used to adjust the parameter space of the CPG model in real time using the predicted CPG model parameter space.

[0109] In this embodiment, the device further includes: a reward module, used to control the movement of the bionic fish based on the parameter space of the adjusted CPG model, and to collect power parameters and motion posture parameters in real time; inputting the real-time collected power parameters and motion posture parameters into a pre-established reward function, and outputting the action evaluation result.

[0110] Example 2

[0111] A biomimetic tuna dorsal fin control device based on strategy search, comprising:

[0112] At least one processor and a memory communicatively connected to said at least one processor;

[0113] The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the method described in one of the embodiments.

[0114] In this embodiment, to better run and process the method, the above method is stored in memory, and the stored method is executed using a processor. It should be noted that the principle and effect of each step have been described above and will not be elaborated upon here.

[0115] Example 4

[0116] A computer-readable storage medium storing a computer program, characterized in that the computer program, when executed by a processor, implements the method described in one of the embodiments.

[0117] In this embodiment, to better operate and use the method, the above method is stored in a computer-readable storage medium and implemented using a processor. It should be noted that the principle and effect of each step have been described above and will not be elaborated upon here.

[0118] The above description is only a part of the embodiments of the present invention and does not limit the scope of protection of the present invention. Any equivalent device or equivalent process transformation made based on the content of the present invention specification and drawings, or direct or indirect application in other related technical fields, are similarly included within the patent protection scope of the present invention.

Claims

1. A biomimetic tuna dorsal fin control method based on strategy search, characterized in that, The dorsal fin is composed of several fin bones and connecting rods forming a four-bar linkage, with the outer layer of the fin bones covered with silicone material; the method includes the following steps: Bionic tuna dorsal and caudal fins controlled based on CPG model; Real-time acquisition of power parameters and motion posture parameters of the dorsal and caudal fins of the bionic fish during movement; The power parameters and motion posture parameters are input into a pre-trained reinforcement learning network model, and the predicted motion parameter space of each CPG model is output; the reinforcement learning network model is constructed based on the policy gradient algorithm. The parameter space of the CPG model is adjusted in real time using the predicted CPG model, and the action results are evaluated through a reward function. The CPG model, specifically: A biomimetic tuna central pattern generator (CPG) model is established, with each CPG unit corresponding to the tail fin control servo and the dorsal fin control servo, respectively. The specific control network equations are as follows: in: Indicates the offset state. Advanced control commands that represent offsets. The representation of a positive number determines converged to speed, Indicates the amplitude state. Advanced control commands indicating amplitude. A positive number determines converges to speed, Indicates phase state, Advanced control commands for representing angular velocity. It is the rotation angle of the servo motor. It is an intermediate variable. The time ratio represents the time ratio between the two stages that form one oscillation cycle; that is, this CPG model has four control parameters: , , , ; The reward function is as follows: in: Indicates average input power. Let g represent the mass of the biomimetic tuna, and g represent the acceleration due to gravity. , This represents the turning angular velocity during the tuna's movement; this formula is used to characterize the energy consumed per unit turning angle during the bionic tuna's turning process, representing the turning efficiency of the bionic tuna. The reinforcement learning network model is constructed based on the policy gradient algorithm, as follows: Using environmental state parameters as input and action selection distribution as output, the expected return is maximized by using policy gradient ascent, and the CPG model control parameters are implemented. , , , Strategy Update: in: It is the learning rate. It is the policy gradient. It is a policy parameter vector; In this reinforcement learning network model, a linear deterministic policy is: In addition, a hyperparameter is introduced. :p( ) Let θ be the prior distribution of the policy parameter θ, and the hyperparameters be... The expected return is expressed as: Policy gradient for: 。 2. The biomimetic tuna dorsal fin control method based on strategy search as described in claim 1, characterized in that, Controlling the dorsal and caudal fins of a biomimetic tuna based on the CPG model also includes: The initial parameter space of the CPG model is determined based on the Monte Carlo method.

3. The biomimetic tuna dorsal fin control method based on strategy search as described in claim 1, characterized in that, The parameter space of the CPG model is adjusted in real time using the predicted CPG model parameter space, and then the following is also included: The biomimetic fish motion is controlled by the parameter space of the adjusted CPG model, and power parameters and motion posture parameters are collected in real time. The real-time acquired power parameters and motion posture parameters are input into a pre-established reward function, which outputs the motion evaluation results.

4. A biomimetic tuna dorsal fin control device based on strategy search, characterized in that, The method based on any one of claims 1-3 includes a CPG module, an acquisition module, a prediction module, and a control module connected in sequence. The CPG module is used to control the dorsal and caudal fins of a biomimetic tuna based on the CPG model. The acquisition module is used to collect the power parameters and motion posture parameters of the dorsal and caudal fins of the bionic fish in real time during movement. The prediction module is used to input power parameters and motion posture parameters into a pre-trained reinforcement learning network model and output the predicted motion parameter space of each CPG model. The control module is used to adjust the parameter space of the CPG model in real time using the predicted CPG model parameter space.

5. The biomimetic tuna dorsal fin control device based on strategy search as described in claim 4, characterized in that, The device further includes a reward module, which controls the movement of the bionic fish based on the parameter space of the adjusted CPG model, and collects power parameters and motion posture parameters in real time; inputs the real-time collected power parameters and motion posture parameters into a pre-established reward function, and outputs the action evaluation result.

6. A biomimetic tuna dorsal fin control device based on strategy search, characterized in that, include: At least one processor and a memory communicatively connected to said at least one processor; The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method as described in any one of claims 1 to 3.

7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 3.

Citation Information

Patent Citations

  • Manta ray type bionic fish control method and device based on reinforcement learning and storage medium

    CN115390573A

  • Soft bionic robotic fish swimming optimization method based on CPG model

    CN116300473A