Unmanned ship and unmanned aerial vehicle system cooperative control method based on hydrodynamic parameters and deep reinforcement learning, electronic equipment and readable storage medium
By proposing a cooperative control method for unmanned surface vessels (USVs) and unmanned aerial vehicles (UAVs) based on hydrodynamic parameters and deep reinforcement learning, the cooperative control problem of USVs and UAVs in complex sea conditions was solved, achieving efficient and adaptive cooperative operation and obstacle avoidance capabilities, and improving the stability and operational efficiency of the system.
Patent Information
- Application Number
- CN202511180408.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2025-10-31
AI Technical Summary
Existing unmanned surface vessel (USV) and unmanned aerial vehicle (UAV) collaborative control systems struggle to accurately describe the motion characteristics of USVs in complex sea conditions. Their control algorithms are unable to handle highly nonlinear and strongly coupled sea-air collaborative control problems, lack adaptive learning capabilities, and thus cannot achieve true collaborative control.
A cooperative control method for unmanned surface vessels and unmanned aerial vehicles (UAVs) based on hydrodynamic parameters and deep reinforcement learning is proposed. This method acquires real-time information by carrying a communication module, constructs a dynamic model and linearly parameterizes it, combines an improved artificial potential field obstacle avoidance algorithm based on deep reinforcement learning, designs a network cooperative algorithm, optimizes potential field parameters and control gain, and uses energy function and input state stability theory for stability analysis.
It improves the stability and efficiency of collaborative control between unmanned surface vessels and unmanned aerial vehicles, enabling adaptive adjustment of control strategies in complex environments, avoiding control conflicts, achieving efficient collaborative operation and obstacle avoidance, and enhancing the system's adaptability and operational efficiency.
Smart Images

Figure CN120872027A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of unmanned system control technology, specifically relating to a collaborative control method, electronic equipment, and readable storage medium for unmanned surface vessels and unmanned aerial vehicles based on hydrodynamic parameters and deep reinforcement learning. Background Technology
[0002] The existing unmanned surface vessel (USV)-unmanned aerial vehicle (UAV) collaborative platform for waterborne operations is regarded as an important technical means to improve the efficiency of mission execution in complex environments. Its collaborative control technology has become a key research direction for intelligent system control. The collaboration between USVs and UAVs requires the precise deployment of missions, involving complex dynamic characteristics and multi-objective coordination optimization, which has become a core issue in intelligent control research. Due to the strong coupling nonlinear dynamic characteristics of USVs and UAVs, as well as the influence of multi-source disturbances in complex environments, traditional control methods are difficult to adapt to dynamically changing mission requirements and achieve the desired collaborative effect.
[0003] Furthermore, in fields such as marine monitoring, maritime rescue, and environmental governance, most collaborative control methods rely on sensors for environmental perception and status monitoring. How to improve the robustness and adaptability of the system by combining dynamic models with intelligent algorithms under the condition of using less data has become a key challenge in realizing the collaborative control of unmanned surface vessels and unmanned aerial vehicles.
[0004] Furthermore, with the increasing demands for ocean development and maritime safety, the collaborative operation of Unmanned Surface Vehicles (USVs) and Unmanned Aerial Vehicles (UAVs) is receiving growing attention. USVs possess advantages such as long endurance and large payload capacity, while UAVs offer superior maneuverability and wide field of vision. Their collaboration can significantly improve the efficiency and safety of maritime missions. However, existing USV / UAV collaborative control systems suffer from several drawbacks: insufficient consideration of hydrodynamic parameters, failing to accurately describe the motion characteristics of USVs in complex sea conditions; predominantly using traditional PID control or model predictive control algorithms, which struggle to address highly nonlinear and strongly coupled sea-air collaborative control problems; a lack of adaptive learning capabilities for the sea and air environment, hindering the dynamic adjustment of control strategies based on actual conditions; and insufficient consideration of the collaborative mechanism between USVs and UAVs, making true collaborative control difficult to achieve. Summary of the Invention
[0005] To address the aforementioned technical problems, this invention provides a collaborative control method for unmanned surface vessels and unmanned aerial vehicles (UAVs) systems based on hydrodynamic parameters and deep reinforcement learning, as well as an electronic device and a readable storage medium.
[0006] To achieve the above objectives, the present invention is implemented through the following technical solution:
[0007] The first aspect of this invention provides a cooperative control method for unmanned surface vessels and unmanned aerial vehicles (UAVs) based on hydrodynamic parameters and deep reinforcement learning, comprising the following steps:
[0008] S1. Each unmanned surface vessel and unmanned aerial vehicle (UAV) is equipped with a communication module for data interaction with neighboring UAVs and UAVs. The module acquires the real-time position, speed and external environment information of each UAV and UAV through sensing devices such as radar and environmental sensors.
[0009] S2. Model the hydrodynamic system of each unmanned surface vessel and unmanned aerial vehicle to obtain a dynamic model.
[0010] S3. Linearize the dynamic equations of each unmanned surface vessel and unmanned aerial vehicle.
[0011] S4. Construct deep reinforcement learning control algorithms for unmanned surface vessels and unmanned aerial vehicles to improve obstacle avoidance in artificial potential fields.
[0012] S5. Construct a network collaboration algorithm for unmanned surface vessels and unmanned aerial vehicles.
[0013] S6. Initialize the hardware modules of each unmanned surface vessel and unmanned aerial vehicle (UAV), and load the network cooperation algorithm into the microcontroller unit of each UAV and UAV. The designed framework optimizes the potential field parameters and control gain through DRL to achieve the balance of multiple objectives, and uses energy function and input state stability theory to perform stability analysis on the torque controller based on hydrodynamic parameters and deep reinforcement learning.
[0014] According to the above-mentioned cooperative control method for unmanned surface vessels and unmanned aerial vehicles (UAVs) based on hydrodynamic parameters and deep reinforcement learning, preferably, in step S2, the modeling method of the dynamic model is as follows:
[0015] The dynamic model of the unmanned surface vessel (USV) is as follows:
[0016]
[0017] The dynamic model of the unmanned aerial vehicle (UAV) is as follows:
[0018]
[0019] Among them, M si (η) and M ai (η) represents the mass matrices corresponding to the USV and UAV, respectively; C si (v) and C ai (v) are the Coriolis force matrices for the USV and UAV, respectively; D si (v) and D ai (v) are the damping matrices for the USV and UAV, respectively; g si (η) and g ai(η) are the gravity and buoyancy matrices, respectively; τ si With τ ai These are the control torques for the USV and UAV, respectively; τ wind,si and τ waves,si The wind and wave forces corresponding to USV; τ wind,ai The wind force effect corresponding to the UAV; i = 1, 2.
[0020] According to the above-mentioned cooperative control method for unmanned surface vessels and unmanned aerial vehicles (UAVs) based on hydrodynamic parameters and deep reinforcement learning, preferably, in step S3, the method for linearly parameterizing the dynamic equations of each UAV is as follows:
[0021]
[0022] Among them, Y si (t), Y ai (t) is the regression vector (a known state function); θ si (t) includes parameters such as the unmanned surface vessel's moment of inertia, Coriolis force, hydrodynamic drag, wind speed, and wave height; θ ai (t) includes parameters such as the UAV's moment of inertia, damping, gravity coefficient, wind speed, and aerodynamic forces.
[0023] Define tracking error:
[0024] e si =η si -η si,d ,e ai =η ai -η ai,d ;
[0025] Among them, e si Let η be the i-th state variable of the system. si Its expected value η si,d deviation, e ai Let η be the i-th attitude variable of the system. ai Its expected value η ai,d The deviation; parameter K si K ai The positive definite matrix is used to adjust the influence of sliding mode surfaces on error convergence. Linearization of dynamic parameters plays a key role in reducing complexity, improving computational efficiency, facilitating control design, and enhancing engineering feasibility in dynamic models. In control scenarios with high real-time requirements, linearization is an effective method. However, model errors caused by linearization and complex external environments require the development of nonlinear compensation techniques in high-precision control or strongly nonlinear scenarios. Among them, the USV contains wave force terms, and the corresponding positions of the UAV are filled with zero terms to maintain the consistency of other dynamic terms.
[0026] According to the above-mentioned cooperative control method for unmanned surface vessels and unmanned aerial vehicles based on hydrodynamic parameters and deep reinforcement learning, preferably, in step S4, the deep reinforcement learning control algorithm includes:
[0027] Calculation of the potential gradient:
[0028]
[0029] in,
[0030]
[0031] The potential gradient describes the repulsive effect of obstacles or other entities on the current system. Let represent the repulsive gradients between different variables in the boat-boat, boat-machine, and machine-machine interactions, respectively. Within the range ρ(q) ≤ ρ0, the potential field gradient is continuous and differentiable, making gradient calculation convenient and suitable for application in real-time control systems. The nonlinear form of the nonlinearly controlled potential field gradient (and...) The square correlation ensures a rapidly increasing repulsive force over short distances, enhancing the system's obstacle avoidance response capability; the range of the potential field gradient is limited by ρ0, and this local characteristic reduces interference with distant obstacles, enhancing the computational efficiency of the control system.
[0032] Establish the value function and policy structure for reinforcement learning, wherein the policy structure is constructed using a softmax-based function policy network π. θ (a|s) transforms policy values into action probability distributions, which helps to strike a balance between exploration and exploitation, and outputs the probability distribution for each action:
[0033] π θ (a|s)=softmax(f θ (s)). (5)
[0034] Among them, f θ (s) is a neural network with parameters θ; it takes state s as input and outputs a scalar score for each action.
[0035] Introducing value function networks Soft update of target network parameters; design value function Q with parameter φ. φ (s,a), the goal is to learn the expected cumulative reward for the corresponding state-action pair; it implements soft parameter updates; φ target ←uφ+(1-u)φ target ,u<<1.
[0036] Optimize the policy using the policy gradient method:
[0037]
[0038] The value function network is a key signal for updating the target network. Policy gradient update optimization enhances the model's adaptability, strengthens the selection of high-reward actions, and is used to achieve fast adaptive control.
[0039] Among them, A π (s,a) is the advantage function, representing the degree to which action a is better or worse than the current policy:
[0040] A π (s,a)=Q φ (s,a)-V π (s); (7)
[0041] Among them, V π (s) is the state value function.
[0042] According to the above-mentioned cooperative control method for unmanned surface vessels (USVs) and unmanned aerial vehicles (UAVs) based on hydrodynamic parameters and deep reinforcement learning, preferably, in step S4, the deep reinforcement learning control algorithm further includes: designing the state vector based on the actual situation of USVs and UAVs as follows:
[0043] s=[η s1 ,η s2 ,η a1 ,η a2 ,v s1 ,v s2 ,v a1 ,v a2 U rep ,η d ] T (8)
[0044] Where, η s1 ,η s2 ,η a1 ,η a2 Location information for USV and UAV respectively; v s1 ,v s2 ,v a1 ,v a2 Speed information for USV and UAV respectively; U rep For collision constraint information of the current environment; η d The target location.
[0045] The action vector is:
[0046] a=[Δη s1 ,Δη s2 ,Δη a1 ,Δη a2 ]; (9)
[0047] Where, Δη s1 ,Δηs2 These represent the displacement direction and magnitude of the USV, respectively; Δη a1 ,Δη a2 These represent the displacement direction and amplitude of the UAV, respectively.
[0048] The obstacle avoidance reward function is: r = w1r c +w2r f (10)
[0049] Collision reward r c r c =-∑i,jmax(0,ρ 0,ij -||η i -η j ||); (11)
[0050] Where, ρ 0,ij The desired safe distance; ||η i -η j || represents the relative distance.
[0051] Formation rewards measure the degree of deviation from the target:
[0052] r f =-||η i -η d ||; (12)
[0053] By using the max function to ensure a safe distance and penalizing excessively close distances, and by increasing the weights w1 and w2 of the collision penalty, mutual interference is prioritized to provide a comprehensive optimization objective for multi-agent systems.
[0054] Based on the correction term of known dynamics, a deep neural network g is used. θ (·) Output adaptive parameter update law: where the correction term based on known dynamics is:
[0055] Γ s,t+1 =Γ s,t +Δt·(Y t θ t +g θ (Γ s,t ,s t ,a t (13)
[0056] Among them, Y t θ t The dynamic model with linearized parameters at time t; g θ (·) is a deep neural network used to compensate for errors in dynamic modeling.
[0057] An experience replay mechanism is employed to improve training efficiency. This is achieved by setting up an experience replay buffer to mitigate issues caused by data correlation. Experience replay rules and the target network are defined, and the experience replay buffer stores recent interaction data.
[0058]
[0059] Data is randomly sampled in batches for training to improve data efficiency;
[0060] The parameters of the target network are updated softly:
[0061] φ target ←uΦ+(1-u)φ target ;
[0062] Where u is the update ratio.
[0063] The final optimization goal is:
[0064]
[0065] It includes a value function update objective in the first term and a policy update objective in the second term, with weights λ; the objective function aims to minimize the predicted action value Q. θ (s,a) and target value The differences between them.
[0066] The experience replay buffer stores a tuple of the state, action, reward, and next state for each time step; the weights θ of the target network... target The system is gradually updated to a smooth version of the current network weights θ; the parameter u is usually a small value (e.g., 0.1) to control the update speed; these update rules are closely related to the learning process and dynamically adjust the system's state, actions, and initial condition parameters through an adaptive function to better adapt to environmental changes and optimize the strategy during reinforcement learning.
[0067] Based on the aforementioned collaborative control method for unmanned surface vessels and unmanned aerial vehicles (UAVs) systems based on hydrodynamic parameters and deep reinforcement learning, preferably, in step S4, the state value function V... π (s) Calculated by weighted average of all actions:
[0068] V π (s)=∑ i,j π θ (a|s)Q φ (s,a).
[0069] According to the above-mentioned cooperative control method for unmanned surface vessels and unmanned aerial vehicles (UAVs) based on hydrodynamic parameters and deep reinforcement learning, preferably, in step S5, the network cooperative algorithm includes the following steps: designing a hybrid control strategy module to integrate the above modules and output the final control command; and constructing a distributed control algorithm using parameter vectors and interactive control laws.
[0070]
[0071]
[0072] in, The gradient of the potential field for obstacle avoidance; η si Let η be the real-time location of the i-th USV; sj v represents the real-time location of the j-th USV; si The velocity information of USV at time i; v sj The velocity information of USV at time j; η ai Let η be the real-time location of the i-th UAV; aj v represents the real-time location of the j-th UAV; ai The velocity information of the UAV at time i; v aj The velocity information of the UAV at time j; a ij b ij The coupling coefficient; capable of receiving η di When the information is (i = 1, 2), ζ si, ζ ai ≠0, conversely ζ si, ζ ai =0; where K η,a ,K v,a ,K as and h is a positive definite adaptive gain matrix used to adjust the interaction strength between the robots; ij With R ij The set formation spacing is used to ensure that the submarine and aircraft maintain a reasonable formation when cooperating.
[0073] According to the above-described cooperative control method for unmanned surface vessels (USVs) and unmanned aerial vehicles (UAVs) based on hydrodynamic parameters and deep reinforcement learning, preferably, in step S5, due to the integration of the adaptive controller of the USV-UAV system with deep reinforcement learning and potential fields, the control protocol of the network cooperative algorithm is as follows:
[0074]
[0075] The parameter update law is:
[0076]
[0077] Among them, Γsi , Γ ai Λ si Λ ai e is a positive definite gain matrix; si e ai For tracking error; T sa T as For interactive mapping matrix; The controller integrates basic adaptive control, distributed formation control, potential field obstacle avoidance, deep reinforcement learning control, and parameter adaptive update methods. The design framework uses DRL to optimize potential field parameters and control gain to achieve multi-objective balance, and uses energy function and input state stability theory to perform stability analysis on the torque controller based on hydrodynamic parameters and deep reinforcement learning.
[0078] A second aspect of the present invention provides an electronic device including at least one control processor and a memory for communicatively connecting to the at least one control processor; the memory stores instructions executable by the at least one control processor, the instructions being executed by the at least one control processor to enable the at least one control processor to execute the computer program to implement any step in the cooperative control method for unmanned surface vessels and unmanned aerial vehicles based on hydrodynamic parameters and deep reinforcement learning as described in the first aspect.
[0079] A third aspect of the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a computer processor, implements any step in the cooperative control method for unmanned surface vessels and unmanned aerial vehicles based on hydrodynamic parameters and deep reinforcement learning as described in the first aspect.
[0080] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0081] 1. This invention utilizes hydrodynamic system modeling and control: the dynamic modeling of unmanned surface vessels (USVs) and unmanned aerial vehicles (UAVs) considers factors such as water flow characteristics, aerodynamics, mass inertia, and structural mechanisms, providing accurate environmental data support for collaborative control. By linearizing the dynamic equations of each system, the calculation process can be simplified, improving the stability and efficiency of the control system; the modeling method provides suitable dynamic support for the designed deep reinforcement learning control algorithm.
[0082] 2. The multi-vessel, multi-machine interactive deep reinforcement learning control algorithm of this invention introduces an improved artificial potential field obstacle avoidance method and combines it with new deep reinforcement learning technology for controller design. The obstacle avoidance task of unmanned surface vessels and unmanned aerial vehicles is completed through agent learning. The agent learns by interacting with surrounding obstacles in the environment and generates the optimal control strategy. The combination of deep reinforcement learning and adaptive algorithms can not only improve the obstacle avoidance effect, but also enable the controller to adapt to complex and dynamically changing underwater environments.
[0083] 3. The obstacle avoidance and collaborative working mechanism of the present invention can not only optimize the obstacle avoidance capability of a single unmanned surface vessel or drone, but also maintain a high degree of adaptability when multiple systems work together. Combined with a deep reinforcement learning strategy, the weights of the controller can be dynamically adjusted according to specific task requirements, enabling unmanned surface vessels and drones to adjust their paths in a timely manner in dynamic environments, achieving collaborative operation and obstacle avoidance, and avoiding the control conflicts and performance degradation problems that may occur in traditional control methods.
[0084] 4. This invention addresses the uncertainties of hydrodynamics through deep learning algorithms and, combined with the characteristics of distributed control algorithms, considers the mutual influence and information transmission characteristics of each system, ensuring the stable operation of each UAV and unmanned surface vessel system in complex environments. By employing a precise physical dynamics model, the deep reinforcement learning control algorithm can dynamically adapt to changing environmental conditions with a relatively small amount of data, ensuring that multiple unmanned surface vessels and UAVs remain stable and achieve predetermined goals when working collaboratively, demonstrating highly efficient dynamic environmental adaptability.
[0085] 5. The adaptive parameter update strategy of this invention introduces a novel interaction matrix, enabling the multi-vessel-machine system to respond in real time to various challenges in the surface and underwater environments when the environment changes, thereby improving operational efficiency and safety. It also ensures basic control performance through dynamics and provides high-precision nonlinear compensation. This control method can significantly improve the operational coordination efficiency and intelligence level of surface and underwater unmanned surface vessels and drones, promoting the widespread application of high-end manufacturing equipment such as unmanned vessels and drones in surface and underwater operations. Its applications in marine exploration, environmental monitoring, and logistics transportation have significant potential social and economic benefits.
[0086] 6. This invention takes hydrodynamic modeling, deep reinforcement learning control algorithms, and distributed control algorithms as its core innovations, which can solve the problems of obstacle avoidance and collaborative work of multiple unmanned surface vessels and unmanned aerial vehicles in complex environments. The control method of this invention can not only improve the system's adaptability and stability and is applicable to various dynamic environmental conditions, but also promote the intelligent development of unmanned systems and generate significant economic and social benefits in multiple fields, with broad application prospects. Attached Figure Description
[0087] Figure 1 This is a control diagram of the present invention. Detailed Implementation
[0088] The present invention will be further described in detail below through specific embodiments, but this does not limit the scope of the present invention.
[0089] Example 1
[0090] A cooperative control method for unmanned surface vessels and unmanned aerial vehicles (UAVs) based on hydrodynamic parameters and deep reinforcement learning, such as... Figure 1 As shown, the specific implementation steps of USV-UAV cooperative control are as follows:
[0091] Structural and dimensional analysis of the dynamic equations;
[0092] The dynamic equation is:
[0093]
[0094] The left and right sides of the entire dynamic equation are 3×1 vectors.
[0095] The dynamic models of unmanned surface vessels and drones are simulated by assigning specific values to each element in the regression matrix and parameter vector;
[0096] The form of the regression model;
[0097] 1. The specific construction of the regression matrix for unmanned surface vessels;
[0098] (1) Inertia term
[0099] The quality matrix is:
[0100]
[0101] Where m: hull mass; Longitudinal added mass; Lateral added mass; Coupled additional mass; I z Moment of inertia about the Z-axis; Additional rotational inertia;
[0102] The moment of inertia is:
[0103]
[0104] The relevant terms of the corresponding regression matrix are:
[0105]
[0106] (2) Coriolis force term C(v)v:
[0107] The Coriolis force matrix is:
[0108]
[0109] Coriolis torque is:
[0110]
[0111] Corresponding terms in the regression matrix:
[0112]
[0113] (3) Damping term D(v)v:
[0114] The damping matrix is:
[0115]
[0116] The damping coefficient is:
[0117]
[0118] The damping torque is:
[0119]
[0120] The relevant terms of the corresponding regression matrix are:
[0121]
[0122] (4) Gravity and buoyancy matrix Y g :
[0123]
[0124] Corresponding parameters φ is the roll angle.
[0125] (5) Wind term τ wind and wave force term τ waves :
[0126] Wind power model:
[0127]
[0128] Where, ρ a Air density, V r Relative wind speed, γ r Relative wind angle, A f : Frontal projected area, A l : Lateral projected area, L: Length of the vessel, C x C y C y Wind force coefficient.
[0129] Wave force model:
[0130]
[0131] Wherein: F w1 F w2 Wave force amplitude, M w Wave torque amplitude, ω e Encounter wave frequency.
[0132] Unify it into matrix form:
[0133]
[0134] Combining all the above components, the overall regression matrix is:
[0135] Y(t) = [Y M Y C Y D Y g Y ext ];
[0136]
[0137] This ensures that the overall dynamic equations remain dimensionally consistent: the regression parameter Θ includes inertia, Coriolis force, damping, restoring torque (gravity and buoyancy), and external disturbance forces (wind and waves).
[0138] 2. The dynamic equations of the unmanned aerial vehicle (UAV);
[0139] The dynamic equations of the drone are described as follows:
[0140]
[0141] The left and right sides of the entire dynamic equation are also 3×1 vectors.
[0142] (1) Inertia matrix M ai for:
[0143]
[0144] Where m is the mass of the drone, I z It is the moment of inertia of the drone about its vertical axis.
[0145] (2) Coriolis force matrix C ai Used to describe coupling effects in motion:
[0146]
[0147] Among them, v y and v zIt is the velocity component of the drone.
[0148] (3) Damping matrix D ai Used to describe the dissipation effects caused by air resistance and other external forces:
[0149]
[0150] Where d1, d2, and d3 are damping coefficients.
[0151] (4) Gravity and buoyancy matrix g ai To account for the effects of gravity and buoyancy on the drone, the following assumptions are made regarding the influence of the drone's attitude on these forces:
[0152]
[0153] Where: m is the mass of the drone, and g is the acceleration due to gravity.
[0154] (5) Wind and external disturbance matrix τ wind,ai and τ ai ;
[0155] The effects of wind and other external disturbances on drones can be represented by an external force matrix:
[0156]
[0157] Where, τ wind,x , τ wind,y , τ wind,z These are the wind forces acting on the drone in the x, y, and z directions; all terms are grouped into three main dynamic equations; the model is assumed to be a simplified translational motion model (e.g., the drone's longitudinal, lateral, and vertical motion):
[0158] (1) Regression matrix Y(t);
[0159] The regression matrix Y(t) contains the relationship between each motion component and velocity, acceleration, etc.; integrating each term yields:
[0160]
[0161] (2) Parameter vector Θ;
[0162] The parameter vector Θ includes all system parameters (mass, moment of inertia, added mass, damping coefficient, etc.):
[0163]
[0164] This embodiment includes two unmanned surface vessels (USV1, USV2) and two unmanned aerial vehicles (UAV1, UAV2), forming a collaborative control system.
[0165] Unmanned surface vessel parameters;
[0166] USV1 / USV2 models: Weight: 45kg, Length: 2.0m, Beam: 0.8m, Draft: 0.2m, Maximum speed: 8 knots, Turning radius: 1.2m.
[0167] Drone parameters
[0168] UAV1 / UAV2 models: Quadrotor UAVs. Takeoff weight: 2kg, payload: 6kg, maximum flight speed: 20m / s, endurance: 45 minutes, hovering accuracy: ±0.1m, flight radius: 500m, working altitude: 1-50m.
[0169] Simulation environment setup;
[0170] Simulation area: 100m×100m; Environmental conditions: Wind speed: 3m / s; Wind direction: 0-360° random; Wave height: 0.3-1m; Initial position: USV1: (0,0), USV2: (10,0), UAV1: (0,10), UAV2: (10,10); Target position: (100,100);
[0171] Control parameters;
[0172]
[0173] Policy network: 2 hidden layers, 64 nodes per layer; Value function network: 2 hidden layers, 128 nodes per layer; The learning rate for both the policy network and the value function network is 3e-4, the discount factor γ = 0.99, the soft update coefficient u = 0.005, the experience replay buffer size = 1,000,000, the batch size = 64; the regularization coefficient λ = 0.1.
[0174] Specific examples;
[0175] The unmanned surface vessel has the following specific parameter values:
[0176] Mass m = 45 kg, additional mass Lateral added mass Moment of inertia I z =10kg·m 2 Coupled additional mass Damping coefficients d1 = 50, d2 = 40, d3 = 30, wind force coefficient C x =1.2,C y =0.8,C n =0.5,
[0177] The drone is assumed to have the following specific parameter values:
[0178] Mass m: Total mass of the UAV, m = 2 kg. Longitudinal added mass. Lateral added mass Additional moment of inertia The moment of inertia of the drone about its vertical axis, I z =0.01kg·m 2 The longitudinal damping d1 = 0.05 N·s / m, the lateral damping d2 = 0.05 N·s / m, and the vertical damping d3 = 0.1 N·s / m. The external forces (wind and other external disturbances) are:
[0179]
[0180] Example 2
[0181] An electronic device includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement any step of the cooperative control method for unmanned surface vessels and unmanned aerial vehicles systems based on hydrodynamic parameters and deep reinforcement learning as described in Embodiment 1.
[0182] Furthermore, the process of the cooperative control method for unmanned surface vessels and unmanned aerial vehicles (UAVs) based on hydrodynamic parameters and deep reinforcement learning described in Embodiment 1 can be implemented as a computer software program. For example, this embodiment includes a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the method. In such an embodiment, the computer program can be downloaded and installed from a network, and / or installed from a removable medium. When the computer program is executed by a processor, it performs the functions defined in the method of this application.
[0183] Example 3
[0184] A computer-readable storage medium storing a computer program that, when executed by a processor, implements any step of a cooperative control method for unmanned surface vessels and unmanned aerial vehicle systems based on hydrodynamic parameters and deep reinforcement learning, as described in Embodiment 1.
[0185] The computer-readable medium described in this application may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable medium other than computer-readable storage media, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0186] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Python and C++, as well as conventional procedural programming languages or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0187] In this embodiment, the computer-readable storage medium can be accelerated using hardware such as a GPU. The parallel computing advantage of the GPU is used to accelerate any step in the implementation of a cooperative control method for unmanned surface vessels and unmanned aerial vehicles based on hydrodynamic parameters and deep reinforcement learning, as described in Embodiment 1.
[0188] In summary, this invention effectively overcomes the shortcomings of the prior art and has high industrial applicability. The above embodiments are intended to illustrate the substantive content of this invention, but are not intended to limit the scope of protection of this invention. Those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of this invention without departing from the essence and scope of protection of this invention.
[0189] The above embodiments are specific implementations of the present invention, but the implementation of the present invention is not limited to the above embodiments. Any other combination, change, modification, substitution, or simplification that does not exceed the design concept of the present invention shall fall within the protection scope of the present invention.
Claims
1. A cooperative control method for unmanned surface vessels and unmanned aerial vehicles (UAVs) based on hydrodynamic parameters and deep reinforcement learning, characterized in that, Includes the following steps: S1. Obtain the real-time position, speed, and external environment information of each unmanned surface vessel and drone; S2. Model the hydrodynamic system of each unmanned surface vessel and unmanned aerial vehicle to obtain a dynamic model; S3. Linearize the dynamic equations of each unmanned surface vessel and unmanned aerial vehicle; S4. Construct deep reinforcement learning control algorithms for unmanned surface vessels and unmanned aerial vehicles to improve obstacle avoidance in artificial potential fields; S5. Construct a network collaboration algorithm for unmanned surface vessels and unmanned aerial vehicles; S6. Initialize the hardware modules of each unmanned surface vessel and drone, and load the network cooperation algorithm into the microcontroller unit of each unmanned surface vessel and drone.
2. The cooperative control method for unmanned surface vessels and unmanned aerial vehicles based on hydrodynamic parameters and deep reinforcement learning according to claim 1, characterized in that, In step S2, the modeling method for the dynamic model is as follows: The dynamic model of the unmanned surface vessel (USV) is: The dynamic model of the unmanned aerial vehicle (UAV) is as follows: Among them, M si (η) and M ai (η) represents the mass matrices corresponding to the USV and UAV, respectively; C si (v) and C ai (v) are the Coriolis force matrices for the USV and UAV, respectively; D si (v) and D ai (v) are the damping matrices for the USV and UAV, respectively; g si (η) and g ai (η) are the gravity and buoyancy matrices, respectively; τ si With τ ai These are the control torques for the USV and UAV, respectively; τ wind,si and τ waves,si The wind and wave forces corresponding to USV; τ wind,ai The wind force effect corresponding to the UAV; i = 1, 2.
3. The cooperative control method for unmanned surface vessels and unmanned aerial vehicles (UAVs) based on hydrodynamic parameters and deep reinforcement learning according to claim 2, characterized in that, In step S3, the method for linearly parameterizing the dynamic equations of each unmanned surface vessel and unmanned aerial vehicle is as follows: Among them, Y si (t), Y ai (t) is the regression vector; θ si (t) includes parameters such as the unmanned surface vessel's moment of inertia, Coriolis force, hydrodynamic drag, wind speed, and wave height; θ ai (t) includes parameters such as the UAV's moment of inertia, damping, gravity coefficient, wind speed, and aerodynamic forces.
4. The cooperative control method for unmanned surface vessels and unmanned aerial vehicles based on hydrodynamic parameters and deep reinforcement learning according to claim 3, characterized in that, In step S4, the deep reinforcement learning control algorithm includes: calculating the potential gradient: in, These represent the gradients of the repulsive effects between different variables in the submarine-submarine, submarine-machine, and machine-machine relationships, respectively. Establish the value function and policy structure for reinforcement learning, wherein the policy structure is constructed using a softmax-based function-policy network π. θ (a|s) converts the policy values into an action probability distribution, outputting the probability distribution for each action: p θ (a|s)=softmax(f θ (s); (5) Among them, f θ (s) is a neural network with parameters θ; it takes state s as input and outputs a scalar score for each action. Introducing value function networks Soft update of target network parameters; Optimize the policy using the policy gradient method: Among them, A π (s,a) is the dominance function: A π (s,a)=Q φ (s,a)-V π (s);(7) Among them, V π (s) is the state value function.
5. The cooperative control method for unmanned surface vessels and unmanned aerial vehicles based on hydrodynamic parameters and deep reinforcement learning according to claim 4, characterized in that, In step S4, the deep reinforcement learning control algorithm further includes: establishing a state vector: s=[η s1 ,or s2 ,or a1 ,or a2 ,v s1 ,v s2 ,v a1 ,v a2 ,U rep ,or d ] T ;(8) Where, η s1 ,η s2 ,η a1 ,η a2 Location information for USV and UAV respectively; v s1 ,v s2 ,v a1 ,v a2 Speed information for USV and UAV respectively; U rep For collision constraint information of the current environment; η d For the target location; The action vector is: a=[Dη s1 ,See s2 ,See a1 ,See a2 ]; (9) Where, Δη s1 ,Δη s2 These represent the displacement direction and magnitude of the USV, respectively; Δη a1 ,Δη a2 These represent the direction and magnitude of the UAV's displacement, respectively. The obstacle avoidance reward function is: r = w1r c +w2r f (10) Collision reward r c r c =-Σi,j(0,ρ 0,ij -||η i -η j ||); (11) Where, ρ 0,ij The desired safe distance; ||η i -η j || represents the relative distance; Formation rewards measure the degree of deviation from the target: r f =-||h i -or d ||; (12) Based on the correction term of known dynamics, a deep neural network g is used. θ (·) Output adaptive parameter update law: where the correction term based on known dynamics is: C s,t+1 =C s,t +Δt·(Y t i t +g θ (C s,t ,s t ,a t )); (13) Among them, Y t θ t The dynamic model with linearized parameters at time t; g θ (·) is a deep neural network used to compensate for errors in dynamic modeling; An experience replay mechanism is adopted, which sets experience replay rules and a target network, and uses an experience replay buffer to store the most recent interaction data. The final optimization goal is: It includes a value function update objective in the first term and a policy update objective in the second term, with weights λ; the objective function aims to minimize the predicted action value Q. θ (s,a) and target value The differences between them.
6. The cooperative control method for unmanned surface vessels and unmanned aerial vehicles (UAVs) based on hydrodynamic parameters and deep reinforcement learning according to claim 4, characterized in that, In step S4, the state value function V π (s) Calculated by weighted average of all actions: V π (s)=∑ i,j π θ (a|s)Q φ (s,a)。 7. The cooperative control method for unmanned surface vessels and unmanned aerial vehicles (UAVs) based on hydrodynamic parameters and deep reinforcement learning according to claim 5 or 6, characterized in that, In step S5, the network cooperation algorithm includes the following steps: designing a hybrid control strategy module to output the final control command; and constructing a distributed control algorithm using parameter vectors and interactive control laws. in, The gradient of the potential field for obstacle avoidance; η si Let η be the real-time location of the i-th USV; sj v represents the real-time location of the j-th USV; si The velocity information of USV at time i; v sj The velocity information of USV at time j; η ai Let η be the real-time location of the i-th UAV; aj v represents the real-time location of the j-th UAV; ai The velocity information of the UAV at time i; v aj The velocity information of the UAV at time j; a ij b ij The coupling coefficient is η; it can receive η. di When the information is ζ si, ζ ai ≠0, conversely ζ si, ζ ai =0; where K η,a ,K v,a ,K as and h is a positive definite adaptive gain matrix. ij With R ij The set formation spacing is used to ensure that the submarine and aircraft maintain a reasonable formation when cooperating.
8. The cooperative control method for unmanned surface vessels and unmanned aerial vehicles based on hydrodynamic parameters and deep reinforcement learning according to claim 7, characterized in that, In step S5, the control protocol of the network cooperation algorithm is as follows: The parameter update law is: Among them, Γ si , Γ ai Λ si Λ ai e is a positive definite gain matrix; si e ai For tracking error; T sa T as For interactive mapping matrix; 9. An electronic device, characterized in that, It includes at least one control processor and a memory for communicatively connecting to the at least one control processor; the memory stores instructions executable by the at least one control processor, which, when executed by the at least one control processor, enable the at least one control processor to execute the computer program to implement any step in the cooperative control method for unmanned surface vessels and unmanned aerial vehicles based on hydrodynamic parameters and deep reinforcement learning as described in any of claims 1-8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a computer processor, implements any step in the cooperative control method for unmanned surface vessels and unmanned aerial vehicles based on hydrodynamic parameters and deep reinforcement learning as described in any of claims 1-8.
Citation Information
Cited By
Unmanned aerial vehicle simulation adaptive management system based on multi-modal neural network
CN121477938A
Cable boat consistency cross-domain cooperative control method and system based on model prediction
CN121957117A
Unmanned aerial vehicle autonomous carrier landing control method based on reinforcement learning and related equipment
CN122308452A