A method, system, and medium for motion planning of a virtual tethered underwater glider
By fusing local flow field information with global flow field prediction data, and combining model predictive control and reinforcement learning, the stability problem of glider arrays in complex ocean flow fields was solved, achieving stable formation and efficient observation of glider arrays.
Patent Information
- Application Number
- CN202510320523.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2045-03-18
AI Technical Summary
Existing motion planning algorithms are inadequate to effectively handle the dynamic changes of virtual moored underwater gliders in complex ocean current fields, causing glider arrays to deviate from their predetermined positions and affecting observation efficiency and stability.
By fusing estimated local flow field information with global flow field prediction data, and combining model predictive control and reinforcement learning methods, the motion planning strategy of the glider is determined, and artificial potential field forces are used to guide the glider to maintain a stable formation in complex flow fields.
It improves the real-time performance and accuracy of flow field prediction, enhances the glider's resistance to flow in complex environments, and ensures the stability and observation efficiency of the glider array.
Smart Images

Figure CN120370984B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of motion planning, and more particularly to a motion planning method, system and medium for a virtual anchored underwater glider. Background Technology
[0002] Virtual moored underwater glider arrays, relying on multi-glider collaborative formation technology, can achieve continuous distributed observation of large-scale sea areas. By sharing tasks, they significantly reduce the energy consumption and mission cycle of individual gliders, thereby improving overall observation efficiency and system robustness. However, due to the highly dynamic and uncertain nature of ocean currents, gliders are prone to deviating from their predetermined positions. Therefore, it is necessary to design anti-current motion planning strategies to maintain the stable stationing capability of the array, i.e., multiple gliders.
[0003] Currently, motion planning is mainly carried out using algorithms such as A* algorithm, fast random search tree algorithm, artificial potential field method, heuristic algorithms (such as genetic algorithm, particle swarm optimization algorithm), and adaptive motion planning algorithm. Although these methods are relatively simple to implement, they are difficult to meet the anti-current requirements of virtual mooring arrays due to their rigidity and inability to cope with the dynamic changes of complex environments. In addition, although deep reinforcement learning (DRL) has been introduced to enhance the adaptability to dynamic environments, its training process relies on unknown flow field interaction and lacks real-time flow field prediction information to optimize control input, resulting in poor local path performance, making it difficult to support multi-glider collaborative anti-current planning, and ultimately affecting the positioning stability of the array.
[0004] Application content
[0005] This application provides a motion planning method, system, and medium for virtual anchored underwater gliders, which performs anti-current motion planning for multiple gliders to ensure the stability of the glider array.
[0006] In a first aspect, this application provides a motion planning method for a virtual moored underwater glider, comprising:
[0007] Based on the actual motion deviations generated by each glider during its floating and sinking process in the preset area, local flow field information is estimated, and the local flow field information is fused with the pre-extracted global flow field prediction data to obtain target flow field information;
[0008] Based on the target flow field information, the first motion planning results of each glider station are determined through a preset model prediction control strategy and a first reinforcement learning method.
[0009] Based on the first motion planning result, the stationing cost and position transfer cost corresponding to each glider are determined respectively. If the stationing cost corresponding to a preset number of gliders is greater than the position transfer cost, the artificial potential field force on each glider is determined, and the first control parameter is determined by the second reinforcement learning method. Based on the first control parameter and the artificial potential field force, the second motion planning result of the glider array is determined, and each glider is controlled to move according to the second motion planning result.
[0010] This application's embodiments obtain relatively accurate sparse flow field information by estimating local flow field information. By fusing the local flow field information with pre-extracted global flow field prediction data, a more accurate target flow field information describing the glider's environment can be obtained, improving the real-time performance, accuracy, and comprehensiveness of flow field prediction and providing a foundation for subsequent motion planning. Through a preset model prediction control strategy and a first reinforcement learning method, the optimal stationing behavior can be found in complex flow fields, while enhancing the glider's resistance to current in complex environments. By comparing the stationing cost and the position transfer cost, it is possible to quantitatively determine whether a stationing adjustment is needed, thereby improving the observation efficiency and adaptability of the glider array. Simultaneously, the artificial potential field method can simulate a virtual force field, guiding the glider to move to the target position in complex flow fields while avoiding obstacles. By adjusting control parameters through reinforcement learning, the motion strategy of the gliders can be dynamically optimized, ensuring that multiple gliders maintain a stable formation during current resistance. Compared with existing technologies, this application can perform current resistance motion planning for multiple gliders, ensuring the stability of the glider array's formation.
[0011] Furthermore, the actual motion deviation includes actual position deviation, actual depth deviation, and actual attitude deviation. The determination of local flow field information based on the actual motion deviation generated during the glider's floating and sinking process in the preset area specifically involves:
[0012] The actual position deviation of the glider during its floating and sinking process in a preset area is obtained, and the horizontal average current velocity at depth is estimated based on the actual position deviation.
[0013] The actual depth deviation and actual attitude deviation during the glider's rise and fall are obtained, and the actual depth deviation and actual attitude deviation are input into a preset state observer to determine the glider's horizontal current velocity depth gradient.
[0014] Local flow field information is determined based on the horizontal average velocity at the depth and the horizontal velocity-depth gradient.
[0015] By estimating local flow field information, relatively accurate sparse flow field information can be obtained.
[0016] Furthermore, the process of fusing the local flow field information with the pre-extracted global flow field prediction data to obtain the target flow field information specifically involves:
[0017] Obtain global flow field prediction data for a preset area;
[0018] The global flow field prediction data and the local flow field information are input into a preset Gaussian process model to dynamically predict the target flow field information in the preset region.
[0019] By fusing the local flow field information with the pre-extracted global flow field prediction data, a more accurate target flow field information describing the environment in which the glider is located can be obtained, improving the real-time performance, accuracy, and comprehensiveness of flow field prediction, and providing a foundation for subsequent motion planning.
[0020] Furthermore, the step of determining the first motion planning result of a single glider based on the target flow field information, through a preset model prediction control strategy and a first reinforcement learning method, specifically involves:
[0021] Based on the target flow field information, the second control parameter is determined using a preset model prediction control strategy;
[0022] The second control parameter is used as a priori guidance for the first reinforcement learning method to determine the third control parameter, and based on the second control parameter and the third control parameter, the first motion planning result for each glider station is determined.
[0023] By using a pre-defined model predictive control strategy and a first reinforcement learning method, the optimal stationary behavior can be found in complex flow fields, while enhancing the glider's resistance to flow in complex environments.
[0024] Furthermore, the calculation formula for determining the second control parameter is as follows:
[0025]
[0026] ||P t -P target ||≤R;
[0027]
[0028] In the formula, J MPC The objective function for predicting the control strategy of the model; The second control parameter is the desired value at time t; It is the second control parameter at time t+i-1; N is the prediction time domain length; P t+i and P t P represents the horizontal position of the glider at times t+i and t, respectively; targetTo achieve the target position of the virtual anchor for the glider; λ is a preset weighting coefficient; f(·) represents the glider motion model; x t v is the glider's position and state vector; t The target flow field information; R is the preset radius of the target station area; u min and u max These are the lower and upper limits of the glider's desired motion control parameters, which are also the lower and upper limits of the motion planning parameters.
[0029] Furthermore, the relevant calculation formula for the artificial potential field force is as follows:
[0030] F total (i)=F att (i)+∑ j≠i F rep (i,j);
[0031] F att (i)=-k att (P i -P i_target );
[0032]
[0033] In the formula, F total (i) represents the artificial potential force acting on glider i; F att (i) represents the gravitational force acting on glider i; F rep (i,j) represents the repulsive force exerted by other gliders j on glider i; k att P is the intensity coefficient of gravity; i P represents the current position of glider i; i_target k represents the virtual mooring target position for glider i. rep P is the intensity coefficient of the repulsive force. j The positions of other gliders j around glider i; d safe To set a safe distance.
[0034] Furthermore, the determination of the first control parameter through the second reinforcement learning method specifically involves:
[0035] A state space for reinforcement learning is constructed based on the glider's current motion state and the force of the artificial potential field.
[0036] Based on the state space, the preset action space, and the preset reward function, a preset second reinforcement learning model is used to perform iterative optimization of the strategy with the goal of maximizing the expected value of the accumulated value of the reward function, and the first control parameters for the glider array to achieve conformity are determined.
[0037] By adjusting control parameters through reinforcement learning, the motion strategy of gliders can be dynamically optimized, ensuring that multiple gliders maintain a stable formation during turbulent conditions.
[0038] Furthermore, the formula for determining the first control parameter using the second reinforcement learning method is as follows:
[0039]
[0040] In the formula, S t_trans Let s be the state space. t_trans Let P be the state vector. i v represents the current position of glider i. i For information about the surrounding flow field, P j For the current position information of other gliders j around glider i, P i_target For the virtual mooring target position of glider i, F total (i) represents the force in the artificial potential field; A t_trans For the action space, Let z be the motion vector, which is the first control parameter for glider i to achieve conformity. i θ i and ψ i These represent the depth of motion, pitch angle, and yaw angle of glider i, respectively; r t_trans Let α be the reward function; t β t and λ t These are the weighting coefficients for each item; c form The cost function for optimizing the formation; N is the number of gliders; d min and d max These represent the minimum and maximum safe distances, respectively; π represents the reinforcement learning strategy. The optimal conservative strategy is formed after strategy optimization; γ t ∈[0,1) represents the discount factor.
[0041] Secondly, this application provides a motion planning system for a virtual moored underwater glider, comprising: an estimation module, a first planning module, and a second planning module;
[0042] The estimation module is used to estimate local flow field information based on the actual motion deviation generated by each glider during the floating and sinking process in the preset area, and to fuse the local flow field information with the pre-extracted global flow field prediction data to obtain the target flow field information.
[0043] The first planning module is used to determine the first motion planning result of each glider station based on the target flow field information, through a preset model prediction control strategy and a first reinforcement learning method.
[0044] The second planning module is used to determine the stationing cost and position transfer cost corresponding to each glider based on the first motion planning result. If the stationing cost corresponding to a preset number of gliders is greater than the position transfer cost, the module determines the artificial potential field force on each glider and determines the first control parameter through the second reinforcement learning method. Based on the first control parameter and the artificial potential field force, the module determines the second motion planning result of the glider array and controls each glider to move according to the second motion planning result.
[0045] This application's embodiments obtain relatively accurate sparse flow field information by estimating local flow field information. By fusing the local flow field information with pre-extracted global flow field prediction data, a more accurate target flow field information describing the glider's environment can be obtained, improving the real-time performance, accuracy, and comprehensiveness of flow field prediction and providing a foundation for subsequent motion planning. Through a preset model prediction control strategy and a first reinforcement learning method, the optimal stationing behavior can be found in complex flow fields, while enhancing the glider's resistance to current in complex environments. By comparing the stationing cost and the position transfer cost, it is possible to quantitatively determine whether a stationing adjustment is needed, thereby improving the observation efficiency and adaptability of the glider array. Simultaneously, the artificial potential field method can simulate a virtual force field, guiding the glider to move to the target position in complex flow fields while avoiding obstacles. By adjusting control parameters through reinforcement learning, the motion strategy of the gliders can be dynamically optimized, ensuring that multiple gliders maintain a stable formation during current resistance. Compared with existing technologies, this application can perform current resistance motion planning for multiple gliders, ensuring the stability of the glider array's formation.
[0046] Thirdly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the motion planning method for a virtual moored underwater glider as described in this application. Attached Figure Description
[0047] Figure 1 This is a flowchart illustrating an embodiment of the motion planning method for a virtual moored underwater glider provided in this application;
[0048] Figure 2 yes Figure 1 A flowchart illustrating step S102;
[0049] Figure 3 This is a schematic diagram of the anti-current motion planning process for a single glider station provided in this application;
[0050] Figure 4 This is a schematic diagram of the anti-current motion planning process for the glider array conformal shape provided in this application;
[0051] Figure 5 This is a flowchart illustrating another embodiment provided in this application;
[0052] Figure 6 This is a schematic diagram of one embodiment of the motion planning system for a virtual anchored underwater glider provided in this application. Detailed Implementation
[0053] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0054] It should be understood that the step numbers used in the text are for ease of description only and are not intended to limit the order in which the steps are performed.
[0055] It should be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0056] The terms “comprising” and “including” indicate the presence of the described feature, whole, step, operation, element and / or component, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or collections thereof.
[0057] The term “and / or” refers to any combination of one or more of the associated listed items, as well as all possible combinations, and includes these combinations.
[0058] Virtual moored underwater glider arrays utilize multi-glider cooperative formation technology for sustainable distributed observation of large-scale sea areas. Task sharing reduces individual glider energy consumption and cycle time, improving observation efficiency and system robustness. However, the strong dynamism and uncertainty of ocean currents can easily cause gliders to deviate from their positions, necessitating the design of current-resistant motion planning strategies to maintain array stability. Currently, motion planning mainly employs A* algorithms, fast random search tree algorithms, artificial potential field methods, heuristic algorithms (such as genetic and particle swarm optimization algorithms), and adaptive motion planning algorithms. While these methods are simple to implement, their algorithms are rigid and difficult to cope with complex dynamic environmental changes. Although deep reinforcement learning is introduced to enhance adaptability to dynamic environments, its training relies on unknown flow field interactions, lacks real-time flow field prediction information to optimize control input, and suffers from suboptimal local path performance, making it difficult to support multi-glider cooperative current-resistant planning and affecting array position stability.
[0059] Next, the terms used in this application will be explained:
[0060] Virtual mooring technology refers to controlling the movement of an underwater glider to maintain a relatively fixed position (i.e., "station") in the water, thereby enabling continuous observation of the hydrological profile of a designated area. Unlike traditional mooring methods, virtual mooring underwater gliders do not require physical anchors; instead, they dynamically maintain their position in the ocean through control algorithms to achieve continuous stationary observation of a designated observation area.
[0061] Model Predictive Control (MPC) is an advanced control method based on system models, widely used in industrial process control, robot motion planning, and autonomous driving. It achieves desired control objectives by predicting the future behavior of the system and optimizing the control inputs.
[0062] The core principle of reinforcement learning (RL) is that an agent continuously adjusts its policy through interaction with the environment to maximize the cumulative reward obtained from the environment. The entire process includes defining the state space and action space, designing the reward function, policy optimization, and value function estimation. The ultimate goal is to learn a policy that can make optimal decisions in a given task, so that the agent achieves optimal long-term rewards.
[0063] Based on this, embodiments of this application provide a motion planning method, system, and medium for virtual anchored underwater gliders, which can perform anti-current motion planning for multiple gliders and ensure the stability of the glider array.
[0064] This application provides a motion planning method, system, and medium for a virtual anchored underwater glider, which will be described in detail through the following embodiments. First, a motion planning method for a virtual anchored underwater glider in this application embodiment is described.
[0065] This application provides a motion planning method for a virtual anchor-tethered underwater glider, relating to the field of motion planning. This method can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing a motion planning method for a virtual anchor-tethered underwater glider, but is not limited to the above forms.
[0066] This application can also be used in numerous general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0067] Example 1
[0068] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the motion planning method for a virtual moored underwater glider provided in this application, including steps S101 to S103.
[0069] Step S101: Based on the actual motion deviation generated by each glider during the floating and sinking process in the preset area, estimate the local flow field information, and fuse the local flow field information with the pre-extracted global flow field prediction data to obtain the target flow field information;
[0070] In some embodiments, the actual motion deviation includes actual position deviation, actual depth deviation, and actual attitude deviation. Determining local flow field information based on the actual motion deviation generated during the glider's floating and sinking process in a preset area includes: acquiring the actual position deviation of the glider during the floating and sinking process in the preset area, and estimating the horizontal average velocity at depth based on the actual position deviation; acquiring the actual depth deviation and the actual attitude deviation during the glider's floating and sinking process, inputting the actual depth deviation and the actual attitude deviation into a preset state observer, and determining the horizontal velocity-depth gradient of the glider; and determining local flow field information based on the horizontal average velocity at depth and the horizontal velocity-depth gradient.
[0071] In some embodiments, the actual position deviation of the glider during its descent and ascent in a preset area is obtained, and the horizontal average current velocity at depth is estimated based on the actual position deviation. Specifically, firstly, after the glider completes each descent and ascent cycle, the descent point position (x1, y1), the ascent point position (x2, y2), and the total time T for descent and ascent are recorded. cycle Secondly, based on the diving point position (x1, y1) and the surfacing point position (x2, y2), the actual displacement of the glider during a single buoyancy maneuver in the actual flow field is determined. Then, the theoretical displacement of the glider under the influence of no flow field can be estimated as follows: in, For the theoretical displacement, It is the average speed of a glider under ideal conditions, T cycle The total time for one gliding cycle; determined by the glider's actual displacement. With theoretical displacement Calculate the horizontal position deviation of a glider during a single dive and ascent caused by the flow field. Then, since the influence of the flow field on the horizontal motion of the glider can be expressed as the average flow velocity at depth level. This can be described by the horizontal position deviation of the glider. (That is, the actual position deviation) is calculated as follows:
[0072] It should be noted that in calculating theoretical displacement... Furthermore, accurate calculations can be made based on glider motion models.
[0073] It should be noted that the preset area is a range of latitude and longitude or a specific geographical area.
[0074] In some embodiments, obtaining the actual depth deviation and the actual attitude deviation during the glider's buoyancy process includes: obtaining actual depth information and actual attitude information during the glider's buoyancy process; calculating theoretical depth information and theoretical attitude information of the glider based on the glider's motion model; determining the actual depth deviation based on the actual depth information and the theoretical depth information; and determining the actual attitude deviation based on the actual attitude information and the theoretical attitude information. Specifically, firstly, during the glider's buoyancy process in a preset area, the glider acquires its own state measurement value Z in real time through navigation and positioning sensors. The measurement value Z includes the glider's actual depth information z. m and glider three-axis attitude θ m ,ψ m That is, the actual attitude information, and based on the acquired information, a measurement model is established as: Z = HX + ν, where Z is the measurement vector: z m This represents the measured depth of the glider. θ m ,ψ m These represent the three-axis attitude; H is the measurement matrix: H = I, where I is the identity matrix, representing the direct measurement of depth and three-axis attitude; ν is the measurement noise. It follows a zero-mean normal distribution, R m To measure the noise covariance matrix and describe the uncertainty of the measurement data; secondly, based on the glider motion model. In the formula, X is the state vector: x, y, z represent the positions of the glider's three axes. θ, ψ are the three-axis attitude angles, u, v, w are the three-axis axial velocities, p, q, r are the three-axis angular velocities, ξ is the vertical gradient of the horizontal flow velocity, which needs to be estimated by designing a state observer; U is the control input. V, r θ and In order, they are net buoyancy control parameters, pitch angle control parameters, and roll angle control parameters or rudder angle control parameters; ω is the process noise. The glider motion model follows a zero-mean normal distribution, and Q is the covariance matrix of the process noise, describing the uncertainty of the state variables. The theoretical depth and attitude information are obtained by solving the glider motion model using numerical integration. Finally, the actual attitude deviation is obtained by comparing the actual depth information with the theoretical depth information, and vice versa.
[0075] In some embodiments, the glider motion model is solved using a numerical integration method to obtain theoretical depth and theoretical attitude information. Specifically, the state vector X is initialized first, including the initial position (x, y, z) and attitude angles. The linear velocity (u, v, w), angular velocity (p, q, r), and vertical gradient of the horizontal flow velocity (ξ) are calculated. Next, the rate of change of state is calculated based on the control input U and the current state vector X. Next, the state vector X is updated using numerical integration methods (such as Euler integral, Runge-Kutta integral, etc.) to obtain the theoretical state at the next moment, including theoretical depth and attitude information. Finally, the above steps are repeated to progressively calculate the theoretical depth and attitude information along the glider's trajectory.
[0076] In some embodiments, the actual depth deviation and the actual attitude deviation are input into a preset state observer to determine the horizontal velocity-depth gradient of the glider; specifically: the actual depth deviation and the actual attitude deviation are input into the preset state observer, and the state observer estimates the horizontal velocity-vertical gradient based on the deviation information, wherein the update formula of the state observer is: In the formula, It is the system state estimated by the state observer, including the vertical gradient of the horizontal flow velocity to be estimated. The actual depth deviation and actual attitude deviation of the glider caused by the flow field are characterized; H is the measurement matrix, here it is the identity matrix I; L is the observer gain matrix, which can be optimized by optimal estimation methods such as Kalman Filter (KF) and its extended variants, Particle Filter (PF), High-Gain Observer (HGO) to ensure that the state estimation error e(t) converges to zero, that is: Solve the above differential equation and update it step by step using numerical integration methods. To obtain an estimate of the system state, thereby determining the horizontal velocity-depth gradient.
[0077] In some embodiments, the horizontal average velocity at depth and the horizontal velocity gradient at depth are combined to construct a local flow field model to analyze the distribution characteristics of the flow field, such as the magnitude, direction and variation of the velocity with depth, and to obtain local flow field information of a preset area. The model can be a parameterized mathematical expression or a numerical simulation model based on a physical mechanism.
[0078] By estimating local flow field information, relatively accurate sparse flow field information can be obtained.
[0079] In some embodiments, the method for obtaining the global flow field prediction data is as follows: by downloading publicly available ocean model flow field prediction data for a preset area, which provides global flow field information for that area. The global flow field prediction data includes, but is not limited to, information such as flow velocity and flow direction. This data, as prior information, can be fused with the estimated local flow field information to further determine the locally accurate global flow field prediction information, i.e., the target flow field information.
[0080] In some embodiments, fusing the local flow field information and pre-extracted global flow field prediction data to obtain target flow field information specifically involves: acquiring global flow field prediction data for a preset region; and inputting the global flow field prediction data and the local flow field information into a preset Gaussian process model to dynamically predict the target flow field information for the preset region. Specifically, after acquiring the global flow field prediction data for the preset region, the global flow field prediction data and the local flow field information need to be input into a Gaussian process model for information fusion to dynamically predict the target flow field information for the preset region, and then outputting the target flow field information for the preset region, including key parameters such as flow velocity and flow direction.
[0081] It should be noted that Gaussian process regression has the following advantages: (1) nonparametric properties, which do not require assumptions about the distribution of data and can directly learn complex nonlinear mapping relationships from the data; (2) probabilistic prediction, which not only outputs the predicted value, but also gives the prediction uncertainty (variance), which is particularly important for flow field estimation in dynamic environments; (3) small sample learning ability, which performs well under small sample conditions and can effectively capture data features.
[0082] It should be noted that, in addition to the Gaussian process model mentioned above, information fusion methods can also include Kalman filtering, Bayesian filtering, deep learning fusion, etc., and this application does not impose any restrictions.
[0083] By fusing the local flow field information with the pre-extracted global flow field prediction data, a more accurate target flow field information describing the environment in which the glider is located can be obtained, improving the real-time performance, accuracy, and comprehensiveness of flow field prediction, and providing a foundation for subsequent motion planning.
[0084] Step S102: Based on the target flow field information, determine the first motion planning result of each glider station by using a preset model prediction control strategy and a first reinforcement learning method;
[0085] In some embodiments, Figure 3 This is a schematic diagram of the anti-current movement planning process for a single glider station, as detailed below. Step S102 may include, but is not limited to, steps S201 to S202:
[0086] Step S201: Based on the target flow field information, determine the second control parameters using a preset model prediction control strategy and a first reinforcement learning method;
[0087] In some embodiments, the target flow field information v t and the glider's position state vector x t (That is, the position within the horizontal plane), in the input model predictive control strategy, the objective function J is minimized within the constraints. MPC In order to solve for the optimal control parameters of the glider under the current flow field, that is, the second control parameters. This is used to guide the short-term movement of a glider, wherein the constraints include position constraints and control input limit constraints, the position constraint being that the glider must remain at the virtual mooring target position P. target Within the nearby circular area, the control input is limited to the upper and lower limits of the glider's desired control input (i.e., the planned motion parameters).
[0088] In some embodiments, the formula for determining the second control parameter is specifically as follows:
[0089] Objective function:
[0090]
[0091] Position constraints:
[0092] ||P t -P target ||≤R;
[0093] Control input limit constraints:
[0094]
[0095] In the formula, J MPC The objective function for predicting the control strategy of the model; The second control parameter is the desired value at time t; It is the second control parameter at time t+i-1; N is the prediction time domain length; P t+i and P t P represents the horizontal position of the glider at times t+i and t, respectively; target To achieve the target position of the virtual anchor for the glider; λ is a preset weighting coefficient; f(·) represents the glider motion model; x t v is the glider's position and state vector; t The target flow field information; R is the preset radius of the target station area; u min and u max These are the lower and upper limits of the glider's desired motion control parameters, which are also the lower and upper limits of the motion planning parameters.
[0096] Step S202: Use the second control parameter as a priori guidance for the first reinforcement learning method to determine the third control parameter, and determine the first motion planning result for each glider station based on the second control parameter and the third control parameter.
[0097] In some embodiments, the second control parameter is used as a priori guidance for the first reinforcement learning method to determine the third control parameter, including: setting a state space, action space, and reward function based on the station-guarding motion planning task requirements; constructing a first reinforcement learning model for a single glider; using the first reinforcement learning model to perform iterative policy optimization with the goal of maximizing the expected value of the accumulated reward function, determining the optimal policy, and then determining the third control parameter for a single glider to achieve station guarding. That is, the first optimal reinforcement learning strategy The action a generated below t_sta This refers to the third control parameter generated by reinforcement learning. (Since reinforcement learning terminology typically uses 'a' to represent the selected action, i.e., the glider control variable (motion planning result) in this application, while robot control typically uses 'u' to represent the control variable, a connection is established between the symbols here, namely...) The third control parameter To achieve the first motion planning result for each glider at the station.
[0098] In some embodiments, the third control parameters for determining a single glider to perform station guarding are described. The calculation formula is as follows:
[0099] S t_sta ={s t_sta |s t_sta =(P t ,v t ,P target )};
[0100] A t_sta ={a t_sta |a t_sta =(z t ,θ t ,ψ t )};
[0101]
[0102] In the formula, S t_sta Let s be the state space of the guard station. t_sta P is the station's state vector; t This represents the glider's current position; v t P represents the flow field velocity. targetA. The virtual mooring target position for the glider; t_sta For the guard action space, a t_sta The stationary action vector; z t ,θ t ,ψ t These are the glider's diving depth, pitch angle, and heading angle, respectively. The second control parameter is determined after the first reinforcement learning method uses the output action of the model-predicted control strategy as a reference for action selection; α s This is an adjustment factor used to balance the second control parameter. And the first reinforcement learning of the action a chosen by itself t_sta The weights; r is the second control parameter; t_sta Let λ be the station-guarding reward function; s These are the weighting coefficients; π represents the reinforcement learning strategy. The optimal guarding strategy formed after strategy optimization is strategy π(a) t_sta |s t_sta In the station's state vector s t_sta Select the guard action vector a t_sta Maximize the accumulated reward function r obtained from the interaction with the environment. t_sta The expected value of the outcome, strategy π(a) t_sta |s t_sta ) represents the state vector s of the station. t_sta The following action vector a is taken to guard the station. t_sta The probability (for stochastic policies) or the direct output action (for deterministic policies); γ s ∈[0,1) represents the discount factor, which controls the rate of decay of future rewards; a smaller γ s It indicates a greater emphasis on short-term rewards; Ε[·] represents the expected value, calculated for different strategies or random environmental state distributions.
[0103] In some embodiments, the first motion planning result for each glider station is determined based on the second control parameter and the third control parameter, specifically by combining the second control parameter. (Short-term optimization) and third control parameter (Long-term optimization) Determine the control parameters u after fusion. t_sta The first motion planning result for a single glider is generated. Short-term and long-term control requirements can be comprehensively considered through weighted averaging or other fusion strategies. The corresponding calculation formula is as follows: In the formula, u t_sta For the final station control parameters; β s The weighting parameter controls the fusion ratio of the two control parameters; This refers to the second control parameter; This refers to the third control parameter.
[0104] It should be noted that the reinforcement learning method mentioned can be deep reinforcement learning, and can be, but is not limited to, DDPG (Deep Deterministic Policy Gradient) or PPO (Proximal Policy Optimization).
[0105] It should be noted that the first control parameter, the second control parameter, and the third control parameter do not indicate an order and can be understood as nouns. These control parameters can be, but are not limited to, parameters such as the pitch angle, heading angle, or diving depth of the glider.
[0106] It should be noted that after generating the first motion planning result for the glider, the reinforcement learning strategy needs to be updated based on the target flow field information to adapt to complex flow field changes.
[0107] It should be noted that this step plans the glider's first motion planning results (including pitch angle, heading angle, and diving depth) so that the glider can move within a circular area (e.g., radius 1 km) centered on the virtual anchor target position. In other words, the glider's position is kept within a certain circular area, rather than being closest to the target position. This control method is more in line with the glider's sawtooth motion characteristics and reduces the difficulty of stationing.
[0108] In this way, the optimal performance of the layout is ensured by using the predictive model, and the actions generated by the predictive model are used as an additional reference for reinforcement learning to enhance the training effect of reinforcement learning, thereby obtaining the global optimal effect. That is, by combining the preset model predictive control strategy and the first reinforcement learning method, the optimal stationing behavior can be found in complex flow fields, while enhancing the glider's resistance to flow in complex environments.
[0109] Step S103: Based on the first motion planning result, determine the stationing cost and position transfer cost corresponding to each glider. If the stationing cost corresponding to a preset number of gliders is greater than the position transfer cost, determine the artificial potential field force on each glider, and determine the first control parameter through the second reinforcement learning method. Based on the first control parameter and the artificial potential field force, determine the second motion planning result of the glider array, and control each glider to move according to the second motion planning result.
[0110] Figure 4 This is a schematic diagram of the anti-current motion planning process for a glider array, as detailed below:
[0111] In some embodiments, the stationing cost C for each glider is determined based on the first motion planning result. stationand location transfer cost C transfer Specifically, the cost of station defense is C. station The cost of maintaining a single glider near a virtual moored target location primarily includes energy consumption C. energy_sta And the penalty for deviating from the target position (i.e., the subsequent position offset cost C) pos Since the glider needs to constantly adjust its pitch angle, heading angle, and depth to resist the flow field during stationing, the energy consumption cost C within a given time T is [amount missing]. energy_sta The calculation formula is: In the formula, C energy_sta As a result of energy consumption, P sta (t) represents the power consumption of the glider at time step t during its stationary operation. It can be estimated from the planned desired control input, i.e., the result of the first motion planning (or by defining a refined energy consumption model). The estimation formula is: P sta (t)=α 1_sta ||z sta (t)‖ 2 +α 2_sta ||θ sta (t)‖ 2 +α 3_sta ||ψ sta (t)‖ 2 In the formula, α 1_sta α 2_sta and α 3_sta The energy consumption correlation coefficient for station maintenance is z. sta (t), θ sta (t) and ψ sta (t) are functions of the glider's diving depth, pitch angle, and heading angle as a function of time during the glider's stationary maneuver; while the position offset cost characterizes the penalty for the glider deviating from the target position, defined as: In the formula, C pos The cost is the location offset, where T is a given time, and P is the cost. t Indicates the glider's current position, P target The virtual anchoring target position for the glider; therefore, considering the energy consumption cost C... energy_sta and position offset cost C pos The cost of guarding the station C station It can be represented as: C station =w1C energy_sta +w2C pos In the formula, w1 and w2 are the weighting coefficients of energy consumption cost and location deviation cost, which can be adjusted according to task priority. The location transfer cost C... transfer The cost of moving a glider from its current target location to another target location mainly includes the path length cost C. length and the energy consumption cost of transfer C energy_trans The glider departs from the current target position P.current Move to the new target location P new Path length cost C length For: C length =‖P current -P new || 2 In the formula, C length P is the cost of the path length. current P represents the glider's current target position. new The new target location for the glider; and the energy cost of the transfer C energy_trans The energy consumption of the glider during the transfer process is represented by the following formula: In the formula, P trans (t) represents the power consumption during the transfer process, where T is a given time, and can be estimated using the formula: P trans (t)=α 1_trans ||z trans (t)‖ 2 +α 2_trans ||θ trans (t)‖ 2 +α 3_trans ||ψ trans (t)‖ 2 In the formula, α 1_trans α 2_trans and α 3_trans z is the energy consumption correlation coefficient for location transfer. trans (t), θ trans (t) and ψ trans (t) are functions of the glider's diving depth, pitch angle, and heading angle as a function of time during the glider's position transfer; finally, considering the path length cost C... length and the energy consumption cost of transfer C energy_trans The cost of the location transfer is: C transfer =w3C length +w4C energy_trans In the formula, w3 and w4 are the weighting coefficients of path length cost and transfer energy cost.
[0112] It should be noted that the stationing cost reflects the cost of a single glider maintaining its position near the virtual mooring target, mainly including the stationing energy cost and the position offset cost; the position transfer cost reflects the cost of a glider moving from its current target position to another target position, mainly including the path length cost and the transfer energy cost.
[0113] It should be noted that the selection of weight coefficients w1, w2, w3, and w4 needs to be adjusted according to the actual application scenario. For example, if the task requires energy saving as a priority, w1 and w4 should be increased appropriately; if the task requires high-precision station guarding, w2 should be increased appropriately; if the path length has a significant impact on the transfer cost, w3 should be increased appropriately.
[0114] In some embodiments, if the station-guarding cost corresponding to a preset number of gliders is greater than the location relocation cost, specifically, when the station-guarding cost C corresponding to each glider is determined... station and the location transfer cost C transfer Subsequently, when the cost of guarding more than a preset number of individual gliders in the glider array exceeds the cost of position relocation (i.e., C... station >C transfer This paper argues that the stationing cost of glider array observation missions outweighs the relocation cost. In this case, the glider array can work collaboratively, using formation change strategies to avoid the influence of strong currents while maintaining its formation. The preset number is a customizable number that can be set as needed, such as one-third of the total number of gliders; this application does not impose any limitations.
[0115] It should be noted that if the condition that "the stationing cost corresponding to the preset number of gliders is greater than the position transfer cost" is not met, then each glider will be controlled to move according to the first motion planning result.
[0116] In some embodiments, the artificial potential field force acting on each glider is determined, specifically by determining the artificial potential field force for each glider based on the artificial potential field method. The artificial potential field force includes an attractive field and a repulsive field, wherein the attractive field is used to attract the glider to move towards the target position, and the repulsive field is used to prevent the glider from colliding with obstacles or other gliders. The relevant calculation formula for the artificial potential field force is as follows:
[0117] F total (i)=F att (i)+∑ j≠i F rep (i,j);
[0118] F att (i)=-k att (P i -P i_target );
[0119]
[0120] In the formula, F total (i) represents the artificial potential force acting on glider i; F att (i) represents the gravitational force acting on glider i; F rep (i,j) represents the repulsive force exerted by other gliders j on glider i; k att P is the intensity coefficient of gravity; i P represents the current position of glider i; i_target k represents the virtual mooring target position for glider i. repP is the intensity coefficient of the repulsive force. j The positions of other gliders j around glider i; d safe To set a safe distance.
[0121] By using the artificial potential field method to construct local gravitational and repulsive fields for each glider, the goal is to ensure that the glider remains within the predetermined virtual mooring area and avoids collisions with other gliders or obstacles.
[0122] In some embodiments, determining the first control parameters using a second reinforcement learning method includes: constructing a reinforcement learning state space based on the glider's current motion state and the artificial potential field force; and using a preset second reinforcement learning model, based on the state space, a preset action space, and a preset reward function, performing iterative policy optimization with the goal of maximizing the expected value of the accumulated reward function, to determine the first control parameters for the glider array to achieve conformity. Specifically: first, based on the glider's current motion state and the artificial potential field force F... total (i) Construct the state space S for reinforcement learning t_trans The state space is as follows: In the formula, s t_trans Let P be the state vector. i v represents the glider's current position. i For information about the surrounding flow field, P j For the current position information of other gliders j around glider i, P i_target For the virtual mooring target position of glider i, F total (i) represents the artificial potential field force. Secondly, based on the state space S... t_trans Preset motion space And a preset reward function, using a preset second reinforcement learning model, to perform policy iterative optimization with the objective of maximizing the expected value of the accumulated value of the reward function, to determine the optimal conformal policy. Wherein, the action space is In the formula, Let z be the action vector. i θ i and ψ i Let represent the depth of motion, pitch angle, and heading angle of glider i, respectively (i.e., the result of the subsequent second motion planning). The formula for determining the optimal conformal strategy is: In the formula, π represents the reinforcement learning strategy. The optimal defensive strategy formed after strategy optimization is the strategy. In the state space vector s t_trans Make a decision Maximize the cumulative conformal reward function r during environmental interaction. t_trans Expected result; γt ∈[0,1) represents the discount factor, which controls the rate of decay of future rewards; a smaller γ t This indicates a greater emphasis on short-term rewards; Ε[·] represents the expected value, calculated for different strategies or stochastic environmental state distributions. Finally, the first control parameters for achieving conformal behavior of the glider array are determined based on the optimal strategy. Because reinforcement learning terminology typically uses 'a' to represent the selected action, i.e., the glider control variable (motion planning result) in this application, while robot control typically uses 'u' to represent the control variable, a connection is established between these symbols here, namely...
[0123] It should be noted that the reinforcement learning method mentioned can be deep reinforcement learning, and can be, but is not limited to, DDPG (Deep Deterministic Policy Gradient) or PPO (Proximal Policy Optimization).
[0124] It should be noted that after generating the second motion planning result for the glider, the reinforcement learning strategy needs to be updated based on the target flow field information to adapt to complex flow field changes.
[0125] In some embodiments, the reward function r t_trans Used for balancing station defense, obstacle avoidance, and formation maintenance, its expression is: In the formula, r t_trans The reward function is P; i P represents the current position of glider i; target For the virtual mooring target position of the glider, α t ,β t ,γ t These are the penalty coefficients for each item. The formation optimization cost function c form In multi-glider cooperative control, it is necessary to consider the global formation and relative positional relationships. This is achieved by defining formation preservation constraints and a formation optimization cost function r. t_trans To achieve formation maintenance, the object maintenance constraint defines the relative distance between gliders i and j as needing to be maintained within a preset range, expressed as: In the formula, d min and d max These are the minimum safe distance and the maximum safe distance, respectively; therefore, the formation optimization cost function c in reinforcement learning is... form The expression is: In the formula, c form The cost function for optimizing the formation is given, where N is the number of gliders and d min and d maxThese are the minimum and maximum safe distances, P. i and P j These are the current positions of gliders i and j, respectively.
[0126] It should be noted that in the reward function expression In the first term, α is used to penalize deviations from the virtual mooring target position. Since cooperative formation control allows each glider to change position within a certain range from the target position to avoid strong current interference, α... t The value should be set slightly smaller; the second term is used to penalize violations of formation preservation, where the formation optimization cost function c form The detailed definition is below; the third item is used to penalize excessive changes in control input and optimize energy consumption.
[0127] It should be noted that, in the formation optimization cost function c form The expression is: In the first max calculation, the goal is to ensure that the relative distance between gliders is not less than the minimum safe distance d. min When the distance between two gliders i and j is ||P i -P j || Less than d min When the distance between the gliders is too close, the max function will output a positive value, indicating that a penalty is needed; if the distance is greater than d, the max function will output a negative value. min If the value is less than or equal to 0, then 0 is returned, and no penalty is incurred. The goal of the second max calculation is to ensure that the relative distance between gliders does not exceed the maximum safe distance d. max When the distance between two gliders i and j is ||P i -P j || Greater than d max When the distance between the gliders is too large, the max function will output a positive value, indicating that a penalty is needed; if the distance is less than d... max If the condition is met, then 0 is returned, and no penalty is imposed.
[0128] In some embodiments, the formula for determining the first control parameter using the second reinforcement learning method specifically includes:
[0129]
[0130] In the formula, S t_trans Let s be the state space. t_trans Let P be the state vector. i v represents the glider's current position. i For information about the surrounding flow field, P j For the current position information of other gliders j around glider i, Pi_target For the virtual mooring target position of glider i, F total (i) represents the force in the artificial potential field; A t_trans For the action space, Let z be the motion vector, which is the first control parameter for glider i to achieve conformity. i θ i and ψ i These represent the depth of motion, pitch angle, and yaw angle of glider i, respectively; r t_trans Let α be the reward function; t β t and λ t These are the weighting coefficients for each item; c form The cost function for optimizing the formation; N is the number of gliders; d min and d max These represent the maximum and minimum distance thresholds, respectively; π represents the reinforcement learning strategy. The optimal conservative strategy is formed after strategy optimization; γ t ∈[0,1) represents the discount factor.
[0131] By adjusting control parameters through reinforcement learning, the motion strategy of gliders can be dynamically optimized, ensuring that multiple gliders maintain a stable formation during turbulent conditions.
[0132] In some embodiments, based on the first control parameters and the artificial potential field force, a second motion planning result for the glider array conformal is determined, specifically: when the determined first control parameters... and the artificial potential field force F total (i) After that, each glider can be updated according to the following formula to obtain the final control parameters, where the relevant formula is: In the formula, u i_trans For the final control parameters, For the first control parameter, F total (i) represents the force in the artificial potential field, λ t These are weight parameters.
[0133] In some embodiments, controlling each glider to move according to the second motion planning result specifically involves: after determining the second motion planning result of the glider array shape, it is necessary to control each glider to move according to the second motion planning parameters.
[0134] Figure 5This is a flowchart of another embodiment provided in this application; first, the local flow field information and the global flow field prediction data are fused by Gaussian process regression. Then, the first motion planning result of each glider station and the second motion planning result of the glider array can be determined. At the same time, the station cost and the position transfer cost can be compared to determine whether to move according to the first motion planning result or the second motion planning result.
[0135] This application provides a simplified summary of the overall steps outlined in Table 1 for ease of understanding:
[0136] Table 1 provides a brief summary of the overall steps involved in this application.
[0137]
[0138] This application embodiment obtains relatively accurate sparse flow field information by estimating local flow field information. By fusing the local flow field information with pre-extracted global flow field prediction data, a more accurate target flow field information describing the environment in which the glider is located can be obtained, improving the real-time performance, accuracy, and comprehensiveness of flow field prediction and providing a foundation for subsequent motion planning. Through a preset model prediction control strategy and a first reinforcement learning method, the optimal stationing behavior can be found in complex flow fields, while enhancing the glider's resistance to current in complex environments. By comparing the stationing cost and the position transfer cost, it is possible to accurately determine whether a stationing adjustment is needed, thereby improving the observation efficiency and adaptability of the glider array. At the same time, by using an artificial potential field to simulate a virtual force field, the glider can be guided to move to the target position in complex flow fields while avoiding obstacles. By adjusting control parameters through reinforcement learning, the motion strategy of the glider can be dynamically optimized, ensuring that multiple gliders maintain a stable formation during the resistance to current. Compared with the prior art, this application can perform resistance motion planning for multiple gliders, ensuring the stability of the glider array's formation.
[0139] Example 2
[0140] like Figure 6 As shown, this application also provides a motion planning system for a virtual moored underwater glider, including: an estimation module 100, a first planning module 200, and a second planning module 300;
[0141] The estimation module 100 is used to estimate local flow field information based on the actual motion deviation generated by each glider during the floating and sinking process in the preset area, and to fuse the local flow field information with the pre-extracted global flow field prediction data to obtain target flow field information.
[0142] The first planning module 200 is used to determine the first motion planning result of each glider station based on the target flow field information, through a preset model prediction control strategy and a first reinforcement learning method.
[0143] The second planning module 300 is used to determine the stationing cost and position transfer cost corresponding to each glider based on the first motion planning result. If the stationing cost corresponding to a preset number of gliders is greater than the position transfer cost, the artificial potential field force on each glider is determined, and the first control parameter is determined through the second reinforcement learning method. Based on the first control parameter and the artificial potential field force, the second motion planning result of the glider array is determined, and each glider is controlled to move according to the second motion planning result.
[0144] The information interaction and execution process between the modules in the motion planning system of the virtual anchored underwater glider described above are based on the same concept as the embodiment of the motion planning method of the virtual anchored underwater glider in the first aspect of the present invention, and the technical effects achieved are basically the same. For details, please refer to the description in the first embodiment of the method of the present invention, which will not be repeated here.
[0145] The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the method in this embodiment, depending on actual needs.
[0146] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the motion planning method for a virtual moored underwater glider as described in this application.
[0147] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0148] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of this application. It should be understood that the above descriptions are merely specific embodiments of this application and are not intended to limit the scope of protection of this application.
[0149] In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application for those skilled in the art.
Claims
1. A motion planning method for a virtual anchor-moored underwater glider, characterized in that, include: Based on the actual motion deviations generated by each glider during its floating and sinking process in the preset area, local flow field information is estimated, and the local flow field information is fused with the pre-extracted global flow field prediction data to obtain target flow field information; Based on the target flow field information, the first motion planning results of each glider station are determined through a preset model prediction control strategy and a first reinforcement learning method. Based on the first motion planning result, the stationing cost and position transfer cost corresponding to each glider are determined respectively. If the stationing cost corresponding to a preset number of gliders is greater than the position transfer cost, the artificial potential field force on each glider is determined, and the first control parameter is determined by the second reinforcement learning method. Based on the first control parameter and the artificial potential field force, the second motion planning result of the glider array is determined, and each glider is controlled to move according to the second motion planning result. The actual motion deviations include actual position deviations, actual depth deviations, and actual attitude deviations. The determination of local flow field information based on the actual motion deviations generated during the glider's buoyancy in the preset area specifically involves: acquiring the actual position deviation of the glider during its buoyancy in the preset area, and estimating the horizontal average velocity at depth based on the actual position deviation; acquiring the actual depth deviation and the actual attitude deviation during the glider's buoyancy, inputting the actual depth deviation and the actual attitude deviation into a preset state observer to determine the horizontal velocity-depth gradient of the glider; and determining local flow field information based on the horizontal average velocity at depth and the horizontal velocity-depth gradient. Specifically, fusing the local flow field information with the pre-extracted global flow field prediction data to obtain the target flow field information involves: acquiring global flow field prediction data for a preset region; and inputting the global flow field prediction data and the local flow field information into a preset Gaussian process model to dynamically predict the target flow field information for the preset region. Specifically, determining the first motion planning result for each glider station based on the target flow field information using a preset model predictive control strategy and a first reinforcement learning method involves: determining a second control parameter based on the target flow field information using a preset model predictive control strategy; using the second control parameter as a priori guide for the first reinforcement learning method to determine a third control parameter; and determining the first motion planning result for each glider station based on the second control parameter and the third control parameter.
2. The motion planning method for a virtual anchor-moored underwater glider according to claim 1, characterized in that, The calculation formula for determining the second control parameter is as follows: ; ; s. t. ; ; ; In the formula, The objective function for predicting the control strategy of the model; In order to be in The second control parameter expected at any given time; It is the first The second control parameter at time; It predicts the length of the time domain; and Glider and The horizontal position at that moment; To achieve the target position of the virtual mooring for the glider; These are preset weighting coefficients; This represents a glider motion model; Let be the position and state vector of the glider; The target flow field information; The radius of the preset target guard area; and These are the lower and upper limits of the glider's desired motion control parameters, which are also the lower and upper limits of the motion planning parameters.
3. The motion planning method for a virtual anchor-moored underwater glider according to claim 1, characterized in that, The relevant calculation formula for the artificial potential field force is as follows: ; ; ; In the formula, For gliders The artificial potential field force it is subjected to; For gliders The gravitational force acting upon it; For other gliders glider The resulting repulsive force; The strength coefficient of gravity; For gliders The current location; For gliders The virtual anchorage target location; The strength coefficient of the repulsive force; For gliders Other gliders around Location; To set a safe distance.
4. The motion planning method for a virtual anchor-moored underwater glider according to claim 1, characterized in that, The determination of the first control parameter through the second reinforcement learning method specifically involves: A state space for reinforcement learning is constructed based on the glider's current motion state and the force of the artificial potential field. Based on the state space, the preset action space, and the preset reward function, a preset second reinforcement learning model is used to perform iterative optimization of the strategy with the goal of maximizing the expected value of the accumulated value of the reward function, and the first control parameters for the glider array to achieve conformity are determined.
5. The motion planning method for a virtual anchor-moored underwater glider according to claim 4, characterized in that, The formula for determining the first control parameter using the second reinforcement learning method is as follows: ; ; ; ; ; ; In the formula, Let the state space be... For state vectors, For gliders Current location For information about the surrounding flow field, For gliders Other gliders around Current location information For gliders The virtual anchorage target location, The artificial potential field force; For the action space, The motion vector is the glider. The first control parameter to achieve conformity They represent gliders The depth of motion, pitch angle, and yaw angle; The reward function; , and These are the weight coefficients for each item; Optimize the cost function for the formation; The number of gliders; and These are the minimum safe distance and the maximum safe distance, respectively. Strategies to enhance learning; This is the optimal defensive strategy formed after strategy optimization; This represents the discount factor.
6. A motion planning system for a virtual anchor-moored underwater glider, characterized in that, include: Estimation module, first planning module, and second planning module; The estimation module is used to estimate local flow field information based on the actual motion deviation generated by each glider during the floating and sinking process in the preset area, and to fuse the local flow field information with the pre-extracted global flow field prediction data to obtain the target flow field information. The first planning module is used to determine the first motion planning result of each glider station based on the target flow field information, through a preset model prediction control strategy and a first reinforcement learning method. The second planning module is used to determine the stationing cost and position transfer cost corresponding to each glider based on the first motion planning result. If the stationing cost corresponding to a preset number of gliders is greater than the position transfer cost, the module determines the artificial potential field force on each glider and determines the first control parameter through the second reinforcement learning method. Based on the first control parameter and the artificial potential field force, the module determines the second motion planning result of the glider array and controls each glider to move according to the second motion planning result. The actual motion deviations include actual position deviations, actual depth deviations, and actual attitude deviations. The determination of local flow field information based on the actual motion deviations generated during the glider's buoyancy in the preset area specifically involves: acquiring the actual position deviation of the glider during its buoyancy in the preset area, and estimating the horizontal average velocity at depth based on the actual position deviation; acquiring the actual depth deviation and the actual attitude deviation during the glider's buoyancy, inputting the actual depth deviation and the actual attitude deviation into a preset state observer to determine the horizontal velocity-depth gradient of the glider; and determining local flow field information based on the horizontal average velocity at depth and the horizontal velocity-depth gradient. Specifically, fusing the local flow field information with the pre-extracted global flow field prediction data to obtain the target flow field information involves: acquiring global flow field prediction data for a preset region; and inputting the global flow field prediction data and the local flow field information into a preset Gaussian process model to dynamically predict the target flow field information for the preset region. Specifically, determining the first motion planning result for each glider station based on the target flow field information using a preset model predictive control strategy and a first reinforcement learning method involves: determining a second control parameter based on the target flow field information using a preset model predictive control strategy; using the second control parameter as a priori guide for the first reinforcement learning method to determine a third control parameter; and determining the first motion planning result for each glider station based on the second control parameter and the third control parameter.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the motion planning method for the virtual anchored underwater glider as described in any one of claims 1-5.
Citation Information
Patent Citations
Method and system for path planning of wave glider
US20210286361A1
Dynamic gliding method and system based on distributed pressure sensors and segmented attitude control
WO2023130691A1