An observation unmanned surface vehicle control method based on risk sensitivity and physical potential field constraint

By constructing a control method for observation unmanned surface vessels based on risk sensitivity and physical potential field constraints, the problems of local minima stagnation, model prediction accuracy decline, and sparse rewards in unmanned surface vessel cooperative observation are solved. This enables unmanned surface vessel formations to safely and quickly observe and avoid collisions with highly dynamic targets, thereby improving the system's safety and real-time response capabilities.

CN121742520BActive Publication Date: 2026-04-24DALIAN MARITIME UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
DALIAN MARITIME UNIVERSITY
Filing Date
2026-03-02
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing unmanned surface vessel (USV) cooperative control technologies suffer from problems such as local minima stagnation, decreased model predictive control accuracy, inefficient training due to sparse rewards, and neglect of long-tail risks in multi-USV cooperative observation missions, making it difficult to meet the real-time response and safe observation requirements of highly maneuverable targets.

Method used

A control method for observational unmanned surface vessels (USVs) based on risk sensitivity and physical potential field constraints is constructed. This method involves building an underactuated USV dynamic model, a dimensionless observation state space, and a composite artificial potential field model. Dense reward functions and joint loss functions are established, and a risk-sensitive multi-agent proximal policy optimization algorithm based on conditional risk value is adopted, combined with a physical vector synthesis mechanism for control.

Benefits of technology

It enables long-term continuous observation and dynamic collision avoidance of highly dynamic targets by unmanned surface vessel formations in complex and dynamic marine environments, improving the safety of observation missions and training convergence speed, reducing the probability of collision damage to observation equipment, and ensuring the robustness and real-time response capability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121742520B_ABST
    Figure CN121742520B_ABST
Patent Text Reader

Abstract

The embodiment discloses a kind of observation unmanned ship control methods based on risk sensitivity and physical potential field constraint, establishes the composite artificial potential field model with smooth transition characteristics, and further establishes the dense reward function based on artificial potential field potential energy difference;And according to reward function, establish joint loss function;Further, using the trained Actor network, according to dimensionless observation state space, the dense reward function based on artificial potential field potential energy difference, the strategy flexible force of observation unmanned ship is acquired and is modified right, the total control force after modification is obtained, and the expected speed and expected bow angle of observation unmanned ship are acquired, combined with the dynamics model of underactuated unmanned ship, realize the control of unmanned ship based on collaborative observation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of unmanned surface vessel (USV) technology, and in particular to a control method for observation USVs based on risk sensitivity and physical potential field constraints. Background Technology

[0002] With the development of marine engineering technology, collaborative observation by multiple unmanned surface vessels (USVs) has become a core scenario for marine scientific research, environmental monitoring, and search and rescue missions. This task requires multiple USVs to maintain specific relative geometric configurations towards highly maneuverable targets in dynamic environments to acquire high-quality continuous data. However, existing collaborative control technologies face the following serious challenges in practical applications:

[0003] First, there are limitations to traditional control methods. While methods based on artificial potential fields are computationally simple, they are prone to getting trapped in local minima in multi-vessel game scenarios, causing unmanned vessels to stagnate and oscillate before reaching the observation position. Model predictive control, on the other hand, relies on precise dynamic models, and its prediction accuracy decreases under wind and wave disturbances. Furthermore, as the number of agents increases, its online solution computational load grows exponentially, making it difficult to meet the requirements for millisecond-level real-time response to high-speed targets.

[0004] Second, there are blind spots in existing reinforcement learning methods. Mainstream multi-agent proximal policy optimization algorithms are based on the risk neutrality assumption, aiming to maximize expected returns, but often neglecting long-tail risks that have low probability of occurrence but serious consequences. In close-range observation, this neglect can easily lead to the destruction of expensive observation equipment due to collisions.

[0005] Third, sparse rewards lead to inefficient training. Cooperative observation tasks are characterized by sparse rewards. Positive feedback is only obtained when multiple submarines simultaneously form a specific configuration, causing the agent to blindly explore in the early stages of training, resulting in extremely slow convergence. Summary of the Invention

[0006] This invention discloses a control method for observation unmanned surface vessels based on risk sensitivity and physical potential field constraints, in order to overcome the above-mentioned technical problems.

[0007] To achieve the above objectives, the technical solution of the present invention is as follows:

[0008] A control method for observational unmanned surface vessels based on risk sensitivity and physical potential field constraints includes the following steps:

[0009] S1: Construct a dynamic model of an underactuated unmanned surface vessel for a collaborative observation scenario;

[0010] S2: Based on the underactuated unmanned surface vessel dynamics model, construct a dimensionless observation state space to obtain the state space vector of the observation unmanned surface vessel;

[0011] S3: Construct a composite artificial potential field model with smooth transition characteristics to obtain the total potential energy of the observed unmanned surface vessel and establish a dense reward function based on the potential energy difference of the artificial potential field.

[0012] S4: Establish a joint loss function based on the dense reward function of the potential energy difference of the artificial potential field; then, train the Critic network and Actor network in the risk-sensitive multi-agent proximal policy optimization algorithm architecture based on the loss function.

[0013] S5: Based on the dimensionless observation state space and the dense reward function based on the potential energy difference of the artificial potential field, the trained Actor network is used to obtain the policy flexibility force of the observation unmanned surface vessel.

[0014] S6: Based on the strategic flexibility of the observed unmanned surface vessel (USV), obtain the corrected total control force to obtain the expected speed and expected heading angle of the USV. Then, combine the underactuated USV dynamics model to realize the control of the USV based on cooperative observation.

[0015] Furthermore, the composite artificial potential field model with smooth transition characteristics is constructed as follows:

[0016]

[0017] In the formula: This represents the current position vector of the observed unmanned surface vessel; Indicates the observed gravitational potential field; It is a repulsive potential field in the environment; The total potential energy function;

[0018] in,

[0019]

[0020] In the formula: This indicates the real-time distance between the observed unmanned surface vessel (USV) and the target vessel. This represents the gravitational gain coefficient, used to adjust the magnitude of gravity. This indicates the optimal orbital observation radius set for the mission;

[0021]

[0022] In the formula: Indicates the distance between the unmanned surface vessel and obstacles or teammates. This represents the maximum saturation repulsion value; For polynomial adjustment coefficients; R Warning radius; r This is the danger radius.

[0023] Furthermore, the dense reward function based on the potential energy difference of the artificial potential field is established as follows:

[0024]

[0025] In the formula: For the total reward function; Rewards for sparse terminals; Rewards are given based on potential energy. Punishment for security risks;

[0026] in,

[0027]

[0028] In the formula: Positive reward; Negative rewards;

[0029]

[0030] In the formula: Indicates the first i An observation unmanned surface vessel in time step t The potential energy guidance reward obtained; For the first i An observation unmanned surface vessel in time step t The state of time; To reinforce the discount factor in learning; Indicates the first i An observation unmanned surface vessel in time step t The total potential energy at position +1; Indicates the first i An observation unmanned surface vessel in time step t The total potential energy at the location.

[0031] Furthermore, the conditional value at risk is constructed as follows:

[0032] S41: Obtain the discounted cumulative reward of the training sampled trajectory of the observed unmanned surface vessel, using the following formula:

[0033]

[0034] In the formula: For the first k The discounted cumulative reward of the training sampling trajectory of the observation unmanned surface vessel; For the index of the time step; This represents the total number of time steps. Discount factor; The index number is used to identify the training sampling trajectory of the unmanned surface vessel. The total number of training sampling trajectories observed from unmanned surface vessels; For the first k The training sampling trajectory of the observation unmanned surface vessel was in the first... t Instant rewards for time steps;

[0035] S42: Based on the discounted cumulative reward of the observed unmanned surface vessel's training sampling trajectory, the risk threshold is obtained as follows:

[0036]

[0037] In the formula: The set of discounted cumulative rewards for training sample trajectories of unmanned surface vessels; This refers to quantile function operations; Confidence level; Risk threshold;

[0038] S43: The conditional value of risk is as follows:

[0039]

[0040] In the formula: Conditional risk value; As expected.

[0041] Furthermore, the joint loss function includes the loss function of the Critic network and the loss function of the Actor network;

[0042] The loss function of the Critic network is expressed as follows:

[0043]

[0044] In the formula: The loss function; The parameter is The state value function estimated by the centralized Critic network; This is the parameter set for the Critic network; For the first The training sampling trajectory of the observation unmanned surface vessel at the time step t The combined state; For the first k The training sampling trajectory of the observation unmanned surface vessel at the time step t Discount return estimate;

[0045] in,

[0046]

[0047]

[0048] In the formula: The amount calculated based on the gap between the current return and the conditional value at risk, i.e., the first... k The magnitude of the reduction in conditional risk value caused by the training sampling trajectory of the observation unmanned surface vessel; For the first k Sample weights for the training sampling trajectories of the observation unmanned surface vessel; As the risk amplification coefficient, ; This is the absolute value of the risk threshold; This is expressed as a minimal constant to prevent the denominator from being zero;

[0049] The loss function of the Actor network is expressed as follows:

[0050]

[0051] In the formula: The objective function is PPO, which is the loss function value of the Actor network; The parameter set of the Actor network; For the first i The expectation of an unmanned observation vessel; It represents the probability ratio; The dominant function; This is a truncation function; This is the PPO trimming factor.

[0052] Furthermore, the formula used to obtain the corrected total control force is as follows:

[0053]

[0054] In the formula: The revised total control force; For strategic flexibility; It has a strong repulsive force on the environment; For task interaction; For strategic flexibility The scaling factor;

[0055] The formulas used to obtain the desired velocity and desired heading angle of the observed unmanned surface vessel are as follows:

[0056]

[0057]

[0058] In the formula: For the desired speed; Desired heading angle; For modulo length calculation; This is the maximum permissible linear velocity for the unmanned surface vessel. All of these are components of the resultant force in the inertial coordinate system.

[0059] Furthermore, the dynamic model of the underactuated unmanned surface vessel is constructed as follows:

[0060]

[0061] in: Indicates the number of the unmanned surface vessel being observed. They represent the first i The longitudinal and lateral position coordinates of the unmanned observation vessel in the inertial coordinate system; Indicates the first i The bow angle of the unmanned observation vessel; They represent the first i The longitudinal and lateral velocities of the unmanned observation vessel in the attached coordinate system; Indicates the forward angular velocity; It represents the first-order differential.

[0062] Furthermore, the dimensional observation state space is represented as follows:

[0063]

[0064] in: For the first i The state space vector of the observation unmanned surface vessel; For the first i The state vector of the unmanned observation vessel; For the first i The target vessel state vector that can be observed by a single unmanned observation vessel; For the first i The state vector of the first cooperative observation vessel among the observation unmanned vessels; For the first i The state vector of the second cooperative observation vessel among the observation unmanned vessels;

[0065] in,

[0066]

[0067] In the formula: W This represents the side length of the square mission area; This indicates the maximum permissible linear velocity of the unmanned surface vessel; This indicates the maximum permissible heading angular velocity of the unmanned surface vessel; These are the cosine and sine values ​​of the heading angle, respectively.

[0068]

[0069] In the formula: They represent the first i An observation unmanned surface vessel and a target vessel were in Relative distance along the axis; Indicates the first i The target vessel that the observation unmanned surface vessel can observe is the first i The Euclidean distance between the observation unmanned vessels; Indicates the first i The target vessel that the observation unmanned surface vessel can observe is relative to the first one. i The relative azimuth angle of the observation unmanned surface vessel.

[0070] Beneficial Effects: This invention provides a control method for observational unmanned surface vessels (USVs) based on risk sensitivity and physical potential field constraints. It establishes a composite artificial potential field model with smooth transition characteristics, and then establishes a dense reward function based on the potential energy difference of the artificial potential field. Based on the reward function, a joint loss function is established. Then, using a trained Actor network, the strategy flexibility force of the observation USV is obtained and modified according to the dimensionless observation state space and the dense reward function based on the potential energy difference of the artificial potential field, resulting in a corrected total control force. This allows the acquisition of the expected velocity and expected heading angle of the observation USV. Combined with the underactuated USV dynamics model, this enables control of the USV based on cooperative observation. This invention addresses the challenge of long-term continuous observation, tracking, and dynamic collision avoidance of a single highly dynamic non-cooperative target by a fixed formation of unmanned surface vessels (USVs) in complex and dynamic marine environments, significantly improving the safety of observation missions. The invention utilizes a composite artificial potential field model with smooth transition characteristics, ensuring that prediction accuracy is not reduced by wind and waves during model predictive control. Furthermore, the invention employs a dense reward function based on the potential energy difference of the artificial potential field to establish a joint loss function, solving the problem of neglecting the low-probability but serious long-tail risk in existing technologies. This reduces the probability of observation equipment being damaged by collisions and also results in faster training convergence. Attached Figure Description

[0071] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0072] Figure 1 This is a flowchart of the observation unmanned surface vessel control method based on risk sensitivity and physical potential field constraints of the present invention;

[0073] Figure 2 This is a diagram of the overall training and control framework in an embodiment of the present invention;

[0074] Figure 3 This is a schematic diagram of the underlying security constraints and action correction mechanism in an embodiment of the present invention;

[0075] Figure 4 This is a simulation trajectory effect diagram in the collaborative observation task according to an embodiment of the present invention. Detailed Implementation

[0076] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0077] This embodiment introduces a control method for an observational unmanned surface vessel based on risk sensitivity and physical potential field constraints, including the following steps: Figure 1 As shown:

[0078] S1: Construct an underactuated unmanned surface vessel dynamics model for a collaborative observation scenario.

[0079] Specifically, an underactuated unmanned surface vessel (USV) dynamic model is established in a simulation environment, including the kinematic equations of three observation USVs and one target USV, considering hydrodynamic damping and the added mass matrix. To adapt to the input requirements of the neural network and eliminate training instability caused by dimensional differences, all input state variables are normalized.

[0080] Specifically, the scenario set in this embodiment includes observation unmanned surface vessel and The non-cooperative target vessel, the first i A three-degree-of-freedom underactuated kinematic model of an observation unmanned surface vessel:

[0081] (1)

[0082] in: Indicates the number of the unmanned surface vessel being observed. They represent the first i The longitudinal and lateral position coordinates of the unmanned observation vessel in the inertial coordinate system; Indicates the first i The bow angle of the unmanned observation vessel; They represent the first i The longitudinal and lateral velocities of the unmanned observation vessel in the attached coordinate system; Indicates the forward angular velocity; It represents the first-order differential.

[0083] S2: Based on the underactuated unmanned surface vessel dynamics model, construct a dimensionless observation state space to obtain the state space vector of the observation unmanned surface vessel;

[0084] Specifically, to eliminate training instability caused by different physical dimensions, all input state variables are physically normalized to construct a dimensionless observation state space. For unmanned surface vessel (USV) formations, observation vectors are constructed using a direct concatenation method. These observation vectors include normalized self-state, target relative state, and teammate relative state, and angle information is processed using trigonometric function encoding to avoid numerical discontinuities caused by angle periodicity. For example, position coordinates... Divide by the side length of the mission area W ,speed Divide by maximum linear velocity The observation vector is constructed using a direct concatenation method. .

[0085] (2)

[0086] in: For the first i The state space vector of the observation unmanned surface vessel; For the first i The state vector of the unmanned observation vessel; For the first i The target vessel state vector that can be observed by a single unmanned observation vessel; For the first i The state vector of the first cooperative observation vessel among the observation unmanned vessels; For the first i The state vector of the second cooperative observation vessel among the observation unmanned vessels;

[0087] Among them, one's own state These are the normalized sine / cosine values ​​of position, velocity, and heading angle, expressed in detail as follows:

[0088] (3)

[0089] In the formula: W This represents the side length of the square mission area; This indicates the maximum permissible linear velocity of the unmanned surface vessel; This indicates the maximum permissible heading angular velocity of the unmanned surface vessel; These are the cosine and sine values ​​of the heading angle, respectively.

[0090] No. i The target vessel state vector that can be observed by an unmanned observation vessel The expression is:

[0091] (4)

[0092] In the formula: They represent the first i An observation unmanned surface vessel and a target vessel were in Relative distance along the axis; Indicates the first i The target vessel that the observation unmanned surface vessel can observe is the first i The Euclidean distance between the observation unmanned vessels; Indicates the first i The target vessel that the observation unmanned surface vessel can observe is relative to the first one. i The relative azimuth of the observation unmanned surface vessel;

[0093] Specifically, the relative condition of teammates Includes relative status information of the two cooperative observation vessels, defined in the same way. .

[0094] In this embodiment, the reward function is defined as follows: , indicating at time step t Immediate rewards for the agent. Actions. It is the intelligent agent based on observation The output control commands have a continuous action space. State This refers to the global or local state of an intelligent agent in the environment, including information such as position, velocity, and azimuth.

[0095] S3: Construct a composite artificial potential field model with smooth transition characteristics to obtain the total potential energy of the observed unmanned surface vessel and establish a dense reward function based on the potential energy difference of the artificial potential field.

[0096] Specifically, a composite physical field is constructed, comprising the gravitational field of the target vessel, the repulsive field of the cooperative observation vessel, the repulsive field of the obstacle, and the collision avoidance potential field for the dynamic target vessel. A smooth transition repulsive force function based on a high-order polynomial is designed to ensure the continuity of the repulsive force and its gradient between the danger radius and the warning radius, providing a smooth physical prior for subsequent reward calculations and safety corrections.

[0097] Specifically, the unmanned surface vessel (USV) is subjected to various virtual forces during its movement. The actions output by the policy network generate additional forces. The combined force of the force and the environmental potential field propels the unmanned surface vessel to the next position. Motion. Define the total potential energy function. The composite artificial potential field model with smooth transition characteristics is constructed as follows:

[0098] (5)

[0099] In the formula: P This represents the current position vector of the observed unmanned surface vessel; Indicates the observed gravitational potential field; It is a repulsive potential field in the environment; The total potential energy function;

[0100] in,

[0101] (6)

[0102] In the formula: d This indicates the real-time distance between the observed unmanned surface vessel (USV) and the target vessel. This represents the gravitational gain coefficient, used to adjust the magnitude of gravity. This indicates the optimal orbital observation radius set for the mission.

[0103] Specifically, set a danger radius for obstacles and teammates. r and warning radius R Traditional artificial potential field repulsive function Often in d=R There is a truncation at this point, resulting in a discontinuity in the potential field force. This embodiment aims to ensure the unmanned surface vessel can operate within the warning radius area. R To ensure the smoothness of the dynamics, a higher-order potential field satisfying specific boundary conditions was constructed. For ease of engineering implementation, the function of the repulsive force amplitude of this potential field with respect to distance is directly given. Its form is:

[0104] (7)

[0105] In the formula: Indicates the distance between the unmanned surface vessel and obstacles or teammates. This represents the maximum saturation repulsive force, which is a negative constant, indicating an extremely strong repulsive effect. These are polynomial adjustment coefficients used to control the rate at which the repulsive force grows; R The warning radius is defined as the distance greater than a certain value. R The repulsive force is 0 at that time; r The danger radius is defined as the distance less than [a certain value]. r The repulsive force reaches its maximum at this time;

[0106] Specifically, functions Corresponding to =R The potential field with zero potential energy and zero gradient ensures that the unmanned surface vessel experiences continuous force the instant it enters the warning zone.

[0107] Specifically, after establishing the composite artificial potential field model, the total potential energy of the observed unmanned surface vessel can be obtained, and then a dense reward function based on the potential energy difference of the artificial potential field can be established.

[0108] Specifically, to address the sparse reward problem, this embodiment introduces a potential energy-based reward reshaping mechanism. The potential energy difference between the unmanned surface vessel (USV) and other time points within the composite artificial potential field is calculated and multiplied by a discount factor, serving as a dense auxiliary reward signal. This mechanism mathematically guarantees the invariance of the optimal policy while guiding the agent to quickly learn risk-avoidance behavior. To balance task completion and training efficiency, a total reward function of the following form is established. :

[0109] (8)

[0110] In the formula: For the total reward function; Rewards for sparse terminals; Rewards are given based on potential energy. Punishment for security risks;

[0111] in,

[0112]

[0113] In the formula: Positive reward; Negative rewards;

[0114] Specifically, sparse terminal rewards A positive reward will be given when the observation unmanned surface vessel successfully forms an encirclement observation configuration around the target vessel. A negative reward is given when the unmanned surface vessel collides (with an obstacle or a teammate) or the target escapes. This ensures that the observation unmanned surface vessels (USVs) clearly define the ultimate objective of their mission. The encirclement observation configuration refers to the distribution of each observation USV on the desired observation circumference centered on the target USV, and the azimuth differences between adjacent observation USVs relative to the target satisfy a specific angular interval constraint (such as three observation USVs being evenly distributed at 120 degrees to each other in a formation).

[0115] Specifically, to address the sparse reward problem in the complex trajectory generation process, this embodiment employs a composite artificial potential field model with smooth transition characteristics for reward reshaping. The potential energy-guided reward at time step t is defined. :

[0116] (9)

[0117] In the formula: Indicates the first i An observation unmanned surface vessel in time step t The potential energy guidance reward obtained; For the first i An observation unmanned surface vessel in time step t The state of time; The discount factor is used in reinforcement learning, and its value ranges from (0, 1]. Indicates the first i An observation unmanned surface vessel in time step t The total potential energy at position +1; Indicates the first i An observation unmanned surface vessel in time step t The total potential energy at the location;

[0118] Specifically, when the unmanned surface vessel moves in the direction of decreasing potential energy, Thus making This provides positive incentives for the intelligent agent.

[0119] Security risk penalties When the distance between the unmanned surface vessel and an obstacle or a teammate is less than a safe threshold, causing a sharp increase in the underlying physical repulsion force, an additional negative penalty is imposed.

[0120] S4: Establish a loss function based on the dense reward function of the potential energy difference of the artificial potential field; then train the Critic network and Actor network in the risk-sensitive multi-agent proximal policy optimization algorithm architecture based on the loss function.

[0121] Specifically, the dense reward function based on the potential energy difference of the artificial potential field provides a non-zero feedback signal for each time step, enabling the observation unmanned surface vessel to obtain a discriminative cumulative reward value for each training sampling trajectory it generates, even in the initial stage before the mission is completed. These differences This constitutes a set of reward distributions that can effectively reflect the systemic risk situation, thus providing the necessary statistical sample basis for calculating quantiles (VaR) and conditional value at risk (CVaR). Without this dense reward, sparse rewards would result in most training sample trajectories having zero returns, making it impossible to calculate an effective risk measure.

[0122] Specifically, this embodiment employs a risk-sensitive multi-agent proximal policy optimization algorithm architecture based on conditional value of risk (VAT) for solving the problem. VAT is introduced during the centralized training phase to measure risk, and the reward distribution is calculated... - Quantiles identify the worst-performing tail trajectories (i.e., samples that collided or lost their targets). A reward-mixing mechanism is employed, incorporating conditional risk value into the advantage function calculation, forcing the policy network to prioritize optimization in extreme risk scenarios and improving the system's robustness.

[0123] Preferably, the conditional value at risk is constructed as follows:

[0124] Specifically, this embodiment introduces a conditional value-at-risk (VAT) metric within a framework of centralized training and distributed execution. This allows the strategy to explicitly focus on high-risk trajectories at the tail of the distribution while optimizing expected returns, thereby improving the system's security and robustness under extreme conditions. In each policy update, the environment is based on the current policy. Generate a batch of lengths T Multi-agent joint trajectory set , of which k Training sampling trajectory The cumulative return on discounts is defined as:

[0125] S41: Obtain the discounted cumulative reward of the training sampled trajectory of the observed unmanned surface vessel:

[0126] (10)

[0127] In the formula: For the first k The discounted cumulative reward of the training sampling trajectory of the observation unmanned surface vessel; For the index of the time step; This represents the total number of time steps. Discount factor; The index number is used to identify the training sampling trajectory of the unmanned surface vessel. The total number of training sampling trajectories observed from unmanned surface vessels; For the first k The training sampling trajectory of the observation unmanned surface vessel was in the first... t Instant rewards for time steps;

[0128] Will Sort by size from smallest to largest, and record the confidence level. The Value at Risk (VaR) is as follows:

[0129] S42: Obtain the risk threshold by accumulating discounted rewards based on the training sampled trajectories of the observed unmanned surface vessel.

[0130] (11)

[0131] In the formula: The set of discounted cumulative rewards for training sample trajectories of unmanned surface vessels; This refers to quantile function operations, which are performed on the input data set. Sort by size from smallest to largest, resulting in the [number]th [rank]. The value at that position, Confidence level; Risk threshold;

[0132] Specifically, In order to make The quantiles, where, The cumulative reward for observing the training sampling trajectory of unmanned surface vessels Less than or equal to the risk threshold The probability of; Confidence level;

[0133] S43: The conditional value at risk is as follows:

[0134] Conditional Value at Risk (CVaR) is defined as the conditional expected value of the following tail returns:

[0135] (12)

[0136] In the formula: Conditional risk value; For expectations;

[0137] Specifically, due to directly optimizing the expected value form Since calculating gradients is difficult, this embodiment transforms it into a sample-based weighted optimization problem. This is achieved by identifying those factors that lead to... By reducing the value of tail risk samples and assigning them higher training weights, the Critic network loss function can explicitly increase the conditional risk value during gradient descent.

[0138] Preferably, the joint loss function includes the loss function of the Critic network and the loss function of the Actor network;

[0139] The loss function of the Critic network is expressed as follows:

[0140] In the update of the centralized Critic network, this embodiment uses risk-sensitive weights derived from Conditional Value at Risk (CVaR) to weight the mean squared error, constructing a weighted mean squared error loss function, as follows:

[0141] (13)

[0142] In the formula: The loss function; The parameter is The state value function estimated by the centralized Critic network; This is the parameter set for the Critic network; For the first The training sampling trajectory of the observation unmanned surface vessel at the time step t The combined state; For the first k The training sampling trajectory of the observation unmanned surface vessel at the time step t Discount return estimate;

[0143] Specifically, to ensure that the policy network focuses on observing the tail trajectory of the unmanned surface vessel, which carries a higher risk, this embodiment designs a weight factor for each training sampling trajectory that is monotonically correlated with the reward magnitude. Its form is:

[0144]

[0145]

[0146] In the formula: The amount calculated based on the gap between the current return and the conditional value at risk, i.e., the first... k The training sampling trajectory of the observation unmanned surface vessel led to The magnitude of the decrease; For the first k Sample weights for the training sampling trajectories of the observation unmanned surface vessel; As the risk amplification coefficient, ; This is the absolute value of the risk threshold; This is a minimal constant used to prevent the denominator from being zero.

[0147] Specifically, by introducing The Critic network will provide higher fitting accuracy for low-reward trajectories at the tail when fitting the value function. Weights The introduction of weighted loss functions in Critic and Actor networks is achieved through this approach. In the Critic network, The fitting accuracy of the reward was adjusted; and in the Actor network, the weighting operation in the PPO loss function helps the policy focus on high-risk tail trajectories.

[0148] Specifically, the policy network takes the state space vector of the observed unmanned surface vessel as input and outputs action commands in a continuous action space. In this embodiment, the action Defined as a two-dimensional continuous vector The physical meaning of is the normalized coefficient of the strategic flexibility force of the unmanned surface vessel in the longitudinal and lateral directions, representing the active exploration intention of the intelligent agent beyond the guidance of the physical potential field.

[0149] In this centralized training architecture, this embodiment aims to simultaneously improve the sensitivity of the Critic network and the Actor network to tail risk. Therefore, for the same training sampling trajectory identified as high-risk, the Critic network is required to have higher weights. To accurately fit its extremely low value, while also requiring the Actor network to use the same weights. The focus is on learning and avoiding actions that lead to this risk. This embodiment introduces weights into the PPO loss of the standard MAPPO. Let the first... i An observation unmanned surface vessel was on the training sampling trajectory. The t The advantage function estimate for the step is: The weighted PPO objective is then written as:

[0150] (14)

[0152] In the formula: The objective function for PPO; The parameter set of the Actor network; For the first i The expectation of an unmanned observation vessel; The probability ratio; The dominant function; This is a truncation function; The PPO cutting factor;

[0153] In actual implementation, the training sampling trajectory obtained from environmental acquisition is first stored in the experience buffer. The advantage function estimate of each step is calculated uniformly through the buffer, and then the parameters of the Critic and Actor networks are updated according to equations (13) and (14).

[0154] S5: Based on the dimensionless observation state space and the dense reward function based on the potential energy difference of the artificial potential field, the trained Actor network is used to obtain the policy flexibility force of the observation unmanned surface vessel.

[0155] In this embodiment, it is assumed that at time step t The unmanned surface vessel was located at the following position. The heading angle is The surrounding area contains static obstacles and intruding targets. To prevent the policy network from outputting dangerous actions that could lead to collisions during the exploration process, this embodiment introduces a physical vector synthesis mechanism in the underlying controller to obtain the action commands output by the Actor policy network based on the current observations, representing the exploration intentions of the unmanned surface vessel.

[0156] S6: Establish a low-level safety constraint and action correction mechanism based on physical vector synthesis to obtain the corrected total control force according to the strategic flexibility force of the observed unmanned surface vessel, so as to obtain the expected speed and expected heading angle of the observed unmanned surface vessel, and then combine it with the underactuated unmanned surface vessel dynamics model to realize the control of the unmanned surface vessel based on cooperative observation.

[0157] Specifically, at the execution end, the flexible motion vector output by the strategy network is dynamically synthesized with the physical force vector generated by the environmental potential field to generate the final control command, ensuring that the unmanned surface vessel is forcibly protected by the physical field when approaching obstacles.

[0158] Define environmental strong repulsion That is, the repulsive potential field of static obstacles in the environment The generated physical repulsive force is directed strictly towards the outward normal direction, away from the nearest obstacle, and its amplitude increases exponentially with decreasing distance. Define the task interaction force. That is, the gravitational potential field of the target, which is formed by the gravitational potential field of the target. It generates a signal that is always directed toward the target, used to maintain continuous tracking and monitoring of the target.

[0159] The vector sum of all the above forces is calculated in real time to obtain the corrected total control force. for:

[0160] (15)

[0161] In the formula: The revised total control force; For strategic flexibility; It has a strong repulsive force on the environment; For task interaction; For strategic flexibility The scaling factor;

[0162] Among them, overall control The angle between the unmanned surface vessel and its current heading is defined as This is used to define the required steering angle.

[0163] When the unmanned surface vessel is in a safe area, the strong repulsive force of the environment ,at this time The system is primarily driven by neural networks, demonstrating its intelligence. When the unmanned surface vessel approaches an obstacle or a teammate, The amplitude will increase dramatically, far exceeding the output of the neural network. Upper limit. According to the vector composition principle, the strong repulsive force of the environment. It will forcibly change the combined force The direction was steered away from the danger zone, and the enormous reverse force generated by the physical field forcibly took over at the bottom layer, ensuring... Always safe.

[0164] The revised total control force Decompose into desired speed and expected heading angle :

[0165] (16)

[0166] In the formula: For the desired speed; Desired heading angle; For modulo length calculation; This is the maximum permissible linear velocity for the unmanned surface vessel. All are components of the resultant force in the inertial coordinate system;

[0167] Specifically, instructions The underactuated kinematic model described in Equation (1) is sent to be executed, thereby at the next time step. t +1 Reach a safe location.

[0168] Specifically, this embodiment is based on closed-loop collaborative observation and execution using a distributed policy network and a security correction layer. The trained and converged policy network is deployed to the observation unmanned surface vessel (USV). Each USV independently senses the environment, generates policy actions, and outputs them to the underlying propulsion system after verification by the local security layer, thereby achieving efficient and safe collaborative observation of dynamic targets.

[0169] This embodiment simulates a realistic distributed control process in a multi-agent simulation environment. After the training phase, the system switches to execution mode. At this point, each observation unmanned surface vessel no longer relies on a centralized Critic network or global state information, but instead makes independent decisions using only the Actor network. The specific closed-loop control process is as follows:

[0170] 1. Extract the parameters of the converged Actor network and load them into the independent decision-making modules representing each unmanned surface vessel. At this point, there is no gradient backpropagation between the agents, and they operate completely independently.

[0171] 2. The simulation environment mimics the functions of airborne sensors, generating local observation vectors for each unmanned surface vessel. .

[0172] 3. Each unmanned surface vessel will transmit its local observation vector. Input the local Actor network. The network performs forward propagation and outputs the original action commands, i.e., the policy flexible force vector. .

[0173] 4. Before executing an action, the underlying control logic calculates the current physical field constraints in parallel. The vector synthesis mechanism described in S6 is applied to calculate... .

[0174] 5. The corrected resultant force The desired acceleration or velocity command is mapped and input into the underactuated kinematic differential equation of the unmanned surface vessel (USV) for integral solution, thereby updating the USV's position and attitude in the next simulation step.

[0175] like Figure 2 As shown, the overall training and control framework of this embodiment includes two parts: a centralized training module and a distributed execution module. During the training phase, the multi-agent simulation environment contains three observation unmanned surface vessels (USVs) and one non-cooperative target vessel, evolving according to the kinematic and observation models in formulas (1) to (4). At each decision step, the Actor network of each USV, based on local observations... Output Action The environment is then updated to the next state, and a reward is returned. ,form Trajectory fragments of the form are stored in the sampling buffer. When a batch of fragments of length accumulating in the buffer... T After the multi-agent joint trajectory is calculated, the trajectory is fed into a centralized Critic network based on CVaR. The Critic estimates the value function on the joint state. And calculate the discounted reward for each trajectory. VaR and CVaR are used to obtain the weights. Subsequently, a weighted mean squared error update is performed on the Critic network, and a CVaR-based weighted PPO update is performed on each Actor network, completing one policy iteration. During the execution phase, Figure 2 The distributed control process is illustrated below. After training convergence, the Actor networks corresponding to the three unmanned surface vessels (USVs) are deployed to their respective onboard computing units, with each USV relying solely on local observations. Forward reasoning yields flexible motion vectors Subsequently, this action, along with the virtual force calculated from the composite artificial potential field, undergoes vector synthesis and command mapping in the safety correction layer, outputting a safety control command. This drives the underlying dynamics model, enabling safe, collaborative, and continuous observation of the target.

[0176] like Figure 3 The diagram illustrates the underlying safety constraints and motion correction mechanism of this embodiment. The obstacle areas in the diagram are shown in gray, representing the restricted areas of the unmanned surface vessel (USV) in the environment. The target location is located at the bottom of the diagram; this is the area the USV needs to continuously track and approach. In this diagram, the arrow points from the current position... Pointing to the position of the next moment , Represents the gravitational pull of the target, used to guide the unmanned surface vessel toward its target location. This represents the repulsive force from obstacles, which is calculated based on the distance between the unmanned surface vessel (USV) and the obstacle, preventing the USV from entering the obstacle area. This represents the flexibility of the strategy, which is the action instructions output by the policy network based on the current observations, representing the agent's willingness to explore. It is a combined force, it is composed of , and The combined effect ultimately determines the expected speed and direction of the unmanned surface vessel.

[0177] Appendix Figure 4 This is an example simulation using Python. The mission scenario includes three unmanned surface vessels (USVs), a static obstacle, and one target vessel. The USVs track the target's path and dynamically adjust according to changes in the environment. Different colors in the figure represent different USV trajectories; the gray trajectory is the defense vessel's trajectory, and the black trajectory is the intruding target's trajectory, showing the relative motion of the three USVs performing a cooperative observation mission. As can be seen in the figure, the trajectories of the three USVs sometimes approach each other and sometimes slightly separate, reflecting their cooperative behavior when performing dynamic collision avoidance and target tracking tasks.

[0178] This embodiment constructs a composite artificial potential field incorporating both target attraction and obstacle repulsion. A dense reward guidance mechanism is designed using potential energy difference to address the cold start problem in training under sparse rewards. During the policy training phase, a conditional risk value algorithm is introduced to improve the multi-agent proximal policy optimization algorithm, enhancing policy robustness by weighting long-tail risk trajectories. In the distributed execution phase, a low-level safety correction mechanism based on physical vector synthesis is designed, dynamically synthesizing the flexible action vectors output by the policy network with the physical force vectors generated by the environmental potential field. This embodiment effectively solves the problems of slow training convergence and lack of security in collaborative observation tasks under complex dynamic environments through the dual protection of algorithmic-level risk aversion and physical-level hard constraints. It is highly feasible and applicable to distributed systems such as unmanned surface vessels and swarms.

[0179] In this embodiment, a composite artificial potential field model with smooth transition characteristics is established, and a dense reward function based on the potential energy difference of the artificial potential field is then established. A joint loss function is then established based on the reward function. A trained Actor network is then used to obtain and modify the policy flexibility force of the observed unmanned surface vessel (USV) based on the dimensionless observation state space and the dense reward function based on the potential energy difference of the artificial potential field, resulting in a corrected total control force. The desired velocity and desired heading angle of the USV are then obtained. Combined with the underactuated USV dynamics model, control of the USV based on cooperative observation is achieved. This invention addresses the long-term continuous observation, tracking, and dynamic collision avoidance of a single highly dynamic non-cooperative target by a fixed formation of unmanned surface vessels (USVs) in complex and dynamic marine environments, significantly improving the safety of observation missions. By establishing a composite artificial potential field model with smooth transition characteristics, the invention ensures that prediction accuracy is not reduced due to wind and wave disturbances during model predictive control. Furthermore, it employs a dense reward function based on the potential energy difference of the artificial potential field to establish a joint loss function, solving the problem of neglecting the low-probability but serious long-tail risk in existing technologies. This reduces the probability of observation equipment being damaged by collisions and also accelerates training convergence. The invention also addresses the cold start problem in training under sparse rewards by utilizing a potential energy-based reward shaping mechanism to transform the physical prior of the artificial potential field into a dense reward signal for each step, guiding the agent to quickly overcome the blind exploration period and reducing the number of training samples required for policy convergence.

[0180] This embodiment achieves dual safety protection by introducing a risk aversion mechanism at the strategy layer and combining it with a physical vector synthesis-based underlying action correction mechanism to provide physical hard constraints at the execution layer, effectively preventing collision damage to unmanned surface vessels in complex sea conditions during mission execution.

[0181] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A control method for an observational unmanned surface vessel based on risk sensitivity and physical potential field constraints, characterized in that, Includes the following steps: S1: Construct a dynamic model of an underactuated unmanned surface vessel for a collaborative observation scenario; S2: Based on the underactuated unmanned surface vessel dynamics model, construct a dimensionless observation state space to obtain the state space vector of the observation unmanned surface vessel; S3: Construct a composite artificial potential field model with smooth transition characteristics to obtain the total potential energy of the observed unmanned surface vessel and establish a dense reward function based on the potential energy difference of the artificial potential field. S4: Establish a joint loss function based on the dense reward function of the potential energy difference of the artificial potential field; then, train the Critic network and Actor network in the risk-sensitive multi-agent proximal policy optimization algorithm architecture based on the loss function. S5: Based on the dimensionless observation state space and the dense reward function based on the potential energy difference of the artificial potential field, the trained Actor network is used to obtain the policy flexibility force of the observation unmanned surface vessel. S6: Based on the strategic flexibility of the observed unmanned surface vessel (USV), obtain the corrected total control force to obtain the expected speed and expected heading angle of the USV. Then, combine the underactuated USV dynamics model to realize the control of the USV based on cooperative observation.

2. The control method for an observation unmanned surface vessel based on risk sensitivity and physical potential field constraints according to claim 1, characterized in that, The composite artificial potential field model with smooth transition characteristics is constructed as follows: In the formula: This represents the current position vector of the observed unmanned surface vessel; Indicates the observed gravitational potential field; It is a repulsive potential field in the environment; The total potential energy function; in, In the formula: This indicates the real-time distance between the observed unmanned surface vessel (USV) and the target vessel. This represents the gravitational gain coefficient, used to adjust the magnitude of gravity. This indicates the optimal orbital observation radius set for the mission; In the formula: Indicates the distance between the unmanned surface vessel and obstacles or teammates. This represents the maximum saturation repulsion value; For polynomial adjustment coefficients; R Warning radius; r This is the danger radius.

3. The control method for an observation unmanned surface vessel based on risk sensitivity and physical potential field constraints according to claim 1, characterized in that, The dense reward function based on the potential energy difference of the artificial potential field is established as follows: In the formula: For the total reward function; Rewards for sparse terminals; Rewards are given based on potential energy. Punishment for security risks; in, In the formula: Positive reward; Negative rewards; In the formula: Indicates the first i An observation unmanned surface vessel in time step t The potential energy guidance reward obtained; For the first i An observation unmanned surface vessel in time step t The state of time; To reinforce the discount factor in learning; Indicates the first i An observation unmanned surface vessel in time step t The total potential energy at position +1; Indicates the first i An observation unmanned surface vessel in time step t The total potential energy at the location.

4. The control method for an observation unmanned surface vessel based on risk sensitivity and physical potential field constraints according to claim 1, characterized in that, The conditional value at risk is constructed as follows: S41: Obtain the discounted cumulative reward of the training sampled trajectory of the observed unmanned surface vessel, using the following formula: In the formula: For the first k The discounted cumulative reward of the training sampling trajectory of the observation unmanned surface vessel; For the index of the time step; This represents the total number of time steps. Discount factor; The index number is used to identify the training sampling trajectory of the unmanned surface vessel. The total number of training sampling trajectories observed from unmanned surface vessels; For the first k The training sampling trajectory of the observation unmanned surface vessel was in the first... t Instant rewards for time steps; S42: Based on the discounted cumulative reward of the observed unmanned surface vessel's training sampling trajectory, the risk threshold is obtained as follows: In the formula: The set of discounted cumulative rewards for training sample trajectories of unmanned surface vessels; This refers to quantile function operations; Confidence level; Risk threshold; S43: The conditional value of risk is as follows: In the formula: Conditional risk value; As expected.

5. The control method for an observation unmanned surface vessel based on risk sensitivity and physical potential field constraints according to claim 4, characterized in that, The joint loss function includes the loss function of the Critic network and the loss function of the Actor network; The loss function of the Critic network is expressed as follows: In the formula: The loss function; The parameter is The state value function estimated by the centralized Critic network; This is the parameter set for the Critic network; For the first The training sampling trajectory of the observation unmanned surface vessel at the time step t The combined state; For the first k The training sampling trajectory of the observation unmanned surface vessel at the time step t Discount return estimate; in, In the formula: The amount calculated based on the gap between the current return and the conditional value at risk, i.e., the first... k The magnitude of the reduction in conditional risk value caused by the training sampling trajectory of the observation unmanned surface vessel; For the first k Sample weights for the training sampling trajectories of the observation unmanned surface vessel; As the risk amplification coefficient, ; This is the absolute value of the risk threshold; This is expressed as a minimal constant to prevent the denominator from being zero; The loss function of the Actor network is expressed as follows: In the formula: The objective function is PPO, which is the loss function value of the Actor network; The parameter set of the Actor network; For the first i The expectation of an unmanned observation vessel; The probability ratio; The dominant function; This is a truncation function; This is the PPO trimming factor.

6. The control method for an observation unmanned surface vessel based on risk sensitivity and physical potential field constraints according to claim 4, characterized in that, The formula used to obtain the corrected total control force is as follows: In the formula: The revised total control force; For strategic flexibility; It has a strong repulsive force on the environment; For task interaction; For strategic flexibility The scaling factor; The formulas used to obtain the desired velocity and desired heading angle of the observed unmanned surface vessel are as follows: In the formula: For the desired speed; Desired heading angle; For modulo length calculation; This is the maximum permissible linear velocity for the unmanned surface vessel. All of these are components of the resultant force in the inertial coordinate system.

7. The control method for an observation unmanned surface vessel based on risk sensitivity and physical potential field constraints according to claim 1, characterized in that, The dynamic model of the underactuated unmanned surface vessel is constructed as follows: in: Indicates the number of the unmanned surface vessel being observed. They represent the first i The longitudinal and lateral position coordinates of the unmanned observation vessel in the inertial coordinate system; Indicates the first i The bow angle of the unmanned observation vessel; They represent the first i The longitudinal and lateral velocities of the unmanned observation vessel in the attached coordinate system; Indicates the forward angular velocity; It represents the first-order differential.

8. The control method for an observation unmanned surface vessel based on risk sensitivity and physical potential field constraints according to claim 7, characterized in that, The dimensional observation state space is represented as follows: in: For the first i The state space vector of the observation unmanned surface vessel; For the first i The state vector of the unmanned observation vessel; For the first i The target vessel state vector that can be observed by a single unmanned observation vessel; For the first i The state vector of the first cooperative observation vessel among the observation unmanned vessels; For the first i The state vector of the second cooperative observation vessel among the observation unmanned vessels; in, In the formula: W This represents the side length of the square mission area; This indicates the maximum permissible linear velocity of the unmanned surface vessel; This indicates the maximum permissible heading angular velocity of the unmanned surface vessel; These are the cosine and sine values ​​of the heading angle, respectively. In the formula: They represent the first i An observation unmanned surface vessel and a target vessel were in Relative distance along the axis; Indicates the first i The target vessel that the observation unmanned surface vessel can observe is the first i The Euclidean distance between the observation unmanned vessels; Indicates the first i The target vessel that the observation unmanned surface vessel can observe is relative to the first one. i The relative azimuth angle of the observation unmanned surface vessel.

Citation Information

Patent Citations

  • Unmanned ship trajectory tracking control method based on Actor-Credit-Advantage network

    CN115793455A

  • Local path planning method and device for unmanned surface vehicle

    CN117519197A