Unmanned ship fleet coordination strategy virtual-real migration method and system based on basin digital twinning
Patent Information
- Application Number
- CN202611318576.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-28
- Publication Date
- 2026-09-29
AI Technical Summary
[0005]本申请提供一种基于流域数字孪生的无人船群协同策略虚实迁移方法及系统,解决现有技术流域无人船群协同策略在虚实迁移过程中,因仿真环境保真度不足、虚实差异不可量化、模型偏差无法闭环修正以及部署缺乏安全机制,导致策略难以从仿真环境可靠迁移至实体船群的问题
本申请通过将训练所得策略下装至试验用实体无人船群进行小样本试运行,并与流域孪生环境同工况回放比对计算虚实差异度,建立了可量化的虚实差距评估手段,使策略迁移前的性能衰减程度可计算、可判定;通过当虚实差异度超阈值时反演修正孪生参数并回写后返回继续训练更新策略,形成了模型偏差的闭环修正机制,使孪生环境能够随实测数据持续对齐,避免虚实偏差累积;通过以连续多个评估回合的虚实差异度与任务成功率双条件作为策略正式迁移部署的判定依据,避免了单回合偶然通过导致的误判,实现了策略可部署性的可靠判定;通过在策略正式迁移部署后对动作施加航速约束、离岸距离约束和最小会遇距离约束,在保留策略决策能力的同时保障了实体无人船群的航行安全,从而建立了从仿真训练到实体部署的可量化、可判定、可回写的完整虚实迁移闭环,使策略部署质量等效于真实环境训练效果。
Smart Images

Figure CN122839862A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of unmanned vessel swarm collaboration technology, and more specifically, to a virtual-to-real migration method and system for unmanned vessel swarm collaboration strategy based on watershed digital twins. Background Technology
[0002] The reliable deployment of unmanned vessel swarm collaborative strategies relies on high-fidelity reproduction of real-world watershed hydrology, pollution diffusion, and vessel dynamics. This is crucial for achieving excellent consistency between virtual and real-world migration and engineering reliability. However, such high-fidelity training currently depends entirely on conducting real-world sea trials and strategy integration in real-world watershed environments. It cannot be directly applied in the early, immature stages of the strategy, otherwise, there are high risks such as collisions and groundings.
[0003] When direct training in a real watershed is not possible, a general simulation environment must be used for strategy training. However, existing simulation environments generally suffer from the following shortcomings: the lack of multiphysics coupling between watershed topography, hydrodynamics, pollution diffusion, and ship dynamics leads to insufficient simulation fidelity; the lack of quantifiable evaluation indicators for virtual-real differences makes it impossible to determine the deployability of strategies before migration; simulation model parameters are fixed once set, making it impossible to perform closed-loop correction of model deviations based on measured data; and the lack of safety constraint mechanisms after strategy deployment to a real fleet makes it difficult to ensure navigation safety. These shortcomings result in a significant performance degradation of strategies trained through simulation when deployed to a real fleet, making it impossible to improve the deployment quality of strategies trained through simulation to a level comparable to that trained in a real environment.
[0004] Therefore, how to construct a high-fidelity river basin twin environment, establish a quantitative assessment mechanism for virtual-real differences and a closed-loop parameter correction mechanism, and deploy safety constraints in a virtual test and training environment without actual ships, so that the migration and deployment quality of the river basin unmanned vessel swarm collaborative strategy from simulation to physical vessel swarm is equivalent to the training effect in the real environment, has become an urgent technical challenge to overcome. Summary of the Invention
[0005] This application provides a virtual-to-real migration method and system for unmanned vessel swarm collaborative strategies based on watershed digital twins. It solves the problem that existing watershed unmanned vessel swarm collaborative strategies are difficult to reliably migrate from the simulation environment to the physical vessel swarm during the virtual-to-real migration process due to insufficient fidelity of the simulation environment, unquantifiable virtual-to-real differences, inability to close-loop correct model deviations, and lack of deployment security mechanisms.
[0006] To solve the above-mentioned technical problems, the technical solution of this application is as follows: Firstly, this application provides a virtual-to-real migration method for unmanned vessel swarm cooperative strategies based on watershed digital twins, including: A watershed twin environment is constructed, which is generated by coupling four layers: topography layer, hydrodynamic layer, pollution diffusion layer and intelligent agent layer. In the watershed twin environment, training scenarios are generated by curriculum-based domain randomization scheduling, and a collaborative strategy for unmanned vessel swarms is trained under the training scenarios using a centralized training and distributed execution framework. The currently trained unmanned vessel swarm cooperative strategy is downloaded to the experimental physical unmanned vessel swarm for small-sample trial operation, and compared with the playback simulation results of the watershed twin environment under the same working conditions to calculate the virtual-real difference degree. When the virtual-real difference exceeds a preset threshold, the twin parameters are inverted and corrected and written back to the watershed twin environment. Training continues to be conducted to update the unmanned vessel swarm cooperative strategy, and the updated unmanned vessel swarm cooperative strategy is re-evaluated. When the virtual-to-real difference is less than the download threshold and the mission success rate is greater than or equal to the success rate threshold for multiple consecutive evaluation rounds, the currently qualified unmanned vessel swarm collaboration strategy will be officially migrated and deployed to the target entity unmanned vessel swarm. When the target entity unmanned vessel swarm executes the unmanned vessel swarm cooperative strategy, speed constraints, offshore distance constraints, and minimum encounter distance constraints are applied to the actions output by the strategy in real time. When the action causes any constraint to be violated, the action is restricted to the boundary value of the corresponding constraint before execution.
[0007] Furthermore, the topographic layer, hydrodynamic layer, pollution diffusion layer, and intelligent agent layer are specifically as follows: The topographic layer is constructed from a digital elevation model (DEM) and measured river cross-section data; the hydrodynamic layer provides an instantaneous velocity field based on the numerical solution of the two-dimensional shallow water equation or its simplified flow field; the pollution diffusion layer calculates pollutant concentration based on the instantaneous velocity field in the form of discrete convolution superposition of continuous and mobile sources; the intelligent agent layer consists of a three-degree-of-freedom dynamic model of the hull, a sensor noise model, and a communication channel model.
[0008] Furthermore, the formula for calculating the pollutant concentration is as follows:
[0009] In the formula, Number the emission sources; The discrete sequence number of the emission release time; This refers to the historical release moment of the pollutants; For the first Each emission source at time The emission rate; The discrete time step; Location of the emission source; The observation time when calculating the concentration; For the first The observation time is at the observation point ( The pollutant concentration at point ( ); D is the diffusion coefficient; The velocity component along the x-direction represents the instantaneous velocity field output by the hydrodynamic layer. This represents the velocity component along the y-direction of the instantaneous velocity field output by the hydrodynamic layer.
[0010] Furthermore, the course-based domain randomization scheduling is implemented using the following formula:
[0011] In the formula, This is the k-th randomization parameter; Nominal value; For training rounds; The maximum randomization amplitude; This is the course time constant.
[0012] Furthermore, the degree of difference between virtual and real is calculated according to the following formula, the specific formula being:
[0013] In the formula, T is the total number of sampling times in a single evaluation round; t is the sampling time number, which takes values from 1 to T. Let t be the position of the unmanned vessel at time t; The position of the virtual ship after playback simulation at time t under the same working conditions in the twin environment of the watershed; Let be the observed concentration of the unmanned surface vessel at time t; The observed concentration of the virtual ship at time t after playback simulation under the same working conditions in the twin environment of the watershed; The total energy consumption per round for the physical unmanned vessel; To determine the total round energy consumption of the virtual ship after replaying the simulation under the same operating conditions in the twin environment of the watershed; As the first normalization benchmark; This serves as the second normalization benchmark. This is the first sensitivity analysis parameter; This is the second sensitivity analysis parameter; This is the third sensitivity analysis parameter.
[0014] Furthermore, the consecutive multiple evaluation rounds refer to K consecutive evaluation rounds; the virtual-to-real difference degree being less than the download threshold and the mission success rate being greater than or equal to the success rate threshold specifically means that the virtual-to-real difference degree is less than the download threshold for all K consecutive evaluation rounds, and the proportion of mission success rounds in the K consecutive evaluation rounds is greater than or equal to the success rate threshold; wherein, in a single evaluation round, when the unmanned vessel swarm completes the location of the pollution emission source within the set round time limit, and none of the vessels violate the safety constraints of the minimum offshore distance and minimum encounter distance in that round, the mission is determined to be successful in that round; otherwise, the mission is determined to be unsuccessful in that round.
[0015] Furthermore, speed constraints, offshore distance constraints, and minimum encounter distance constraints are applied to the actions output in real time by the strategy. When an action causes any constraint to be violated, the action is restricted to the boundary value of the corresponding constraint before execution. Specifically, the speed constraint is that the speed of each unmanned vessel in the unmanned vessel swarm is less than or equal to the upper limit of a first set threshold; the offshore distance constraint is that the offshore distance of each unmanned vessel is greater than or equal to the lower limit of a second set threshold; the minimum encounter distance constraint is that the minimum encounter distance between any unmanned vessel and other unmanned vessels in the unmanned vessel swarm is greater than or equal to the lower limit of a third set threshold. If an action output by the strategy causes any constraint to be violated, the action is restricted to the boundary value of the corresponding constraint before execution.
[0016] Furthermore, the inversion correction twin parameters are specifically as follows: With the objective of minimizing the residual between the actual observation data of the experimental unmanned vessel fleet and the simulated observation data generated by the watershed twin environment playback simulation, parameters to be corrected are inverted. The parameters to be corrected include at least one of diffusion coefficient, propulsion efficiency, and sensor noise variance. The inverted and corrected twin parameters are solved using least squares or Bayesian optimization methods, and an iteration upper limit and convergence tolerance are set.
[0017] Furthermore, the centralized training distributed execution framework includes a feature extraction layer, a policy branch, and a value branch; the output of the feature extraction layer is connected to the inputs of both the policy branch and the value branch; the feature extraction layer includes a first fully connected layer, a first ReLU activation function, a second fully connected layer, and a second ReLU activation function connected in sequence, used to receive the local observation state of a single unmanned vessel and output a feature vector; the policy branch is used to receive the feature vector and output the probability distribution of the unmanned vessel's actions; the value branch is used to receive the feature vector and output a value estimate of the unmanned vessel's state.
[0018] Secondly, this application provides a virtual-to-real migration system for unmanned vessel swarm cooperative strategies based on watershed digital twins, comprising: The twin environment construction module is used to construct a watershed twin environment, which is generated by coupling four layers: topography layer, hydrodynamic layer, pollution diffusion layer and intelligent agent layer. The collaborative strategy training module is used to generate training scenarios in the watershed twin environment according to the curriculum-based domain randomization scheduling, and train the collaborative strategy of the unmanned vessel swarm under the training scenarios using a centralized training and decentralized execution framework. The virtual-real consistency evaluation module is used to download the currently trained unmanned vessel swarm cooperative strategy to the experimental real unmanned vessel swarm for small-sample trial operation, and compare it with the playback simulation results of the watershed twin environment under the same working conditions to calculate the virtual-real difference degree. The twin parameter identification and write-back module is used to invert and correct the twin parameters and write them back to the watershed twin environment when the virtual-real difference is greater than a preset threshold, return to continue training to update the unmanned vessel swarm cooperative strategy, and re-evaluate the updated unmanned vessel swarm cooperative strategy. The strategy determination download module is used to formally migrate and deploy the currently satisfied unmanned vessel swarm collaborative strategy to the target entity unmanned vessel swarm when the virtual-real difference degree is less than the download threshold and the task success rate is greater than or equal to the success rate threshold in multiple consecutive evaluation rounds. The safety constraint execution module is used to apply speed constraints, offshore distance constraints, and minimum encounter distance constraints to the actions output in real time by the unmanned vessel swarm of the target entity when the unmanned vessel swarm executes the unmanned vessel swarm cooperative strategy. When the action causes any constraint to be violated, the action is restricted to the boundary value of the corresponding constraint before execution.
[0019] Compared with the prior art, the beneficial effects of the technical solution of this application are: This application establishes a quantifiable method for evaluating the virtual-real gap by downloading the trained strategy to a swarm of experimental unmanned surface vessels (USVs) for small-sample trial operation and comparing it with the simulated environment of the watershed under the same operating conditions. This allows the performance degradation before strategy migration to be calculated and determined. By inverting and correcting the USV parameters when the virtual-real gap exceeds a threshold and writing them back to continue training and updating the strategy, a closed-loop correction mechanism for model bias is formed, enabling the USV environment to continuously align with the measured data and avoiding the accumulation of virtual-real bias. By using the virtual-real gap and mission success rate of multiple consecutive evaluation rounds as the dual criteria for determining the formal migration and deployment of the strategy, misjudgments caused by accidental passage in a single round are avoided, and the deployability of the strategy is reliably determined. By imposing speed constraints, offshore distance constraints, and minimum encounter distance constraints on the actions after the formal migration and deployment of the strategy, the navigation safety of the USV swarm is ensured while retaining the strategy's decision-making ability. This establishes a complete virtual-real migration closed loop that is quantifiable, determineable, and write-backable from simulation training to physical deployment, making the quality of strategy deployment equivalent to the training effect in the real environment. Attached Figure Description
[0020] Figure 1 A flowchart illustrating a virtual-to-real migration method for an unmanned vessel swarm cooperative strategy based on a watershed digital twin, provided for embodiments of this application; Figure 2 This is a schematic diagram of the structure of a virtual-to-real migration system for an unmanned vessel swarm cooperative strategy based on a watershed digital twin, provided as an embodiment of this application. Detailed Implementation
[0021] To make the objectives and advantages of this application clearer, the application will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining this application and are not intended to limit this application.
[0022] The prefixes such as "first" and "second" used in this application embodiment are merely for distinguishing different descriptive objects and do not limit the position, order, priority, quantity, or content of the described objects. The use of ordinal numbers and other prefixes used to distinguish descriptive objects in this application embodiment does not constitute a restriction on the described objects. The description of the described objects is given in the context of the embodiments, and the use of such prefixes should not constitute unnecessary restrictions.
[0023] Preferred embodiments of this application will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of this application and are not intended to limit the scope of protection of this application.
[0024] Please see Figure 1 This is a flowchart illustrating a virtual-to-real migration method for an unmanned vessel swarm cooperative strategy based on a watershed digital twin, provided in an embodiment of this application. The embodiment of this application provides a virtual-to-real migration method for an unmanned vessel swarm cooperative strategy based on a watershed digital twin, comprising: Step S1: Construct a watershed twin environment, which is generated by coupling four layers: topography layer, hydrodynamic layer, pollution diffusion layer, and intelligent agent layer. Step S2: In the watershed twin environment, training scenarios are generated by curriculum-based domain randomization scheduling, and the unmanned vessel swarm cooperative strategy is trained under the training scenarios using a centralized training and distributed execution framework. Step S3: Download the currently trained unmanned vessel swarm cooperative strategy to the experimental physical unmanned vessel swarm for small-sample trial operation, and compare it with the playback simulation results of the watershed twin environment under the same working conditions to calculate the virtual-real difference degree. Step S4: When the virtual-real difference is greater than a preset threshold, the twin parameters are inverted and corrected and written back to the watershed twin environment. Then, the training continues to update the unmanned vessel swarm cooperative strategy, and the updated unmanned vessel swarm cooperative strategy is re-evaluated. Step S5: When the virtual-real difference is less than the download threshold and the mission success rate is greater than or equal to the success rate threshold for multiple consecutive evaluation rounds, the currently qualified unmanned vessel swarm collaboration strategy is officially migrated and deployed to the target entity unmanned vessel swarm. Step S6: When the target entity unmanned vessel swarm executes the unmanned vessel swarm cooperative strategy, apply speed constraints, offshore distance constraints, and minimum encounter distance constraints to the actions output in real time by the strategy. When the action causes any constraint to be violated, the action is restricted to the boundary value of the corresponding constraint before execution.
[0025] In one embodiment, a method for virtual-real migration of unmanned vessel swarm collaborative strategy based on watershed digital twin is provided. A scheme is established that includes twin environment construction, strategy training, small sample trial operation and virtual-real difference assessment, parameter write-back closed loop, formal migration deployment, and security constraint execution. Taking the collaborative source tracing task of pollution sources in a certain watershed section as an example, a typical navigable watershed section of about 10 km in length is selected to construct a watershed twin environment generated by coupling four layers: topography layer, hydrodynamic layer, pollution diffusion layer, and intelligent agent layer.
[0026] Topographic layer: The three-dimensional topography and shoreline boundary of the river section are constructed using digital elevation model (DEM) and measured river cross-section data to determine the range of navigable waters and the distribution of static obstacles such as shoals and bridge piers, providing a geometric basis for hydrodynamic solutions and obstacle avoidance by the ship.
[0027] Hydrodynamic layer: The flow field of the river section is numerically solved using two-dimensional shallow water equations, and the instantaneous velocity components Ux and Uy of each grid node are output. The assimilation and calibration are performed using measured velocity cross-sectional data to ensure that the twin flow field is consistent with the real flow field.
[0028] Pollution diffusion layer: Based on the instantaneous flow field provided by the hydrodynamic layer, the pollutant concentration is calculated using a discrete convolutional superposition of continuous and moving sources. The formula for calculating the pollutant concentration is:
[0029] In the formula, Number the emission sources; The discrete sequence number of the emission release time; This refers to the historical release moment of the pollutants; For the first Each emission source at time The emission rate; The discrete time step; Location of the emission source; The observation time when calculating the concentration; For the first The observation time is at the observation point ( The pollutant concentration at point ( ); D is the diffusion coefficient; The velocity component along the x-direction represents the instantaneous velocity field output by the hydrodynamic layer. This represents the velocity component along the y-direction of the instantaneous velocity field output by the hydrodynamic layer.
[0030] By analyzing emission rates With source location The settings are configured to access instantaneous sources ( Non-zero only at a single moment), continuous source ( (constant during emission periods) and mobile sources (source location varies) (Changes) Three emission conditions, velocity component is taken from the instantaneous field output of the hydrodynamic layer along the trajectory. By integrating to t, the evolution of plumes with continuous and moving emissions can be characterized in an unsteady flow field.
[0031] It should be noted that the above formula for calculating pollutant concentrations is derived from the superposition principle of the linear convection-diffusion equation: when there is only a single emission source (m=1) and the emission is concentrated at a single moment ( =0 and = ), constant flow velocity (integral term reduced to t and When t), the superposition formula, after simplification term by term, degenerates into the classical two-dimensional instantaneous point source analytical solution, and the two are mathematically consistent; and The dimension of is mass, which is the same as the total mass of pollutants released in the classical solution, making the formula consistent. Therefore, while maintaining consistency with the classical solution, the diffusion layer further supports continuous sources, mobile sources, and unsteady flow fields, significantly expanding the configurable training conditions.
[0032] The intelligent agent layer consists of a three-degree-of-freedom (sway, roll, bow) dynamic model of the hull, a sensor noise model, and a communication channel model. The sensor noise model superimposes zero-mean Gaussian noise on the concentration observation and positioning observation, and the communication channel model imposes time delay and packet loss on the information interaction between multiple ships. The channel parameters correspond to the bandwidth-limited training conditions and are used to reproduce the real perception and communication constraints in the twin environment.
[0033] The collaborative strategy training module generates training scenarios in the constructed watershed twin environment using a course-based domain randomization scheduling method, and trains the collaborative strategy of the unmanned vessel swarm using a centralized training and decentralized execution framework under the training scenarios.
[0034] Curriculum-based randomized scheduling is achieved through the following formula:
[0035] In the formula, This is the k-th randomization parameter; Nominal value; For training rounds; The maximum randomization amplitude; This is the course time constant.
[0036] Training Start =0, meaning training with nominal parameters to ensure convergence stability; as training progresses... Monotonically increasing and gradually saturating to This allows the strategy to learn across a gradually expanding parameter distribution, thereby improving its cross-environment generalization ability. The randomization parameter is determined based on the fluctuation range of the real historical data of each randomization parameter, for example, taking the relative amplitude corresponding to the 5th to 95th percentile of the empirical distribution; The randomization is set according to a certain proportion of the total training rounds, so that randomization approaches saturation in the later stages of training.
[0037] The centralized training and distributed execution framework includes a feature extraction layer, a policy branch, and a value branch. The output of the feature extraction layer is connected to the inputs of both the policy branch and the value branch. The feature extraction layer consists of a first fully connected layer, a first ReLU activation function, a second fully connected layer, and a second ReLU activation function, connected sequentially. This layer receives the local observation state of a single unmanned surface vessel (USV) and outputs a feature vector. The policy branch receives the feature vector and outputs the probability distribution of USV actions. The value branch receives the feature vector and outputs a value estimate of the USV's state. One feasible network structure is as follows: the first layer of the feature extraction layer is a fully connected layer with an input dimension equal to the USV's state dimension and an output dimension of 64, introduced with ReLU activation for non-linearity; the second layer is also a fully connected layer with an input dimension of 64 and an output dimension of 64, again activated with ReLU; subsequently, the policy branch and value branch are separated. The policy branch receives the 64-dimensional output of the feature extraction layer and outputs the probability distribution of actions, while the value branch receives the same 64-dimensional output and outputs a 1-dimensional scalar representing the value estimate of the state. The USV's state can be (…). , , The coordinates of the ship's position and the current observed concentration are used; the action can be the desired course increment and the desired speed increment (Δψ, Δv), and the action output is scaled and mapped to the physical range before execution.
[0038] The virtual-real consistency assessment module works after the strategy training is completed: it downloads the currently trained unmanned vessel swarm collaborative strategy to the experimental real unmanned vessel swarm for small-sample trial operation, and compares it with the playback results of the watershed twin environment under the same working conditions to calculate the virtual-real difference degree.
[0039] Specifically, real-world navigation data of the experimental unmanned surface vessel swarm was collected, including vessel position trajectory sequences, observed concentration sequences, and total round energy consumption. Simultaneously, a replay simulation was performed in a watershed twin environment using the same initial conditions and control sequences as the trial run, obtaining twin simulation data. The virtual-real difference was calculated using the following formula:
[0040] In the formula, T is the total number of sampling moments in a single evaluation round, that is, the number of sampling points that participate in the statistics after the actual flight data and the playback simulation data are aligned in time. Its value is equal to the ratio of the round duration to the sampling period. The entity side and the twin side use the same sampling period and the same round duration, so the same T value is taken on both sides; t is the sampling moment number, and t takes the values from 1 to T in sequence; 1 / T means that the deviation of each sampling moment is averaged over time, so that the virtual-real difference obtained under different round durations is comparable; Let t be the position of the unmanned vessel at time t; The position of the virtual ship after playback simulation at time t under the same working conditions in the twin environment of the watershed; Let be the observed concentration of the unmanned surface vessel at time t; The observed concentration of the virtual ship at time t after playback simulation under the same working conditions in the twin environment of the watershed; The total energy consumption per round for the physical unmanned vessel; To determine the total round energy consumption of the virtual ship after replaying the simulation under the same operating conditions in the twin environment of the watershed; As the first normalization benchmark; This serves as the second normalization benchmark. This is the first sensitivity analysis parameter; This is the second sensitivity analysis parameter; This is the third sensitivity analysis parameter. Gap ≥ 0, and Gap = 0 if and only if the trajectories, observations, and energy consumption on both sides are completely consistent.
[0041] The twin parameter identification and write-back module operates when the virtual-real difference exceeds a preset threshold: It aims to minimize the residual between the actual observation data of the experimental unmanned surface vessel (USV) swarm and the simulated observation data generated by the watershed twin environment replay simulation. It then inverts the parameters to be corrected, which include at least one of the diffusion coefficient, propulsion efficiency, and sensor noise variance. The inversion solution employs least squares or Bayesian optimization methods, with an iteration upper limit and convergence tolerance set to ensure convergence. The identified corrected parameters are written back to the watershed twin environment, updating the corresponding parameters of the pollution diffusion layer and the agent layer, ensuring continuous alignment between the twin environment and the measured data. Training continues in the updated twin environment to update the USV swarm's collaborative strategy, and the virtual-real consistency evaluation module re-evaluates the updated strategy. This process of trial run—evaluation—write-back—retraining is repeated until the virtual-real difference converges to within the threshold.
[0042] The determination is based on two conditions: the difference between virtual and real data and the task success rate. The virtual-real difference is determined after K consecutive evaluation rounds. All are less than the download threshold Furthermore, the proportion of successful rounds in the K consecutive evaluation rounds is greater than or equal to the success rate threshold. Only when the conditions are met is the current unmanned surface vessel (USV) swarm collaborative strategy allowed to be formally migrated and deployed to the target USV swarm; if any condition is not met, training continues or parameters are rewritten. The use of a dual-condition judgment method over K consecutive rounds effectively avoids misjudgments caused by accidental success in a single round, improving the reliability of the migration decision. K and the threshold... , Set according to the task acceptance requirements.
[0043] In this embodiment, the success criterion for a single evaluation round is as follows: the unmanned vessel swarm completes the localization of the pollution emission source within the set round time limit, that is, at least one unmanned vessel in the swarm arrives within a neighborhood centered on the actual emission source location and with a set localization error threshold as the radius, and none of the vessels violate the safety constraints of the minimum offshore distance and minimum encounter distance within that round. Otherwise, the mission is considered a failure. The round time limit, localization error threshold, and success rate threshold are defined as follows: All settings are configured according to the task acceptance requirements.
[0044] The safety constraint enforcement module operates when the target entity's unmanned surface vessel (USV) swarm executes the formally deployed USV swarm coordination strategy: it applies speed constraints, offshore distance constraints, and minimum encounter distance constraints to the actions output by the strategy in real time. The speed constraint requires that the speed of each USV in the swarm is less than or equal to a set upper limit; the offshore distance constraint requires that the offshore distance of each USV is not less than a set lower limit; and the minimum encounter distance constraint requires that the minimum encounter distance between any USV and other USVs in the swarm is not less than a set lower limit. When an action output by the strategy results in a violation of any constraint, the action is restricted to the boundary value of the corresponding constraint, and the target entity's USV swarm executes the constraint-processed action. Thus, while preserving the strategy's decision-making capabilities, the navigation safety of the entity's USV swarm is ensured from the ground up.
[0045] In one embodiment, training of unmanned vessel swarm coordination strategies is conducted in a twin environment, providing a set of specific, implementable settings.
[0046] Configure the navigable area of the river section and the discharge source, and set up a continuous discharge source at (50, 50) with a discharge rate of A constant value is taken during the emission period; the velocity vector is set to (0.1, 0.0), which means that the flow field flows at a speed of 0.1 along the x direction, and the instantaneous velocity is updated in each grid according to the numerical solution of the hydrodynamic layer; 3 to 5 unmanned vessels are configured, and each vessel is initialized at a different position near the river inlet.
[0047] The feature extraction layer uses two 64-dimensional fully connected layers with a ReLU activation function. The policy branch outputs the probability distribution of actions, and the value branch outputs the state value estimate. The state space dimension is set to 3, and the state is represented as ( , , The action space has a dimension of 2, representing the expected heading increment and the expected speed increment (Δψ, Δv); hyperparameters such as learning rate, discount factor γ (0≤γ<1), pruning range ε and batch size are set. The discount factor is used to balance the importance of current rewards and future rewards.
[0048] The total number of training rounds is set to 500. At the start of each round, the environment is reset and the initial state of each ship is obtained. The state list, action list, reward list, next state list, and termination flag list for storing trajectories are initialized. Each ship selects an action based on the policy branch output. The action range is initially limited to [-1, 1], then scaled to the physical range of the desired heading increment and desired speed increment before execution. The next state, reward, and termination flag are obtained, and the current state, action, reward, next state, and termination flag are stored in the corresponding lists. Rewards are designed based on tracing objectives such as approaching the target, suppressing ineffective movements, and reducing energy consumption, encouraging the fleet to complete tasks efficiently and with low energy consumption. After each round, the policy network and value network are updated using the trajectory data collected in that round. The update logic includes calculating the advantage function, pruning the policy update magnitude according to the pruning objective function, calculating the policy loss and value loss, and performing backpropagation. The training progress is output every 50 rounds, and the total number of training rounds is adjusted according to the task complexity to ensure sufficient model convergence. Curriculum-based randomization is applied synchronously during training. Take the relative amplitude of the historical fluctuations of each parameter from the 5th to the 95th percentile. A certain proportion of the total number of training rounds is taken so that the randomization amplitude approaches saturation in the later stages of training.
[0049] After training, the pre-download evaluation and parameter write-back are performed according to the aforementioned virtual-real transfer closed loop: the trained strategy is downloaded to the experimental unmanned vessel fleet and tested with a small sample (e.g., dozens of trajectories), and the position trajectory, observed concentration, and round energy consumption of the experimental unmanned vessel fleet are collected; at the same time, the playback simulation is performed in the watershed twin environment with the same initial conditions and the same control sequence to obtain twin simulation data; the virtual-real difference degree is calculated according to the aforementioned virtual-real difference degree formula, and the difference degree is obtained by weighting the normalized three terms: trajectory error, concentration error, and energy consumption error.
[0050] Set download threshold Success rate threshold With the number of consecutive rounds K: when the difference between the real and virtual values in a certain evaluation is greater than At that time, the twin parameter identification and write-back module uses least squares or Bayesian optimization to invert the diffusion coefficient, propulsion efficiency, and sensor noise variance within the set iteration upper limit and convergence tolerance. The identification results are then written back to the watershed twin environment to update the corresponding parameters of the pollution diffusion layer and the agent layer. Subsequently, training and evaluation continue in the updated twin environment. This process of trial run-evaluation-write-retraining is repeated until the virtual-real difference is less than 0.5% for K consecutive evaluation rounds. And the task success rate is greater than or equal to Once the strategy is deemed to meet the migration and deployment conditions, it is formally migrated and deployed to the target entity's unmanned vessel swarm. When the formally deployed strategy is executed by the target entity's unmanned vessel swarm, hard constraints such as upper speed limits, lower offshore distance limits, and lower minimum encounter distance limits are imposed on its real-time output actions. If any constraint is violated by an action, the strategy is restricted to the constraint boundary before execution. Therefore, the entire process, from building the twin environment and training the unmanned vessel swarm's collaborative strategy to small-sample trial operation, virtual-real difference assessment, parameter write-back, and safe execution, can be implemented step-by-step according to the above settings, and converges when the download conditions are met.
[0051] Please see Figure 2 This is a schematic diagram of the structure of a virtual-to-real migration system for an unmanned vessel swarm cooperative strategy based on a watershed digital twin, provided in an embodiment of this application. This embodiment of the application provides a virtual-to-real migration system for an unmanned vessel swarm cooperative strategy based on a watershed digital twin, used to implement the virtual-to-real migration method for an unmanned vessel swarm cooperative strategy based on a watershed digital twin provided in the above embodiment, including: The twin environment construction module is used to construct a watershed twin environment, which is generated by coupling four layers: topography layer, hydrodynamic layer, pollution diffusion layer and intelligent agent layer. The collaborative strategy training module is used to generate training scenarios in the watershed twin environment according to the curriculum-based domain randomization scheduling, and train the collaborative strategy of the unmanned vessel swarm under the training scenarios using a centralized training and decentralized execution framework. The virtual-real consistency evaluation module is used to download the currently trained unmanned vessel swarm cooperative strategy to the experimental real unmanned vessel swarm for small-sample trial operation, and compare it with the playback simulation results of the watershed twin environment under the same working conditions to calculate the virtual-real difference degree. The twin parameter identification and write-back module is used to invert and correct the twin parameters and write them back to the watershed twin environment when the virtual-real difference is greater than a preset threshold, return to continue training to update the unmanned vessel swarm cooperative strategy, and re-evaluate the updated unmanned vessel swarm cooperative strategy. The strategy determination download module is used to formally migrate and deploy the currently satisfied unmanned vessel swarm collaborative strategy to the target entity unmanned vessel swarm when the virtual-real difference degree is less than the download threshold and the task success rate is greater than or equal to the success rate threshold in multiple consecutive evaluation rounds. The safety constraint execution module is used to apply speed constraints, offshore distance constraints, and minimum encounter distance constraints to the actions output in real time by the unmanned vessel swarm of the target entity when the unmanned vessel swarm executes the unmanned vessel swarm cooperative strategy. When the action causes any constraint to be violated, the action is restricted to the boundary value of the corresponding constraint before execution.
[0052] The technical solutions of this application have been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of this application is obviously not limited to these specific embodiments. Without departing from the principles of this application, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of this application.
[0053] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A virtual-to-real migration method for unmanned vessel swarm cooperative strategy based on watershed digital twin, characterized in that, include: A watershed twin environment is constructed, which is generated by coupling four layers: topography layer, hydrodynamic layer, pollution diffusion layer and intelligent agent layer. In the watershed twin environment, training scenarios are generated by curriculum-based domain randomization scheduling, and a collaborative strategy for unmanned vessel swarms is trained under the training scenarios using a centralized training and distributed execution framework. The currently trained unmanned vessel swarm cooperative strategy is downloaded to the experimental physical unmanned vessel swarm for small-sample trial operation, and compared with the playback simulation results of the watershed twin environment under the same working conditions to calculate the virtual-real difference degree. When the virtual-real difference exceeds a preset threshold, the twin parameters are inverted and corrected and written back to the watershed twin environment. Training continues to be conducted to update the unmanned vessel swarm cooperative strategy, and the updated unmanned vessel swarm cooperative strategy is re-evaluated. When the virtual-to-real difference is less than the download threshold and the mission success rate is greater than or equal to the success rate threshold for multiple consecutive evaluation rounds, the currently qualified unmanned vessel swarm collaboration strategy will be officially migrated and deployed to the target entity unmanned vessel swarm. When the target entity unmanned vessel swarm executes the unmanned vessel swarm cooperative strategy, speed constraints, offshore distance constraints, and minimum encounter distance constraints are applied to the actions output by the strategy in real time. When the action causes any constraint to be violated, the action is restricted to the boundary value of the corresponding constraint before execution.
2. The method according to claim 1, characterized in that, The topographic layer, hydrodynamic layer, pollution diffusion layer, and intelligent agent layer are specifically as follows: The topographic layer is constructed from a digital elevation model (DEM) and measured river cross-section data; the hydrodynamic layer provides an instantaneous velocity field based on the numerical solution of the two-dimensional shallow water equation or its simplified flow field; the pollution diffusion layer calculates pollutant concentration based on the instantaneous velocity field in the form of discrete convolution superposition of continuous and mobile sources; the intelligent agent layer consists of a three-degree-of-freedom dynamic model of the hull, a sensor noise model, and a communication channel model.
3. The method according to claim 2, characterized in that, The formula for calculating the concentration of the pollutants is as follows: In the formula, Number the emission sources; The discrete sequence number of the emission release time; This refers to the historical release moment of the pollutants; For the first Each emission source at time The emission rate; The discrete time step; Location of the emission source; The observation time when calculating the concentration; For the first The observation time is at the observation point ( The pollutant concentration at point ( ); D is the diffusion coefficient; The velocity component along the x-direction represents the instantaneous velocity field output by the hydrodynamic layer. This represents the velocity component along the y-direction of the instantaneous velocity field output by the hydrodynamic layer.
4. The method according to claim 1, characterized in that, The course-based domain randomization scheduling is implemented using the following formula: In the formula, This is the k-th randomization parameter; Nominal value; For training rounds; The maximum randomization amplitude; This is the course time constant.
5. The method according to claim 1, characterized in that, The degree of difference between virtual and real data is calculated according to the following formula: In the formula, T is the total number of sampling times in a single evaluation round; t is the sampling time number, which takes values from 1 to T. Let t be the position of the unmanned vessel at time t; The position of the virtual ship after playback simulation at time t under the same working conditions in the twin environment of the watershed; Let be the observed concentration of the unmanned surface vessel at time t; The observed concentration of the virtual ship at time t after playback simulation under the same working conditions in the twin environment of the watershed; The total energy consumption per round for the physical unmanned vessel; To determine the total round energy consumption of the virtual ship after replaying the simulation under the same operating conditions in the twin environment of the watershed; As the first normalization benchmark; This serves as the second normalization benchmark. This is the first sensitivity analysis parameter; This is the second sensitivity analysis parameter; This is the third sensitivity analysis parameter.
6. The method according to claim 1, characterized in that, The consecutive evaluation rounds refer to K consecutive evaluation rounds; the virtual-to-real difference degree being less than the download threshold and the mission success rate being greater than or equal to the success rate threshold are specifically defined as follows: the virtual-to-real difference degree being less than the download threshold for all K consecutive evaluation rounds, and the proportion of mission success rounds in the K consecutive evaluation rounds being greater than or equal to the success rate threshold; wherein, in a single evaluation round, when the unmanned vessel swarm completes the location of the pollution emission source within the set round time limit, and none of the vessels violate the safety constraints of the minimum offshore distance and minimum encounter distance within that round, the mission is deemed successful for that round; otherwise, the mission is deemed to have failed for that round.
7. The method according to claim 1, characterized in that, The actions output in real time by the strategy are subject to speed constraints, offshore distance constraints, and minimum encounter distance constraints. When an action causes any constraint to be violated, the action is restricted to the boundary value of the corresponding constraint before execution. Specifically: the speed constraint is that the speed of each unmanned vessel in the unmanned vessel swarm is less than or equal to the upper limit of a first set threshold; the offshore distance constraint is that the offshore distance of each unmanned vessel is greater than or equal to the lower limit of a second set threshold; the minimum encounter distance constraint is that the minimum encounter distance between any unmanned vessel and other unmanned vessels in the unmanned vessel swarm is greater than or equal to the lower limit of a third set threshold; if the action output by the strategy causes any constraint to be violated, the action is restricted to the boundary value of the corresponding constraint before execution.
8. The method according to claim 1, characterized in that, The inversion correction twin parameters are specifically as follows: With the objective of minimizing the residual between the actual observation data of the experimental unmanned vessel fleet and the simulated observation data generated by the watershed twin environment playback simulation, parameters to be corrected are inverted. The parameters to be corrected include at least one of diffusion coefficient, propulsion efficiency, and sensor noise variance. The inverted and corrected twin parameters are solved using least squares or Bayesian optimization methods, and an iteration upper limit and convergence tolerance are set.
9. The method according to claim 1, characterized in that, The centralized training and distributed execution framework includes a feature extraction layer, a policy branch, and a value branch. The output of the feature extraction layer is connected to the inputs of both the policy branch and the value branch. The feature extraction layer includes a first fully connected layer, a first ReLU activation function, a second fully connected layer, and a second ReLU activation function connected in sequence. It is used to receive the local observation state of a single unmanned vessel and output a feature vector. The policy branch is used to receive the feature vector and output the probability distribution of the unmanned vessel's actions. The value branch is used to receive the feature vector and output a value estimate of the unmanned vessel's state.
10. A virtual-to-real migration system for unmanned vessel swarms based on a watershed digital twin, characterized in that, include: The twin environment construction module is used to construct a watershed twin environment, which is generated by coupling four layers: topography layer, hydrodynamic layer, pollution diffusion layer and intelligent agent layer. The collaborative strategy training module is used to generate training scenarios in the watershed twin environment according to the curriculum-based domain randomization scheduling, and train the collaborative strategy of the unmanned vessel swarm under the training scenarios using a centralized training and decentralized execution framework. The virtual-real consistency evaluation module is used to download the currently trained unmanned vessel swarm cooperative strategy to the experimental real unmanned vessel swarm for small-sample trial operation, and compare it with the playback simulation results of the watershed twin environment under the same working conditions to calculate the virtual-real difference degree. The twin parameter identification and write-back module is used to invert and correct the twin parameters and write them back to the watershed twin environment when the virtual-real difference is greater than a preset threshold, return to continue training to update the unmanned vessel swarm cooperative strategy, and re-evaluate the updated unmanned vessel swarm cooperative strategy. The strategy determination download module is used to formally migrate and deploy the currently satisfied unmanned vessel swarm collaborative strategy to the target entity unmanned vessel swarm when the virtual-real difference degree is less than the download threshold and the task success rate is greater than or equal to the success rate threshold in multiple consecutive evaluation rounds. The safety constraint execution module is used to apply speed constraints, offshore distance constraints, and minimum encounter distance constraints to the actions output in real time by the unmanned vessel swarm of the target entity when the unmanned vessel swarm executes the unmanned vessel swarm cooperative strategy. When the action causes any constraint to be violated, the action is restricted to the boundary value of the corresponding constraint before execution.