An escort method for an unmanned underwater vehicle and related products
By constructing a diverse data system and strategy scenario mapping relationship, and combining deep reinforcement learning and adversarial behavior rules to optimize the escort strategy of the UUV simulation agent, the problem of insufficient robustness and adaptability of UUV formation control in complex underwater environments has been solved, achieving rapid response and efficient task execution.
Patent Information
- Application Number
- CN202511649224.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-12
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2045-11-12
AI Technical Summary
Existing unmanned underwater vehicle (UUV) formation control lacks robustness and adaptability in complex and unknown underwater environments, making it difficult to quickly adapt to environmental changes and unforeseen disturbances through reasonable morphological changes, resulting in unstable mission execution.
By acquiring real-time scenario state data and multiple sets of raw simulation data, and combining them with random simulation data, intelligent escort strategies are generated. A strategy-scenario mapping relationship is constructed, and deep reinforcement learning and adversarial behavior rules are used to optimize the generation of escort strategies for UUV simulation agents, thereby achieving a diversified data system and rapid response.
It significantly improves the robustness and response speed of UUVs in complex and unknown environments, ensures the safety of protected targets and the success rate of missions, and solves the problem that formations have difficulty adapting to environmental changes quickly.
Smart Images

Figure CN121115519B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, specifically to an escort method for an unmanned underwater vehicle and related products. Background Technology
[0002] An unmanned underwater vehicle (UUV) is an unmanned system capable of autonomous or remote control to perform various tasks in underwater environments. It is typically equipped with navigation, sensing, communication, and execution mechanisms, enabling it to perform multiple functions such as ocean exploration, environmental monitoring, and underwater operations without direct human intervention. Due to the complexity of the underwater environment and limited communication, UUVs need to possess a high degree of autonomy and intelligence to cope with changing currents, obstacles, and unknown terrain, achieving stable and reliable underwater navigation and mission execution.
[0003] In UUV operations, to ensure the designated target can efficiently and accurately complete its mission objectives, several UUVs are typically deployed in coordinated operations. UUVs work closely together with the designated UUV through formation control strategies to jointly address various challenges in complex underwater environments. During this coordination process, UUVs achieve dynamic coordination and resource sharing within the formation based on cooperative control algorithms and real-time information exchange, thereby effectively improving the overall mission execution success rate and ensuring the designated UUV successfully completes its predetermined objectives.
[0004] Reinforcement learning, as an intelligent control method based on trial and error and feedback mechanisms, provides new ideas and means for UUV formation control. In formation control, reinforcement learning enables individual UUVs to autonomously learn efficient collaborative strategies through continuous interaction with the environment and other members, thereby achieving dynamic adaptation and optimization of mission objectives. Especially in multi-UUV scenarios, reinforcement learning can help form flexible and robust formation behavior, improving the overall UUV coordination and anti-interference capabilities.
[0005] However, the application of reinforcement learning in UUV formation control is currently limited, and UUVs exhibit poor robustness in complex and unknown underwater environments. For example, they cannot quickly adapt to sudden environmental changes or unforeseen disturbances through reasonable formation configuration changes.
[0006] Therefore, how to improve the robustness and adaptability of UUV formation control in escort operations against complex and unknown underwater environments, so as to achieve more flexible and reliable collaborative operations, is a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0007] To address the aforementioned issues, this application provides an escort method and related products for unmanned underwater vehicles, which can enhance the robustness and adaptability of UUV formation control in the escort process to cope with complex and unknown underwater environments, and achieve more flexible and reliable collaborative operations.
[0008] The embodiments of this application disclose the following technical solutions:
[0009] An escort method based on an unmanned underwater vehicle (UUV) is applied to a UUV simulation agent. The UUV simulation agent is obtained through multi-strategy fusion training. The method includes:
[0010] The system acquires real-time scene status data and task objectives, and obtains multiple sets of raw simulation data from the target simulation environment based on the task objectives. The task objectives include assisting the protected target in escaping the tracking and interference of multiple interference sources through the coordinated escort of multiple UUVs, ensuring the protected target safely reaches the target location. The target simulation environment is a platform integrating multi-agent dynamic simulation function and underwater physical environment simulation function. One set of raw simulation data includes a set of scene status simulation data. The set of scene status simulation data includes a set of virtual poses of the protected target, a set of virtual poses and a set of virtual detection data for each of the multiple UUVs, a set of virtual poses and a set of virtual detection data for each of the multiple interference sources, and a set of underwater virtual environment parameters.
[0011] Multiple sets of random simulation data are randomly generated based on the aforementioned multiple sets of original simulation data;
[0012] Based on the multiple sets of random simulation data, the multiple sets of original simulation data, and the real-time scene state data, escort strategies are generated to obtain a variety of intelligent escort strategies; each intelligent escort strategy corresponds to a set of state data, and the state data includes the random simulation data, the original simulation data, or the real-time scene state data.
[0013] Based on the state data corresponding to each of the various intelligent escort strategies, multiple virtual scenarios are generated;
[0014] The various intelligent escort strategies are matched and associated with their respective virtual scenarios to obtain a strategy-scenario mapping relationship. The strategy-scenario mapping relationship is then distributed to each UUV so that each UUV can perform intelligent escort in accordance with the strategy-scenario mapping relationship.
[0015] In one possible implementation, the multi-policy fusion training process of the UUV simulation agent includes:
[0016] A training dataset is obtained, and a deep reinforcement learning algorithm is used to perform transfer learning on the agent in combination with the training dataset to obtain a basic escort strategy model; the training dataset includes multiple sets of historical scene state data labeled with escort strategy tags;
[0017] Obtain pre-defined adversarial behavior rules and input them into the basic escort strategy model to obtain a rule enhancement model; the adversarial behavior rules are behavioral guidelines formulated for scenarios with clear solutions in all adversarial scenarios between the protected target and the multiple interference sources.
[0018] The UUV simulation agent is obtained by performing game-theoretic training on the rule enhancement model based on centralized training and distributed execution CTDE technology.
[0019] The UUV simulation agent generates escort strategies based on the adversarial behavior rules and the neural network of the UUV simulation agent.
[0020] In one possible implementation, the step of training the rule-enhancing model using CTDE technology to obtain the UUV simulation agent includes:
[0021] Multiple training agents are acquired; the multiple training agents include multiple full-scale defect training agents and multiple specialized defect training agents with different adversarial styles; the full-scale defect training agents are designed to target the global weaknesses of the rule enhancement model; the specialized defect training agents are designed to target a single or some specific weaknesses of the rule enhancement model.
[0022] Based on the CTDE technology, the multiple training agents are included in the adversarial training pool. In the adversarial training pool, the multiple training agents conduct multiple rounds of game training on the rule enhancement model until the training cutoff condition is reached, at which point the game training stops, and the UUV simulation agent is obtained.
[0023] In each round of game training, a centralized controller integrates the observation data and action feedback from the multiple training agents and the rule enhancement model to obtain global state information. Based on the global state information, a reward and punishment signal based on the degree of weakness improvement is calculated, and the model parameters of the rule enhancement model are optimized based on the reward and punishment signal. In the reward and punishment signal, the rule enhancement model's successful avoidance of weaknesses is recorded as a positive reward, and the rule enhancement model's failure to avoid weaknesses is recorded as a negative punishment.
[0024] In one possible implementation, the intelligent escort strategy includes pre-set identification information for each UUV and corresponding task assignment instructions; the task assignment instructions are used to guide each UUV to cooperate and form an effective escort formation for the protected target.
[0025] The process of intelligent escort by each UUV in combination with the strategy scene mapping relationship includes: for each UUV, the UUV selects the intelligent escort strategy corresponding to the current scene in the strategy scene mapping relationship based on the virtual scene matching the current scene to obtain the target escort strategy, and calls the corresponding control command to execute action planning and collaborative operation according to the division of labor instruction corresponding to the identification information of the target escort strategy and the identification information of the UUV, so as to realize the escort of the protected target.
[0026] In one possible implementation, the escort strategy is generated based on the multiple sets of random simulation data, the multiple sets of original simulation data, and the real-time scenario state data, resulting in various intelligent escort strategies, including:
[0027] Based on the multiple sets of random simulation data, the multiple sets of original simulation data, and the real-time data of the scene state, corresponding escort strategies are generated respectively, resulting in multiple initial escort strategies.
[0028] Based on the aforementioned multiple initial escort strategies, iterative derivative strategies are generated to obtain multiple derivative escort strategies;
[0029] Wherein, both the initial escort strategy and the derived escort strategy are the intelligent escort strategy;
[0030] The iterative derivation strategy generation process includes: in the target simulation environment, acquiring multiple sets of new original simulation data based on the current batch escort strategy to obtain the corresponding batch of derivation simulation data; randomly generating multiple sets of random derivation simulation data based on the corresponding batch of derivation simulation data; generating a derivation escort strategy based on each derivation simulation data in the corresponding batch, and generating a derivation escort strategy based on each random derivation simulation data in the corresponding batch, to obtain a new batch of derivation escort strategies; repeating the steps of acquiring derivation simulation data, generating random derivation simulation data, and generating derivation escort strategies until the iteration stops when the iteration termination condition is met;
[0031] In the first iteration, the escort strategy for the current batch is the multiple initial escort strategies, and in subsequent iterations, it is the derived escort strategy generated in the previous batch.
[0032] In one possible implementation, the state data further includes derived simulation data or randomized derived simulation data;
[0033] The process involves generating multiple virtual scenarios based on the state data corresponding to each of the various intelligent escort strategies, including:
[0034] For each intelligent escort strategy, the virtual poses of the protected target, UUV virtual poses and detection data, interference source virtual poses and detection data, and underwater virtual environment parameters contained in the random simulation data, original simulation data, real-time scene state data, derived simulation data, or randomly derived simulation data corresponding to the intelligent escort strategy are extracted. Then, a three-dimensional virtual scene is constructed according to the preset scene modeling rules by combining the virtual poses of the protected target, UUV virtual poses and detection data, interference source virtual poses and detection data, and underwater virtual environment parameters contained in the random simulation data, original simulation data, real-time scene state data, derived simulation data, or randomly derived simulation data corresponding to the intelligent escort strategy.
[0035] A UUV simulation intelligent agent, the UUV simulation intelligent agent comprising:
[0036] The integrated acquisition unit is used to acquire real-time scene state data and task objectives, and to acquire multiple sets of raw simulation data from the target simulation environment based on the task objectives. The task objectives include assisting the protected target in escaping the tracking and interference of multiple interference sources through the coordinated escort of multiple UUVs, ensuring that the protected target safely reaches the target location. The target simulation environment is a platform that integrates multi-agent dynamic simulation function and underwater physical environment simulation function. One set of raw simulation data includes a set of scene state simulation data. The set of scene state simulation data includes a set of virtual poses of the protected target, a set of virtual poses and a set of virtual detection data for each of the multiple UUVs, a set of virtual poses and a set of virtual detection data for each of the multiple interference sources, and a set of underwater virtual environment parameters.
[0037] The simulation data generation unit is used to randomly generate multiple sets of random simulation data based on the multiple sets of original simulation data.
[0038] The escort strategy generation unit is used to generate escort strategies based on the multiple sets of random simulation data, the multiple sets of original simulation data, and the real-time scenario state data, to obtain a variety of intelligent escort strategies; each intelligent escort strategy corresponds to a set of state data, and the state data includes the random simulation data, the original simulation data, or the real-time scenario state data.
[0039] The virtual scene generation unit is used to generate multiple virtual scenes based on the state data corresponding to each of the various intelligent escort strategies.
[0040] The intelligent escort unit is used to match and associate the various intelligent escort strategies with their respective corresponding virtual scenarios to obtain a strategy-scenario mapping relationship, and then distribute the strategy-scenario mapping relationship to each UUV so that each UUV can perform intelligent escort in combination with the strategy-scenario mapping relationship.
[0041] In one possible implementation, the UUV simulation agent further includes:
[0042] The basic model training unit is used to acquire a training dataset and use a deep reinforcement learning algorithm to perform transfer learning on the agent in combination with the training dataset to obtain a basic escort strategy model; the training dataset includes multiple sets of historical scene state data labeled with escort strategy tags;
[0043] The enhanced model acquisition unit is used to acquire pre-set adversarial behavior rules and input the adversarial behavior rules into the basic escort strategy model to obtain a rule enhancement model; the adversarial behavior rules are behavioral guidelines formulated for the coordinated actions between various UUVs and the adversarial strategies between the protected target and the multiple interference sources.
[0044] The game training unit is used to train the rule enhancement model using CTDE technology to obtain the UUV simulation agent.
[0045] An escort device for an unmanned underwater vehicle includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the escort method for the unmanned underwater vehicle as described above.
[0046] A computer-readable storage medium storing instructions that, when executed on a terminal device, cause the terminal device to perform the escort method for an unmanned underwater vehicle as described above.
[0047] Compared with the prior art, this application has the following beneficial effects:
[0048] This application provides an escort method and related products for unmanned underwater vehicles (UUVs). Specifically, when executing the escort method for UUVs provided in this application, firstly, real-time scene state data and mission objectives are acquired. Based on the mission objectives, multiple sets of original simulation data are obtained from the target simulation environment. Each set of original simulation data includes a set of scene state simulation data, which includes a set of virtual poses for the protected target, a set of virtual poses and virtual detection data for each of the multiple UUVs, a set of virtual poses and virtual detection data for each of the multiple interference sources, and a set of underwater virtual environment parameters. Next, multiple sets of random simulation data are randomly generated based on the multiple sets of original simulation data. Escort strategies are then generated based on the multiple sets of random simulation data, the multiple sets of original simulation data, and the real-time scene state data, resulting in various intelligent escort strategies. Then, multiple virtual scenarios are generated according to the state data corresponding to each of the various intelligent escort strategies. Finally, the various intelligent escort strategies are matched and associated with their corresponding virtual scenarios to obtain a strategy-scenario mapping relationship. This strategy-scenario mapping relationship is then distributed to each UUV so that each UUV can perform intelligent escort in accordance with the strategy-scenario mapping relationship. This application utilizes raw simulation data provided by the target simulation environment, combined with randomly generated simulation data and real-time data from actual scenarios, to construct a diverse data system covering multiple elements including the protected target, UUVs, interference sources, and the underwater environment. The resulting intelligent escort strategies overcome the limitations of single UUV agent reinforcement learning training, which relies heavily on training data, significantly expanding the applicable scenarios for UUV agents. Furthermore, by mapping intelligent escort strategies to corresponding virtual scenarios, UUVs can quickly invoke strategies adapted to the current scenario when facing sudden environmental changes or unforeseen interference, without relying on the trial-and-error optimization process of reinforcement learning. This significantly improves the robustness and response speed of UUVs in unknown and complex environments, effectively solving the problem of multiple UUVs struggling to quickly adapt to the environment through reasonable morphological changes, thus ensuring the safety of the protected target and the mission success rate in collaborative escort. Attached Figure Description
[0049] To more clearly illustrate the technical solutions in this embodiment or the prior art, the drawings used in the description of the embodiment or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0050] Figure 1 A flowchart illustrating an escort method for an unmanned underwater vehicle provided in this application embodiment;
[0051] Figure 2A flowchart illustrating the implementation of a multi-strategy fusion training method for an intelligent agent, as provided in an embodiment of this application;
[0052] Figure 3 A flowchart illustrating the implementation of an intelligent agent game training method provided in this application embodiment;
[0053] Figure 4 A flowchart illustrating the implementation of an escort strategy generation method provided in this application embodiment;
[0054] Figure 5 This is a schematic diagram of the structure of a UUV simulation intelligent agent provided in an embodiment of this application. Detailed Implementation
[0055] To facilitate understanding of the technical solutions provided in the embodiments of this application, the background technology involved in the embodiments of this application will be described below.
[0056] A UUV (Unmanned Underwater Vehicle) is an unmanned system capable of autonomously or remotely controlling tasks in underwater environments. It possesses navigation, sensing, communication, and execution mechanisms, and is suitable for various functions such as ocean exploration and environmental monitoring. In complex underwater environments, UUVs require a high degree of autonomy and intelligence to cope with currents, obstacles, and unknown terrain. To improve mission efficiency, UUVs are typically deployed in cooperative operations, achieving dynamic coordination and resource sharing through formation control strategies. Reinforcement learning, as an intelligent control method, enables each UUV to autonomously learn efficient cooperative strategies, enhancing the coordination and anti-interference capabilities of multiple UUVs.
[0057] Although reinforcement learning has attracted widespread attention due to its ability to continuously optimize strategies through interaction with the environment, its application in UUV formation control still faces certain limitations. Firstly, reinforcement learning models typically require large amounts of training data and repeated trials to obtain effective control strategies. However, the complexity of the underwater environment and the high cost of experimentation limit the practical implementation of sufficient data collection and training. This results in a relatively narrow scope of application for reinforcement learning methods in real-world scenarios, particularly in underwater operations characterized by variable environments and diverse tasks.
[0058] Secondly, the underwater environment is highly uncertain and dynamic, including factors such as changing ocean currents, low visibility, communication delays or even interruptions, and the appearance of unforeseen obstacles or interference sources. These complex and unknown factors place extremely high demands on the robustness of UUV formations. However, existing reinforcement learning algorithms often lack sufficient generalization ability and stability when faced with sudden environmental changes or abnormal interference, easily leading to policy failures or significant performance degradation. In other words, models trained through reinforcement learning perform poorly in new environments outside the training environment, making it difficult to guarantee real-time response and reliable safety in formation control.
[0059] Furthermore, the lack of pre-defined adversarial behavior rules in UUV formation control leads to poor interpretability of reinforcement learning strategies. Existing reinforcement learning models are mostly "black box" decision-making processes; their generated formation control commands (such as UUV speed adjustments and heading changes) rely on complex parameter mappings of neural networks rather than explicit logical rules. When formation coordination deviations occur (such as formation shifts or response delays), technicians struggle to trace the key influencing factors in the decision-making process, making it difficult to quickly pinpoint the root cause of the problem or optimize strategies for specific scenarios. For example, determining that "a certain UUV needs to shift 50 meters to the left" to maintain escort formation cannot explain whether this decision is based on "the detection threat of interference sources" or "the influence of water currents on the formation's center of gravity." This not only increases the difficulty of strategy debugging but also reduces the level of trust in the safety of formation control in practical applications.
[0060] To address this issue, this application provides an escort method and related products for unmanned underwater vehicles (UUVs). First, real-time scene state data and mission objectives are acquired. Based on this mission, multiple sets of raw simulation data are collected from the target simulation environment. Each set of raw simulation data contains a complete set of scene state simulation data, specifically including a set of virtual poses for the protected target, a set of virtual poses and detection data for each of the multiple UUVs, the virtual poses and detection data for each of the multiple interference sources, and a set of underwater virtual environment parameters. Subsequently, multiple sets of random simulation data are randomly generated using these raw simulation data. Next, by combining the multiple sets of random simulation data, the multiple sets of raw simulation data, and the real-time scene state data, escort strategies are generated, producing various intelligent escort strategies. Multiple virtual scenarios are constructed based on the state data corresponding to each intelligent escort strategy, and each escort strategy is matched and associated with its corresponding virtual scenario to form a strategy-scenario mapping relationship. Finally, this mapping relationship is distributed to each UUV, enabling it to perform intelligent escort based on the mapping information. This application leverages the rich raw data and randomly generated data provided by the target simulation environment, integrating multiple elements such as the protected target, UUVs, interference sources, and the underwater environment to construct a diverse and comprehensive data system. This overcomes the limitation of traditional reinforcement learning's reliance on a single training data, significantly expanding the applicable scenarios. Simultaneously, through the mapping and association between policies and virtual scenarios, UUVs can quickly invoke the most suitable policy for the current scenario when facing sudden environmental changes or unforeseen interference, eliminating the need for repeated trial and error processes in reinforcement learning. This significantly improves the robustness and response speed of UUVs in complex and unknown underwater environments, effectively solving the problem of formations struggling to quickly adapt to the environment through reasonable morphological changes, and ensuring the safety of the protected target and the mission success rate during collaborative escort operations.
[0061] Furthermore, step S202, by inputting adversarial behavior rules into the basic escort strategy model, can obtain a rule-enhanced model, thereby effectively compensating for the problems of unstable decision-making or poor interpretability that may occur when pure deep reinforcement learning models face complex, volatile, or extreme scenarios.
[0062] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0063] See Figure 1 The figure is a flowchart of an escort method for an unmanned underwater vehicle provided in an embodiment of this application. The escort method for the unmanned underwater vehicle is applied to a UUV simulation agent, which is obtained by multi-strategy fusion training based on the agent.
[0064] like Figure 1 As shown, the escort method for this unmanned underwater vehicle may include steps S101-S105:
[0065] S101: Obtain real-time data of the scene status and the task objective, and obtain multiple sets of raw simulation data from the target simulation environment based on the task objective.
[0066] To achieve efficient escort of the protected target, the UUV simulation agent first acquires real-time status data of the scene and the specific target of the current task through sensors and communication equipment. The main task objective is to utilize multiple UUVs to coordinate escort, helping the protected target escape tracking and interference from multiple sources of interference, thereby ensuring that the protected target can safely reach the predetermined target location. To support this task, the UUV simulation agent acquires multiple sets of raw simulation data from a highly integrated platform (i.e., a target simulation environment that includes multi-agent dynamic simulation capabilities and underwater physical environment simulation capabilities) based on the task objective.
[0067] Each set of raw simulation data contains a complete set of scenario state simulation data, which details various aspects of the underwater combat environment. Specifically, each set of raw simulation data corresponds to a complete set of scenario state simulation data, covering all key information of the escort scenario: it includes not only a set of virtual poses of the protected target (the virtual poses specifically include virtual position reflecting spatial location, virtual heading indicating direction of travel, and virtual speed reflecting speed of movement, the three of which together define the dynamic state of the protected target), but also core data from multiple UUVs. Each UUV corresponds to a set of virtual poses (consistent with the virtual pose dimensions of the protected target to ensure the consistency of the formation's coordinated state) and a set of virtual detection data (the virtual detection data includes information obtained by the UUV through simulated equipment such as sonar and electromagnetic sensors, such as the distance and azimuth detection results of interference sources, the contour data of surrounding underwater obstacles, and the real-time relative position of the protected target). The simulation data includes: feedback mechanisms to provide a basis for UUVs to determine the timing of cooperation; a set of virtual poses for multiple interference sources (used to simulate the movement trajectory and real-time position of the interference sources, supporting the simulation of adversarial scenarios); and a set of virtual detection data (such as the signal detection range of the interference sources, the type and intensity of the released interference signals, providing interference scenario references for strategy generation). In addition, a set of underwater virtual environment parameters is required, which specifically covers key environmental factors affecting UUV navigation and detection, such as water flow velocity reflecting water dynamics, temperature and salinity reflecting water characteristics, and topographic simulation data simulating seabed topography (such as reef distribution, trench location, etc.). These data together construct a realistic underwater scenario, ensuring that the original simulation data can provide comprehensive and accurate scenario support for subsequent strategy generation.
[0068] S102: Generate multiple sets of random simulation data based on the multiple sets of original simulation data.
[0069] To enhance data diversity and cover more possible underwater environments and combat scenarios, the UUV simulation agent generates multiple sets of randomized simulation data based on multiple sets of original simulation data. Specifically, while maintaining the core features of the original data, the virtual poses (including position, heading, and speed) of the protected target, UUV, and interference sources, as well as detection data and underwater environmental parameters (water flow velocity, temperature, salinity, terrain structure, etc.), are randomly perturbed or recombined within a certain range to generate new simulation scenarios. For example, in the original simulation data, the virtual position of the protected target is located near a fixed coordinate. The UUV simulation agent can randomly adjust its position within a certain radius of that coordinate, or change the speed and heading angle of the UUV, simulating the response of different cooperative strategies to environmental changes. At the same time, environmental factors such as water flow velocity and salinity can also be perturbed according to a preset probability distribution to simulate more complex and variable underwater conditions. This random generation method enriches the training dataset, improves the adaptability of intelligent escort strategies to unknown and dynamic environments, and avoids the risk of overfitting caused by relying solely on limited original data.
[0070] It should be noted that the data structure of the random simulation data is similar to that of the original simulation data. That is, a set of random simulation data includes a set of scene state random simulation data, a set of scene state random simulation data includes another set of virtual poses of the protected target, another set of virtual poses and another set of virtual detection data for each of the multiple UUVs, another set of virtual poses and another set of virtual detection data for each of the multiple interference sources, and another set of underwater virtual environment parameters.
[0071] S103: Based on the multiple sets of random simulation data, the multiple sets of original simulation data, and the real-time data of the scene status, an escort strategy is generated to obtain a variety of intelligent escort strategies.
[0072] In order to develop a more comprehensive and robust escort plan, the UUV simulation agent comprehensively utilizes multiple sets of random simulation data, multiple sets of raw simulation data, and real-time data of scene status as inputs to generate intelligent escort strategies.
[0073] One of the intelligent escort strategies corresponds to a set of state data, which includes the random simulation data, the original simulation data, or the real-time scenario state data.
[0074] S104: Generate multiple virtual scenarios based on the state data corresponding to each of the various intelligent escort strategies.
[0075] Based on various intelligent escort strategies and their corresponding state data, the UUV simulation agent transforms this state data into virtual scenarios with specific spatial and environmental characteristics, thereby providing a targeted simulation testing environment for each strategy. A virtual scenario refers to a three-dimensional underwater environment model constructed in a computer simulation platform. This model realistically reproduces the position and movement of the protected target, UUV, and interference sources, as well as underwater environmental parameters such as water flow velocity, temperature, salinity, and terrain, enabling the simulation of complex and ever-changing combat conditions. By generating virtual scenarios, the UUV simulation agent can comprehensively evaluate and optimize the performance of different escort strategies under various environmental conditions.
[0076] Specifically, when generating a virtual scenario, the UUV simulation agent first extracts the virtual poses (including position, heading, and speed) of the protected target and the UUVs, the distribution and dynamic information of interference sources, and the physical parameters of the underwater environment from the state data. This information is then mapped onto the simulation platform, forming a dynamically interactive virtual sea area by setting three-dimensional spatial coordinates, motion trajectories, and environmental attributes. For example, if a set of state data reflects that the protected target is in a strong current area, surrounded by multiple UUVs and multiple moving interference sources, the UUV simulation agent will construct a corresponding current vector field in the virtual scenario, arranging the initial positions and motion patterns of the UUVs and interference sources to realistically simulate the actual combat environment. Another set of state data may represent static interference sources and steady currents; in this case, the virtual scenario focuses on showcasing the cooperative formation and real-time response capabilities of the UUVs.
[0077] The virtual scenarios generated in this way not only enhance the realism of training and testing, but also provide a rich and diverse verification environment for intelligent escort strategies, helping to improve the autonomous decision-making and adaptability of unmanned underwater vehicles in complex underwater missions.
[0078] S105: Match and associate the various intelligent escort strategies with their respective corresponding virtual scenarios to obtain a strategy-scenario mapping relationship, and send the strategy-scenario mapping relationship to each UUV so that each UUV can perform intelligent escort in combination with the strategy-scenario mapping relationship.
[0079] To achieve intelligent and precise UUV escort, the UUV simulation agent first matches and associates various intelligent escort strategies with their corresponding virtual scenarios, forming a clear strategy-scenario mapping relationship. This mapping relationship describes the application conditions and execution methods of each escort strategy under specific underwater environments and situations, ensuring that the UUV can select the most suitable collaborative solution based on the current environmental state. Subsequently, the UUV simulation agent distributes this strategy-scenario mapping relationship to each UUV, enabling them to dynamically invoke the corresponding intelligent escort strategy in actual missions, combining their own environmental information and mission requirements, thus achieving highly autonomous and collaborative escort operations.
[0080] By establishing and applying this strategy-scenario mapping mechanism, UUVs not only possess the ability to respond quickly to complex and ever-changing environments, but also significantly improve the intelligence level and execution effect of UUV simulation agents.
[0081] Based on the content of S101-S105, the process first acquires real-time scene state data and mission objectives, and collects multiple sets of raw simulation data from the simulation environment, including parameters of the protected target, UUV, interference sources, and underwater environment. Then, by randomly expanding these data and combining them with real-time data, various intelligent escort strategies are generated. Next, based on the state data corresponding to each strategy, multiple virtual scenarios are constructed, and a mapping relationship is established between the strategies and scenarios. These scenarios are then distributed to each UUV, enabling it to intelligently select the appropriate escort strategy according to the current environment. This application utilizes a diverse data system, overcoming the limitations of single reinforcement learning on training data, improving the adaptability and response speed of UUVs in complex underwater environments, effectively ensuring the safety of the protected target, and increasing the mission success rate of the protected target.
[0082] In one possible implementation, this application also provides a process for multi-strategy fusion training of a UUV simulation agent, see [link to relevant documentation]. Figure 2 , Figure 2 A flowchart illustrating the implementation of a multi-policy fusion training method for intelligent agents provided in this application embodiment can be specifically implemented through steps S201-S203:
[0083] S201: Obtain the training dataset, and use a deep reinforcement learning algorithm to perform transfer learning on the agent in combination with the training dataset to obtain a basic escort strategy model.
[0084] In the initial stage of agent training, a training dataset containing multiple sets of historical scene state data labeled with escort strategies is first acquired. This data records in detail key information such as the pose, velocity, and distribution of interference sources of the protected target and UUV under different underwater environments, as well as the corresponding optimal escort strategy selection. By combining deep reinforcement learning algorithms and utilizing this rich and diverse historical data for transfer learning of the agent, the training process of the model can be effectively accelerated, and its generalization ability can be significantly improved. Transfer learning enables the agent to learn from successful experiences in existing scenarios, quickly master basic escort decision-making capabilities, and avoid the inefficiency and blindness of learning from scratch.
[0085] S202: Obtain the pre-set adversarial behavior rules, input the adversarial behavior rules into the basic escort strategy model, and obtain the rule enhancement model.
[0086] After the basic escort strategy model is constructed, pre-defined adversarial behavior rules are introduced. These rules are explicit behavioral guidelines formulated for various typical adversarial scenarios between the protected target and multiple interference sources. They include expert experience and tactical guidance, providing the agent with clear solutions. Inputting these rules into the basic escort strategy model yields a rule-enhanced model, effectively mitigating the instability or poor interpretability that pure deep reinforcement learning models may exhibit when facing complex, variable, or extreme scenarios. The rule-enhanced model combines the autonomous learning capabilities of neural networks with the logical constraints of rules to generate more robust, reasonable, and practically operational escort strategies.
[0087] In one possible implementation, the adversarial behavior rules specifically include, but are not limited to, strategies for rationally allocating the use of decoys and jammers in different interference scenarios. For example, when the protected target detects acoustic interference sources from multiple directions, the rules explicitly instruct the UUV to prioritize releasing decoys to attract the interference sources' attention, thereby protecting the protected target from direct threats. Simultaneously, after the decoys are released, the UUV must also activate jammers to conduct electronic jamming on key frequency bands based on the dynamic changes of the interference sources, weakening the detection and tracking capabilities of the interference sources. The rules also specify the timing of decoy and jammer releases, spatial layout, and energy management strategies to ensure effective resource utilization and continuous protection capabilities. By inputting these specific decoy release and jammer deployment rules into the basic escort strategy model, the agent can coordinate the use of hardware and software means in complex adversarial environments to achieve multi-layered, multi-faceted proactive defense, significantly improving the effectiveness of the adversarial efforts.
[0088] S203: The UUV simulation agent is obtained by game-theoretic training of the rule enhancement model based on CTDE technology.
[0089] The rule-enhancing model is trained using a game-theoretic approach based on Centralized Training Distributed Execution (CTDE). This approach aims to continuously optimize the model's policy performance through dynamic interactions among multiple agents. During training, the centralized training phase aggregates the state and action information of all agents, comprehensively capturing the adversarial relationship between the rule-enhancing model and the training agents. This guides the rule-enhancing model to adjust and improve its decision-making capabilities in complex and ever-changing interference environments. Simultaneously, distributed execution ensures that each agent can operate independently based on local observations during actual deployment, achieving efficient real-time responses. Through repeated game-theoretic training, the rule-enhancing model gradually adapts to and overcomes the diverse challenges posed by the training agents, generating escort strategies more robustly and accurately, ultimately resulting in a high-performance UUV simulation agent with practical value.
[0090] In one possible implementation, the UUV simulation agent achieves efficient generation of escort strategies by combining adversarial behavior rules with a neural network model. The combination methods mainly include three types:
[0091] Horizontal integration: refers to processing adversarial behavior rules and neural network outputs in parallel at the same decision-making level, and fusing the strategies or action suggestions generated independently by the rules and the network to obtain more robust and diverse decision results.
[0092] Vertical integration: This refers to embedding adversarial behavior rules as prior knowledge or constraints into the training and inference process of neural networks, guiding the network to learn rule-compliant policy performance, and enhancing the interpretability and security of the model.
[0093] Hybrid approach: Combining the advantages of horizontal and vertical integration, rules and neural networks are flexibly applied at different stages or modules. For example, rules can first filter out a reasonable action space, and then neural networks can refine the decision-making, or vice versa, to form a dynamic and collaborative escort strategy generation mechanism.
[0094] Through the above multi-dimensional combination, the UUV simulation intelligent agent can give full play to the clear guiding role of rules and the powerful learning ability of neural networks, so as to achieve intelligent and flexible response to diverse interferences in complex underwater environments, and effectively improve the overall performance and reliability of escort missions.
[0095] Steps S201-S203, through a hierarchical learning and optimization process, not only improve the agent's decision-making and anti-interference capabilities, but also enhance its adaptability and robustness in complex environments, providing a more reliable and efficient escort strategy for UUVs.
[0096] In one possible implementation, this application also provides a game training process for a rule-enhancing model, see [link to relevant documentation]. Figure 3 , Figure 3 This is a flowchart illustrating the implementation of an intelligent agent game-playing training method provided in this application embodiment. Accordingly, in step S203, the UUV simulated intelligent agent is obtained by game-playing training the rule enhancement model based on CTDE technology, which can be specifically implemented through steps S301-S302:
[0097] S301: Acquire multiple training agents.
[0098] During the game-playing training process, multiple training agents are first acquired. These agents are divided into two categories: one is a full-scale defect training agent, designed to comprehensively challenge the overall weaknesses of the rule-enhancing model, aiming to reveal common shortcomings of the model in various scenarios, thereby promoting the overall performance improvement of the model; the other is a specialized defect training agent with different adversarial styles. These agents are specifically designed to target single or specific weaknesses of the rule-enhancing model, simulating diverse and targeted adversarial strategies to effectively strengthen the model in its weak areas. By introducing these two types of training agents, the training process can cover a wide range of defects while meticulously refining key weaknesses, providing a solid guarantee for the continuous optimization of the rule-enhancing model.
[0099] S302: Based on the CTDE technology, the multiple training agents are included in the adversarial training pool. In the adversarial training pool, the multiple training agents conduct multiple rounds of game training on the rule enhancement model until the training cutoff condition is reached, at which point the game training stops, and the UUV simulation agent is obtained.
[0100] Based on CTDE technology, multiple training agents are integrated into an adversarial training pool. In this environment, a rule-enhancing model is trained through multiple rounds of game theory. During training, the multiple training agents continuously compete against the rule-enhancing model, promoting the continuous improvement of the model's strategy until a preset training cutoff condition is met, at which point training ends, ultimately forming a high-performance UUV simulation agent.
[0101] In each round of training, the centralized controller integrates the observation data and action feedback from all training agents and the rule-enhancing model to construct a complete global state. Based on this global state, reward and penalty signals are calculated to guide the optimization and adjustment of model parameters based on the degree of improvement in weaknesses: a positive reward is given when the rule-enhancing model successfully avoids a weakness, and a negative penalty is imposed otherwise. Through this continuous iteration based on reward and penalty signals, the rule-enhancing model can continuously correct its own defects, achieving robust policy enhancement and a significant improvement in overall capabilities.
[0102] In one possible implementation, the training cutoff condition can be set as, but is not limited to, the weakness coverage of the rule-enhancing model reaching a pre-set standard, or the reward / penalty signal showing a clear convergence trend. Furthermore, it can be judged based on the continuous performance of the mean weakness improvement: when this mean is higher than a set threshold for several consecutive rounds (e.g., 5 to 10 rounds), and the rule-enhancing model's adversarial win rate against all training agents in the adversarial training pool remains stable within the expected range (e.g., 85% to 95%), the model is considered to have reached the ideal performance level, and training is terminated. Through these multi-dimensional judgment criteria, it is possible to effectively ensure that the training process both fully explores and repairs model weaknesses while avoiding overtraining, thereby improving the practical adaptability and robustness of the final UUV simulated agent.
[0103] Steps S301-S302 introduce diverse training agents to specifically detect and improve the shortcomings of the rule-enhancing model. Combined with CTDE technology, global information sharing and distributed execution are achieved, effectively promoting continuous improvement of the model and enhancing its practical adaptability.
[0104] In one possible implementation, the intelligent escort strategy includes pre-set identification information for each UUV and corresponding task assignment instructions; the task assignment instructions are used to guide the various UUVs to cooperate and form an effective escort formation for the protected target.
[0105] Specifically, the intelligent escort strategy pre-sets unique identification information for each UUV and assigns specific task instructions to each identification information. These task instructions clearly guide the responsibilities and actions of each UUV, achieving collaborative cooperation among multiple UUVs through reasonable allocation and coordination. With the help of these task instructions, each UUV can form an effective escort formation around the protected target according to a predetermined plan, thereby enhancing overall protection capabilities and improving the safety of the protected target and the success rate of missions in complex environments.
[0106] In one possible implementation, the process of intelligent escort by each UUV in conjunction with the strategy scene mapping relationship includes: for each UUV, the UUV selects an intelligent escort strategy corresponding to the current scene in the strategy scene mapping relationship based on a virtual scene that matches the current scene to obtain a target escort strategy, and calls corresponding control instructions to execute action planning and collaborative operation according to the division of labor instructions corresponding to the identification information of the target escort strategy and the UUV, so as to achieve escort of the protected target.
[0107] Specifically, each UUV conducts intelligent escort operations based on a strategy-scenario mapping relationship. The process includes: for each UUV, matching a corresponding virtual scenario based on the current operational environment; then selecting the intelligent escort strategy corresponding to that virtual scenario from the strategy-scenario mapping relationship to obtain a specific target escort strategy. Subsequently, the UUV, based on its own identification information and corresponding task instructions, invokes the appropriate control commands to execute action planning and collaborative operations. Through this strategy execution based on dynamic scenario selection and task guidance, effective collaboration among multiple UUVs is achieved, ensuring precise escort of the protected target and maximizing protection capabilities.
[0108] In one possible implementation, this application also provides a method for generating an escort strategy, see [link to relevant documentation]. Figure 4 , Figure 4 This is a flowchart illustrating the implementation of an escort strategy generation method provided in this application embodiment. Accordingly, step S103 generates escort strategies based on the multiple sets of random simulation data, the multiple sets of original simulation data, and the real-time scene state data, resulting in various intelligent escort strategies. Specifically, this can be implemented through steps S401-S402.
[0109] S401: Based on the multiple sets of random simulation data, the multiple sets of original simulation data, and the real-time data of the scene state, generate corresponding escort strategies respectively to obtain multiple initial escort strategies.
[0110] To build a foundational strategy library covering diverse scenarios, corresponding escort strategies can be independently generated based on multiple sets of random simulation data, multiple sets of raw simulation data, and real-time scenario state data. Specifically, analyzing and utilizing random simulation data can capture uncertainties and diversity in the environment; raw simulation data reflects the basic performance and patterns of UUVs, protected targets, and interference sources; and real-time scenario state data reflects the dynamic changes in the current environment. By combining data from these three different sources, corresponding initial escort strategies can be constructed, thereby obtaining multiple targeted and diverse initial solutions, providing a rich strategy foundation for subsequent iterative optimization.
[0111] S402: Based on the multiple initial escort strategies, iterative derivative strategies are generated to obtain multiple derivative escort strategies.
[0112] To effectively cover diverse escort needs in complex environments, iterative derivative strategies can be generated based on multiple initial escort strategies. This iterative generation process dynamically expands the strategy library's coverage, ultimately forming a complete intelligent escort strategy system together with the initial strategies. This iterative process follows a closed-loop logic of "strategy execution - data update - strategy regeneration," with the specific mechanism as follows:
[0113] First, in the target simulation environment, based on the current batch of escort strategies, simulations are run to obtain multiple sets of new original simulation data, which constitute the derived simulation data for the corresponding batch. Then, for these derived simulation data, multiple sets of random derived simulation data are generated through methods such as random perturbation or transformation, increasing the diversity and coverage of the data. Next, based on each derived simulation data and each random derived simulation data, corresponding derived escort strategies are generated, thus forming a new batch of derived escort strategy sets. This process continuously optimizes and enriches the escort strategy by repeatedly acquiring derived simulation data, generating random derived data, and generating derived escort strategies, until a preset iteration termination condition is met. In the first iteration, the escort strategy for the current batch is the multiple initial escort strategies obtained in the previous stage. In subsequent iterations, the derived escort strategies generated in the previous batch are used as input to continuously advance and improve the strategy. Through this iterative derivation mechanism, not only can the adaptability and robustness of the intelligent escort strategy be effectively improved, but it can also better cope with complex and ever-changing simulation environments and real-world application scenarios.
[0114] Steps S401-S402 continuously enrich and optimize the escort strategy library, improving the adaptability and robustness of intelligent escort strategies.
[0115] In one possible implementation, the state data further includes derived simulation data or randomized derived simulation data.
[0116] In one possible implementation, generating multiple virtual scenarios based on the state data corresponding to each of the multiple intelligent escort strategies includes:
[0117] For each intelligent escort strategy, the virtual poses of the protected target, UUV virtual poses and detection data, interference source virtual poses and detection data, and underwater virtual environment parameters contained in the random simulation data, original simulation data, real-time scene state data, derived simulation data, or randomly derived simulation data corresponding to the intelligent escort strategy are extracted. Then, a three-dimensional virtual scene is constructed according to the preset scene modeling rules by combining the virtual poses of the protected target, UUV virtual poses and detection data, interference source virtual poses and detection data, and underwater virtual environment parameters contained in the random simulation data, original simulation data, real-time scene state data, derived simulation data, or randomly derived simulation data corresponding to the intelligent escort strategy.
[0118] Specifically, for each intelligent escort strategy, key elements are first extracted from its corresponding random simulation data, original simulation data, real-time scene state data, derived simulation data, or randomly derived simulation data. These elements include the virtual pose of the protected target, the virtual pose and detection data of the UUV, the virtual pose and detection data of the interference source, and underwater virtual environment parameters. Then, combining this extracted information with pre-defined scene modeling rules, a 3D virtual scene conforming to the actual operating environment is constructed. Through this process, the complex dynamic environment associated with different intelligent escort strategies is realistically and accurately reproduced, providing a solid simulation foundation for subsequent strategy verification and optimization.
[0119] See Figure 5 , Figure 5 This is a schematic diagram of the structure of a UUV simulation intelligent agent provided in an embodiment of this application. Figure 5 As shown, the UUV simulation intelligent agent includes:
[0120] The integrated acquisition unit 501 is used to acquire real-time scene state data and task objectives, and to acquire multiple sets of raw simulation data from the target simulation environment based on the task objectives. The task objectives include assisting the protected target in escaping the tracking and interference of multiple interference sources through the coordinated escort of multiple UUVs, ensuring that the protected target safely reaches the target location. The target simulation environment is a platform that integrates multi-agent dynamic simulation function and underwater physical environment simulation function. One set of raw simulation data includes a set of scene state simulation data. The set of scene state simulation data includes a set of virtual poses of the protected target, a set of virtual poses and a set of virtual detection data for each of the multiple UUVs, a set of virtual poses and a set of virtual detection data for each of the multiple interference sources, and a set of underwater virtual environment parameters.
[0121] The simulation data generation unit 502 is used to randomly generate multiple sets of random simulation data based on the multiple sets of original simulation data;
[0122] The escort strategy generation unit 503 is used to generate escort strategies based on the multiple sets of random simulation data, the multiple sets of original simulation data, and the real-time scene state data, to obtain a variety of intelligent escort strategies; each intelligent escort strategy corresponds to a set of state data, and the state data includes the random simulation data, the original simulation data, or the real-time scene state data.
[0123] The virtual scene generation unit 504 is used to generate multiple virtual scenes based on the state data corresponding to each of the multiple intelligent escort strategies.
[0124] The intelligent escort unit 505 is used to match and associate the various intelligent escort strategies with their respective corresponding virtual scenes to obtain a strategy-scene mapping relationship, and to send the strategy-scene mapping relationship to each UUV so that each UUV can perform intelligent escort in combination with the strategy-scene mapping relationship.
[0125] In one possible implementation, the UUV simulation agent further includes:
[0126] The basic model training unit is used to acquire a training dataset and use a deep reinforcement learning algorithm to perform transfer learning on the agent in combination with the training dataset to obtain a basic escort strategy model; the training dataset includes multiple sets of historical scene state data labeled with escort strategy tags;
[0127] The enhanced model acquisition unit is used to acquire pre-set adversarial behavior rules and input the adversarial behavior rules into the basic escort strategy model to obtain a rule enhancement model; the adversarial behavior rules are behavioral guidelines formulated for the coordinated actions between various UUVs and the adversarial strategies between the protected target and the multiple interference sources.
[0128] The game training unit is used to train the rule enhancement model using CTDE technology to obtain the UUV simulation agent.
[0129] In one possible implementation, the game training unit is specifically used for:
[0130] Multiple training agents are acquired; the multiple training agents include multiple full-scale defect training agents and multiple specialized defect training agents with different adversarial styles; the full-scale defect training agents are designed to target the global weaknesses of the rule enhancement model; the specialized defect training agents are designed to target a single or some specific weaknesses of the rule enhancement model.
[0131] Based on the CTDE technology, the multiple training agents are included in the adversarial training pool. In the adversarial training pool, the multiple training agents conduct multiple rounds of game training on the rule enhancement model until the training cutoff condition is reached, at which point the game training stops, and the UUV simulation agent is obtained.
[0132] In each round of game training, a centralized controller integrates the observation data and action feedback from the multiple training agents and the rule enhancement model to obtain global state information. Based on the global state information, a reward and punishment signal based on the degree of weakness improvement is calculated, and the model parameters of the rule enhancement model are optimized based on the reward and punishment signal. In the reward and punishment signal, the rule enhancement model's successful avoidance of weaknesses is recorded as a positive reward, and the rule enhancement model's failure to avoid weaknesses is recorded as a negative punishment.
[0133] In one possible implementation, the intelligent escort strategy includes pre-set identification information for each UUV and corresponding task assignment instructions; the task assignment instructions are used to guide the various UUVs to cooperate and form an effective escort formation for the protected target.
[0134] In one possible implementation, the process of intelligent escort by each UUV in conjunction with the strategy scene mapping relationship includes: for each UUV, the UUV selects an intelligent escort strategy corresponding to the current scene in the strategy scene mapping relationship based on a virtual scene that matches the current scene to obtain a target escort strategy, and calls corresponding control instructions to execute action planning and collaborative operation according to the division of labor instructions corresponding to the identification information of the target escort strategy and the UUV, so as to achieve escort of the protected target.
[0135] In one possible implementation, the escort strategy generation unit 503 is specifically used for:
[0136] Based on the multiple sets of random simulation data, the multiple sets of original simulation data, and the real-time data of the scene state, corresponding escort strategies are generated respectively, resulting in multiple initial escort strategies.
[0137] Based on the aforementioned multiple initial escort strategies, iterative derivative strategies are generated to obtain multiple derivative escort strategies;
[0138] Wherein, both the initial escort strategy and the derived escort strategy are the intelligent escort strategy;
[0139] The iterative derivation strategy generation process includes: in the target simulation environment, acquiring multiple sets of new original simulation data based on the current batch escort strategy to obtain the corresponding batch of derivation simulation data; randomly generating multiple sets of random derivation simulation data based on the corresponding batch of derivation simulation data; generating a derivation escort strategy based on each derivation simulation data in the corresponding batch, and generating a derivation escort strategy based on each random derivation simulation data in the corresponding batch, to obtain a new batch of derivation escort strategies; repeating the steps of acquiring derivation simulation data, generating random derivation simulation data, and generating derivation escort strategies until the iteration stops when the iteration termination condition is met;
[0140] In the first iteration, the escort strategy for the current batch is the multiple initial escort strategies, and in subsequent iterations, it is the derived escort strategy generated in the previous batch.
[0141] In one possible implementation, the state data further includes derived simulation data or randomized derived simulation data;
[0142] In one possible implementation, the virtual scene generation unit 504 is specifically used for:
[0143] For each intelligent escort strategy, the virtual poses of the protected target, UUV virtual poses and detection data, interference source virtual poses and detection data, and underwater virtual environment parameters contained in the random simulation data, original simulation data, real-time scene state data, derived simulation data, or randomly derived simulation data corresponding to the intelligent escort strategy are extracted. Then, a three-dimensional virtual scene is constructed according to the preset scene modeling rules by combining the virtual poses of the protected target, UUV virtual poses and detection data, interference source virtual poses and detection data, and underwater virtual environment parameters contained in the random simulation data, original simulation data, real-time scene state data, derived simulation data, or randomly derived simulation data corresponding to the intelligent escort strategy.
[0144] In addition, this application embodiment also provides an escort device for an unmanned underwater vehicle, including: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the escort method for the unmanned underwater vehicle as described above.
[0145] In addition, this application embodiment also provides a computer-readable storage medium storing instructions that, when executed on a terminal device, cause the terminal device to perform the escort method for the unmanned underwater vehicle as described above.
[0146] This application utilizes raw simulation data and randomly generated data provided by the target simulation environment, combined with real-time scene state information, to construct a rich data system covering multiple elements such as the protected target, UUVs, interference sources, and the underwater environment. Based on this, various intelligent escort strategies are generated, effectively breaking through the dependence of traditional reinforcement learning on single training data and significantly expanding the applicability of UUVs. Simultaneously, by matching these escort strategies with corresponding virtual scenarios and establishing mapping relationships, UUVs can quickly select the most suitable strategy in the event of sudden environmental changes or unexpected interference, without relying on the repeated trial-and-error process in reinforcement learning. This not only significantly improves the robustness and response speed of UUVs in complex and unknown environments but also effectively solves the problem of formations struggling to quickly adapt to the environment through reasonable morphological adjustments, ensuring the safety and success rate of collaborative escort missions for protected targets.
[0147] The foregoing provides a detailed description of the escort method and related products for an unmanned underwater vehicle (UUV) provided in this application. The various embodiments in the specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. Regarding the UUV simulation intelligent agent disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and relevant parts can be referred to in the method section. It should be noted that those skilled in the art can make several improvements and modifications to this application without departing from the principles of this application, and these improvements and modifications also fall within the protection scope of the claims of this application.
[0148] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0149] It should also be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
Claims
1. A method for escorting vessels based on unmanned underwater vehicles, characterized in that, An intelligent agent for simulating unmanned underwater vehicles (UUVs) is provided, wherein the UUV simulation agent is obtained through multi-strategy fusion training based on the agent. The method includes: The system acquires real-time scene status data and task objectives, and obtains multiple sets of raw simulation data from the target simulation environment based on the task objectives. The task objectives include assisting the protected target in escaping the tracking and interference of multiple interference sources through the coordinated escort of multiple UUVs, ensuring the protected target safely reaches the target location. The target simulation environment is a platform integrating multi-agent dynamic simulation function and underwater physical environment simulation function. One set of raw simulation data includes a set of scene status simulation data. The set of scene status simulation data includes a set of virtual poses of the protected target, a set of virtual poses and a set of virtual detection data for each of the multiple UUVs, a set of virtual poses and a set of virtual detection data for each of the multiple interference sources, and a set of underwater virtual environment parameters. Multiple sets of random simulation data are randomly generated based on the aforementioned multiple sets of original simulation data; Based on the multiple sets of random simulation data, the multiple sets of original simulation data, and the real-time scene state data, escort strategies are generated to obtain a variety of intelligent escort strategies; each intelligent escort strategy corresponds to a set of state data, and the state data includes the random simulation data, the original simulation data, or the real-time scene state data. Based on the state data corresponding to each of the various intelligent escort strategies, multiple virtual scenarios are generated; The various intelligent escort strategies are matched and associated with their respective virtual scenarios to obtain a strategy-scenario mapping relationship. The strategy-scenario mapping relationship is then distributed to each UUV so that each UUV can perform intelligent escort in accordance with the strategy-scenario mapping relationship.
2. The method according to claim 1, characterized in that, The multi-strategy fusion training process of the UUV simulation agent includes: A training dataset is obtained, and a deep reinforcement learning algorithm is used to perform transfer learning on the agent in combination with the training dataset to obtain a basic escort strategy model; the training dataset includes multiple sets of historical scene state data labeled with escort strategy tags; Obtain pre-defined adversarial behavior rules and input them into the basic escort strategy model to obtain a rule enhancement model; the adversarial behavior rules are behavioral guidelines formulated for scenarios with clear solutions in all adversarial scenarios between the protected target and the multiple interference sources. The UUV simulation agent is obtained by performing game-theoretic training on the rule enhancement model based on centralized training and distributed execution CTDE technology. The UUV simulation agent generates escort strategies based on the adversarial behavior rules and the neural network of the UUV simulation agent.
3. The method according to claim 2, characterized in that, The process of training the rule-enhancing model using CTDE technology to obtain the UUV simulation agent includes: Multiple training agents are acquired; the multiple training agents include multiple full-scale defect training agents and multiple specialized defect training agents with different adversarial styles; the full-scale defect training agents are designed to target the global weaknesses of the rule enhancement model; the specialized defect training agents are designed to target the single or partial weaknesses of the rule enhancement model. Based on the CTDE technology, the multiple training agents are included in the adversarial training pool. In the adversarial training pool, the multiple training agents conduct multiple rounds of game training on the rule enhancement model until the training cutoff condition is reached, at which point the game training stops, and the UUV simulation agent is obtained. In each round of game training, a centralized controller integrates the observation data and action feedback from the multiple training agents and the rule enhancement model to obtain global state information. Based on the global state information, a reward and punishment signal based on the degree of weakness improvement is calculated, and the model parameters of the rule enhancement model are optimized based on the reward and punishment signal. In the reward and punishment signal, the rule enhancement model's successful avoidance of weaknesses is recorded as a positive reward, and the rule enhancement model's failure to avoid weaknesses is recorded as a negative punishment.
4. The method according to claim 1, characterized in that, The intelligent escort strategy includes pre-set identification information for each UUV and corresponding task assignment instructions. The task assignment instructions are used to guide the various UUVs to cooperate and form an effective escort formation for the protected target. The process of intelligent escort by each UUV in combination with the strategy scene mapping relationship includes: for each UUV, the UUV selects the intelligent escort strategy corresponding to the current scene in the strategy scene mapping relationship based on the virtual scene matching the current scene to obtain the target escort strategy, and calls the corresponding control command to execute action planning and collaborative operation according to the division of labor instruction corresponding to the identification information of the target escort strategy and the identification information of the UUV, so as to realize the escort of the protected target.
5. The method according to claim 1, characterized in that, The escort strategy is generated based on the multiple sets of random simulation data, the multiple sets of original simulation data, and the real-time scenario state data, resulting in various intelligent escort strategies, including: Based on the multiple sets of random simulation data, the multiple sets of original simulation data, and the real-time data of the scene state, corresponding escort strategies are generated respectively, resulting in multiple initial escort strategies. Based on the aforementioned multiple initial escort strategies, iterative derivative strategies are generated to obtain multiple derivative escort strategies; Wherein, both the initial escort strategy and the derived escort strategy are the intelligent escort strategy; The iterative derivation strategy generation process includes: in the target simulation environment, acquiring multiple sets of new original simulation data based on the current batch escort strategy to obtain the corresponding batch of derivation simulation data; randomly generating multiple sets of random derivation simulation data based on the corresponding batch of derivation simulation data; generating a derivation escort strategy based on each derivation simulation data in the corresponding batch, and generating a derivation escort strategy based on each random derivation simulation data in the corresponding batch, to obtain a new batch of derivation escort strategies; repeating the steps of acquiring derivation simulation data, generating random derivation simulation data, and generating derivation escort strategies until the iteration stops when the iteration termination condition is met; In the first iteration, the escort strategy for the current batch is the multiple initial escort strategies, and in subsequent iterations, it is the derived escort strategy generated in the previous batch.
6. The method according to claim 5, characterized in that, The state data also includes derived simulation data or randomly derived simulation data; The process involves generating multiple virtual scenarios based on the state data corresponding to each of the various intelligent escort strategies, including: For each intelligent escort strategy, the virtual poses of the protected target, UUV virtual poses and detection data, interference source virtual poses and detection data, and underwater virtual environment parameters contained in the random simulation data, original simulation data, real-time scene state data, derived simulation data, or randomly derived simulation data corresponding to the intelligent escort strategy are extracted. Then, a three-dimensional virtual scene is constructed according to the preset scene modeling rules by combining the virtual poses of the protected target, UUV virtual poses and detection data, interference source virtual poses and detection data, and underwater virtual environment parameters contained in the random simulation data, original simulation data, real-time scene state data, derived simulation data, or randomly derived simulation data corresponding to the intelligent escort strategy.
7. A UUV simulation intelligent agent, characterized in that, The UUV simulation intelligent agent includes: The integrated acquisition unit is used to acquire real-time scene state data and task objectives, and to acquire multiple sets of raw simulation data from the target simulation environment based on the task objectives. The task objectives include assisting the protected target in escaping the tracking and interference of multiple interference sources through the coordinated escort of multiple UUVs, ensuring that the protected target safely reaches the target location. The target simulation environment is a platform that integrates multi-agent dynamic simulation function and underwater physical environment simulation function. One set of raw simulation data includes a set of scene state simulation data. The set of scene state simulation data includes a set of virtual poses of the protected target, a set of virtual poses and a set of virtual detection data for each of the multiple UUVs, a set of virtual poses and a set of virtual detection data for each of the multiple interference sources, and a set of underwater virtual environment parameters. The simulation data generation unit is used to randomly generate multiple sets of random simulation data based on the multiple sets of original simulation data. The escort strategy generation unit is used to generate escort strategies based on the multiple sets of random simulation data, the multiple sets of original simulation data, and the real-time scenario state data, to obtain a variety of intelligent escort strategies; each intelligent escort strategy corresponds to a set of state data, and the state data includes the random simulation data, the original simulation data, or the real-time scenario state data. The virtual scene generation unit is used to generate multiple virtual scenes based on the state data corresponding to each of the various intelligent escort strategies. The intelligent escort unit is used to match and associate the various intelligent escort strategies with their respective corresponding virtual scenarios to obtain a strategy-scenario mapping relationship, and then distribute the strategy-scenario mapping relationship to each UUV so that each UUV can perform intelligent escort in combination with the strategy-scenario mapping relationship.
8. The UUV simulation intelligent agent according to claim 7, characterized in that, The UUV simulation intelligent agent also includes: The basic model training unit is used to acquire a training dataset and use a deep reinforcement learning algorithm to perform transfer learning on the agent in combination with the training dataset to obtain a basic escort strategy model; the training dataset includes multiple sets of historical scene state data labeled with escort strategy tags; The enhanced model acquisition unit is used to acquire pre-set adversarial behavior rules and input the adversarial behavior rules into the basic escort strategy model to obtain a rule enhancement model; the adversarial behavior rules are behavioral guidelines formulated for the coordinated actions between various UUVs and the adversarial strategies between the protected target and the multiple interference sources. The game training unit is used to train the rule enhancement model using CTDE technology to obtain the UUV simulation agent.
9. An escort device for an unmanned underwater vehicle, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the escort method for an unmanned underwater vehicle as described in any one of claims 1-6.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed on a terminal device, cause the terminal device to perform the escort method for an unmanned underwater vehicle as described in any one of claims 1-6.
Citation Information
Patent Citations
Unmanned underwater vehicle formation control simulation method, system and equipment
CN116841208A
Unmanned aerial vehicle cluster confrontation strategy acquisition method based on reinforcement learning and related equipment
CN117572893A