Cooperative hunting method and device for multiple underwater vehicles and computer equipment
By building submarine and target body models, establishing a hydroacoustic channel and hidden communication model, and optimizing the coordinated roundup path of multi-submarines, the adaptability and communication concealment of coordinated roundup of multi-submarines in complex environments is solved, and the roundup efficiency and success rate are improved.
Patent Information
- Application Number
- CN202510275951.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-07-29
AI Technical Summary
The existing multi-submarine collaborative roundup technology is insufficient in adaptability and robustness in complex dynamic environments, and lacks system design for communication concealment and collaboration, resulting in low roundup efficiency.
By obtaining the operation information of the submarine and target body, constructing the submarine and target body model, establishing a water acoustic channel model and a hidden communication model, generating a coordinated roundup path, and optimizing trajectory planning to improve the efficiency of coordinated roundup.
Under the satisfaction of hidden communication constraints, the convergence speed and success rate of multi-submarine coordinated roundup are improved, and the efficiency of coordinated roundup in complex task scenarios is improved.
Smart Images

Figure CN120386371A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of artificial intelligence and multi-agent collaboration technologies, and particularly to a collaborative hunting method, device, and computer equipment for multiple underwater vehicles. Background Art
[0002] The collaborative hunting of autonomous underwater vehicles (AUVs) is one of the key tasks in unmanned swarm intelligence collaboration. Traditional rule-based control methods have limited adaptability and robustness in complex dynamic environments, while multi-agent reinforcement learning (MARL) improves decision-making efficiency and task success rate through adaptive learning. However, existing methods have insufficient capabilities in coordinating multiple agents and modeling dynamic environments, and it is difficult to solve the high data requirements and communication concealment problems of online reinforcement learning. Therefore, how to improve the collaborative hunting efficiency and hunting accuracy of multiple underwater vehicles is the current research focus.
[0003] In the target hunting task, the intelligent collaboration of multiple AUVs requires efficient trajectory planning and dynamic coordination. However, existing technologies still face some challenges in practical applications. First, some online reinforcement learning-based methods highly rely on interactions with the real-time environment, with low data utilization. In complex dynamic environments, such as in the presence of obstacles and water flow interference, training is extremely unstable and the actual training cost is too high. Second, although some technologies improve collaboration efficiency by optimizing the trajectory planning of multiple agents, they do not fully consider the possible eavesdropping capabilities of the target in the actual scenario. Once the communication of the hunter AUV is intercepted, the target can use the obtained information to adjust its escape strategy, thus significantly reducing the success rate of the hunting task. In addition, existing multi-agent frameworks lack a systematic design between communication concealment and collaboration, and cannot maintain high task performance while taking into account the need for covert communication. These problems make the adaptability and effectiveness of existing technologies insufficient in actual complex task scenarios, resulting in low efficiency of multi-agent collaborative hunting in complex task scenarios. Summary of the Invention
[0004] Based on this, it is necessary to provide a collaborative hunting method, device, and computer equipment for multiple underwater vehicles to address the above technical problems.
[0005] In a first aspect, this application provides a collaborative hunting method for multiple underwater vehicles, including:
[0006] Obtain the operation information of multiple underwater vehicles, the operation information of the target object, the transmission channel information of each underwater vehicle, and the underwater environment information, and construct the underwater vehicle model of each underwater vehicle and the target object model of the target object based on the operation information of each underwater vehicle and the operation information of the target object;
[0007] Construct an underwater acoustic channel model based on the underwater environment information, and construct a covert communication model between each underwater vehicle through a covert communication modeling strategy based on the transmission channel information of each underwater vehicle;
[0008] Collect the observation information of each underwater vehicle and the action information of each underwater vehicle, and generate the cooperative encirclement path of each underwater vehicle through each underwater vehicle model, the target object model, the underwater acoustic channel model, and the covert communication model based on the observation information of each underwater vehicle and the action information of each underwater vehicle;
[0009] Take the cooperative encirclement path of each underwater vehicle as the target cooperative encirclement plan for multiple underwater vehicles.
[0010] Optionally, the constructing the underwater vehicle model of each underwater vehicle and the target object model of the target object based on the operation information of each underwater vehicle and the operation information of the target object includes:
[0011] For each underwater vehicle, identify the condition information of each motion condition of the underwater vehicle based on the operation information of the underwater vehicle;
[0012] Construct the underwater vehicle model of the underwater vehicle through a kinematic modeling strategy based on the condition information of each motion condition;
[0013] Replace the operation information of the underwater vehicle with the operation information of the target object, and return to execute the step of identifying the condition information of each motion condition of the underwater vehicle based on the operation information of the underwater vehicle to obtain the target object model of the target object.
[0014] Optionally, the constructing the underwater acoustic channel model based on the underwater environment information includes:
[0015] Split the underwater environment information into underwater propagation information and underwater noise information, and identify underwater propagation characteristic information based on the underwater propagation information;
[0016] Generate the underwater propagation loss formula through a path loss algorithm based on the underwater propagation characteristic information;
[0017] Identify the noise eigenvalue of each underwater noise angle based on the underwater noise information, and identify the total noise level formula through the noise level algorithm corresponding to each noise angle;
[0018] Based on the underwater propagation loss formula and the noise level formula, calculate the acoustic path loss at each time step and the underwater noise power at each time step, and use the acoustic path loss at each time step and the underwater noise power at each time step as the underwater acoustic channel model.
[0019] Optionally, based on the transmission channel information of each submersible, constructing a covert communication model between each submersible through a covert communication modeling strategy includes:
[0020] Based on the transmission channel information of each submersible, generate the signal detection information of each submersible of the target at each time step through a signal detection strategy;
[0021] Based on the signal detection information of each submersible at each time step, identify the optimal threshold power of the target, and based on the optimal threshold power and the preset constraint conditions for the change of the communication state of the submersible, construct a covertness constraint equation between each submersible;
[0022] Use the covertness constraint equation between each submersible as the covert communication model between each submersible.
[0023] Optionally, after collecting the observation information of each submersible and the action information of each submersible, it further includes:
[0024] Based on the observation information of each submersible, identify the detection content of each detection condition type detected by each submersible, and based on each detection content, construct and generate the state space of each submersible, and based on the action information of each submersible, generate the action space of each submersible through a kinematic algorithm;
[0025] Based on the action space of each submersible and the observation information of each submersible, identify the state transition probability information of each submersible, and collect the reward condition information of each reward type;
[0026] Based on the reward condition information of each reward type, generate the reward constraint function of each reward type through the reward constraint strategy of each reward type, and based on all reward constraint functions, generate the comprehensive reward function of each submersible;
[0027] Based on the state space of each submersible, the action space of each submersible, the state transition probability information of each submersible, and the comprehensive reward function of each submersible, construct the encirclement modeling information of each submersible through an encirclement decision generation strategy.
[0028] Optionally, generating the cooperative hunting paths of the submarines based on the observation information of each submarine and the action information of each submarine through each submarine model, the target model, the underwater acoustic channel model, and the covert communication model includes:
[0029] Based on the underwater acoustic channel model and the covert communication model, generating the initial hunting trajectories of each submarine at each time step through the noise prediction network and the hunting modeling information of each submarine;
[0030] Based on each submarine model and the target model, adjusting the initial hunting trajectories of each submarine at each time step through the inverse dynamics model to obtain the optimized hunting trajectories of each submarine at each time step;
[0031] Taking the optimized hunting trajectories of each submarine at each time step as the cooperative hunting paths of the submarines.
[0032] In a second aspect, the present application also provides a cooperative hunting device for multiple submarines, including:
[0033] An acquisition module, configured to acquire the operation information of multiple submarines, the operation information of the target, the transmission channel information of each submarine, and the underwater environment information, and construct the submarine model of each submarine and the target model of the target based on the operation information of each submarine and the operation information of the target;
[0034] A construction module, configured to construct an underwater acoustic channel model based on the underwater environment information, and construct a covert communication model between the submarines through a covert communication modeling strategy based on the transmission channel information of each submarine;
[0035] A generation module, configured to collect the observation information of each submarine and the action information of each submarine, and generate the cooperative hunting paths of the submarines through each submarine model, the target model, the underwater acoustic channel model, and the covert communication model based on the observation information of each submarine and the action information of each submarine;
[0036] A determination module, configured to take the cooperative hunting paths of the submarines as the target cooperative hunting scheme of the multiple submarines.
[0037] Optionally, the acquisition module is specifically configured to:
[0038] For each submarine, based on the operation information of the submarine, identify the condition information of each motion condition of the submarine;
[0039] Based on the condition information of each motion condition, construct the submarine model of the submarine through a kinematic modeling strategy;
[0040] Replace the operation information of the submersible with the operation information of the target object, and return to execute the step of identifying the condition information of each motion condition of the submersible based on the operation information of the submersible, to obtain the target object model of the target object.
[0041] Optionally, the construction module is specifically configured to:
[0042] Split the underwater environment information into underwater propagation information and underwater noise information, and identify underwater propagation feature information based on the underwater propagation information;
[0043] Generate the underwater propagation loss formula based on the underwater propagation feature information through the path loss algorithm;
[0044] Identify the noise eigenvalue of each underwater noise angle based on the underwater noise information, and identify the total noise level formula through the noise level algorithm corresponding to each noise angle;
[0045] Calculate the acoustic path loss of each time step and the underwater noise power of each time step based on the underwater propagation loss formula and the noise level formula, and use the acoustic path loss of each time step and the underwater noise power of each time step as the underwater acoustic channel model.
[0046] Optionally, the construction module is specifically configured to:
[0047] Generate the signal detection information of each submersible of the target object at each time step based on the transmission channel information of each submersible through a signal detection strategy;
[0048] Identify the optimal threshold power of the target object based on the signal detection information of each submersible at each time step, and construct a concealment constraint equation between each submersible based on the optimal threshold power and a preset submersible communication state change constraint condition;
[0049] Use the concealment constraint equation between each submersible as the covert communication model between each submersible.
[0050] Optionally, the device further includes:
[0051] A space generation module, configured to identify the detection content of each detection condition type detected by each submersible based on the observation information of each submersible, construct and generate the state space of each submersible based on each detection content, and generate the action space of each submersible through a kinematic algorithm based on the action information of each submersible.
[0052] The acquisition module is used to identify the state transition probability information of each submersible based on the action space of each submersible and the observation information of each submersible, and collect the reward condition information of each reward type;
[0053] The function generation module is used to generate the reward constraint functions of each reward type through the reward constraint strategies of each reward type based on the reward condition information of each reward type, and generate the comprehensive reward function of each submersible based on all the reward constraint functions;
[0054] The modeling module is used to construct the encirclement and capture modeling information of each submersible through the encirclement and capture decision-making generation strategy based on the state space of each submersible, the action space of each submersible, the state transition probability information of each submersible, and the comprehensive reward function of each submersible.
[0055] Optionally, the generation module is specifically used for:
[0056] Based on the underwater acoustic channel model and the covert communication model, generate the initial encirclement and capture trajectories of each submersible at each time step through the noise prediction network and the encirclement and capture modeling information of each submersible;
[0057] Based on each submersible model and the target model, adjust the initial encirclement and capture trajectories of each submersible at each time step through the inverse dynamics model to obtain the optimized encirclement and capture trajectories of each submersible at each time step;
[0058] Use the optimized encirclement and capture trajectories of each submersible at each time step as the collaborative encirclement and capture paths of each submersible.
[0059] In a third aspect, the present application provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the method described in any one of the first aspects are implemented.
[0060] In a fourth aspect, the present application provides a computer-readable storage medium. A computer program is stored thereon, and when the computer program is executed by a processor, the steps of the method described in any one of the first aspects are implemented.
[0061] In a fifth aspect, the present application provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps of the method described in any one of the first aspects are implemented.
[0062] The above-mentioned cooperative encirclement method, device and computer equipment for multiple submersibles obtain the operation information of multiple submersibles, the operation information of the target, the transmission channel information of each submersible, and the underwater environment information, and construct a submersible model for each submersible and a target model for the target based on the operation information of each submersible and the operation information of the target; construct an underwater acoustic channel model based on the underwater environment information, and construct a covert communication model between each submersible through a covert communication modeling strategy based on the transmission channel information of each submersible; collect the observation information of each submersible and the action information of each submersible, and generate a cooperative encirclement path for each submersible through each submersible model, the target model, the underwater acoustic channel model, and the covert communication model based on the observation information of each submersible and the action information of each submersible; use the cooperative encirclement path of each submersible as the target cooperative encirclement plan for multiple submersibles. This solution models the submersibles, the target, the underwater environment, and the shielding communication process, thereby not only improving the comprehensiveness and accuracy of the research on the underwater cooperative encirclement of multiple submersibles, but also avoiding the problems of target escape and encirclement failure caused by target eavesdropping and signal monitoring. Then, in the actual encirclement process, this solution improves the convergence speed and encirclement success rate of multi-agent cooperative encirclement, and improves the adaptive encirclement success rate of multiple submersibles by means of multi-agent cooperative encirclement and covert information communication while satisfying the covert communication constraint conditions, thereby comprehensively improving the efficiency of multi-agent cooperative encirclement in complex task scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0064] Figure 1 It is a schematic flowchart of the cooperative encirclement method for multiple submersibles in an embodiment;
[0065] Figure 2 It is a schematic flowchart of the processing flow of the AMADP network architecture in an embodiment;
[0066] Figure 3 It is a parameter list of the AMADP algorithm in an embodiment;
[0067] Figure 4 It is a schematic diagram of a scenario where the target is successfully surrounded using the AMADP algorithm in an embodiment;
[0068] Figure 5 The numerical distribution diagram of the values for each time step in an embodiment;
[0069] Figure 6 The schematic comparison diagram of AMADP and the latest offline multi-agent reinforcement learning (MARL) algorithm in an embodiment;
[0070] Figure 7 The schematic flow diagram of the cooperative hunting example of multiple underwater vehicles in an embodiment;
[0071] Figure 8 The structural block diagram of the cooperative hunting device of multiple underwater vehicles in an embodiment;
[0072] Figure 9 The internal structure diagram of a computer device in an embodiment. Detailed implementation manners
[0073] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.
[0074] The cooperative hunting method for multiple underwater vehicles provided by the embodiments of the present application can be applied to the application environment of the cooperative hunting of multiple underwater vehicles. This method can be applied to a terminal, a server, or a system including a terminal and a server, and is realized through the interaction between the terminal and the server. Among them, the terminal can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, etc. Among them, the terminal models the underwater vehicle, the target, the underwater environment, and the shielding communication process, thereby not only improving the comprehensiveness and accuracy of the research on the underwater cooperative hunting of multiple underwater vehicles, but also avoiding the problems of target escape and hunting failure caused by target eavesdropping and signal monitoring. Then, in the actual hunting process, this solution uses the method of multi-agent cooperative hunting and covert information communication to improve the convergence speed and hunting success rate of multi-agent cooperative hunting while meeting the constraints of covert communication, and improves the adaptive hunting success rate of multiple underwater vehicles, thereby comprehensively improving the efficiency of multi-agent cooperative hunting in complex task scenarios.
[0075] [[ID=3`1]]In an exemplary embodiment, as Figure 1 shown, a cooperative hunting method for multiple underwater vehicles is provided. Taking this method applied to a terminal as an example, it includes the following steps S101 to S104. Among them:
[0076] Step S101: Obtain the operation information of multiple underwater vehicles, the operation information of the target object, the transmission channel information of each underwater vehicle, and the underwater environment information, and based on the operation information of each underwater vehicle and the operation information of the target object, construct the underwater vehicle model of each underwater vehicle and the target object model of the target object.
[0077] In this embodiment, in response to the information uploading operation of the staff, the terminal obtains the operation information of each underwater vehicle and the operation information of the target object. Here, both the underwater vehicle and the target object are autonomous underwater vehicles (AUVs). Among them, the above operation information includes various motion condition information, for example, surge velocity, sway velocity, heave velocity, rotation angle, and also includes the inertia matrix including added mass, the Coriolis centrifugal force matrix, and the damping matrix. Then, the terminal obtains the transmission channel information of each underwater vehicle. Among them, each transmission channel information includes information such as the channel transmission signal between each underwater vehicle (or hunter AUV). Then, the terminal obtains the underwater environment information related to the underwater environment. This underwater environment information includes underwater propagation information and underwater noise information. Among them, the underwater propagation information is used to characterize the diffusion factor, communication frequency, and communication distance of the geometric characteristics in the shallow water communication process. And the underwater noise information includes turbulent noise, shipping noise, wind-driven wave noise, and thermal noise, etc. Finally, the terminal constructs the underwater vehicle model of each underwater vehicle and the target object model of the target object based on the operation information of each underwater vehicle and the operation information of the target object. Among them, the above underwater vehicle model and target object model are both dynamic models and can be expressed as a simplified three-degree-of-freedom model, which describes the motion of the AUV in the horizontal plane. The specific construction process will be described in detail later.
[0078] Step S102: Based on the underwater environment information, construct an underwater acoustic channel model, and based on the transmission channel information of each underwater vehicle, through a covert communication modeling strategy, construct a covert communication model between each underwater vehicle.
[0079] In this embodiment, the terminal constructs an underwater acoustic channel model based on the underwater environment information, and based on the transmission channel information of each underwater vehicle, through a covert communication modeling strategy, constructs a covert communication model between each underwater vehicle. Among them, this underwater acoustic channel model is a model used to characterize the acoustic path loss and underwater noise power at different time steps. The specific construction process will be described in detail later. And the covert communication model is a communication model between each underwater vehicle corresponding to the covert constraint conditions under the signal detection conditions of the target object. The specific construction process will be described in detail later.
[0080] Step S103: Collect the observation information of each submersible and the action information of each submersible, and generate the cooperative encirclement paths of each submersible based on the observation information of each submersible and the action information of each submersible through each submersible model, target model, underwater acoustic channel model, and covert communication model.
[0081] In this embodiment, the terminal collects the observation information of each submersible and the action information of each submersible, and generates the cooperative encirclement paths of each submersible based on the observation information of each submersible and the action information of each submersible through each submersible model, target model, underwater acoustic channel model, and covert communication model. The specific generation process will be described in detail later.
[0082] Step S104: Use the cooperative encirclement paths of each submersible as the target cooperative encirclement scheme for multiple submersibles.
[0083] In this embodiment, the terminal uses the cooperative encirclement paths of each submersible as the target cooperative encirclement scheme for multiple submersibles.
[0084] Based on the above scheme, by modeling the submersibles, the target, the underwater environment, and the shielding communication process, not only the comprehensiveness and accuracy of the research on the underwater cooperative encirclement of multiple submersibles are improved, but also the problems of the target escaping and the encirclement failing due to the target being able to eavesdrop and monitor signals are avoided. Then, in the actual encirclement process, this scheme improves the convergence speed and encirclement success rate of the multi-agent cooperative encirclement, and improves the adaptive encirclement success rate of multiple submersibles by means of multi-agent cooperative encirclement and covert information communication while meeting the constraints of covert communication, thus comprehensively improving the efficiency of multi-agent cooperative encirclement in complex mission scenarios.
[0085] Optionally, based on the operation information of each submersible and the operation information of the target, construct the submersible model of each submersible and the target model of the target, including: for each submersible, based on the operation information of the submersible, identify the condition information of each motion condition of the submersible; based on the condition information of each motion condition, construct the submersible model of the submersible through the kinematic modeling strategy; replace the operation information of the submersible with the operation information of the target, and return to execute the step of identifying the condition information of each motion condition of the submersible based on the operation information of the submersible to obtain the target model of the target.
[0086] In this embodiment, the terminal, for each submersible, based on the operation information of the submersible, identifies the condition information of each motion condition of the submersible. Among them, each motion condition includes but is not limited to surge speed, sway speed, heave speed, and rotation angle.
[0087] Based on the condition information of each motion condition, the terminal constructs a submersible model of the submersible through a kinematic modeling strategy. Finally, the terminal replaces the operation information of the submersible with the operation information of the target body, returns to execute the step of identifying the condition information of each motion condition of the submersible, and obtains the target body model of the target body. That is, the terminal re-models the target body in the same way as modeling each submersible. Among them, both the target body model and the submersible model are dynamic models.
[0088] Specifically, the terminal is fixed to the body coordinate system and the earth reference system where and represent the surge velocity, sway velocity, heave velocity, and rotation angle respectively. Constrained by the maximum velocity to satisfy , where is the horizontal axis moving distance, is the vertical axis moving distance. Similarly, the terminal assumes that the velocity of the target is restricted by , that is . The hunter AUV (i.e., the submersible) and the target (i.e., the target body) use the same dynamic model, which can be expressed as:
[0089]
[0090] where represent the inertia matrix including added mass, the Coriolis centrifugal force matrix, and the damping matrix respectively. represents the combined matrix of gravity and buoyancy. The control input and environmental disturbance are represented by and respectively. In addition is the transformation matrix, which can be expressed as:
[0091]
[0092] In the underwater target hunting task, the hunter AUV shows better hunting ability through cooperation. Therefore, the accelerations of the submersible and the target body satisfy .
[0093] Based on the above scheme, by constructing the kinematic models of each submersible and the target body respectively, the comprehensiveness and accuracy of the analysis of each submersible and the target body are improved.
[0094] Optionally, based on the underwater environment information, an underwater acoustic channel model is constructed, including: splitting the underwater environment information into underwater propagation information and underwater noise information, and identifying underwater propagation characteristic information based on the underwater propagation information; generating an underwater propagation loss formula through a path loss algorithm based on the underwater propagation characteristic information; identifying the noise eigenvalue at each underwater noise angle based on the underwater noise information, and identifying the total noise level formula through the noise level algorithm corresponding to each noise angle; calculating the acoustic path loss at each time step and the underwater noise power at each time step based on the underwater propagation loss formula and the noise level formula, and taking the acoustic path loss at each time step and the underwater noise power at each time step as the underwater acoustic channel model.
[0095] In this embodiment, the terminal splits the underwater environment information into underwater propagation information and underwater noise information, and identifies underwater propagation characteristic information based on the underwater propagation information; generates an underwater propagation loss formula through a path loss algorithm based on the underwater propagation characteristic information. Among them, the path loss formula in the shallow water communication environment is formulated as:
[0096]
[0097] Among them The spreading factor representing the propagation geometric characteristics, is the communication frequency, is the distance. represents the attenuation coefficient per kilometer at kHz frequency.
[0098] The optimized target path loss formula is given by Thorp's empirical formula:
[0099]
[0100] Then, the terminal identifies the noise eigenvalue at each underwater noise angle based on the underwater noise information, and identifies the total noise level formula through the noise level algorithm corresponding to each noise angle. Among them, each noise angle includes but is not limited to angles such as turbulent noise, shipping noise, wind-driven wave noise, and thermal noise.
[0101] Then, the terminal calculates the acoustic path loss at each time step and the underwater noise power at each time step based on the underwater propagation loss formula and the noise level formula, and takes the acoustic path loss at each time step and the underwater noise power at each time step as the underwater acoustic channel model.
[0102] Among them, the underwater environmental noise includes turbulent noise , shipping noise , wind-driven wave noise and thermal noise According to the underwater noise model, the total noise level can be expressed as:
[0103]
[0104] where s and w represent the shipping activity factor and the wind speed respectively. These noise sources are modeled as Gaussian processes, and the total underwater ambient noise power spectral density is given by in dB relative to per Hz.
[0105] where the acoustic path loss and the underwater noise power at the th time step are represented as and respectively. The terminal takes the target path loss formula and the formula corresponding to the total underwater ambient noise power spectral density as the underwater acoustic channel model.
[0106] Based on the above scheme, by analyzing the noise characteristics and path loss, the comprehensiveness and accuracy of the communication signal analysis for underwater communication are improved.
[0107] Optionally, based on the transmission channel information of each submersible vehicle, a covert communication model between each submersible vehicle is constructed through a covert communication modeling strategy, including: based on the transmission channel information of each submersible vehicle, through a signal detection strategy, generating the signal detection information of each submersible vehicle for the target object at each time step; based on the signal detection information of each submersible vehicle at each time step, identifying the optimal threshold power of the target object, and based on the optimal threshold power and the preset constraint conditions for the change of the communication state of the submersible vehicle, constructing a covertness constraint equation between each submersible vehicle; taking the covertness constraint equation between each submersible vehicle as the covert communication model between each submersible vehicle.
[0108] In this embodiment, the terminal generates the signal detection information of each submersible vehicle for the target object at each time step based on the transmission channel information of each submersible vehicle through a signal detection strategy. Specifically, it is assumed that the channel state is constant within each time step, but independently changes from one time step to the next. To describe the underwater time-varying channel state, the terminal uses a block fading channel model, considering each channel to be independent. The hunter AUV transmits signals through multiple channels during the communication process, expressed as where L represents the number of channel uses. At the same time, the surrounded target detects whether the hunter AUVs are communicating by monitoring the signal strength. The signal received by the target on the l-th channel within the th time step (i.e., the signal detection information of each submersible vehicle for the target object at each time step) is:
[0109]
[0110] where it is assumed that Indicates that communication is taking place between hunter AUVs, while Indicates no communication. And Respectively represent the Transmission power and transmitted signal on the Channel at the Time step. In addition, Where Represents the total Gaussian noise observed by the target within the t-th time step,
[0111] In this case, the target body infers the communication state between hunter AUVs by measuring the received signal power and applying a hypothesis testing strategy to assist its decision-making in the encirclement mission. Specifically, to detect the existence of covert communication, the target body needs to determine whether one hunter AUV is sending information to other hunters (i.e., the identification method for the target body to judge whether there is communication between submersibles). During the encirclement process, the target body uses the likelihood ratio test (LRT) as the optimal detection method to determine whether the hunter AUVs are communicating while minimizing its detection error. This process can be simplified as:
[0112]
[0113] Where Represents the average signal power received by the target, Is the total power received by the target over the entire block duration, Is the Threshold at the Time step. If the received power is below the threshold, the target tends to assume ; Otherwise, it tends to
[0114] Then, the terminal identifies the optimal threshold power of the target body based on the signal detection information of each submersible at each time step. Specifically, assuming that the target knows the transmission power And the noise power Of the hunter AUV, then the optimal threshold At the target (i.e., the optimal threshold power) can be expressed as:
[0115]
[0116] Where Is equal to .
[0117] Finally, the terminal constructs the covertness constraint equation between each submersible based on the optimal threshold power and the preset constraint conditions for the change of the communication state of the submersible. Among them, let Represent the false alarm probability, that is, when the actual state is When the target tends to assume probability. Similarly, represents the probability of missed detection, that is, when the actual state is When the target tends to assume probability. To ensure covert communication, the following constraints must be satisfied:
[0118]
[0119] where represents the acceptable level of concealment. However, directly calculating is challenging. For this reason, the terminal can transform the constraint equation into:
[0120]
[0121] Hypothesis and The total divergence between the probability distributions under is subject to the following constraint:
[0122]
[0123] where represents the error tolerance, is the Kullback-Leibler (KL) divergence of the time distribution under different hypotheses. and represent the probability densities under hypotheses and respectively.
[0124] The terminal assumes that the transmission power between the hunter AUVs and the communication frequency as well as the jammer power against the target remain unchanged. Therefore, the covert communication constraint only depends on the distance between the hunter and the target. The goal of the terminal is to optimize the trajectory of the hunter formation while satisfying the covert communication constraint to complete the encirclement mission.
[0125] Finally, the terminal optimizes the above constraint information to obtain (i.e., the covertness constraint equation between the submersibles):
[0126]
[0127] where represents the probability of successful encirclement when all hunter AUVs are within the target capture range . Enforce the covert communication constraint, and are the maximum speeds of the hunter and the target respectively, is the minimum safety distance required to avoid collisions.
[0128] Finally, the terminal uses the concealment constraint equation between each submersible as the covert communication model between each submersible.
[0129] Based on the above solution, by considering the information leakage problem existing in the process of submersible hunting, a collaborative hunting framework based on covert communication technology is proposed. The communication security protection effect of the obtained shielding communication model between each submersible on the target and the anti-signal acquisition effect are improved.
[0130] Optionally, after collecting the observation information of each submersible and the action information of each submersible, it further includes: based on the observation information of each submersible, identifying the detection content of each detection condition type detected by each submersible, and based on each detection content, constructing and generating the state space of each submersible, and based on the action information of each submersible, through the kinematic algorithm, generating the action space of each submersible; based on the action space of each submersible and the observation information of each submersible, identifying the state transition probability information of each submersible, and collecting the reward condition information of each reward type; based on the reward condition information of each reward type, through the reward constraint strategy of each reward type, generating the reward constraint function of each reward type, and based on all reward constraint functions, generating the comprehensive reward function of each submersible; based on the state space of each submersible, the action space of each submersible, the state transition probability information of each submersible, and the comprehensive reward function of each submersible, through the hunting decision generation strategy, constructing the hunting modeling information of each submersible.
[0131] In this embodiment, the terminal, based on the observation information of each submersible, identifies the detection content of each detection condition type detected by each submersible, and based on each detection content, constructs and generates the state space of each submersible, and based on the action information of each submersible, through the kinematic algorithm, generates the action space of each submersible.
[0132] Among them, the state space: The observation information of each agent can be defined as:
[0133]
[0134] where and respectively represent the speed and position information of agent i, represents the position of the obstacle, represents the position of other agents.
[0135] Action space: The action of each agent includes the moving direction and speed , which is expressed as:
[0136]
[0137] Based on the action space of each submersible and the observation information of each submersible, the terminal identifies the state transition probability information of each submersible and collects the reward condition information of each reward type.
[0138] Then, based on the reward condition information of each reward type, the terminal generates the reward constraint functions of each reward type through the reward constraint strategies of each reward type, and generates the comprehensive reward function of each submersible based on all the reward constraint functions.
[0139] Reward function: In order to achieve effective encirclement and meet specific communication constraints under dynamic and uncertain underwater conditions, the reward function is designed to encourage the hunter AUV to form an optimal encirclement around the target while maintaining specific communication constraints.
[0140] a. Encirclement reward: This reward is used to encourage the hunter to form a uniformly distributed encirclement around the target.
[0141]
[0142] b. Collision avoidance reward: This reward punishes collisions between hunter AUVs or collisions with landmarks, helping to maintain a safe distance during the pursuit mission.
[0143]
[0144] c. Covert communication constraint reward: This reward is designed to ensure that the pursuit process complies with the covert communication constraints.
[0145]
[0146] The total reward of each hunter AUV is the sum of the above rewards:
[0147]
[0148] The above rewards ensure that the hunter AUV effectively completes the pursuit mission while avoiding collisions and meeting the covert communication requirements.
[0149] Based on the state space of each submersible, the action space of each submersible, the state transition probability information of each submersible, and the comprehensive reward function of each submersible, the terminal generates a strategy through the pursuit decision and constructs the pursuit modeling information of each submersible. Specifically, the terminal models the pursuit process as a partially observable Markov decision process (POMDP), where each AUV and the target are regarded as an agent:
[0150]
[0151] in Represents the observation information of each agent, is the action taken by each agent, determines the probability of transitioning to the next state, and is the reward information. In this context, the hunter AUV receives the same reward.
[0152] Based on the above scheme, by analyzing the observation information of each intelligent agent, thus starting with the motion space, state space, reward function, and prediction of the next state, the capture modeling information of each submersible is constructed, which improves the accuracy and comprehensiveness of the capture modeling information of each submersible.
[0153] Optionally, based on the observation information of each submersible and the action information of each submersible, a collaborative capture path for each submersible is generated through each submersible model, target body model, underwater acoustic channel model, and covert communication model, including: based on the underwater acoustic channel model and the covert communication model, through the noise prediction network and the capture modeling information of each submersible, generating the initial capture trajectory of each submersible at each time step; based on each submersible model and the target body model, through the inverse dynamics model, adjusting the initial capture trajectory of each submersible at each time step to obtain the optimized capture trajectory of each submersible at each time step; using the optimized capture trajectory of each submersible at each time step as the collaborative capture path of each submersible.
[0154] In this embodiment, the terminal generates an initial capture trajectory for each submersible at each time step based on the underwater acoustic channel model and the covert communication model, using a noise prediction network and the capture modeling information for each submersible. Finally, based on each submersible model and the target model, the terminal uses an inverse dynamics model to adjust the initial capture trajectory for each submersible at each time step, resulting in an optimized capture trajectory for each submersible at each time step. The terminal uses this optimized capture trajectory for each submersible at each time step as the collaborative capture path for each submersible.
[0155] Specifically, the terminal assumes that the detection range and attack range of the hunter AUV are Specifically, when the distance between the target and the i-th hunter AUV is When , the hunter AUV obtains the target’s location information and shares it within the formation. If all M hunter AUVs are located within a certain distance from the target If all the , or if the conditions for successful capture are not met within this time, the capture mission is considered a failure.
[0156] Then, the terminal algorithm flowchart is as follows Figure 2 As shown, the entire AUV formation shares a noise prediction network, which is based on a modified U-Net architecture (U-Net: Convolutional Networks for Biomedical ImageSegmentation, a deep learning model for image segmentation), and contains three downsampling convolutional layers and three upsampling convolutional layers, with a bottleneck layer connected in the middle. The downsampling layer gradually compresses the state trajectory features while integrating conditional information. The upsampling layer reconstructs the state trajectory by concatenating the intermediate features of the corresponding downsampling layer through skip connections. After the modified U-Net predicts the state trajectories for each hunter AUV, these trajectories are input into their respective inverse dynamics models, enabling each hunter to determine the optimal actions required at each time step to complete the target encirclement.
[0157] To coordinate the formation among multiple hunter AUVs, the terminal introduces an adaptive attention mechanism between the initial downsampling layer and the upsampling layer of the U-Net network. Different from the traditional self-attention mechanism that only focuses on global information, the adaptive attention mechanism is designed to dynamically adjust the attention weights according to global and agent-specific inputs. Formally, the adaptive attention can be expressed as:
[0158]
[0159] where is the query of the i-th AUV, and represent the global key and value information respectively. represents the attention weight of each hunter AUV. The attention weight att reflects the global attention information and combines the contributions of each hunter during the encirclement process. By dynamically adjusting these attention weights, the model can better capture the interactions among the hunter AUVs and coordinate the encirclement formation to achieve effective target encirclement.
[0160] Among them, the training process of this noise prediction network is as follows:
[0161] Unlike the terminal and online reinforcement learning (Online RL), which need to interact with the environment in real time for continuous policy updates, offline reinforcement learning (Offline RL) relies on a pre-collected static dataset D to learn policies, thereby improving data utilization. Considering that the challenges brought by obstacles and unstable communication make real-time interaction in the underwater environment extremely difficult, AMADP (Mobile Application Development Platform) adopts an offline RL training method. In addition, offline RL based on diffusion models is suitable for solving cooperative games. Therefore, the AMADP algorithm proposed in this paper generates the encirclement strategy for the entire hunter AUV formation, while the escape strategy of the target adopts a pre-trained traditional RL method - deep deterministic policy gradient (DDPG).
[0162] Among them, the diffusion model is conditioned guided (which includes the current state, the obtained reward, and the current time step) to generate future trajectories. Considering the requirements of the real encirclement scenario, AMADP adopts a centralized training and decentralized execution (CTDE) framework, accessing global information during the training process, but each hunter AUV makes decisions based on local observations during the execution process. To simplify the representation of the decision-making process in the diffusion model, the terminal defines a learning state sequence:
[0163]
[0164] where represents the state of the i-th hunter AUV at time step and H represents the time step of forward planning.
[0165] For each hunter AUV, the terminal defines an inverse dynamics model , which predicts the action of the i-th AUV at time step t according to the current state and the next state :
[0166]
[0167] Combining the DDPM loss, the inverse dynamics model loss, and the classifier-free guidance mechanism, the overall training loss function can be expressed as:
[0168]
[0169] in Sampling from Bernoulli distribution can balance the diffusion process of conditional guidance and unconditional guidance and improve the generalization ability of the algorithm.
[0170] During the experiment, the AUV formation starts from the center point (500, 500), and the position of the target is randomly initialized. The system model and the parameters of the AMADP algorithm are listed in Figure 3 middle.
[0171] Figure 4 The scene of successfully encircling a target using the AMADP algorithm is demonstrated, in which the AUV formation successfully surrounds the target and forms an encirclement while avoiding obstacles without any collision.
[0172] The terminal measured the The value is taken as the average value of each time step, such as Figure 5 As shown. The results show that the AMADP algorithm always satisfies the covert communication constraints throughout the entire roundup process. To demonstrate the superior performance of the proposed algorithm, the terminal compared AMADP with the latest offline multi-agent reinforcement learning (MARL) algorithms, including multi-agent conservative Q-learning (MACQL), OMAR (Offline MARL, offline multi-agent reinforcement learning method) and MADIFF (Motion-Aware Mamba Diffusion Models, a model for hand trajectory prediction). The experimental results are shown in Figure 6 As shown, unlike traditional multi-agent offline reinforcement learning algorithms based on time-delayed learning (TD-Learning) such as MACQL and OMAR, AMADP significantly outperforms other algorithms in terms of the success rate of round-up tasks by relying on the powerful strategy generation capabilities of the diffusion model and the adaptive attention mechanism to adjust the round-up formation. Furthermore, due to its simpler structure, this algorithm converges faster.
[0173] Based on the above solution, by first considering the eavesdropping ability of the target during the collaborative encirclement process, a framework for encirclement with information security guaranteed by using covert communication is proposed. Meanwhile, to improve the generalization ability, data efficiency, and trajectory diversity of traditional collaborative encirclement algorithms, the terminal proposes AMADP, an offline multi-agent reinforcement learning (MARL) algorithm. This algorithm uses a diffusion model to generate diverse encirclement trajectories and an adaptive attention mechanism to adjust the hunter formation under complex constraints. Experimental results show that AMADP satisfies the covert communication constraints and outperforms existing state-of-the-art algorithms in terms of success rate and convergence speed.
[0174] This application also provides a collaborative encirclement example of multiple underwater vehicles, as Figure 7 shown. The specific processing process includes the following steps:
[0175] Step S701, obtain the operation information of multiple underwater vehicles, the operation information of the target, the transmission channel information of each underwater vehicle, and the underwater environment information.
[0176] Step S702, for each underwater vehicle, based on the operation information of the underwater vehicle, identify the conditional information of each motion condition of the underwater vehicle.
[0177] Step S703, based on the conditional information of each motion condition, through a kinematic modeling strategy, construct an underwater vehicle model of the underwater vehicle.
[0178] Step S704, replace the operation information of the underwater vehicle with the operation information of the target, and return to execute the step of identifying the conditional information of each motion condition of the underwater vehicle based on the operation information of the underwater vehicle to obtain a target model of the target.
[0179] Step S705, split the underwater environment information into underwater propagation information and underwater noise information, and based on the underwater propagation information, identify the underwater propagation feature information.
[0180] Step S706, based on the underwater propagation feature information, through a path loss algorithm, generate an underwater propagation loss formula.
[0181] Step S707, based on the underwater noise information, identify the noise eigenvalue of each underwater noise angle, and through the noise level algorithm corresponding to each noise angle, identify the total noise level formula.
[0182] Step S708, based on the underwater propagation loss formula and the noise level formula, calculate the acoustic path loss of each time step and the underwater noise power of each time step, and use the acoustic path loss of each time step and the underwater noise power of each time step as the underwater acoustic channel model.
[0183] Step S709: Based on the transmission channel information of each submersible vehicle, generate the signal detection information of each submersible vehicle for the target at each time step through a signal detection strategy.
[0184] Step S710: Based on the signal detection information of each submersible vehicle at each time step, identify the optimal threshold power of the target, and construct the concealment constraint equations between each submersible vehicle based on the optimal threshold power and the preset constraints on the change of the communication state of the submersible vehicle.
[0185] Step S711: Use the concealment constraint equations between each submersible vehicle as the covert communication model between each submersible vehicle.
[0186] Step S712: Collect the observation information of each submersible vehicle and the action information of each submersible vehicle.
[0187] Step S713: Based on the observation information of each submersible vehicle, identify the detection content of each type of detection condition detected by each submersible vehicle, construct the state space of each submersible vehicle based on each detection content, and generate the action space of each submersible vehicle through a kinematic algorithm based on the action information of each submersible vehicle.
[0188] Step S714: Based on the action space of each submersible vehicle and the observation information of each submersible vehicle, identify the state transition probability information of each submersible vehicle, and collect the reward condition information of each type of reward.
[0189] Step S715: Based on the reward condition information of each type of reward, generate the reward constraint functions of each type of reward through the reward constraint strategies of each type of reward, and generate the comprehensive reward function of each submersible vehicle based on all the reward constraint functions.
[0190] Step S716: Based on the state space of each submersible vehicle, the action space of each submersible vehicle, the state transition probability information of each submersible vehicle, and the comprehensive reward function of each submersible vehicle, construct the encirclement modeling information of each submersible vehicle through an encirclement decision-making generation strategy.
[0191] Step S717: Based on the underwater acoustic channel model and the covert communication model, generate the initial encirclement trajectory of each submersible vehicle at each time step through a noise prediction network and the encirclement modeling information of each submersible vehicle.
[0192] Step S718: Based on each submersible vehicle model and the target model, adjust the initial encirclement trajectory of each submersible vehicle at each time step through an inverse dynamics model to obtain the optimized encirclement trajectory of each submersible vehicle at each time step.
[0193] Step S719: Use the optimized encirclement trajectory of each submersible vehicle at each time step as the cooperative encirclement path of each submersible vehicle.
[0194] Step S720: Use the cooperative encirclement path of each submersible vehicle as the target cooperative encirclement plan for multiple submersible vehicles.
[0195] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0196] Based on the same inventive concept, the embodiments of the present application further provide a cooperative encirclement device for multiple submersible vehicles for implementing the cooperative encirclement method for multiple submersible vehicles involved above. The solution provided by this device for solving problems is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the cooperative encirclement device for multiple submersible vehicles provided below can refer to the limitations on the cooperative encirclement method for multiple submersible vehicles in the above text, and will not be repeated here.
[0197] In an exemplary embodiment, as Figure 8 shown, a cooperative encirclement device for multiple submersible vehicles is provided, including: an acquisition module 810, a construction module 820, a generation module 830, and a determination module 840, where:
[0198] The acquisition module 810 is configured to acquire the operation information of multiple submersible vehicles, the operation information of the target, the transmission channel information of each submersible vehicle, and the underwater environment information, and construct a submersible vehicle model for each submersible vehicle and a target model for the target based on the operation information of each submersible vehicle and the operation information of the target;
[0199] The construction module 820 is configured to construct an underwater acoustic channel model based on the underwater environment information, and construct a covert communication model between the submersible vehicles through a covert communication modeling strategy based on the transmission channel information of each submersible vehicle;
[0200] A generating module 830, configured to collect the observation information of each submersible and the action information of each submersible, and generate the cooperative encirclement paths of the submersibles based on the observation information of each submersible and the action information of each submersible through each submersible model, the target body model, the underwater acoustic channel model, and the covert communication model;
[0201] A determining module 840, configured to use the cooperative encirclement paths of the submersibles as the target cooperative encirclement scheme for multiple submersibles.
[0202] Optionally, the obtaining module 810 is specifically configured to:
[0203] For each submersible, based on the operation information of the submersible, identify the condition information of each motion condition of the submersible;
[0204] Based on the condition information of each motion condition, construct the submersible model of the submersible through a kinematic modeling strategy;
[0205] Replace the operation information of the submersible with the operation information of the target body, and return to execute the step of identifying the condition information of each motion condition of the submersible based on the operation information of the submersible to obtain the target body model of the target body.
[0206] Optionally, the constructing module 820 is specifically configured to:
[0207] Split the underwater environment information into underwater propagation information and underwater noise information, and identify underwater propagation characteristic information based on the underwater propagation information;
[0208] Based on the underwater propagation characteristic information, generate the underwater propagation loss formula through a path loss algorithm;
[0209] Based on the underwater noise information, identify the noise characteristic values of each underwater noise angle, and identify the total noise level formula through the noise level algorithm corresponding to each noise angle;
[0210] Based on the underwater propagation loss formula and the noise level formula, calculate the acoustic path loss at each time step and the underwater noise power at each time step, and use the acoustic path loss at each time step and the underwater noise power at each time step as the underwater acoustic channel model.
[0211] Optionally, the constructing module 820 is specifically configured to:
[0212] Based on the transmission channel information of each submersible, generate the signal detection information of each submersible for the target body at each time step through a signal detection strategy;
[0213] Based on the signal detection information of each submersible at each time step, identify the optimal threshold power of the target object, and based on the optimal threshold power and the preset constraints on the change of the communication state of the submersibles, construct the concealment constraint equations between each of the submersibles;
[0214] Take the concealment constraint equations between each of the submersibles as the covert communication model between each of the submersibles.
[0215] Optionally, the device further includes:
[0216] A space generation module, configured to identify the detection content of each type of detection condition detected by each submersible based on the observation information of each submersible, and based on each detection content, construct and generate the state space of each submersible, and based on the action information of each submersible, generate the action space of each submersible through a kinematic algorithm;
[0217] An acquisition module, configured to identify the state transition probability information of each submersible based on the action space of each submersible and the observation information of each submersible, and acquire the reward condition information of each type of reward;
[0218] A function generation module, configured to generate the reward constraint functions of each type of reward through the reward constraint strategies of each type of reward based on the reward condition information of each type of reward, and generate the comprehensive reward function of each submersible based on all the reward constraint functions;
[0219] A modeling module, configured to construct the encirclement and capture modeling information of each submersible through an encirclement and capture decision generation strategy based on the state space of each submersible, the action space of each submersible, the state transition probability information of each submersible, and the comprehensive reward function of each submersible.
[0220] Optionally, the generation module 830 is specifically configured to:
[0221] Based on the underwater acoustic channel model and the covert communication model, generate the initial encirclement and capture trajectories of each time step of each submersible through a noise prediction network and the encirclement and capture modeling information of each submersible;
[0222] Based on each submersible model and the target object model, adjust the initial encirclement and capture trajectories of each time step of each submersible through an inverse dynamics model to obtain the optimized encirclement and capture trajectories of each time step of each submersible;
[0223] Take the optimized encirclement and capture trajectories of each time step of each submersible as the collaborative encirclement and capture paths of each of the submersibles.
[0224] Each module in the above-mentioned collaborative hunting device for multiple submersibles can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above modules can be embedded in or independent of a processor in a computer device in the form of hardware, or stored in a memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.
[0225] In an exemplary embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 9 shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner. The wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC), or other technologies. The computer program, when executed by the processor, implements a collaborative hunting method for multiple submersibles. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, a touchpad, or a mouse, etc.
[0226] Those skilled in the art can understand that Figure 9 the structure shown in
[0227] is only a block diagram of a part of the structure related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. A specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.
[0228] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps corresponding to the collaborative hunting method for multiple underwater vehicles are implemented.
[0229] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps corresponding to the collaborative hunting method for multiple underwater vehicles are implemented.
[0230] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0231] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.
[0232] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in the present application.
[0233] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A cooperative encirclement method for multiple submersibles, characterized in that, The method includes: Obtaining the operation information of multiple submersibles, the operation information of the target object, the transmission channel information of each submersible, and the underwater environment information, and constructing a submersible model for each submersible and a target object model for the target object based on the operation information of each submersible and the operation information of the target object; Constructing an underwater acoustic channel model based on the underwater environment information, and constructing a covert communication model between the submersibles through a covert communication modeling strategy based on the transmission channel information of each submersible; Collecting the observation information of each submersible and the action information of each submersible, and generating a cooperative encirclement path for each submersible through each submersible model, the target object model, the underwater acoustic channel model, and the covert communication model based on the observation information of each submersible and the action information of each submersible; Taking the cooperative encirclement path of each submersible as the target cooperative encirclement scheme for multiple submersibles.
2. The method according to claim 1, wherein The constructing a submersible model for each submersible and a target object model for the target object based on the operation information of each submersible and the operation information of the target object includes: For each submersible, identifying the condition information of each motion condition of the submersible based on the operation information of the submersible; Constructing a submersible model of the submersible through a kinematic modeling strategy based on the condition information of each motion condition; Replacing the operation information of the submersible with the operation information of the target object, and returning to execute the step of identifying the condition information of each motion condition of the submersible based on the operation information of the submersible to obtain the target object model of the target object.
3. The method according to claim 1, wherein The constructing an underwater acoustic channel model based on the underwater environment information includes: Splitting the underwater environment information into underwater propagation information and underwater noise information, and identifying underwater propagation characteristic information based on the underwater propagation information; Generating the underwater propagation loss formula through a path loss algorithm based on the underwater propagation characteristic information; Identifying the noise eigenvalue of each underwater noise angle based on the underwater noise information, and identifying the total noise level formula through the noise level algorithm corresponding to each noise angle; Calculating the acoustic path loss of each time step and the underwater noise power of each time step based on the underwater propagation loss formula and the noise level formula, and taking the acoustic path loss of each time step and the underwater noise power of each time step as the underwater acoustic channel model.
4. The method according to claim 1, wherein The constructing a covert communication model between the submersibles through a covert communication modeling strategy based on the transmission channel information of each submersible includes: Generating signal detection information of each submersible of the target object at each time step through a signal detection strategy based on the transmission channel information of each submersible; Identifying the optimal threshold power of the target object based on the signal detection information of each submersible at each time step, and constructing a covertness constraint equation between the submersibles based on the optimal threshold power and a preset submersible communication state change constraint condition; Take the concealment constraint equation between each submersible as the covert communication model between each submersible.
5. The method according to claim 1, characterized in that, After collecting the observation information of each submersible and the action information of each submersible, it further includes: Based on the observation information of each submersible, identify the detection content of each type of detection condition detected by each submersible, and based on each detection content, construct the state space of each submersible, and based on the action information of each submersible, through the kinematic algorithm, generate the action space of each submersible; Based on the action space of each submersible and the observation information of each submersible, identify the state transition probability information of each submersible, and collect the reward condition information of each type of reward; Based on the reward condition information of each type of reward, through the reward constraint strategy of each type of reward, generate the reward constraint function of each type of reward, and based on all reward constraint functions, generate the comprehensive reward function of each submersible; Based on the state space of each submersible, the action space of each submersible, the state transition probability information of each submersible, and the comprehensive reward function of each submersible, through the encirclement decision-making generation strategy, construct the encirclement modeling information of each submersible.
6. The method according to claim 5, characterized in that, The generating the collaborative encirclement path of each submersible through each submersible model, the target body model, the underwater acoustic channel model, and the covert communication model based on the observation information of each submersible and the action information of each submersible includes: Based on the underwater acoustic channel model and the covert communication model, through the noise prediction network and the encirclement modeling information of each submersible, generate the initial encirclement trajectory of each time step of each submersible; Based on each submersible model and the target body model, through the inverse dynamics model, adjust the initial encirclement trajectory of each time step of each submersible to obtain the optimized encirclement trajectory of each time step of each submersible; Take the optimized encirclement trajectory of each time step of each submersible as the collaborative encirclement path of each submersible.
7. A collaborative hunting device for multiple underwater vehicles, characterized in that, The device includes: An acquisition module, configured to acquire the operation information of multiple submersibles, the operation information of the target body, the transmission channel information of each submersible, and the underwater environment information, and based on the operation information of each submersible and the operation information of the target body, construct the submersible model of each submersible and the target body model of the target body; A construction module, configured to construct an underwater acoustic channel model based on the underwater environment information, and construct a covert communication model between each submersible through a covert communication modeling strategy based on the transmission channel information of each submersible; A generation module, configured to collect the observation information of each submersible and the action information of each submersible, and based on the observation information of each submersible and the action information of each submersible, generate the collaborative encirclement path of each submersible through each submersible model, the target body model, the underwater acoustic channel model, and the covert communication model; A determination module, configured to take the collaborative encirclement path of each submersible as the target collaborative encirclement scheme of multiple submersibles.
8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method described in any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method described in any one of claims 1 to 6.
Citation Information
Cited By
Underwater vehicle cooperative detection method based on multi-agent reinforcement learning
CN121254280A
Cooperative detection method for underwater vehicle formation and related products
CN122172858A
Generative sound-scene-imitating underwater covert communication method based on message latent variable injection
CN122293217A