Situation information-based unmanned ship cluster control method, equipment and medium
By generating a situational threat feature map and inputting the target strategy model, the problem of low control efficiency of unmanned boat clusters is solved, and fast and accurate unmanned boat cluster control is achieved.
Patent Information
- Application Number
- CN202510618828.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-08-15
AI Technical Summary
In the prior art, unmanned boat cluster control efficiency is low, making it difficult to achieve rapid and precise control in a complex and changeable maritime combat environment.
By obtaining the operation data and observation data of the unmanned boat, a situation threat feature map is generated, and input it into the agent of the pre-trained target strategy model, generating control actions, and using situation information to control the unmanned boat cluster.
It realizes fast and precise control of unmanned boat clusters, solving the problems of low efficiency and poor collaboration in cluster combat scenarios.
Smart Images

Figure CN120491663A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of unmanned boat control technology, and in particular to a control method, device and medium for an unmanned boat cluster based on situation information. Background Art
[0002] As naval warfare environments become increasingly complex and maritime combat styles become increasingly diverse, the development of swarm warfare is becoming increasingly important. In swarm confrontation scenarios (or simulated swarm confrontation scenarios), bioinspired methods (such as bee swarm algorithms) are used to control swarms of intelligent agents. However, due to the large number of devices on both sides, the high dimensionality of the observation space, and the complex and ever-changing offensive and defensive situations, bioinspired methods are less efficient in controlling swarms. Summary of the Invention
[0003] In response to the above-mentioned problems and technical needs, the applicant has proposed a control method, equipment and medium for an unmanned boat cluster based on situation information, so as to solve the problem of low control efficiency in the existing technology when controlling an unmanned boat cluster, and to achieve rapid and accurate control of the unmanned boat cluster.
[0004] The embodiment of the application provides a method for controlling a swarm of unmanned boats based on situation information, the method comprising:
[0005] Obtaining operational data of each unmanned boat in the first target cluster in a preset area, and observation data of each unmanned boat observing target devices of the second target cluster;
[0006] Determining a first distribution of unmanned boats based on the operational data, and determining a second distribution of target devices based on the observation data;
[0007] Generating a situation threat feature map based on the first distribution and the second distribution;
[0008] For each unmanned vehicle, the situation threat feature map and the observation data corresponding to the unmanned vehicle are input into the intelligent agent of the pre-trained target strategy model to obtain the control action output by the target strategy model;
[0009] Among them, the total number of unmanned boats is the same as the total number of intelligent agents, and the observation data corresponding to one unmanned boat corresponds to one intelligent agent;
[0010] Among them, the target strategy model is trained based on operation data samples, observation data samples and control action samples.
[0011] According to the control method of the unmanned watercraft swarm based on situation information provided in an embodiment of the present application, a situation threat characteristic map is generated based on the first distribution situation and the second distribution situation, including:
[0012] Divide the preset area into N sub-areas of equal size, where N is an integer greater than or equal to 2;
[0013] Determine a first number of unmanned boats and a second number of target devices in each sub-area based on the first distribution and the second distribution;
[0014] The situation threat level of each sub-region is calculated based on the first quantity and the second quantity, and a situation threat feature map for representing each situation threat level is generated.
[0015] According to the control method of the unmanned boat swarm based on situation information provided in an embodiment of the present application, the situation threat level of each sub-area is calculated based on the first quantity and the second quantity, including:
[0016] Inputting the first quantity and the second quantity into a predetermined situation threat degree calculation formula to obtain a situation threat degree output by the situation threat degree calculation formula;
[0017] The situation threat degree calculation formula includes:
[0018]
[0019] Among them, T i Indicates the situation threat level corresponding to the i-th sub-region, represents the first quantity, Indicates the second quantity.
[0020] According to the control method of the unmanned vehicle swarm based on situation information provided in an embodiment of the present application, before obtaining observation data of each unmanned vehicle in the first target cluster observing the target device of the second target cluster in the preset area, the method further includes:
[0021] Obtaining operation data samples, observation data samples and control action samples;
[0022] Generate situation threat feature map samples and total threat degree value samples based on operation data samples and observation data samples;
[0023] For each unmanned boat sample, a reward function is created based on the total threat value sample and the observation data sample corresponding to the unmanned boat sample;
[0024] The target strategy model is trained based on the situation threat feature map samples, control action samples and reward function until the number of training times reaches a preset threshold, and the training of the target strategy model is determined to be completed.
[0025] According to the control method of the unmanned watercraft swarm based on situation information provided in an embodiment of the present application, based on the operation data samples and the observation data samples, a situation threat feature map sample and a total threat degree value sample are generated, including:
[0026] Generate situation threat degree samples based on operation data samples and observation data samples;
[0027] Generate a situation threat feature map sample based on the situation threat degree sample, and obtain a total threat degree value sample based on a preset total threat degree calculation formula;
[0028] The total threat calculation formula includes:
[0029]
[0030] Among them, T 总 Indicates the total threat value sample, represents the situation threat level sample corresponding to the i-th sub-region, and N represents the number of sub-regions.
[0031] According to the control method of the unmanned boat swarm based on situation information provided in an embodiment of the present application, a reward function is created based on the total threat value samples and the observation data samples corresponding to the unmanned boat samples, including:
[0032] Calculate the difference between the total threat value samples at two adjacent moments to obtain the situation change reward;
[0033] For each unmanned boat, calculate individual rewards based on the operation status of the unmanned boat;
[0034] Create a reward function based on situation change rewards and individual rewards.
[0035] According to the control method of the unmanned boat swarm based on situation information provided in an embodiment of the present application, determining the second distribution of target devices based on observation data includes:
[0036] The observation data corresponding to each unmanned boat is processed based on a centralized information fusion algorithm to obtain target observation data;
[0037] Based on the target observation data, the total number of target devices and the position and second number of each target device in the preset area are determined.
[0038] According to the control method of the unmanned vehicle swarm based on situation information provided in an embodiment of the present application, before obtaining the operating data of each unmanned vehicle in the first target swarm in the preset area and the observation data of each unmanned vehicle observing the target device of the second target swarm, the method further includes:
[0039] Obtain a target policy model where the total number of agents is the same as the total number of unmanned boats.
[0040] An embodiment of the present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the program, the steps of the control method of the unmanned boat cluster based on situation information as described in any one of the above items are implemented.
[0041] An embodiment of the present application also provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of the control method of the unmanned boat cluster based on situation information as described in any one of the above items are implemented.
[0042] The control method, device and medium of the unmanned boat cluster based on situation information provided by the embodiment of the present application obtain the operating data of each unmanned boat in the first target cluster in the preset area, and the observation data of each unmanned boat observing the target equipment of the second target cluster, to obtain the first distribution of the unmanned boats and the second distribution of the target equipment. The present application processes the observation data to obtain the second distribution to reduce the data dimension; based on the first distribution and the second distribution, a situation threat feature map is generated. It can be seen that the present application obtains situation information based on the distribution, which also effectively reduces the data dimension and provides an effective data basis for the accurate prediction of the subsequent model; for each unmanned boat Boat, the situation threat feature map and the observation data corresponding to the unmanned boat are input into the intelligent agent of the pre-trained target strategy model to obtain the control action output by the target strategy model; wherein, the total number of unmanned boats is the same as the total number of intelligent agents, and the observation data corresponding to one unmanned boat corresponds to one intelligent agent. The present application inputs the observation data monitored by different unmanned boats into different intelligent agents, ensuring the rapid and accurate output of the control action of each unmanned boat in the cluster combat scenario, realizing the rapid and accurate completion of the unmanned boat cluster control, and solving the problems in the prior art of low control action output efficiency and poor collaboration between the unmanned boats due to the complexity of the cluster combat scenario or the large number of equipment. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0044] Figure 1 This is one of the flow charts of the control method of the unmanned boat swarm based on situation information provided in the embodiment of the present application;
[0045] Figure 2 This is a schematic diagram of the distribution of unmanned boats and target equipment provided in an embodiment of the present application;
[0046] Figure 3 This is the second flow chart of the control method of the unmanned boat swarm based on situation information provided in an embodiment of the present application;
[0047] Figure 4 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0048] To make the purpose, technical solutions, and advantages of the embodiments of the present application more clear, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0049] Specifically, conventional cluster adversarial methods rely on pre-programmed rules or centralized control, making them difficult to adapt to sudden changes in dynamic environments. Furthermore, these methods suffer from slow computational convergence and low collaborative efficiency. Furthermore, existing technologies lack quantitative assessment of real-time situational awareness and proactive guidance mechanisms, resulting in irrational cluster resource allocation and poor adversarial effectiveness.
[0050] In order to solve the above series of problems, the embodiment of the present application provides a control method for an unmanned boat cluster based on situation information. The method can be applied to smart terminals, servers, and controllers of unmanned boats. This application uses the application of the method in the controller of an unmanned boat as an example to illustrate. This is an example and is not used to limit the scope of protection of this application. Some other descriptions in the embodiments are also examples and will not be described one by one later. Figure 1 As shown, the method includes:
[0051] Step 101: Acquire the operating data of each unmanned boat in a first target cluster in a preset area, and the observation data of each unmanned boat observing the target equipment in a second target cluster.
[0052] Among them, the operating data includes the operating position, operating speed, operating direction, etc. of the unmanned boat in the preset area.
[0053] Among them, the total number of unmanned boats had been obtained before the battle.
[0054] Step 102: determining a first distribution of unmanned boats based on the operating data, and determining a second distribution of target devices based on the observation data.
[0055] Among them, the first distribution situation includes: the operating position of each unmanned boat in the preset area.
[0056] Specifically, the operating position of each unmanned boat in the preset area can be extracted from the operating data.
[0057] Step 103: Generate a situation threat feature map based on the first distribution and the second distribution.
[0058] Step 104: For each unmanned boat, the situation threat feature map and the observation data corresponding to the unmanned boat are input into the intelligent agent of the pre-trained target strategy model to obtain the control action output by the target strategy model.
[0059] Among them, the control actions include controlling the steering angle and propeller speed of the unmanned boat.
[0060] The steering gear angle is represented by δ, and its value range is δ∈[-δ max ,δ max ], the propeller speed is represented by n, and its value range is n∈[0,n max ]. Among them, δ max Indicates the maximum steering angle of the unmanned boat, n max Indicates the maximum propeller speed of the unmanned boat.
[0061] Specifically, the present application sets a value range for the servo angle and propeller speed, and limits the servo angle and propeller speed output by the target control model. If they exceed the corresponding value range, the limit value corresponding to the value range is used as the final servo angle and / or propeller speed.
[0062] Among them, the total number of unmanned boats is the same as the total number of intelligent agents, and the observation data corresponding to one unmanned boat corresponds to one intelligent agent.
[0063] Among them, the target strategy model is trained based on operation data samples, observation data samples and control action samples.
[0064] Specifically, the operational and observational data acquired by this application is current data, while the target strategy model predicts the control actions for the next moment. This entire process is iteratively executed as the UAV's combat cycle progresses. This application enables real-time monitoring of the UAV and target equipment, enabling timely response to various sudden changes in a dynamic environment and ensuring efficient collaboration between the UAVs.
[0065] Specifically, the agent corresponding to each set of situational threat feature maps and observation data corresponding to the unmanned boat can be matched based on the situational threat level contained in the situational threat feature maps. Unmanned boats in the same sub-region have the same situational threat level. There can be one or more agents corresponding to the same situational threat level.
[0066] In addition, this application can also be used in the simulation scenario of cluster confrontation of unmanned boats.
[0067] The control method of the unmanned boat cluster based on situation information provided by the embodiment of the present application obtains the operating data of each unmanned boat in the first target cluster in the preset area, and the observation data of each unmanned boat observing the target equipment of the second target cluster, to obtain the first distribution of the unmanned boats and the second distribution of the target equipment. The present application processes the observation data to obtain the second distribution to reduce the data dimension; based on the first distribution and the second distribution, a situation threat feature map is generated. It can be seen that the present application obtains situation information based on the distribution, which also effectively reduces the data dimension and provides an effective data basis for the accurate prediction of the subsequent model; for each unmanned boat, The situation threat characteristic map and the observation data corresponding to the unmanned boat are input into the intelligent agent of the pre-trained target strategy model to obtain the control action output by the target strategy model; wherein, the total number of unmanned boats is the same as the total number of intelligent agents, and the observation data corresponding to one unmanned boat corresponds to one intelligent agent. The present application inputs the observation data monitored by different unmanned boats into different intelligent agents, ensuring the rapid and accurate output of the control action of each unmanned boat in the cluster combat scenario, realizing the rapid and precise completion of the unmanned boat cluster control, and solving the problems in the existing technology of low control action output efficiency and poor collaboration between unmanned boats due to the complexity of the cluster combat scenario or the large number of equipment.
[0068] In a specific embodiment, before obtaining the operating data of each unmanned boat in the first target cluster in the preset area, and the observation data of each unmanned boat observing the target device of the second target cluster, a target strategy model is obtained in which the total number of intelligent agents is the same as the total number of unmanned boats.
[0069] Specifically, in order to ensure that the target strategy model can output control actions accurately and quickly, this application adopts the same total number of intelligent agents as the total number of unmanned boats, so that the observation data corresponding to one unmanned boat is matched with one intelligent agent.
[0070] Specifically, this application pre-trains multiple target strategy models, and different target strategy models correspond to different numbers of agents.
[0071] In a specific embodiment, determining the second distribution of target devices based on the observed data includes:
[0072] The observation data corresponding to each unmanned boat is processed based on a centralized information fusion algorithm to obtain target observation data; based on the target observation data, the total number of target devices and the position and second number of each target device in the preset area are determined.
[0073] Specifically, the observation data obtained by each unmanned boat is fused and processed to achieve real-time fusion, with high processing precision and flexible algorithm, which improves the accuracy of data processing.
[0074] In a specific embodiment, the specific implementation of generating a situation threat feature map based on the first distribution situation and the second distribution situation includes:
[0075] A preset area is divided into N sub-areas of equal size; a first number of unmanned boats and a second number of target devices in each sub-area are determined based on a first distribution and a second distribution; a situation threat degree of each sub-area is calculated based on the first number and the second number, and a situation threat characteristic map is generated for characterizing each situation threat degree.
[0076] Wherein, N is an integer greater than or equal to 2.
[0077] Specifically, a fixed size can be set for the sub-area, and the user can also set a size corresponding to the actual situation according to the actual situation. This application does not impose any restrictions.
[0078] Specifically, the operating positions and first number of unmanned boats within each sub-area are determined based on the first distribution, and the positions and second number of target devices within each sub-area are determined based on the second distribution. This allows the first number and operating positions of unmanned boats, the positions and second number of target devices within each sub-area to be determined, and the positional relationship between the unmanned boats and the target devices to be determined.
[0079] Specifically, the situation threat feature map includes an RGB image, which includes multiple colors, and each color corresponds to a different situation threat level.
[0080] The first and second distributions can be found in Figure 2 .
[0081] Among them, Figure 2 Each quadrilateral in represents a sub-area, and 25 sub-areas are used as an example for illustration. This is only an example and is not used to limit the scope of protection of this application. Users can divide it according to their actual needs.
[0082] Among them, Figure 2 The colors in the image represent the threat level. Darker colors indicate a higher threat level for that sub-region. Later models can also determine the threat level of a sub-region by identifying the depth of the color. Of course, a mapping relationship between color depth and threat level is pre-set.
[0083] The present application generates a situation threat feature map for prediction of subsequent control actions, which reduces data dimensions and improves data processing speed.
[0084] In a specific embodiment, the specific implementation of calculating the situation threat level of each sub-region based on the first quantity and the second quantity includes:
[0085] The first quantity and the second quantity are input into a predetermined situation threat degree calculation formula to obtain a situation threat degree output by the situation threat degree calculation formula.
[0086] The calculation formula of situation threat degree is shown in formula (1):
[0087]
[0088] Among them, T i Indicates the situation threat level corresponding to the i-th sub-region, represents the first quantity, Indicates the second quantity.
[0089] In a specific embodiment, before obtaining the observation data of each unmanned boat in the first target cluster observing the target device of the second target cluster in the preset area, the target strategy model is trained. The specific implementation of the training process is as follows: Figure 3 As shown:
[0090] Step 301: Acquire operation data samples, observation data samples, and control action samples.
[0091] Step 302: Generate situation threat feature map samples and total threat degree value samples based on the operation data samples and the observation data samples.
[0092] Step 303: For each unmanned boat sample, a reward function is created based on the total threat value sample and the observation data sample corresponding to the unmanned boat sample.
[0093] Step 304 : training the target policy model based on the situation threat feature map samples, the control action samples, and the reward function until the number of training times reaches a preset threshold, and determining that the training of the target policy model is completed.
[0094] Specifically, historical data corresponding to the unmanned boat in the historical cluster confrontation scenario is obtained and used as training sample data, including operation data samples, observation data samples and control action samples.
[0095] In a specific embodiment, the specific implementation of generating a situation threat feature map sample and a total threat degree value sample based on the operation data sample and the observation data sample includes:
[0096] Based on the operation data samples and the observation data samples, a situation threat degree sample is generated; based on the situation threat degree sample, a situation threat feature map sample is generated; and based on the preset total threat degree calculation formula, a total threat degree value sample is obtained.
[0097] The total threat degree calculation formula is shown in formula (2):
[0098]
[0099] Among them, T 总 Indicates the total threat value sample, represents the situation threat level sample corresponding to the i-th sub-region, and N represents the number of sub-regions.
[0100] Specifically, a first distribution sample is obtained based on the operational data sample, and a second distribution sample is obtained based on the observed data sample. The regional sample is divided into N sub-region samples of equal size; based on the first distribution sample and the second distribution sample, a first number sample of unmanned boats and a second number sample of target equipment within each sub-region sample are determined; based on the first number sample and the second number sample, a situation threat level sample is calculated for each sub-region sample, and a situation threat level feature map sample is generated for characterizing each situation threat level.
[0101] In a specific embodiment, the specific implementation of creating a reward function based on the total threat value sample and the observation data sample corresponding to the unmanned boat sample includes:
[0102] The difference between the total threat value samples at two adjacent moments is calculated to obtain the situation change reward. For each unmanned boat, the individual reward is calculated based on the operation status of the unmanned boat. A reward function is created based on the situation change reward and the individual reward.
[0103] Specifically, the situation change reward is obtained based on formula (3):
[0104] R st =T″ 总 -T′ 总 …………………………(3)
[0105] Among them, R st represents the situation change reward, T″ 总 Represents the total threat value sample corresponding to the previous moment, T′ 总 Indicates the total threat value sample corresponding to the current moment.
[0106] Specifically, the individual reward is obtained based on formula (4):
[0107]
[0108] Among them, R id represents individual rewards, n at Indicates the number of hits in a single-step operation of the unmanned boat, n ij Indicates the number of times the unmanned boat is hit by the target equipment during a single-step operation.
[0109] Among them, the operating status of the unmanned boat includes: the number of hits in a single-step operation and the number of hits by the target equipment in a single-step operation.
[0110] Specifically, the reward function is shown in formula (5):
[0111] R=ωR st +(1-ω)R id ………………(5)
[0112] Among them, R represents the total reward obtained using the reward function, and ω represents the weight coefficient, which ranges from 0 to 1.
[0113] Among them, the reward function is used to evaluate the total reward of the current single-step job.
[0114] In a specific embodiment, the specific implementation of training the target policy model based on the situation threat feature map samples, the control action samples and the reward function includes:
[0115] A target policy model for the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm is established. This target policy model constructs a pair of actor-critic deep neural networks (agents) for each unmanned vehicle. The actor is a policy generation network that takes as input a sample of a situational threat feature map and observation data, and outputs a predicted control action. The critic is a value assessment network that takes as input a sample of a situational threat feature map and a predicted control action, and outputs the value of the action corresponding to the unmanned vehicle.
[0116] Specifically, all agents are trained based on the training sample data, and the network parameters of each agent are updated until a preset threshold is reached, and the target strategy model training is determined to be completed.
[0117] Ultimately, the target policy model is used to predict control actions. However, in the actual application stage, it is only necessary to input the situation threat feature map and observation data into the Actor, and use the Actor to predict the control action at the next moment.
[0118] In addition, the environmental information corresponding to the preset area can be obtained, and the situation threat feature map, the observation data corresponding to the unmanned boat and the environmental information can be input into the intelligent agent of the pre-trained target strategy model to obtain the control action output by the target strategy model.
[0119] Its training and application stages both correspond to matching environmental information to improve the accuracy of control action output.
[0120] This application preprocesses observation data, reducing its dimensionality and increasing its processing speed. It also introduces a reward function for situation assessment, using situational information to guide reinforcement learning training. This achieves a closed-loop process of environmental perception and policy learning, improving the model's convergence efficiency and enhancing the cluster's collaborative capabilities based on the model's output.
[0121] Figure 4 An example of a physical structure diagram of an electronic device is shown below. Figure 4 As shown, the electronic device may include: a processor 401, a communication interface 402, a memory 403, and a communication bus 404. The processor 401, the communication interface 402, and the memory 403 communicate with each other via the communication bus 404. The processor 401 may call logic instructions in the memory 403 to execute a control method for an unmanned vehicle swarm based on situation information.
[0122] In addition, the logic instructions in the above-mentioned memory 403 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0123] On the other hand, the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the control method of the unmanned boat cluster based on situation information provided by the above methods.
[0124] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the control method of the unmanned boat cluster based on situation information provided in the above embodiments.
[0125] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0126] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.
[0127] Finally, it should be noted that the above description is merely a preferred embodiment of the present application and the present application is not limited to the above embodiments. It is understood that other improvements and variations directly derived or conceived by those skilled in the art without departing from the spirit and concept of the present application should be considered to be included within the scope of protection of the present application.
Claims
1. A control method for an unmanned boat swarm based on situation information, characterized in that: The method comprises: Obtaining operational data of each unmanned boat in the first target cluster in a preset area, and observation data of each unmanned boat observing target devices of the second target cluster; Determining a first distribution of unmanned boats based on the operational data, and determining a second distribution of target devices based on the observation data; Generating a situation threat feature map based on the first distribution and the second distribution; For each unmanned vehicle, the situation threat feature map and the observation data corresponding to the unmanned vehicle are input into the intelligent agent of the pre-trained target strategy model to obtain the control action output by the target strategy model; Among them, the total number of unmanned boats is the same as the total number of intelligent agents, and the observation data corresponding to one unmanned boat corresponds to one intelligent agent; Among them, the target strategy model is trained based on operation data samples, observation data samples and control action samples.
2. The control method of the unmanned boat swarm based on situation information according to claim 1 is characterized in that: Based on the first distribution and the second distribution, a situation threat feature map is generated, including: Divide the preset area into N sub-areas of equal size, where N is an integer greater than or equal to 2; Determine a first number of unmanned boats and a second number of target devices in each sub-area based on the first distribution and the second distribution; The situation threat level of each sub-region is calculated based on the first quantity and the second quantity, and a situation threat feature map for representing each situation threat level is generated.
3. The control method of the unmanned boat swarm based on situation information according to claim 2, characterized in that: Calculating the situation threat level of each sub-area based on the first quantity and the second quantity includes: Inputting the first quantity and the second quantity into a predetermined situation threat degree calculation formula to obtain a situation threat degree output by the situation threat degree calculation formula; The situation threat degree calculation formula includes: Among them, T i Indicates the situation threat level corresponding to the i-th sub-region, represents the first quantity, Indicates the second quantity.
4. The control method of the unmanned boat swarm based on situation information according to any one of claims 1 to 3, characterized in that: Before obtaining observation data of each unmanned boat in the first target cluster in the preset area observing the target device of the second target cluster, the method further includes: Obtaining operation data samples, observation data samples and control action samples; Generate situation threat feature map samples and total threat degree value samples based on operation data samples and observation data samples; For each unmanned boat sample, a reward function is created based on the total threat value sample and the observation data sample corresponding to the unmanned boat sample; The target strategy model is trained based on the situation threat feature map samples, control action samples and reward function until the number of training times reaches a preset threshold, and the training of the target strategy model is determined to be completed.
5. The control method of the unmanned boat swarm based on situation information according to claim 4 is characterized in that: Based on the operation data samples and observation data samples, the situation threat feature map samples and the total threat value samples are generated, including: Generate situation threat degree samples based on operation data samples and observation data samples; Generate a situation threat feature map sample based on the situation threat degree sample, and obtain a total threat degree value sample based on a preset total threat degree calculation formula; The total threat calculation formula includes: Among them, T 总 Indicates the total threat value sample, represents the situation threat level sample corresponding to the i-th sub-region, and N represents the number of sub-regions.
6. The control method of the unmanned boat swarm based on situation information according to claim 4 is characterized in that: Create a reward function based on the total threat value sample and the observation data sample corresponding to the unmanned boat sample, including: Calculate the difference between the total threat value samples at two adjacent moments to obtain the situation change reward; For each unmanned boat, calculate individual rewards based on the operation status of the unmanned boat; Create a reward function based on situation change rewards and individual rewards.
7. The control method of the unmanned boat swarm based on situation information according to claim 2, characterized in that: Determining a second distribution of target devices based on the observed data includes: The observation data corresponding to each unmanned boat is processed based on a centralized information fusion algorithm to obtain target observation data; Based on the target observation data, the total number of target devices and the position and second number of each target device in the preset area are determined.
8. The control method of the unmanned boat swarm based on situation information according to any one of claims 1 to 3, characterized in that: Before obtaining the operation data of each unmanned boat in the first target cluster in the preset area and the observation data of each unmanned boat observing the target equipment of the second target cluster, the method further includes: Obtain a target policy model where the total number of agents is the same as the total number of unmanned boats.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the method for controlling an unmanned boat cluster based on situation information as described in any one of claims 1 to 8 are implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for controlling a swarm of unmanned boats based on situation information as described in any one of claims 1 to 8 are implemented.