A heterogeneous cooperative exploration trajectory planning method and device
Patent Information
- Application Number
- CN202311088162.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-28
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2043-08-28
AI Technical Summary
通过构建并训练初始探测模型,解决了现有技术中环境发生变化时难以保证实时性、可扩展性差的问题
本发明通过分别构建至少一个飞机类平台和/或至少一个舰船类平台的能力模型,实现兼容异构的多个探测平台的协同探测规划。通过引入确定度图,分别获取所述飞机类平台和/或所述舰船类平台的初始局部观测,降低探测对环境中运动时敏目标的遗漏概率。通过构建并训练初始探测模型,实现实时规划探测轨迹,提高了多平台自主协同探测的实时性和可扩展性。
Smart Images

Figure CN117109589B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of autonomous decision-making and planning technology, and in particular to a heterogeneous collaborative detection trajectory planning method and apparatus. Background Technology
[0002] Target detection, as a crucial means of understanding the operational situation on the maritime battlefield, is the foundation and prerequisite for combat command and firepower strikes, and is of great significance for achieving battlefield transparency and the integration of combat forces and operations. With the continuous development of advanced technologies and multi-platform joint combat systems, modern maritime warfare faces a battlefield situation characterized by diversified targets, complex electromagnetic environments, and ever-changing tactics. Using only a single platform for detection presents challenges such as vulnerability to reconnaissance and attack, susceptibility to jamming, and limited coverage areas. Therefore, multi-platform collaborative detection is necessary. Utilizing multiple platforms, including aircraft and ships distributed across the sea and air, for collaborative detection has greater applicability to modern maritime combat scenarios and holds broad research potential.
[0003] Multi-platform autonomous collaborative detection targets large-area mission areas with specific and unknown environments. Limited multi-platform detection resources, through collaborative cooperation, achieve rapid and efficient coverage of the target area and detect targets within it. Traditional methods for multi-platform autonomous collaborative detection decompose collaborative detection into two sub-problems based on a hierarchical task approach: multi-platform task area decomposition and detection trajectory planning. Task area decomposition, based on the target area, initial platform position, and performance, uses performance constraints and minimum width principles to divide the task area, allocating non-overlapping task areas that achieve full coverage for each platform. Detection trajectory planning generates the optimal path that completely covers the task area. Considering energy consumption during platform turns, common trajectory planning methods include figure-eight, inner spiral, and parallel line patterns. However, two problems exist: first, the area is usually pre-defined and a fixed route is planned, making it difficult to guarantee real-time performance when the environment changes, resulting in poor scalability, and the planned path still requires further smoothing in practical applications; second, time-sensitive moving targets in the environment may move from previously searched areas to already searched areas, leading to a higher probability of detection omissions.
[0004] Therefore, overcoming the shortcomings of the existing technology is an urgent problem to be solved in this technical field. Summary of the Invention
[0005] The technical problem this invention aims to solve is to provide a heterogeneous collaborative detection trajectory planning method and apparatus. By constructing capability models for heterogeneous detection platforms, it achieves collaborative detection planning for multiple heterogeneous detection platforms. Introducing a determinism map reduces the probability of missing moving, time-sensitive targets in the environment. By constructing and training an initial detection model, it addresses the problems of poor real-time performance and scalability in existing technologies when the environment changes.
[0006] The present invention adopts the following technical solution: In a first aspect, the present invention provides a heterogeneous cooperative detection trajectory planning method, comprising: Construct capability models for at least one aircraft platform and / or at least one ship platform; Initialize the overall detection area and construct the initial detection model; Based on the initial detection model, initial determination maps of the aircraft platform and / or the ship platform are obtained respectively. Based on the capability model and the initial determination maps, initial local observations of the aircraft platform and / or the ship platform are obtained respectively. Based on the initial local observations, the initial detection model is trained to obtain the target detection model; The actual local observations of at least one aircraft platform and / or at least one ship platform are respectively input into the target detection model to obtain the optimal detection trajectory.
[0007] Furthermore, the capability models for constructing at least one aircraft platform and / or at least one ship platform include: The maneuverability of the aircraft platform and / or the ship platform is abstracted as maintaining a constant speed at a preset speed. The endurance of the aircraft platform and / or the ship platform is characterized as the farthest sailing distance, and the distance is the first distance; The detection range of the aircraft platform and / or the ship platform is abstracted as a square with a preset side length to obtain a local detection area; The probability of local detection within the local detection region is selectively abstracted to 0 or 1; The capability model is obtained based on the preset speed, the first distance, the local detection area, and the local detection probability.
[0008] Furthermore, the initialization of the overall detection area and the construction of the initial detection model include: Determine the first greatest common divisor of the preset side lengths of the local detection areas of the aircraft platform and / or the ship platform; set the overall detection area as a two-dimensional rectangular area; rasterize the overall detection area into at least one square with a side length of the first greatest common divisor; establish a foot force coordinate system of the overall detection area with the first greatest common divisor as the unit length; the foot force coordinate system is used to determine the training local observation or the actual local observation. Each of the aircraft-type platforms and / or the ship-type platforms is treated as an agent. The training local observations and environmental information of all the agents in the foot-force coordinate system are modeled as a distributed decision process. Based on the distributed decision process, the initial detection model is constructed, which includes the structure of an Actor network and the structure of a Critic network.
[0009] Further, the step of obtaining initial determination maps for the aircraft-type platform and / or the ship-type platform based on the initial detection model, and obtaining initial local observations for the aircraft-type platform and / or the ship-type platform based on the capability model and the initial determination maps, includes: Based on the initial detection model, a determination map for each of the aircraft-type platforms and / or the ship-type platforms is initialized, and the determination of each grid of the determination map is set to a preset initial value to obtain the initial determination map including multiple grids; wherein, the unit length of the grid of the grid is the first greatest common divisor; Based on the corresponding capability model, the given position coordinates of the aircraft platform and / or the ship platform in the overall detection area and the given heading in the overall detection area are obtained respectively. The initial local observations are obtained based on the initial determination map, the given position coordinates, and the given heading.
[0010] Further, the step of training the initial detection model based on the initial local observations to obtain the target detection model includes: The initial local observation is used as the training local observation at time t; the training local observation is input into the Actor network of the initial detection model, and the Actor network outputs the local action at time t after making a decision; based on the local action, the joint action of all the aircraft-type platforms and / or the ship-type platforms at time t is obtained; The combined action is input to the environment of the overall detection area, and the environment outputs the training local observation at time t+1 and the environment reward at time t based on the corresponding determination map. The training local observations of all detection platforms at time t are combined into a global state, and the global state is input into the Critic network of the initial detection model to obtain the current action value function. The training local observation at time t, the local action, the global state, the current action value function, the training local observation at time t+1, and the environmental reward are combined to form the current experience tuple; the current experience tuple is stored in the experience pool, and when the experience pool is full, the current experience tuple is selected to be retained in the experience pool based on importance sampling. Based on at least one selected current empirical tuple, an advantage function is obtained, and based on the advantage function, the target detection model is obtained.
[0011] Furthermore, in the determinism map, the raster map of a given position (x, y) at time t is: ; If the grid at a given location (x, y) is not covered by the corresponding local detection region at time t, the certainty of the corresponding certainty map continuously decreases; wherein, the certainty decreases according to the decay factor η, resulting in... The raster image at time +1 is ; If the grid at a given location (x, y) is covered by the corresponding local detection region at time t, then the certainty of the corresponding determination map continuously increases, resulting in... The raster image at time +1 is .
[0012] Furthermore, the step of inputting the trained local observations into the Actor network of the initial detection model, and the Actor network outputting the local action at time t after making a decision, includes: The training local observation is to maintain the original speed and direction and continue to move forward, and the obtained local action is the first preset value; The training local observation is obtained by turning left along the Dubins path and maintaining straight-line motion, and the local action is the second preset value. The training local observation is obtained by turning right along the Dubins path and maintaining straight-line motion, and the local action is the third preset value. The training local observation is obtained when the local action is the fourth preset value after turning left on the Dubins path and moving in a straight line in the opposite direction while maintaining the original speed. The training local observation is obtained when the local action is the fifth preset value after turning right on the Dubins path and maintaining the original speed while moving in a straight line in the opposite direction.
[0013] Furthermore, the environmental reward includes a certainty reward and a range reward, wherein: Based on the corresponding certainty map, the sum of the increments of certainty corresponding to all grids within the overall detection area covered at time t during the execution of the joint action is obtained, and the certainty reward is obtained. Based on the first distance and the corresponding preset speed, a range bonus is selectively awarded. ;in, For preset speed, This represents the furthest possible sailing distance. The environmental reward is obtained by adding the product of the certainty reward and the corresponding certainty weight, and the product of the range reward and the corresponding range weight.
[0014] Further, obtaining the advantage function based on at least one selected current experience tuple, and obtaining the target detection model based on the advantage function, includes: Based on at least one selected current experience tuple, at least one environmental reward for the current experience tuple is obtained; using a preset advantage estimation algorithm, the long-term discounted reward for the corresponding local action is obtained based on the environmental reward; Based on the long-term discount reward and the action value function of the current experience tuple, the advantage function of the corresponding local action relative to the average reward of all local actions is obtained; Calculate the loss functions of the Actor network and the Critic network based on the advantage function, and update the network parameters of the Actor network and the Critic network through backpropagation; When the parameters of the Actor network and the Critic network converge or reach a preset number of rounds, the parameters of the Actor network and the Critic network are saved to obtain the target detection model.
[0015] Secondly, the present invention also provides a heterogeneous cooperative detection trajectory planning device for implementing the heterogeneous cooperative detection trajectory planning method described in the first aspect, the heterogeneous cooperative detection trajectory planning device comprising: At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the processor for performing the heterogeneous cooperative detection trajectory planning method described in the first aspect.
[0016] Thirdly, the present invention also provides a non-volatile computer storage medium storing computer-executable instructions, which are executed by one or more processors to perform the heterogeneous cooperative detection trajectory planning method described in the first aspect.
[0017] Unlike existing technologies, the present invention has at least the following beneficial effects: This invention enables collaborative detection planning for multiple heterogeneous detection platforms by constructing capability models for at least one aircraft platform and / or at least one ship platform. By introducing a determinism map, initial local observations of the aircraft platform and / or ship platform are obtained, reducing the probability of missing time-sensitive moving targets in the environment. Furthermore, by constructing and training the initial detection model, real-time trajectory planning is achieved, improving the real-time performance and scalability of multi-platform autonomous collaborative detection.
[0018] Furthermore, by designing the detection trajectory to consist only of straight lines and turning arcs (Dubins paths) parallel to the grid when designing the motion space, the planned path can be made without further smoothing. Attached Figure Description
[0019] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments of the present invention will be briefly described below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.
[0020] Figure 1 This is a schematic diagram of the overall process of the heterogeneous cooperative detection trajectory planning method according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating the implementation of the heterogeneous cooperative detection trajectory planning method according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the specific process of step 10 in an embodiment of the present invention; Figure 4 This is a schematic diagram of the specific process of step 20 in an embodiment of the present invention; Figure 5 This is a schematic diagram of the overall detection area after being rasterized and placed in a foot-force coordinate system, provided by an embodiment of the present invention; Figure 6 This is a schematic diagram of the specific process of step 30 in an embodiment of the present invention; Figure 7 This is a framework diagram of the heterogeneous cooperative detection algorithm according to an embodiment of the present invention; Figure 8 This is a schematic diagram of the specific process of step 40 provided in an embodiment of the present invention; Figure 9 This is a schematic diagram of the specific process of step 401 in an embodiment of the present invention; Figure 10 This is a schematic diagram of the reinforcement learning action space definition according to an embodiment of the present invention; Figure 11 This is a schematic diagram of the parallel line coverage method according to an embodiment of the present invention; Figure 12 This is a schematic diagram of the specific process of step 405 in an embodiment of the present invention; Figure 13 This is a schematic diagram of the architecture of a heterogeneous collaborative detection trajectory planning device provided in an embodiment of the present invention. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0022] In the description of this invention, the terms "inner", "outer", "longitudinal", "lateral", "upper", "lower", "top", "bottom", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and do not require that this invention must be constructed and operated in a specific orientation. Therefore, they should not be construed as limiting this invention.
[0023] In this invention, the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined with "first," "second," etc., may explicitly or implicitly include one or more of that feature. In the description of this application, unless otherwise stated, "a plurality of" means two or more.
[0024] In this application, unless otherwise expressly specified and limited, the term "connection" should be interpreted broadly. For example, "connection" can be a fixed connection, a detachable connection, or an integral part; it can be a direct connection or an indirect connection through an intermediate medium. Furthermore, the term "coupled" can refer to an electrical connection that enables signal transmission.
[0025] Furthermore, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0026] Example 1: Traditional methods for multi-platform autonomous collaborative detection suffer from poor scalability and real-time performance when the environment changes due to pre-defined areas and fixed routes. Furthermore, the planned paths require further smoothing in practical applications, and the probability of missing moving, time-sensitive targets in the environment cannot be reduced. While existing technologies offer solutions to improve scalability, they only enable collaborative detection for homogeneous detection platforms and do not reduce the probability of missing targets; the planned paths still require further smoothing.
[0027] To address the issues of poor real-time performance and scalability in multi-platform autonomous collaborative detection, this invention, after modeling the environment, uses reinforcement learning to train the detection platform to achieve real-time planning of coverage paths. This eliminates the need for pre-segmentation of the task region and improves scalability. By employing an environment certainty graph from information graph theory to describe the state space of reinforcement learning, the probability of missing moving, time-sensitive targets in the environment is reduced. Specifically, as... Figure 1 and Figure 2As shown, this embodiment of the invention provides a heterogeneous cooperative detection trajectory planning method, including: Step 10: Construct capability models for at least one aircraft platform and / or at least one ship platform.
[0028] The heterogeneous cooperative detection trajectory planning method of this invention primarily considers the cooperative detection of two heterogeneous detection platforms—aircraft and ships—in the maritime battlefield detection platform. Those skilled in the art can also implement cooperative detection of other heterogeneous detection platforms using the heterogeneous cooperative detection trajectory planning method of this invention without creative effort; this is not limited here. Since this invention requires compatibility with heterogeneous detection platforms, the parameters of the networks corresponding to the heterogeneous detection platforms need to be trained separately. The capabilities of at least one heterogeneous aircraft platform and / or at least one heterogeneous ship platform are modeled to obtain detection platform data.
[0029] Step 20: Initialize the overall detection area and build the initial detection model.
[0030] The entire detection area is rasterized and a two-dimensional rectangular coordinate system is established to obtain available detection resources, namely data from aircraft and ship detection platforms. Initial given position coordinates and initial given headings are set for each detection platform in the overall detection area.
[0031] Step 30: Based on the initial detection model, obtain the initial determination maps of the aircraft platform and / or the ship platform respectively. Based on the capability model and the initial determination maps, obtain the initial local observations of the aircraft platform and / or the ship platform respectively.
[0032] The initial local observations are given initialization data. The initial local observations of each detection platform at a certain moment consist of three parts: given position coordinates, given heading, and initial determination map (the determination map contains the environmental determination matrix). The determination map is the environmental determination map of the environment in which each detection platform is located. In this embodiment of the invention, the determination map is used to represent the detection platform's grasp of the situation information of the corresponding local detection area.
[0033] Step 40: Based on the initial local observations, train the initial detection model to obtain the target detection model.
[0034] In this invention, for heterogeneous detection platforms, the initial detection model has the same network structure, but the parameters of each platform are trained independently. This embodiment employs multi-agent reinforcement learning, using infographics to describe the agents' local observations and the global state of the environment. This improves the scalability of collaborative detection for heterogeneous platforms and reduces the probability of omissions.
[0035] Step 50: Input the actual local observations of at least one aircraft platform and / or at least one ship platform into the target detection model to obtain the optimal detection trajectory.
[0036] In this context, the actual local observations refer to the local observations acquired by the detection platform when actually using the target detection model of this embodiment. When using the target detection model for online planning of heterogeneous collaborative detection trajectories, each detection platform inputs the actual local observations into the Actor network corresponding to the target detection model. The target detection model then makes online decisions to generate the optimal detection trajectory, i.e., the optimal action, which is then applied to the detection area to carry out detection.
[0037] The heterogeneous cooperative detection trajectory planning method of this invention realizes cooperative detection planning for multiple heterogeneous detection platforms by constructing capability models for at least one aircraft platform and / or at least one ship platform. By introducing a determinism map, initial local observations of the aircraft platform and / or ship platform are obtained, reducing the probability of missing time-sensitive moving targets in the environment. By constructing and training the initial detection model, real-time detection trajectory planning is achieved, improving the real-time performance and scalability of multi-platform autonomous cooperative detection.
[0038] This invention first models the cooperative detection process of multiple heterogeneous detection platforms, including capability models of aircraft and / or ship-type detection platforms, and a distributed partially observable Markov decision process based on reinforcement learning. To better illustrate the heterogeneous cooperative detection trajectory planning method of this invention, step 10 of the heterogeneous cooperative detection trajectory planning method in this embodiment is further refined below. Specifically, as follows... Figure 3 As shown, step 10 includes: Step 101: Abstract the maneuverability of the aircraft platform and / or the ship platform to maintain a constant speed at a preset speed. The speed of the aircraft platform is represented as... The speed of a ship-type platform is expressed as .
[0039] Step 102: Characterize the endurance of the aircraft platform and / or the ship platform as the maximum range, where the range is a first distance. The maximum range of the aircraft platform is expressed as... The farthest sailing distance of naval platforms is expressed as .
[0040] Step 103: Abstract the detection range of the aircraft-type platform and / or the ship-type platform into a square with a preset side length, thus obtaining a local detection area. Specifically, the detection range of the aircraft-type platform is abstracted into a square with a side length of... The shape is square; the detection range of the ship-type platform is abstracted as having a side length of... It is a square.
[0041] Step 104: Selectively abstract the local detection probability within the local detection area to either 0 or 1. For both aircraft and ship detection platforms, a binary perception model is used. If a platform can be detected at any location within the local detection area, the corresponding local detection probability is 1; if it cannot be detected outside the local detection area, the corresponding local detection probability is 0.
[0042] Step 105: Obtain the capability model based on the preset speed, the first distance, the local detection area, and the local detection probability. This model is based on the movement speed of the aircraft-type platform. Maximum flight distance of aircraft platforms The side length is Based on the square-shaped local detection area and the corresponding local detection probability, the capability model of aircraft-type platforms is obtained; based on the motion speed of ship-type platforms... The furthest sailing distance of ship-type platforms The side length is The capability model of a ship-type platform is obtained by defining the local detection area of a square and the corresponding local detection probability. The specific motion speed, maximum sailing distance, and side length mentioned above are set by those skilled in the art according to the specific application scenario.
[0043] After modeling the cooperative detection process, it is also necessary to standardize the detection area and initialize the detection resources. For example... Figure 4 As shown, step 20 includes: Step 201: Determine the first greatest common divisor of the preset side length of the local detection area of the aircraft platform and / or the ship platform; set the overall detection area as a two-dimensional rectangular area, rasterize the overall detection area into at least one square with a side length of the first greatest common divisor, and establish a foot force coordinate system of the overall detection area with the first greatest common divisor as the unit length. The foot force coordinate system is used to determine the training local observation or the actual local observation.
[0044] Among them, the first greatest common divisor is the preset side length of the local detection area of the heterogeneous detection platform (aircraft platform and ship platform). and The greatest common divisor of the given information is used, and the side length R of the grid is taken as the first greatest common divisor. The detection area is normalized to a two-dimensional rectangle with length L and width W. Cooperative detection aims to achieve full coverage of this two-dimensional rectangle. To intuitively describe the state of the target detection area and reduce the problem size, the overall detection area is rasterized, with the grid set to squares. For example... Figure 5 As shown, a coordinate system is established with the lower left corner of the overall detection area as the origin, the length and width directions as the x-axis and y-axis, and R as the unit length. Detection resources are initialized by representing the number of available aircraft and ship-type detection platforms as follows: and Then, aircraft platform i in The initial given position coordinates and initial given heading at time are: Ship-type platforms i in The initial given position coordinates and initial given heading at time are: Where, ori is the heading of the detection platform, and in an optional embodiment, it can take the value 0, 1, 2, 3, representing heading up, down, left, and right, respectively.
[0045] Step 202: Treat each of the aircraft-type platforms and / or the ship-type platforms as an agent, model the training local observations and environmental information of all the agents in the foot-force coordinate system as a distributed decision process, and construct the initial detection model based on the distributed decision process. The initial detection model includes the structure of the Actor network and the structure of the Critic network.
[0046] The distributed decision-making process is a distributed partially observable Markov decision-making process. In this embodiment, after modeling the environment, reinforcement learning is used to train the detection platform to achieve real-time planning of the coverage path, eliminating the need for pre-segmentation of the task region and improving scalability. Leveraging the trial-and-error mechanism of reinforcement learning, the optimal policy is learned through continuous interaction with the environment. The definition of the environment is the same as that in reinforcement learning. Multi-agent reinforcement learning algorithms can be divided into centralized and decentralized approaches. Centralized approaches directly extend and apply single-agent algorithms, training multiple agents as a whole, but they are not good at providing individual agent decisions. Decentralized approaches treat each agent as a separate entity, independently generating behavior and optimizing its own policy, which may face non-stationary states of the environment, and the resulting policy is not a global policy. Therefore, this embodiment uses a compromise method: centralized training and decentralized execution. It typically utilizes an Actor-Critic framework, using a centralized Critic network to evaluate distributed Actors, achieving the goal of optimizing multi-agent policies, and exhibiting good algorithm robustness and anti-interference capabilities.
[0047] To address the collaborative detection problem of multiple heterogeneous detection platforms, this invention utilizes the MAPPO (Multi-agent Proximal Policy Optimization) algorithm, which employs centralized training and distributed execution, to independently construct and train Actor-Critic networks for aircraft and / or ship platforms. The Actor-Critic algorithm is a reinforcement learning method combining policy gradient and temporal difference learning, comprising two parts: the Actor and the Critic. The Actor refers to the policy function. In essence, it learns a strategy to maximize rewards. It's used to generate actions and interact with the environment. Critic refers to the value function. The value function of the current policy is estimated, which evaluates the actor's performance and guides its actions in the next stage. Using the value function, the actor-critic algorithm can update parameters step-by-step, without waiting for the end of each iteration. Homogeneous probe platforms (agents) share the network, while heterogeneous probe platforms have consistent network structures but independent parameters. The parameterization policy of the Actor network of the probe platform is... The shared parameters for aircraft platforms are The Actor network and parameters are as follows The Critic network shares parameters for ship-type platforms. The Actor network and parameters are as follows The Critic network is described in this invention. This embodiment uses the actor-critic algorithm to construct the network model and the MAPPO algorithm to improve the real-time performance and scalability of multi-platform autonomous collaborative detection.
[0048] To address the issue that a single agent within a cluster cannot obtain precise state information of other agents and overall environmental information, this invention employs a distributed partially observable Markov decision process for modeling, represented as a seven-tuple. Let I be a finite set of n detection platforms; A represents the overall state information of the detection area; A represents the action space shared by all detection platforms. The actions selected for the i-th detection platform constitute a multi-platform joint action. ; It is the local observation of the detection platform i under the global state s; r=(s,a) is the global reward signal shared by the detection platform; Let be the state transition function of the environment, representing the transition from state s to the next state after performing a global action 'a'. The probability of; This is a discount factor. The parts of the heterogeneous cooperative detection trajectory planning method in this embodiment of the invention that involve prior art can be implemented by those skilled in the art according to specific application scenarios, and are not limited here.
[0049] Since the heterogeneous cooperative detection trajectory planning method in this embodiment of the invention adopts a multi-agent reinforcement learning method, it is necessary to design multi-agent reinforcement learning elements. Regarding the design of multi-agent reinforcement learning elements, the following section introduces the use of infographics to describe the local observations of agents and the global state of the environment during the determination map acquisition process in step 30. In the training process in step 40, the section introduces the design of the action space based on parallel line trajectory features and the design of environmental reward rewards with the goal of full coverage of the detection area and reducing the probability of missing time-sensitive moving targets.
[0050] To better illustrate the heterogeneous cooperative detection trajectory planning method of the present invention, step 30 of the heterogeneous cooperative detection trajectory planning method of the present invention will be further refined below. Specifically, as follows: Figure 6 As shown, step 30 includes: Step 301: Based on the initial detection model, initialize the determination map of each of the aircraft-type platforms and / or the ship-type platforms, and set the determination of each grid of the determination map to a preset initial value to obtain the initial determination map including multiple grids; wherein, the unit length of the grid of the grid is the first greatest common divisor.
[0051] Step 302: Based on the corresponding capability model, obtain the given position coordinates of the aircraft platform and / or the ship platform in the overall detection area and the given heading in the overall detection area.
[0052] Step 303: Based on the initial degree of determination map, the given position coordinates, and the given heading, obtain the initial local observation.
[0053] The dimension of the determinism map is consistent with the dimension of the grid in the overall detection area, and the unit length of the grid is the first greatest common divisor. This facilitates the correspondence between the determinism map of the training local observations and the overall detection area, i.e., the detection trajectory. Since the detection platform uses a binary sensing model, it can detect any position within the local detection area, so the corresponding local detection probability is 1. If it cannot detect anything outside the local detection area, the corresponding local detection probability is 0. Therefore, the determinism value range of the grid map at a given position (x, y) at time t in the corresponding determinism map is [0, 1]. 0 indicates that the detection platform has no control at all, and 1 indicates that the detection platform has complete control. In an optional embodiment, the preset initial value of the certainty of each grid map is 0.5, but those skilled in the art can also set it according to the needs of specific application scenarios. As time progresses and the detection activities of the detection platform continue, the grid map in the certainty map is updated after each local action and interaction between the detection platform and the environment. Based on the certainty map of each detection platform at a certain moment, and the preset speed, first distance, local detection area, and local detection probability of each detection platform's capability model, the corresponding training local observation is calculated.
[0054] The framework diagram of the heterogeneous cooperative detection algorithm in this embodiment of the invention is as follows: Figure 7 As shown, the number of aircraft-type and ship-type detection platforms is represented as follows: and The training local observations of aircraft-type and ship-type detection platforms i are respectively represented as: and The local actions of the detection platform i for aircraft and ships are respectively represented as: and Actor networks for aircraft platforms ( Figure 7 The parameters shared by the Actor_a network are: Actor networks for ship-type platforms ( Figure 7 The shared parameters of the Actor_s network are Local movements of the detection platform and After interacting with the environment, the various platforms at time t+1 are obtained. Local detection results , and environmental rewards , Critic network outputs to various platforms Corresponding current action value function , .
[0055] To better illustrate the heterogeneous cooperative detection trajectory planning method of the present invention, step 40 of the heterogeneous cooperative detection trajectory planning method of the present invention will be further refined below. Specifically, as follows: Figure 8 As shown, step 40 includes: Step 401: Use the initial local observation as the training local observation at time t; input the training local observation into the Actor network of the initial detection model, and output the local action at time t after the Actor network makes a decision; based on the local action, obtain the joint action of all the aircraft-type platforms and / or the ship-type platforms at time t.
[0056] Among them, the local actions of each detection platform i at time t Combining actions at time t Since this embodiment of the invention uses a network model training method to achieve heterogeneous cooperative detection trajectory planning, training data exists. The training local observations are the training data used by the detection platform for training. The initial training local observations are the initial local observations, which are the data given during initialization. Subsequent training local observations are output by the environment based on the corresponding determinism map after interaction with the environment. The Actor network receives a state each time, i.e., the training local observation input to detection platform i at time t. The selection probability of local action selection is obtained. The selection probability is sampled to generate and select the corresponding local action. .
[0057] Step 402: Input the joint action into the environment volume of the overall detection area. The environment volume outputs the training local observation at time t+1 and the environment reward at time t based on the corresponding determination map.
[0058] The environment is a data-driven model (or system) of the actual overall detection area, abstracted from real-world data. The environment transitions to the next state based on its dynamic model. In optional embodiments, during practical application, the local actions of each detection platform are input into the environment to improve real-time performance.
[0059] Based on the training data, the parameters of the Actor network are updated using gradient ascent, and the parameters of the Critic network are updated using gradient descent. The network is trained based on the MAPPO algorithm, iteratively optimizing the network parameters. Those skilled in the art can set the training parameters according to the actual application scenario, including setting the maximum number of training episodes and the number of training steps per episode. The parameters of the Actor network for aircraft-type platforms are initialized. Parameters of the Critic network Parameters of Actor networks for ship-type platforms Parameters of the Critic network Set the learning rate .
[0060] Step 403: Combine the training local observations of all detection platforms at time t into a global state, and input the global state into the Critic network of the initial detection model to obtain the current action value function.
[0061] The global state is the result of training local observations that merge all heterogeneous detection platforms.
[0062] Initialize the global state in each training epoch (episodes). and local observations from each detection platform The Actor network is trained based on local observations from each probe at time t. Output the local action at time t. Joint operations with all detection platforms The coordinated operation of all detection platforms The global state at time t+1 is obtained after interacting with the environment. Various detection platforms Local detection results and environmental rewards The Critic network is based on the global state at time t. Output the current action value function This invention, through learning the optimal current action value function, determines how the agent selects the optimal action (the action corresponding to the optimal current action value function) under different states, thus obtaining the optimal detection trajectory.
[0063] Step 404: The training local observation at time t, the local action, the global state, the current action value function, the training local observation at time t+1, and the environmental reward are combined to form the current experience tuple; the current experience tuple is stored in the experience pool; when the experience pool is full, the current experience tuple is selected to be retained in the experience pool based on importance sampling.
[0064] The current empirical tuple also includes the global state at time t, that is, the current empirical tuple is composed of the global state at time t. Training local observations at time t Local action at time t Global state at time t+1 The training local observations at time t+1 and environmental reward at time t Composition. Importance sampling is a commonly used Monte Carlo integral method for estimating the expectation of a probability density function.
[0065] Each update calculation requires a set of current experience tuples. An experience pool is set up to store interaction information, which is used to store sample sequences generated by the interaction between local observations and the environment, and provided to the Actor network and Critic network for training the corresponding network parameters.
[0066] Step 405: Obtain the advantage function based on at least one selected current empirical tuple, and obtain the target detection model based on the advantage function.
[0067] Here, the advantage function expresses the advantage of a local action 'a' relative to the average of all local actions in state S. During the training of the initial probe model, a centralized Critic network with global information is trained to estimate the state value and advantage function of each agent.
[0068] To better illustrate the heterogeneous cooperative detection trajectory planning method of the present invention, the following will further refine step 301 of the heterogeneous cooperative detection trajectory planning method of the present invention. Specifically, as follows: Figure 9 As shown, in step 401, the trained local observations are input into the Actor network of the initial detection model. The local actions output by the Actor network at time t after making a decision include: Step 4011: When the training local observation maintains the original speed and direction and continues to move forward, the local action is obtained as the first action value.
[0069] Step 4012: The training local observation is obtained as the second action value when the local action is obtained after turning left along the Dubins path and maintaining straight-line movement.
[0070] Step 4013: The training local observation is obtained as the third action value when the local action is obtained after turning right on the Dubins path and maintaining straight-line movement.
[0071] Step 4014: The training local observation is obtained when the local action is the fourth action value after turning left on the Dubins path and maintaining the original speed while moving in a straight line in the opposite direction.
[0072] Step 4015: The training local observation is obtained when the local action is the fifth action value after turning right on the Dubins path and maintaining the original speed while moving in a straight line in the opposite direction.
[0073] The first mean is obtained by averaging the length and width of the overall detection area. The execution time of each action is obtained by averaging the first mean and the first greatest common divisor. The Dubins path is the shortest path connecting two points in a two-dimensional plane (i.e., the xy plane) under the condition of satisfying curvature constraints and specified tangent directions at the beginning and end.
[0074] The first action value, second action value, third action value, fourth action value, and fifth action value are used to define the action selected by the detection platform i in the action space A at time t. That is, the detection trajectory, because the agent selects the action. To obtain the global state at time t .like Figure 10 As shown, in an optional embodiment, the first action value is 1, the second action value is 2, the third action value is 3, the fourth action value is 4, and the fifth action value is 5. The action space A of this embodiment of the invention is A={1, 2, 3, 4, 5}. It is worth noting that, as... Figure 10 As shown, if the starting point of the position coordinates is close to the edge of the overall detection area, such as only one unit length R away from the edge, when selecting a local action, the position coordinates are turned around with a radius of one unit length R according to the Dubins path, so that the detection trajectory returns to the overall detection area as quickly as possible.
[0075] Traditional detection trajectory planning methods commonly use figure-eight, inner spiral, and parallel line patterns. Among these, the figure-eight and inner spiral patterns occur less frequently than the parallel line pattern, and the parallel line pattern does not match the motion trajectory of aircraft, ships, and other equipment in actual applications. Therefore, the planned detection trajectory needs further smoothing. Figure 11 As shown, this embodiment of the invention adaptively improves the detection trajectory for different moving speeds of heterogeneous detection platforms by constructing a reinforcement learning action space. The detection trajectory in this embodiment consists only of straight lines parallel to the grid and turning arcs (Dubins paths), which can be used directly in practical applications without further smoothing, thus improving efficiency. This embodiment of the invention combines the physical law that the detection platform consumes a lot of energy when turning with the commonly used parallel line trajectory planning method to design an action space based on the Dubins path, reducing the number of turns and making the obtained detection trajectory easy to use in practical applications.
[0076] This invention employs a determinism graph from information graph theory to establish local observations (training local observations during the training phase and actual local observations during the practical application phase), using these local observations as the state space corresponding to the agent. The determinism graph guides the detection platform to perform detection in directions with high determinism gradients, thereby minimizing the uncertainty of the environment as quickly as possible. To better illustrate the heterogeneous cooperative detection trajectory planning method of this invention, the determinism graph in step 402 of the heterogeneous cooperative detection trajectory planning method of this embodiment will be further explained below. Specifically, in the determinism graph, the grid diagram at a given position (x, y) at time t is... .
[0077] In this context, due to the presence of moving, time-sensitive targets, the environment is actually dynamic. While the certainty of a grid previously detected by the platform increases at the time of detection, its certainty gradually decreases if it is not subsequently covered by the platform. This invention addresses the problem of high detection misses caused by moving, time-sensitive targets potentially moving from unsearched areas to already searched areas by introducing a certainty map and setting specific conditions for attenuation and / or increase through gridding.
[0078] If the grid at a given location (x, y) is not covered by the corresponding local detection region at time t, the certainty of the corresponding certainty map continuously decreases; wherein, the certainty decreases according to the decay factor η, resulting in... The raster image at time +1 is .
[0079] If the grid at a given location (x, y) is covered by the corresponding local detection region at time t, then the certainty of the corresponding determination map continuously increases, resulting in... The raster image at time +1 is .
[0080] This invention addresses the problem of missed detection of time-sensitive moving targets when the target location is unknown. It introduces environmental volume certainty by using an infographic to describe local observations and global states. The certainty of the detected areas decreases and / or increases over time. This allows the detection platform to detect undetected areas and guides the platform to revisit areas that have not been detected for a long time but may contain time-sensitive moving targets, thereby reducing the probability of missed target detection.
[0081] To better illustrate the heterogeneous cooperative detection trajectory planning method of the present invention, the environmental reward in step 402 of the heterogeneous cooperative detection trajectory planning method of the present invention will be further explained below. Specifically, the environmental reward includes a certainty reward and a range reward, wherein: Based on the corresponding certainty map, the sum of the incremental certainty of all grids within the overall detection area covered at time t during the joint action execution is obtained, thus yielding the certainty reward. The certainty reward is... Set to perform local actions During the process, the overall detection area covered at time t The sum of the increments of the determinism of all the grid cells within the range, i.e. When certainty decays, the increment becomes negative.
[0082] Based on the first distance and the corresponding preset speed, a range bonus is selectively awarded. ;in, For preset speed, This represents the maximum range. The range bonus is set as the score for the cumulative maneuvering distance of the detection platform after performing the action. If the cumulative maneuvering distance exceeds the maximum range, the value is reduced by 1.
[0083] The environment reward is obtained by adding the product of the certainty reward and its corresponding certainty weight, and the product of the range reward and its corresponding range weight. The detection platform i performs a local action at time t. The environmental rewards obtained later are recorded as The certainty reward The product of the corresponding certainty weights The aforementioned range reward With corresponding range weights The products are added together to obtain environmental rewards. ,Right now + .
[0084] This invention provides an embodiment that guides the detection platform to achieve coverage of unknown local detection areas as quickly as possible, given its limited capabilities, by setting certainty rewards and range rewards.
[0085] To better illustrate the heterogeneous cooperative detection trajectory planning method of the present invention, step 405 of the heterogeneous cooperative detection trajectory planning method of the present invention will be further refined below. Specifically, as follows: Figure 12 As shown, step 405 includes: Step 4051: Based on the selected at least one current experience tuple, obtain at least one environmental reward for the current experience tuple; using a preset advantage estimation algorithm, obtain the long-term discounted reward of the corresponding local action based on the environmental reward.
[0086] The pre-defined advantage estimation algorithm is the Generalized Advantage Estimator (GAE). This is based on the discount factor in the distributed decision-making process. The GAE algorithm estimates the advantage function and obtains the corresponding long-term discounted return based on the accumulated environmental reward in each round.
[0087] Step 4052: Based on the long-term discount reward and the action value function of the current experience tuple, obtain the advantage function of the corresponding local action relative to the average reward of all local actions.
[0088] Step 4053: Calculate the loss functions of the Actor network and the Critic network based on the advantage function, and update the network parameters of the Actor network and the Critic network through backpropagation.
[0089] Step 4054: When the parameters of the Actor network and the Critic network converge or reach a preset number of rounds, save the parameters of the Actor network and the Critic network to obtain the target detection model.
[0090] The preset number of rounds is set by those skilled in the art based on actual usage scenarios and is not limited here.
[0091] The heterogeneous cooperative detection trajectory planning method of this invention constructs capability models for at least one aircraft platform and / or at least one ship platform, respectively. Using a reinforcement learning Actor-Critic framework, the parameters of the Actor and Critic networks for the heterogeneous detection platforms are trained, enabling cooperative detection planning for multiple heterogeneous detection platforms. By introducing a determinism graph to describe the state space of reinforcement learning, training local observations for the aircraft platform and / or ship platform are obtained based on the corresponding determinism graph. The determinism is incorporated into the setting of the reinforcement learning reward function, increasing the revisit probability of already detected areas and reducing the probability of missing moving, time-sensitive targets in the environment. By constructing and training an initial detection model based on the MAPPO algorithm, real-time detection trajectory planning is achieved, improving the real-time performance and scalability of multi-platform autonomous cooperative detection. By designing the action space in conjunction with the structure of the actual detection path, the detection trajectory is designed to consist only of straight lines parallel to the grid and turning arcs (Dubins paths), eliminating the need for further smoothing of the planned path (detection trajectory).
[0092] Example 2: like Figure 13 The diagram shown is a schematic representation of the architecture of a heterogeneous cooperative detection trajectory planning device according to an embodiment of the present invention. The heterogeneous cooperative detection trajectory planning device according to an embodiment of the present invention includes one or more processors 21 and a memory 22. Wherein, Figure 13 Take a processor 21 as an example.
[0093] Processor 21 and memory 22 can be connected via a bus or other means. Figure 13 Taking the example of a connection between China and Israel via a bus.
[0094] The memory 22, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs and non-volatile computer-executable programs, such as the heterogeneous cooperative detection trajectory planning method in Embodiment 1. The processor 21 executes the heterogeneous cooperative detection trajectory planning method by running the non-volatile software programs and instructions stored in the memory 22.
[0095] Memory 22 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, memory 22 may optionally include memory remotely located relative to processor 21, which can be connected to processor 21 via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0096] The program instructions / modules are stored in the memory 22. When executed by one or more processors 21, they perform the heterogeneous cooperative detection trajectory planning method described in Embodiment 1 above, for example, the method described above. Figures 1-4 , Figure 6 , Figures 8-9 and Figure 12 The steps shown.
[0097] It is worth noting that the information interaction and execution process between the modules and units in the above-mentioned device and system are based on the same concept as the processing method embodiment of the present invention. For details, please refer to the description in the method embodiment of the present invention, and will not be repeated here.
[0098] Those skilled in the art will understand that all or part of the steps in the various methods of the embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, which may include: read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc.
[0099] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A heterogeneous cooperative detection trajectory planning method, characterized in that, include: Construct capability models for at least one aircraft platform and / or at least one ship platform; Initialize the overall detection area and construct the initial detection model; This includes: determining the first greatest common divisor of the preset side lengths of the local detection areas of the aircraft-type platform and / or the ship-type platform; setting the overall detection area as a two-dimensional rectangular area, rasterizing the overall detection area into at least one square with a side length of the first greatest common divisor, establishing a foot-force coordinate system of the overall detection area with the first greatest common divisor as the unit length, the foot-force coordinate system being used to determine training local observations or actual local observations; treating each of the aircraft-type platform and / or the ship-type platform as an agent, modeling the training local observations and environmental information of all the agents in the foot-force coordinate system as a distributed decision-making process, and constructing the initial detection model based on the distributed decision-making process, the initial detection model including the structure of an Actor network and the structure of a Critic network; Based on the initial detection model, initial determination maps of the aircraft platform and / or the ship platform are obtained respectively. Based on the capability model and the initial determination maps, initial local observations of the aircraft platform and / or the ship platform are obtained respectively. Based on the initial local observations, the initial detection model is trained to obtain the target detection model; wherein, the initial local observations are used as training local observations at time t; the training local observations are input into the Actor network of the initial detection model, and the Actor network outputs the local action at time t after making a decision; The actual local observations of at least one aircraft platform and / or at least one ship platform are respectively input into the target detection model to obtain the optimal detection trajectory; The step of inputting the trained local observations into the Actor network of the initial detection model, and the Actor network outputting the local action at time t after making a decision, includes: The training local observation is to maintain the original speed and direction and continue to move forward, and the obtained local action is the first preset value; The training local observation is obtained by turning left along the Dubins path and maintaining straight-line motion, and the local action is the second preset value. The training local observation is obtained by turning right along the Dubins path and maintaining straight-line motion, and the local action is the third preset value. The training local observation is obtained when the local action is the fourth preset value after turning left on the Dubins path and moving in a straight line in the opposite direction while maintaining the original speed. The training local observation is obtained when the local action is the fifth preset value after turning right on the Dubins path and maintaining the original speed while moving in a straight line in the opposite direction.
2. The heterogeneous cooperative detection trajectory planning method according to claim 1, characterized in that, The capability models for constructing at least one aircraft platform and / or at least one ship platform include: The maneuverability of the aircraft platform and / or the ship platform is abstracted as maintaining a constant speed at a preset speed. The endurance of the aircraft platform and / or the ship platform is characterized as the farthest sailing distance, and the distance is the first distance; The detection range of the aircraft platform and / or the ship platform is abstracted as a square with a preset side length to obtain a local detection area; The probability of local detection within the local detection region is selectively abstracted to 0 or 1; The capability model is obtained based on the preset speed, the first distance, the local detection area, and the local detection probability.
3. The heterogeneous cooperative detection trajectory planning method according to claim 1, characterized in that, The step of obtaining initial determination maps for the aircraft-type platform and / or the ship-type platform based on the initial detection model, and obtaining initial local observations for the aircraft-type platform and / or the ship-type platform based on the capability model and the initial determination maps, includes: Based on the initial detection model, a determination map for each of the aircraft-type platforms and / or the ship-type platforms is initialized, and the determination of each grid of the determination map is set to a preset initial value to obtain the initial determination map including multiple grids; wherein, the unit length of the grid of the grid is the first greatest common divisor; Based on the corresponding capability model, the given position coordinates of the aircraft platform and / or the ship platform in the overall detection area and the given heading in the overall detection area are obtained respectively. The initial local observations are obtained based on the initial determination map, the given position coordinates, and the given heading.
4. The heterogeneous cooperative detection trajectory planning method according to claim 1, characterized in that, The step of training the initial detection model based on the initial local observations to obtain the target detection model includes: Based on the local actions, the combined actions of all the aircraft-type platforms and / or the ship-type platforms at time t are obtained; The combined action is input to the environment of the overall detection area, and the environment outputs the training local observation at time t+1 and the environment reward at time t based on the corresponding determination map. The training local observations of all detection platforms at time t are combined into a global state, and the global state is input into the Critic network of the initial detection model to obtain the current action value function. The training local observation at time t, the local action, the global state, the current action value function, the training local observation at time t+1, and the environmental reward are combined to form the current experience tuple; the current experience tuple is stored in the experience pool, and when the experience pool is full, the current experience tuple is selected to be retained in the experience pool based on importance sampling. Based on at least one selected current empirical tuple, an advantage function is obtained, and based on the advantage function, the target detection model is obtained.
5. The heterogeneous cooperative detection trajectory planning method according to claim 4, characterized in that, In the determinism map, the raster image of a given position (x, y) at time t is: ; If the grid at a given location (x, y) is not covered by the corresponding local detection region at time t, the certainty of the corresponding certainty map continuously decreases; wherein, the certainty decreases according to the decay factor η, resulting in... The raster image at time +1 is ; If the grid at a given location (x, y) is covered by the corresponding local detection region at time t, then the certainty of the corresponding determination map continuously increases, resulting in... The raster image at time +1 is .
6. The heterogeneous cooperative detection trajectory planning method according to claim 4, characterized in that, The environmental rewards include certainty rewards and range rewards, wherein: Based on the corresponding certainty map, the sum of the increments of certainty corresponding to all grids within the overall detection area covered at time t during the execution of the joint action is obtained, and the certainty reward is obtained. Based on the first distance and the corresponding preset speed, a range bonus is selectively awarded. ;in, For preset speed, This represents the furthest possible sailing distance. The environmental reward is obtained by adding the product of the certainty reward and the corresponding certainty weight, and the product of the range reward and the corresponding range weight.
7. The heterogeneous cooperative detection trajectory planning method according to claim 4, characterized in that, The step of obtaining the advantage function based on at least one selected current experience tuple, and obtaining the target detection model based on the advantage function, includes: Based on at least one selected current experience tuple, at least one environmental reward for the current experience tuple is obtained; using a preset advantage estimation algorithm, the long-term discounted reward for the corresponding local action is obtained based on the environmental reward; Based on the long-term discount reward and the action value function of the current experience tuple, the advantage function of the corresponding local action relative to the average reward of all local actions is obtained; Calculate the loss functions of the Actor network and the Critic network based on the advantage function, and update the network parameters of the Actor network and the Critic network through backpropagation; When the parameters of the Actor network and the Critic network converge or reach a preset number of rounds, the parameters of the Actor network and the Critic network are saved to obtain the target detection model.
8. A heterogeneous cooperative detection trajectory planning device, characterized in that, It includes at least one processor and a memory, which are connected via a data bus. The memory stores instructions that can be executed by the at least one processor. After being executed by the processor, the instructions are used to complete the heterogeneous cooperative detection trajectory planning method according to any one of claims 1-7.