Cluster collaborative pursuit decision optimization method and device

By building a knowledge graph and relationship matrix and optimizing the pursuit strategy, the problem of collaborative pursuit decisions in complex environments is solved, and efficient pursuit in remote areas is achieved.

CN120373468APending Publication Date: 2025-07-25NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510643282.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Existing methods cannot achieve collaborative pursuit and escape decisions in complex environments, especially in complex environments in remote areas. Traditional prediction methods are difficult to accurately predict the path of escape entities, and traditional methods fail to effectively combine environmental factors and multi-party knowledge, resulting in unreasonable selection of pursuit strategies.

Method used

By constructing the initial knowledge graph and relationship matrix, updating entity attribute characteristics, coding relationship weights based on the key subgraphs and relational logic graphs of escape entities, and combining the autonomous reward function, the pursuit strategy is optimized to achieve efficient encirclement.

Benefits of technology

Under the condition of small sample size, the accuracy of entity relationship prediction is improved, the calculation loss is reduced, and the flexibility and pursuit success rate of the cluster system in different scenarios is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373468A_ABST
    Figure CN120373468A_ABST
Patent Text Reader

Abstract

The invention discloses a cluster collaborative pursuit decision optimization method and device, and relates to the field of local security and protection. The problem that an existing method cannot realize collaborative pursuit decision in a complex environment is solved. The method comprises the following steps: initially obtaining initial attribute characteristics of a plurality of entities in a pursuit environment within a set time according to a task instruction, obtaining an initial knowledge graph and an initial relation matrix, and taking a first escape entity as a center node in the pursuit environment to construct a sub-problem about the first escape entity; if it is determined that the first pursuit entity and the second pursuit entity complete surrounding of the first escape entity according to the updated position of the first pursuit entity, the updated position of the second pursuit entity and the surrounding circle condition, it is determined that pursuit of the first escape entity is completed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of on-site security, and more particularly to a method and device for optimizing cluster collaborative pursuit and escape decision-making. Background Art

[0002] With the development and maturity of related research technologies of logical reasoning in the application of unmanned systems, while the application fields and scopes of entities are constantly expanding, the potential threats brought to society are also increasing. It is very necessary to intercept illegal entities efficiently and accurately. However, it should be noted that when illegal entities escape, they will go to remote areas with more complex and hidden terrain, far from urban areas, thus increasing the probability of escape. The reason why it is difficult for entities to form an efficient encirclement in these complex environments is that the situation is complex and changeable, and the sample size is small. Traditional training methods are difficult to cover comprehensively, resulting in entities being difficult to learn effective coping knowledge. Therefore, researching a collaborative pursuit and escape decision optimization system in complex environments will play a crucial role.

[0003] In the current research on entity pursuit and escape decision-making, the generally adopted idea is that the pursuing entity first predicts the path of the escaping entity, and learns the pursuit strategy based on the current position of the escaping entity and the predicted position. It can be seen that the accuracy of the prediction directly affects the choice of the pursuit strategy. In data-driven prediction methods, the law is usually learned based on a large amount of data, and similar information in the current state information and historical information is used to achieve prediction. However, due to the small amount of data in some remote areas, it is difficult to learn useful information from historical data, further increasing the difficulty of prediction. At the same time, traditional prediction methods such as least squares and neural network fitting only consider the historical state of the escaping entity, and do not consider the influence of environmental factors on it. For example, it is impossible to fly over mountainous areas. Therefore, how to couple multiple elements to achieve accurate prediction directly affects the success or failure of the task.

[0004] In the process of strategy selection of the pursuing entity after prediction, the traditional method sets a reward function for each entity, and selects the optimal reward by exploring the rewards obtained under different strategies. However, it should be noted that the current states of each pursuing entity are different, and it is unreasonable to set the reward function to be the same. For example, the pursuing entity closer to the escaping entity needs to complete the encirclement formation as soon as possible, while the pursuing entity farther from the escaping entity needs to accelerate to catch up with the escaping entity and then form the encirclement formation. Therefore, how to combine multiple knowledge sources to enable entities to have the ability to infer autonomous reward functions can better play the advantages of unmanned systems. Summary of the Invention

[0005] Aiming at the problem that the existing methods cannot achieve collaborative pursuit and escape decision-making in complex environments, the embodiments of the present invention provide a method and device for optimizing cluster collaborative pursuit and escape decision-making.

[0006] The embodiments of the present invention provide a method for optimizing the decision-making of cluster collaborative pursuit and escape, including:

[0007] Obtaining the initial attribute features of multiple entities in the pursuit and escape environment within a set time according to the task instruction, and obtaining the initial knowledge graph and the initial relationship matrix; at the current moment, updating the initial knowledge graph and the initial relationship matrix according to the current attribute features of the multiple entities;

[0008] Determining a first escaping entity as the central node in the pursuit and escape environment, constructing a sub-problem regarding the first escaping entity, and successively determining the relationship weights related to the first escaping entity and the logical feature encoding of the first escaping entity based on the key sub-graph and the relationship logic graph including the first escaping entity; obtaining the final information encoding of the sub-problem based on the logical feature encoding of the first escaping entity, the final time information encoding, and the final coupling information encoding; predicting the most probable position of the first escaping entity at the next moment according to the final information encoding of the sub-problem;

[0009] Determining the revenue strategies of the first pursuer entity and the second pursuer entity according to the revenue coefficient matrix and the reward and punishment conditions of the first pursuer entity and the second pursuer entity, and updating the target positions of the first pursuer entity and the second pursuer entity with the highest revenue strategy; if it is determined that the first pursuer entity and the second pursuer entity complete the encirclement of the first escaping entity according to the updated position of the first pursuer entity, the updated position of the second pursuer entity, and the encirclement condition, it is determined that the pursuit of the first escaping entity is completed.

[0010] The embodiments of the present invention provide a device for optimizing the decision-making of cluster collaborative pursuit and escape, including:

[0011] An updating unit, configured to obtain the initial attribute features of multiple entities in the pursuit and escape environment within a set time according to the task instruction, and obtain the initial knowledge graph and the initial relationship matrix; at the current moment, update the initial knowledge graph and the initial relationship matrix according to the current attribute features of the multiple entities;

[0012] A predicting unit, configured to determine a first escaping entity as the central node in the pursuit and escape environment, construct a sub-problem regarding the first escaping entity, and successively determine the relationship weights related to the first escaping entity and the logical feature encoding of the first escaping entity based on the key sub-graph and the relationship logic graph including the first escaping entity; obtain the final information encoding of the sub-problem based on the logical feature encoding of the first escaping entity, the final time information encoding, and the final coupling information encoding; predict the most probable position of the first escaping entity at the next moment according to the final information encoding of the sub-problem;

[0013] The pursuit unit is used to determine the revenue strategies of the first pursuit entity and the second pursuit entity according to the revenue coefficient matrix and the reward and punishment conditions of the first pursuit entity and the second pursuit entity, update the target positions of the first pursuit entity and the second pursuit entity with the highest revenue strategy; if it is determined that the first pursuit entity and the second pursuit entity have completed the encirclement of the first escape entity according to the updated position of the first pursuit entity, the updated position of the second pursuit entity and the encirclement condition, it is determined that the pursuit of the first escape entity is completed.

[0014] An embodiment of the present invention provides a computer device, which includes a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor executes the cluster collaborative pursuit and escape decision optimization method described in any one of the above.

[0015] An embodiment of the present invention provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor executes the cluster collaborative pursuit and escape decision optimization method described in any one of the above.

[0016] An embodiment of the present invention provides a cluster collaborative pursuit and escape decision optimization method and device. By introducing temporal features, the method deeply mines the implicit information contained in the time series, effectively reducing the dependence on other environmental data, enabling the prediction of future relationships between entities under small sample conditions; furthermore, based on weight screening, key subgraphs are extracted, and a relational logic graph system is constructed to depict the relationships between entities. By converting the relational logic graph into a mathematical expression, the relationship scores between different entities are calculated quantitatively. For specific sub-problems, the entity with the highest relationship score is selected as the answer to fill in, significantly improving the accuracy of predicting future relationships between entities and reducing unnecessary computational losses; further, an autonomous reward function reasoning mechanism is designed, which has the ability to autonomously select a more suitable revenue for itself, thereby improving the flexibility of the cluster system in dealing with different scenarios. It solves the problem that existing methods cannot achieve collaborative pursuit and escape decisions in complex environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0018] Figure 1 It is a flowchart of the cluster collaborative pursuit and escape decision optimization method provided by an embodiment of the present invention;

[0019] Figure 2ASchematic diagram of the full-background entity graph in the multi-layer knowledge graph provided by the embodiment of the present invention;

[0020] Figure 2B Schematic diagram of the key sub-graph in the multi-layer knowledge graph provided by the embodiment of the present invention;

[0021] Figure 2C Schematic diagram of the relationship logic in the multi-layer knowledge graph provided by the embodiment of the present invention;

[0022] Figure 3 Schematic diagram of the link prediction result of the temporal knowledge graph provided by the embodiment of the present invention;

[0023] Figure 4A Schematic diagram of the selection of the discrete policy space of the agent provided by the embodiment of the present invention;

[0024] Figure 4B Schematic diagram of the selection of the continuous policy space of the agent provided by the embodiment of the present invention;

[0025] Figure 5 Schematic diagram of a cluster pursuit and escape provided by the embodiment of the present invention;

[0026] Figure 6 Schematic diagram of a successful cluster pursuit and escape provided by the embodiment of the present invention;

[0027] Figure 7 Schematic diagram of the structure of the cluster collaborative pursuit and escape decision optimization device provided by the embodiment of the present invention. Detailed implementation manners

[0028] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0029] The following will be described in detail by taking Figures 1 to 6 as an example, the cluster collaborative pursuit and escape decision optimization method provided by the embodiment of the present invention. As Figure 1 shown, the method includes the following steps:

[0030] Step 101, obtain the initial attribute features of multiple entities in the pursuit and escape environment within a set time according to the task instruction and obtain the initial knowledge graph and the initial relationship matrix; at the current moment, update the initial knowledge graph and the initial relationship matrix according to the current attribute features of the multiple entities

[0031] Step 102: Determine the first escaping entity as the central node in the escape and pursuit environment, construct sub-problems regarding the first escaping entity, and successively determine the relationship weights related to the first escaping entity and the logical feature encoding of the first escaping entity based on the key sub-graph and the relationship logic graph including the first escaping entity; obtain the final information encoding of the sub-problem based on the logical feature encoding of the first escaping entity, the final time information encoding, and the final coupling information encoding; predict the most probable position of the first escaping entity at the next moment according to the final information encoding of the sub-problem.

[0032] Step 103: Determine the revenue strategies of the first pursuer entity and the second pursuer entity according to the revenue coefficient matrix and the reward and punishment conditions of the first pursuer entity and the second pursuer entity, and update the target positions of the first pursuer entity and the second pursuer entity with the highest revenue strategy; if it is determined that the first pursuer entity and the second pursuer entity complete the encirclement of the first escaping entity according to the updated position of the first pursuer entity, the updated position of the second pursuer entity, and the encirclement condition, it is determined that the pursuit of the first escaping entity is completed.

[0033] It should be noted that in the method provided by the embodiments of the present invention, a three-dimensional example of cluster escape and pursuit is Figure 5 as shown. The figure includes pursuer entities shown as solid circles, escaping entities shown as hollow circles, and reference entity. Among them, the reference entity includes roads with black dotted lines at the edges, mountains, rivers, lakes, etc. shown as black closed solid lines. The encirclement is a black solid curve formed by the pursuer entities around the escaping entity, as Figure 6 shown. For the cluster cooperative escape and pursuit decision optimization method provided by the embodiments of the present invention, starting from a random initial state, as Figure 5 shown, successfully surrounding the escaping entity means the task is successful, as Figure 6 shown.

[0034] In step 101, according to the task instruction, initially obtain the initial attribute features of multiple entities within the set time, such as Figure 5 the initial attribute features of the pursuer entities, the initial attribute features of the escaping entities, and the initial attribute features of the reference entities included. According to the initial attribute features of the multiple entities, an initial knowledge graph can be obtained, and then based on the initial knowledge graph, the relationship matrix among the pursuer entities, the escaping entities, and the reference entities is determined.

[0035] For example, as Figure 5 shown, assume there are m pursuer entities, n escaping entities, and the reference entities include roads, mountains, rivers, lakes, stores, houses, etc. Specifically, at the current time step t, the numbers of the pursuer entities are respectively R = {r 1 , r 2 ,..., r m}, and their state position coordinates are The escape entity numbers are respectively B = {b 1 , b 2 ,..., b n}, and the state position coordinates are The number of the road is road k , the numbers of mountains, rivers, and lakes are terrain l , and the numbers of stores and houses are house s . Among them, m represents the total number of pursuit entities, R represents the set of numbers of all pursuit entities, r i represents the number of the i-th pursuit entity, i ∈ R; n represents the total number of escape entities, B represents the set of numbers of all escape entities, bj represents the number of the j-th escape entity, j ∈ B; k represents the set of road numbers, k = {1,..., q}, q represents the total number of roads; l represents the set of special terrain numbers, l = {1,..., p}, p represents the total number of special terrains; s represents the set of special building numbers, s = {1,..., g}, g represents the total number of special buildings.

[0036] In practical applications, the initial attribute features of pursuit entities and escape entities at least include numbers, initial state coordinates, and actions.

[0037] Furthermore, based on the initial attribute features of multiple entities in the initially obtained environment, an initial knowledge graph that can reflect each entity and its relationships is constructed. Specifically, the initially obtained pursuit entities (such as R = {r 1 , r 2 ,..., r m}), escape entities (such as B = {b 1 , b 2 ,..., b n}) and the initial characteristic attributes of the reference entity are stored as entities in the knowledge graph, and their initial attribute features, such as position coordinates, numbers, lengths, etc., are recorded. An initial knowledge graph of entities and their initial attribute features as shown in Table 1 can be obtained:

[0038] Table 1: Initial Knowledge Graph

[0039] entity attribute feature escaping entity number, status (x, y, z) coordinates, time t, action pursuing entity number, status (x, y, z) coordinates, time t, action road road ID, starting position (x, y, z), length, direction (degrees) traffic facilities (traffic lights, parking lots) facility ID, position (x, y, z) buildings (shops, residences) building ID, position (x, y, z) special terrains (rivers, mountains, lakes) terrain ID, position (x, y, z) … …

[0040] Furthermore, according to the initial knowledge graph, a relationship matrix between pursuit entities, escape entities, and reference entities is determined. In the embodiments of the present invention, various relationships between pursuit entities and escape entities, pursuit entities and reference entities, and escape entities and reference entities can be obtained according to the initial knowledge graph, and a relationship matrix is obtained according to the various relationships included therein. Among them, the various relationships at least include any one or more combinations of adjacency relationships, pursuit-evasion relationships, passing relationships, and inclusion relationships.

[0041] Specifically, the adjacency relationship can be determined by calculating the spatial distances and spatial thresholds between the pursuing entity and the escaping entity, between the pursuing entity and the reference entity, and between the escaping entity and the reference entity. In practical applications, if there is a wall between two entities, the spatial distance represents the shortest spatial distance along the wall between the two entities to reach the opposite target; if the spatial distance between two entities is less than the spatial threshold then these two are considered [adjacent], otherwise they are defined as [non - adjacent]. Here, the specific value of the spatial threshold is not limited.

[0042] The pursuit - escape relationship is used to determine the relationship between the pursuing entity and the escaping entity. If a certain pursuing entity pursues a certain escaping entity, this one - to - one pursuit - escape relationship can be expressed as the pursuing entity [pursuing] the escaping entity, and the escaping entity [being pursued] by the pursuing entity.

[0043] The passing - through relationship is used to describe whether there is a pursuing entity or an escaping entity on a road. If a pursuing entity or an escaping entity passes through this road, traffic facility, building, or special terrain, it is recorded as [passing through], otherwise it is recorded as [not passing through]; if a certain traffic facility is on a certain road, it is recorded that this road [contains] this traffic facility, and this facility [is contained] in this road.

[0044] Furthermore, the multiple relationships determined among the above - mentioned multiple entities are added to the relationship matrix, thereby obtaining the relationship matrix shown in Table 2:

[0045] Table 2: Relationship Matrix

[0046]

[0047] In practical applications, task instructions will be received within a set time. Here, there may be multiple task instructions. Correspondingly, the set time can be a fixed time period or an unfixed time period. That is, in the embodiments of the present invention, the processor will continuously send task instructions to the pursuing entity according to the situations of multiple entities in the environment. At the current moment, if a task instruction is received again, the current attribute characteristics of the pursuing entity and the current attribute characteristics of the escaping entity in the environment are obtained according to the task instruction.

[0048] Furthermore, according to the current attribute characteristics of the pursuing entity and the current attribute characteristics of the escaping entity, the initial knowledge graph is updated and the relationship matrix is updated to select new action information for the pursuing entity at the next moment.

[0049] Specifically, as time goes by, if the attributes of the chasing entity and the escaping entity change due to movement, the distances between the chasing entity and the escaping entity, the chasing entity and the reference entity, and the escaping entity and the reference entity may change, and their attributes need to be updated in the initial knowledge graph. The relationship matrix includes any one or more of the spatial adjacency relationship, the chasing relationship, and the passing relationship. The relationship trigger conditions between entities are detected in real time. If the distance between two entities is less than the spatial threshold, a [adjacent] relationship is added. If the distance between two entities is greater than the spatial threshold, the [adjacent] relationship is removed.

[0050] Furthermore, motion prediction and action selection can also be performed. That is, based on the current speed and acceleration of the escaping entity, a motion prediction model is established based on uniformly accelerated linear motion, and the possible position of the escaping entity at the next moment is obtained. Then, according to the possible position of the escaping entity at the next moment, in the knowledge graph, a [possible position] relationship is established between the escaping entity and the new position entity, and a relationship matrix is established between the new position entity and the reference entity. Furthermore, the relationship matrix between the escaping entity and the reference entity is determined, the distance between the chasing entity and the escaping entity is calculated, and then the next action is selected for the chasing entity according to the probability vectors corresponding to acceleration, encirclement, collision avoidance, and cooperation included in the action table shown in Table 3.

[0051] Table 3: Probability vectors of acceleration, encirclement, collision avoidance, and cooperation

[0052] acceleration [ρ, 0, 1 - ρ, 0] enclosure [0, ρ, 1 - ρ, 0] anti - collision [0,0,1,0] cooperation [0, 0, 1 - ρ, ρ]

[0053] For example, at the current moment t, if the first embedding feature of each chasing entity is [number i, t, action], where action = {1, 2, 3, 4} represents the 4 actions of acceleration, encirclement, collision avoidance, and cooperation respectively, the meaning represented by the first embedding feature is that the task success probability is the highest when the chasing entity number i selects the action action at the moment t. In Figure 3 , there are 3 first embedding features of the entity represented by the solid circle, that is, the action selection of the solid circle at the next moment includes 3. Among them, the acceleration in the action selection represents one position, the encirclement and collision represent one position, and the collision avoidance represents one position. In Figure 3 , the entity represented by the solid circle is the chasing entity. Therefore, the entity represented by the hollow circle is the escaping entity.

[0054] In practical applications, an environment may include multiple escaping entities and multiple chasing entities. In the embodiments of the present invention, in order to more clearly introduce the method provided by the embodiments of the present invention, the following takes one escaping entity and one or two chasing entities chasing the escaping entity as an example to introduce the cluster collaborative pursuit and escape decision optimization method in detail.

[0055] In step 102, based on the first escaping entity, sub-problems regarding the first escaping entity and sub-problems of the pursuing entities that pursue the first escaping entity are constructed. Since the number of pursuing entities that pursue the first escaping entity can be one or more, when constructing the sub-problems of the first escaping entity, correspondingly, the sub-problems of the pursuing entities that pursue the first escaping entity also need to be constructed. Here, no specific limit is imposed on the number of sub-problems of the constructed pursuing entities.

[0056] For example, the sub-problem of the first escaping entity is: [Number b 1 , the maximum probability position coordinates at the next moment are,?, t]; the sub-problem of the pursuing entity r 1 : [Number r 1 , what action should be taken at the next moment to maximize the likelihood of mission success,?, t].

[0057] Furthermore, a full background entity graph is obtained based on the first escaping entity, the pursuing entities related to the first escaping entity, the motion prediction and selection results of the first escaping entity, the reference entity, and the updated information. For example, the pursuing entity moves to (15, 20, 5) and selects "accelerate", the current position of the first escaping entity is (35, 38, 3), and the predicted possible positions at the next moment are (39, 42, 4) and (38, 41, 3.5), etc. Figure 2A For the Figure 5 corresponding schematic diagram of the full background entity, the full background entity graph can include the pursuing entity, the first escaping entity, and the predicted position information of the first escaping entity.

[0058] Furthermore, key sub-graphs related to the first escaping entity are extracted from the full background entity graph. For example, Figure 2B the hollow circle No. 1 in it can be determined as the first escaping entity, Figure 2B and the content shown is the key sub-graph of the first escaping entity.

[0059] In practical applications, according to the full background entity graph, key sub-graphs corresponding to the full background entity graph need to be extracted and weights calculated. Here, the weights refer to the weights between entities related to the first escaping entity extracted from the full background entity graph. For example, Figure 2B the relationship weights between the hollow circle No. 1 and the black circle No. 1, the relationship weights between the hollow circle No. 1 and the black circle No. 2, etc.

[0060] In the embodiments of the present invention, the relationship weights between two entities are determined by the following formula taking the entity e d as an example:

[0061]

[0062] Among them, represents the entity ed The central node of the key subgraph and its neighbor nodes The link between them df The weight of which represents the entity e d The central node of the key subgraph of entity e, t d represents the entity e d The timestamp of entity e, e d represents the number of the d-th entity represents the middle node Its neighbor nodes, t f represents the entity e f The timestamp of entity e, e f represents the number of the f-th entity and f ≠ d; represents the entity e d All the neighbor nodes connected to the central node in the key subgraph of entity e , t F represents the relationship between entity e d and all its neighbor nodes in the key subgraph N d represents the set of all neighbor nodes of the central node The key subgraph of entity e, t represents the entity e d The key subgraph of entity e, t T represents the current query time, θ represents the relationship weight, link df represents the entity e d and the entity e f The link between them

[0063] In practical applications, θ is a weight parameter, taking negative values, and its value can be adjusted according to the actual task to adapt to different distributions

[0064] It should be noted that if the entity e in formula (1) d represents the first escaping entity, since t d represents the timestamp of entity e d So represents the central node of the key subgraph of the first escaping entity e d Accordingly, e f can represent the tracking entity represents the middle node Its neighbor nodes, the link between the first escaping entity and the tracking entity is link df . Then, according to formula (1), the relationship weights between the first escaping entity and the reference entity, and between the first escaping entity and the pursuing entity can be determined respectively

[0065] For example, when for sub-question Q = (e′, link T ,?, t T ), the meaning of sub-question Q is: at time t T , the entity related to entity 'e′' by link T is '?', where entity '?' is what needs to be obtained in the embodiments of the present invention. If the finally calculated answer is entity 'e t ', then the output answer is that at time t T , the entity related to entity 'e′' by link T is 'e t '. More specifically, when link T represents 'pass through', 'e′' represents a certain escaping entity j, and 'e t ' represents a certain position, the question answering becomes: escaping entity j will 'pass through' a certain position 'e T ' at time t t .

[0066] Furthermore, in order to further explore the inference logic between relationships in the relationship graph, the relationships in the key sub-graph are further refined using a relationship logic graph. For example, assume that b 1 has a 'pass through' relationship (denoted as r1) with road 2 , and has an 'adjacent' relationship (denoted as r2) with house 1 , etc. Using these relationships as nodes, determine the relationship types between b 1 and related entities (such as the master-master relationship ι3, and refine the relationship logic graphs of other entities.

[0067] In the embodiments of the present invention, link df in formula (1) represents the edge connecting entity e d and entity e f , that is, the relationship between two entities. There are 6 definitions for the edge between two entities, which successively include two unidirectional connection edges (ι1, ι2) and 4 bidirectional connection edges (ι3, ι4, ι5, ι6). Among them, ι1 represents the master-object relationship, ι2 represents the object-master relationship, ι3 represents the master-master relationship, ι4 represents the object-object relationship, ι5 represents that two nodes act as the subject and object of each other, and ι6 represents that one node is fixed as the subject and the other node is fixed as the object. Specifically, reference can be made to Figure 2C the relationship logic graph shown.

[0068] Further, after obtaining the relationship logic graph, the relationship nodes in the key subgraph can be modeled to construct a relationship learning module, and the structural information of the multi-hop neighborhood of the relationship nodes is used to capture the logical correlation between relationships. In the embodiments of the present invention, a K-layer graph convolutional neural network (GNN, Graph Neural Network) is used to perform logical learning on the constructed relationship graph. After K-layer information transmission is completed, the logical feature encoding of the central node can be determined by the following formula. In the following formula, v d is the abbreviation of the central node , F is the abbreviation of , v f is 's abbreviation.

[0069]

[0070] Among them, represents the logical feature encoding of the central node v d after passing through the K-layer network. N d represents the set of all neighbor nodes of the central node v d . |N d | represents the number of elements in the set N d . represents the logical feature encoding of the neighbor node v f after passing through the K-1 layer network. represents the weight matrix of the Kth layer convolutional neural network. σ represents the non-linear activation function, usually ReLU is selected. K represents a constant, which can be adjusted according to the actual application effect; h d′ represents the semantic information encoding of the edge between the central node v d and the neighbor node v f , that is, any one of the 6 relationships included in the edge link df between two entities. As shown in Figure 2C , each relationship corresponds to a semantic encoding, that is, (ι1, ι2, ι3, ι4, ι5, ι6). d' represents the edge between the central node v d and the neighbor node v f . represents the logical feature encoding of the central node v d after passing through the K-1 layer network.

[0071] It should be noted that if the entity e d in formula (1) is the first escaping entity, then further according to formula (2), the logical feature encodings of the first escaping entity and the pursuing entity can be determined.

[0072] Further, in order to better understand the time information, the embodiments of the present invention can simultaneously encode the time information, and the encoding of the time information includes two aspects, short-term encoding and periodic encoding, where the short-term encoding and the periodic encoding can be represented by the following formulas respectively:

[0073]

[0074] Among them, represents the short-term encoding in the encoding of time t, represents the periodic encoding in the encoding of time t, Dimen represents the feature dimension of the current entity, ω i represents the weight in the neural network, which is a learnable parameter, represents the offset of the neural network, which is a learnable parameter.

[0075] Further, the final time information encoding can be obtained according to the short-term encoding and the periodic encoding, specifically:

[0076]

[0077] Among them, W t represents the weight matrix of the network for encoding time, α represents the weighting parameter, and the value ranges from 0 to 1, h t represents the final time information encoding.

[0078] Further, by coupling the logical feature encoding of the central node after passing through K layers of the network and the semantic information encoding of the connection edge between the central node and the neighbor node v f the initial coupling information encoding of the central node as shown below can be obtained:

[0079]

[0080] Among them, represents the initial coupling information encoding of the central node, W C represents the weight parameter, σ represents the activation function, usually ReLU is selected, h d′ represents the semantic information encoding of the connection edge between the central node v d and the neighbor node v f represents the semantic information encoding of the connection edge between the central node v d d and the neighbor node v

[0081] Further, by embedding the final time information encoding into the initial coupling information encoding of the central node, the final coupling information encoding of the central node can be obtained:

[0082]

[0083] Among them, h d,d′,t represents the encoded final coupling information of the central node, represents the encoded initial coupling information of the central node, h t represents the encoded final time information, || represents vector concatenation, and T represents the transpose operation of a matrix or vector.

[0084] Furthermore, according to the encoded final coupling information of the central node, it is input into a graph convolutional neural network with L layers for feature fusion. After passing through L layers of the network, the encoded initial information of sub-problem Q can be represented by the following formula:

[0085]

[0086] Among them, represents the encoded initial information of sub-problem Q after passing through L layers of the network, represents the weight matrix of the network encoding for sub-problem Q, represents the coupling weight between the central node and its neighbor nodes after passing through L layers of the network, h d,d′,t represents the encoded final coupling information of the central node, represents entity e f the logical feature encoding of sub-problem Q after passing through L - 1 layers of the network.

[0087] In the embodiments of the present invention, can be determined by the following formula:

[0088]

[0089] Among them, represents the coupling weight between the central node and its neighbor nodes after passing through L layers of the network, σ represents the activation function, represents for and for the coupling weights, represents the encoded final coupling information of the central node after passing through L layers of the network, represents entity e f the logical feature encoding of sub-problem Q after passing through L - 1 layers of the network, represents the semantic information encoding of the edge between the central node v d and the neighbor node v f after passing through L layers of the network, || represents vector concatenation.

[0090] Furthermore, in order to effectively learn the sequentiality between entities, a gated recurrent unit (GRU) is introduced. After passing through L layers of the graph convolutional neural network, the encoded initial information of sub-problem Q can obtain the encoded final information of sub-problem Q through the gate control unit, as shown specifically below:

[0091]

[0092] Among them, represents the initial information encoding of sub-problem Q after passing through L layers of the network, represents the final information encoding of sub-problem Q after passing through L layers of the network, which can also be understood as the final information encoding of sub-problem Q obtained by processing the initial information encoding of sub-problem Q through a gated recurrent unit.

[0093] The entity here is the central node v d , that is, the first escaping entity. In the embodiment of the present invention, if sub-problem Q is to predict the position of the first escaping entity b 1 at the next moment, it may include information such as the current position, speed, acceleration of b 1 , and the relationships between it and the surrounding roads, buildings, and pursuing entities. For the sub-problem of predicting the position of b 1 with the sub-problem number b 1 , the maximum probability position coordinates at the next moment are?, t, it will combine the movement conditions of b 1 in the past few time steps to give a position prediction feature encoding that is more in line with the actual situation, providing a more reliable basis for determining the 1 maximum probability position of b at the next moment.

[0094] Furthermore, according to the "?" in the sub-problem of the first escaping entity constructed in step 102 [sub-problem number b 1 , the maximum probability position coordinates at the next moment are?, t], which represents scoring all entities. For example Figure 2B scoring the 2nd hollow circle, 1st solid circle, and 2nd solid circle that are all possible connections to the 1st hollow circle in (x′, y′, z′), (x″, y″, z″), (x″′, y″′, z″′), and then the one with the highest score is the answer to "the maximum probability position coordinates at the next moment are" in the sub-problem, that is Figure 3 the dashed edge connection shown in.

[0095] In the embodiment of the present invention, the score of entity e d in sub-problem Q can be determined by the following formula:

[0096]

[0097] Among them, W T represents the transpose of the weight matrix W, e d represents the number of the d-th entity, which represents the entity for which the score needs to be calculated in this formula, represents entity e fAfter passing through the L-1 layer network, the logical feature encoding of sub-problem Q, f(e d ,Q) represents the score of entity e d in sub-problem Q.

[0098] As described before step 102, steps 102 and 103 take an escaping entity and one or two pursuing entities tracking the escaping entity as examples. In the following steps, the focus is on describing two pursuing entities tracking the escaping entity.

[0099] In step 103, action a represents the specific actions that the pursuing entity can take at the current moment. The reward coefficient matrix D is related to the actions selected by the pursuing entity. The action selection and the corresponding reward coefficient matrix provided in the embodiments of the present invention for the pursuing entity are shown in Table 3. Among them, the reward coefficient matrix D is the probability vector of acceleration, encirclement, collision prevention, and cooperation, which reflects the weight distribution of different types of rewards and punishments in the overall reward. ρ is a constant that can be adjusted according to the actual pursuit and escape effect. This matrix will be used for subsequent reward calculation to measure the advantages and disadvantages of different strategies.

[0100] In the embodiments of the present invention, the reward for the pursuing entity to take a specific action is calculated according to the following formula:

[0101] reward = D × [U1,U2,U3,U4]′ (7)

[0102] Among them, reward represents the reward function of the pursuing entity. The higher the reward, the better the strategy. In formula (7), [U1,U2,U3,U4]′ represents the transpose of [U1,U2,U3,U4]. U1 represents the guiding reward, U2 represents the task success reward, U3 represents the collision prevention penalty, and U4 represents the cooperation reward, which are respectively shown as follows:

[0103]

[0104] Among them, U1 represents the guiding reward, which means that the closer the pursuing entity is to the first escaping entity, the better. U2 represents the task success reward, which means that the pursuing entity has formed an encirclement. represents the radius of the encirclement, δ represents the encirclement error, and ε represents the minimum safety distance. These three parameters are all constants that can be adjusted according to the actual situation. When the distance between the pursuing entity and the first escaping entity meets specific conditions, that is, an effective encirclement is formed, the pursuing entity can obtain a reward of 2, otherwise the reward is 0. U3 is the collision penalty, which means that if a collision occurs between entities, a penalty will be received. Denote all entities except the \(i\)-th entity, i.e., \(i' \neq i\); if the distance between the pursuer entity and other entities is less than the minimum safety distance \(\varepsilon\), it indicates a collision, and at this time, the pursuer entity will receive a penalty of -2, otherwise the penalty is 0. \(U_4\) is the cooperation reward, representing the collaborative benefit between entities. The higher this benefit, the better the collaboration. \(m\) represents the number of pursuer entities. By calculating the average of the sum of squared deviations of the distances between all pursuer entities and the first escape entity and the radius of the encirclement, it encourages the pursuer entities to collaborate together to form an effective encirclement. The inner \(|\cdot|\) in \(\|\cdot\|\) represents the modulus of the vector, and the outer \(|\cdot|\) represents taking the absolute value. The \(|\cdot|\) in it represents the modulus of the vector.

[0105] In the embodiment of the present invention, taking the first pursuer entity and the second pursuer entity as examples, the strategies adopted by the two pursuer entities when pursuing the first escape entity are introduced in detail. Specifically, the benefits of the first pursuer entity and the second pursuer entity can be determined according to Table 3 and formula (7) respectively; that is, according to the benefit coefficient matrix of the first pursuer entity, the benefit coefficient matrix of the second pursuer entity, the reward and punishment conditions of the first pursuer entity, and the reward and punishment conditions of the second pursuer entity, the benefits of the first pursuer entity and the second pursuer entity can be determined respectively, and the strategy with the highest benefit is selected in the strategy space of the entities. The reward and punishment conditions here refer to the guiding reward shown in formula (7 - 1), the task success reward shown in formula (7 - 2), the collision penalty shown in formula (7 - 3), and the cooperation reward shown in formula (7 - 4).

[0106] In practical applications, if it is a discrete action space, directly select the action with the highest corresponding benefit, as Figure 4A shown; if it is a continuous space, the strategy is represented by the angle \(\theta\), as Figure 4B shown.

[0107] Furthermore, based on the selected strategy with the highest benefit, update the target positions of the first pursuer entity and the second pursuer entity. If it is determined that the first pursuer entity and the second pursuer entity complete the encirclement of the first escape entity according to the updated position of the first pursuer entity, the updated position of the second pursuer entity, and the encirclement condition, it is determined that the pursuit of the first escape entity is completed.

[0108] Specifically, it is determined whether all pursuer entities have completed the encirclement through the following formula, that is, satisfying the inequality:

[0109]

[0110] where \(\delta\) represents the encirclement error, represents the position of the pursuer entity \(r\) i at time \(t\), represents the escape entity \(b\) jThe position at time t, j represents the number of fleeing entities, and i represents the number of pursuing entities. In the inner layer, |·| represents the modulus of a vector, and in the outer layer, |·| represents taking the absolute value.

[0111] In the embodiment of the present invention, if this inequality is satisfied, it indicates that the first pursuing entity and the second pursuing entity have successfully surrounded the first fleeing entity, the task is completed, and the process ends; otherwise, return to step 101 to update the initial knowledge graph, re-execute the above steps, and continue to execute the pursuit and escape task.

[0112] It should be noted that the first pursuing entity and the second pursuing entity here are only two of the multiple pursuing entities that pursue the first fleeing entity, that is, the first pursuing entity and the second pursuing entity may each include only one pursuing entity, or may include multiple pursuing entities. The specific number of the first pursuing entity and the second pursuing entity is not limited here.

[0113] The embodiment of the present invention provides a method for optimizing the decision-making of cluster collaborative pursuit and escape. In order to play a role in a small sample size environment, a temporal knowledge graph is designed to mine the hidden information in the time series, thereby reducing the demand for other data information, expanding the application scope of the system, and enhancing the adaptability of the decision-making system; during the pursuit and escape process, a large number of entities will appear, and the decision-making method needs to learn the relationships between these different entities, so as to infer the most likely relationships in the future. Since the traditional method learns all entities and the relationships among them from a global perspective, the efficiency is relatively low. The method for optimizing the decision-making of cluster collaborative pursuit and escape first extracts key subgraphs based on weights, reducing unnecessary computational losses. At the same time, a logical subgraph is constructed and encoded, which can learn the features related to the task faster and more directly, thereby optimizing the decision-making efficiency in the process of cluster collaborative pursuit and escape; furthermore, in order to avoid unexpected task situations during the process of the cluster executing tasks, compared with the traditional method of modeling the unified payoff matrix for agents, this method designs an autonomous reward function inference mechanism, which has the ability to independently select a more suitable payoff for itself, thereby improving the flexibility of the cluster system in dealing with different scenarios.

[0114] To more clearly introduce the method for optimizing the decision-making of cluster collaborative pursuit and escape provided by the embodiment of the present invention, the following explains this method in combination with a specific scenario. Suppose in a certain pursuit and escape scenario, there are 3 pursuing entities r 1 , r 2 , r 3 , 2 fleeing entities b 1 , b 2 and reference entities. The above-mentioned multiple entities are stored as entities in the initial knowledge graph, and their initial attribute features, such as position coordinates, numbers, lengths, etc., are recorded. Among them, the reference entity includes 5 roads road 1 , road2 , road 3 , road 4 , road 5 , and 2 special terrains (such as teerrain 1 , teerrain 2 ), 4 buildings (such as numbered house 1 , house 2 , house 3 , house 4 ). The current time step t = 3, and the time interval Δt = 1 second.

[0115] Step 201, obtain the initial attribute features of the pursuit entity, the initial attribute features of the escape entity, and the initial attribute features of the reference entity in the environment according to the task instructions, and obtain the initial knowledge graph based on the multiple initial attribute features and determine the relationship matrix among the pursuit entity, the escape entity, and the reference entity.

[0116] Specifically, m = 3 pursuit entities, numbered r 1 , r 2 , r 3 . At time step t = 0, their position coordinates are randomly assigned. For example, the position of r 1 is (10, 20, 5), which represents its initial position in three-dimensional space; n = 2 escape entities, numbered b 1 , b 2 . At t = 0, the position coordinates of b 1 are (30, 40, 3); the initial attributes of the reference entity. Since the reference entity includes road entities, special terrain entities, and building entities, the acquisition of the following multiple initial attribute features can be divided. The total number of roads q = 5, numbered road 1 , road 2 , road 3 , road 4 , road 5 . Taking road 1 as an example, in addition to having a number, it also has attributes such as the starting position (x, y, z), length len, and direction (degrees). For example, the starting position is (5, 5, 0), the length is 100 meters, and the direction is 90 degrees (i.e., the north-south direction). The total number of special terrains p = 2, numbered terrain 1 , terrain 2 . Assume that terrain 1 is a lake, and its central position coordinates are (25, 35, 0). Each special terrain is determined by attributes such as the number and the central position to determine its position and characteristics in the scene. The total number of buildings g = 4, numbered house 1, house 2 , house 3 , house 4 。For example, house 1 has a position coordinate of (15, 15, 0), and the building is also characterized by attributes such as number and position.

[0117] It should be noted that the initial attribute features of the pursuit entity and the escape entity at least include number, initial state coordinate, and action.

[0118] Furthermore, an initial knowledge graph is obtained based on multiple initial attribute features, and the relationship matrix among the pursuit entity, the escape entity, and the reference entity is determined.

[0119] Specifically, the obtained pursuit entity (r 1 , r 2 , r 3 ), the escape entity (b 1 , b 2 ), and the initial feature attributes of the reference entity are stored as entities in the knowledge graph, and their initial attribute features, such as position coordinates, numbers, lengths, etc., are recorded. An initial knowledge graph of some entities and their initial attribute features in the temporal knowledge graph as shown in Table 1 can be obtained.

[0120] Determine the relationship matrix among the pursuit entity, the escape entity, and the reference entity according to the initial knowledge graph. Among them, the relationships at least include any one or more combinations of adjacency relationship, pursuit - escape relationship, passing - through relationship, and inclusion relationship.

[0121] Specifically, if the distance between the pursuit entity r 1 and the building house 2 is less than the distance threshold of 10 meters, then in the knowledge graph, record that r 1 and house 2 are in an [adjacent] relationship; if it is greater than 10 meters, record it as [not adjacent]. For the calculation of the spatial distance, if there are obstacles (such as walls) between two entities, the spatial shortest distance along the wall to reach the opposite target is calculated. If the pursuit entity r 2 is pursuing the escape entity b 1 , this one - to - one relationship is recorded in the knowledge graph as r 2 [pursuing] b 1 , and at the same time, b 1 [being pursued] by r 2 . If the escape entity b 2 has passed through the road road 3 , record in the knowledge graph that b 2 [has passed through] road 3; If not passed, record as [not passed]. If there is a traffic facility (such as a traffic light) on the road road 4 then record road 4 [contains] this traffic facility, and this traffic facility [is contained] in road 4 . For other similar inclusion relationships, such as a building within a certain area, etc., they are also defined and recorded in a similar manner.

[0122] Furthermore, add the various relationships determined among the above-mentioned multiple entities to the relationship matrix, thereby obtaining the relationship matrix shown in Table 2.

[0123] Step 202, at the current moment, based on the current attribute characteristics of the pursuing entity, the escaping entity, and the reference object, obtain an updated knowledge graph and an updated relationship matrix, and select new action information for the pursuing entity at the next moment.

[0124] Step 202-1, update the initial knowledge graph. As time goes by and the pursuing entity and the escaping entity move, various information in the environment will change. At the current moment, collect the latest attribute characteristics of each entity in the environment in real time, and update the initial knowledge graph and the relationship matrix.

[0125] Specifically, if the position of the pursuing entity r 1 moves from (10, 20, 5) to (12, 22, 5), its speed may change from 0 to 5 m / s, and its acceleration may change from 0 to 2 m / s 2 . The changes in the attribute values such as the latest position, speed, and acceleration of the pursuing entity and the escaping entity need to be parsed and matched with the pursuing entity r 1 in the initial knowledge graph, and its attribute values in the initial graph need to be updated in a timely manner. Similarly, if the attribute characteristics of the escaping entity and the reference object entity change, they also need to be updated in the initial knowledge graph.

[0126] Step 202-2, adjust the relationship dynamics of the relationship matrix, monitor the triggering conditions of various relationships among each entity in real time, and update the relationships among each entity.

[0127] Step 202-3, perform motion prediction for the next moment. For each escaping entity, based on the current speed, acceleration, and environmental constraints of the escaping entity, establish a motion prediction model. For example, the escaping entity b 2 , as well as its current position coordinates, speed, and acceleration, and through the updated knowledge graph, learn about the reference object entity (obstacle information) in the cycle of the escaping entity b 2 , such as the position and shape of the nearest building.

[0128] According to the kinematic formula and environmental constraints, the escaping entity b 2Possible positions at the next moment. Assume there are no obstacles blocking the escaping entity b 2 In the x, y, and z directions of movement, then in the x direction, the possible position at the next moment may be where represents the escaping entity b 2 The velocity in the x direction, represents the escaping entity b 2 The acceleration in the x direction, Δt represents the time interval. Similarly, the positions of the escaping entity b 2 at the next moment in the y and z directions can be calculated. In practical applications, considering that the escaping entity b 2 may change its direction of movement, so multiple possible position coordinates will be obtained at the next moment.

[0129] Furthermore, take these calculated possible positions as new position entities, attach a time stamp (such as the t+1 moment) to each position entity, and then add these position entities to the temporal knowledge graph through incremental update. At the same time, establish the "possible position" relationship between these new position entities and the escaping entity b 2 in the knowledge graph, indicating that these are the possible positions that the escaping entity b 2 may reach at the t+1 moment.

[0130] Step 202-4, perform the prediction of the action selection at the next moment. For each pursuing entity, analyze the state information in the updated knowledge graph (the current knowledge graph). For example, for the pursuing entity r 1 , check the distance and pursuit-evasion relationship between the pursuing entity r 1 and the escaping entity b 1 , and the relationship with environmental elements such as surrounding roads and buildings. It is found that the distance between the pursuing entity r 1 and the escaping entity b 1 is relatively far, and the surrounding roads are unobstructed without obstacles affecting its movement. 1 The distance between the pursuing entity r

[0131] and the escaping entity b 1 is relatively far. Refer to the action table shown in Table 3 (assuming the action table contains actions such as acceleration, encirclement, collision prevention, cooperation, etc. and their corresponding probability representations, such as the probability vector corresponding to the acceleration action is [ρ, 0, 1-ρ, 0]), and combine the current graph state to select the next action for the escaping entity b 1 Since the pursuing entity r 1 is far from the escaping entity b 1 , it may choose the "acceleration" action to approach the escaping entity b 1 . Generate a new action entity for the "acceleration" action and add it to the updated knowledge graph. At the same time, establish the association relationship between this action entity and the pursuing entity r1 The "acceleration" action is selected at the current moment for subsequent decision-making analysis, link prediction and other operations.

[0132] For the three pursuit entities r in the above-mentioned pursuit and escape scenario 1 ,r 2 ,r 3 ,and two escape entities b 1 ,b 2 ,in the following steps, taking the escape entity b 1 as an example, this embodiment will be described.

[0133] Step 203, based on the escape entity b 1 ,construct a sub-problem regarding the escape entity b 1 ,and obtain a full-background entity graph according to the three pursuit entities r 1 ,r 2 ,r 3 ,the escape entity b 1 ,new location information, new action information, and the predicted location information of the escape entity b 1 ;extract the key sub-graph related to the escape entity b 1 from the full-background entity graph and determine the relationship weight and relationship logic graph; predict the most probable location of the escape entity b 1 at the next moment, and the three pursuit entities r 1 that have a pursuit relationship with the escape entity b 1 ,r 2 ,r 3 select actions at the next moment.

[0134] Specifically, based on the escape entity b 1 ,construct a sub-problem regarding the escape entity b 1 ,the sub-problem of the escape entity b 1 : [Number b 1 ,the most probable location coordinates at the next moment are,?, t]; the sub-problem of the pursuit entity r 1 : [Number r 1 ,what action should be taken at the next moment to maximize the likelihood of mission success,?, t], the sub-problem of the pursuit entity r 2 : [Number r 2 ,what action should be taken at the next moment to maximize the likelihood of mission success,?, t], the sub-problem of the pursuit entity r 3 : [Number r 3 ,what action should be taken at the next moment to maximize the likelihood of mission success,?, t].

[0135] Step 203-1, according to the three pursuit entities r 1 ,r 2 ,r3 , escape entity b 1 , new position information, new action information, and escape entity b 1 The predicted position information is used to obtain the full background entity map.

[0136] In step 201 above, motion prediction and motion selection have been performed on both the pursuit entity and the escape entity. Specifically, when the pursuit entity r 1 moves to the next position (15, 20, 5) and selects the "accelerate" action; the pursuit entity r 2 moves to the next position (18, 22, 4) and selects the "encircle" action; the pursuit entity r 3 moves to the next position (20, 25, 3) and selects the "cooperate" action; the escape entity b 1 The current position is (35, 38, 3), the speed is 4 m / s, and the acceleration is 1 m / s 2 , predicting the escape entity b 1 The possible positions at the next moment are (39, 42, 4) and (38, 41, 3.5); the escape entity b 2 The current position is (40, 45, 2), the speed is 5 m / s, and the acceleration is 0.8 m / s 2 , predicting the escape entity b 2 The possible positions at the next moment are (45, 50, 3) and (44, 49, 2.8).

[0137] At the same time, the attributes and relationships of each map element have also been updated according to the actual situation. Integrating all these entities (entities, map elements, new positions, new actions) and their relationships, the full background entity map is obtained.

[0138] Step 203-2: According to the full background entity map, extract the corresponding key subgraph and calculate the weights.

[0139] In this embodiment, a key subgraph related to the escape entity b 1 is extracted from the full background entity map. Assuming only the escape entity b 1 is concerned with the surrounding road road 2 , road 3 and the building house 1 , house 2 's relationship.

[0140] For the escape entity b 1 and the building house 1 , determine the relationship weight between them. Set the escape entity b 1 as e d , and the surrounding building house 1 as ef Set the weight parameter θ = 0.5, and the current query time t q = 4. For b 1 (e d ) and house 1 (e f ), assuming t j = 2, according to the relationship weight formula between entities, that is, formula (1), first calculate the relationship weight between the escaping entity b 1 and the surrounding building house 1 , that is

[0141] Furthermore, according to the method of determining the relationship weight between the escaping entity b 1 and the surrounding building house 1 , the relationship weight between the escaping entity b 1 and road 2 can be obtained respectively

[0142] Step 203-3, improve the relationship logic diagram. Assume that there is a "passing by" relationship between the escaping entity b 1 and the road road 2 , then record the relationship between the two as r1; there is an "adjacent" relationship between the escaping entity b 1 and house 1 , then record the relationship between the two as r2; furthermore, in the relationship logic diagram, use r1 and r2 as nodes. Since the escaping entity b 1 is the subject in these relationships, the main-main relationship (l3) between the escaping entity and road 2 , house 1 can be determined

[0143] Step 203-4, based on the graph convolutional neural network, predict the most probable position of the escaping entity b 1 at the next moment. Assume there are K = 2 layers of graph convolutional neural networks. Using the relationships r1 and r2 between the escaping entity b 1 and the reference entity as nodes, construct a relationship learning module, use the structural information of the multi-hop neighborhood of the relationship nodes to capture the logical correlation between the relationships, and determine the logical feature encoding of the escaping entity b 1 through formula (2)

[0144] In the previous formula (1), when calculating the relationship weight between the escaping entity b 1 and the surrounding building house 1 , the central node in formula (1) 1 represents the escaping entity b dDenote the escaping entity b 1 .

[0145] Specifically, the neighbor node set N of the escaping entity b 1 , includes other relationship nodes related to road d , (assuming there is a relationship node related to the road direction and a relationship node related to the road length). Let the transformation matrices of the first and second layer convolutional neural networks be: 2 σ is a non-linear activation function. σ is a non-linear activation function.

[0146] Among them, if the central node v in formula (2) d 's neighbor nodes are v f1 and v f2 , the initial feature encodings of the neighbor nodes v f1 and v f2 are respectively and h d′ denotes the semantic information encoding of the edge connecting the central node v d and the neighbor node v f , and can also be called the semantic information encoding of the corresponding relationship type between the central node v d and the neighbor node v f . Here, the corresponding relationship type between the central node v d and the neighbor node v f can be ι3, Through formula (2), the logical feature encoding of the central node v d after passing through one layer of the network can be obtained Assume that the neighbor nodes v f1 and v f2 become after the first layer update, then through formula (2), the logical feature encoding of the central node v d after passing through two layers of the network can be obtained

[0147] Furthermore, assume the entity feature dimension Dimen = 3, the weights in the neural network are ω1 = 0.2, ω2 = 0.3, ω3 = 0.3 in turn; the offsets of the neural network are The weighting parameter is α = 0.2, and the weight matrix of the network encoding for time is When t = 4, the short-term encoding can be determined by formula (3-1) [-0.4339]. The periodic encoding is determined by formula (3-2)

[0148] Furthermore, the final time information encoding h is determined through formula (3-3). t :

[0149] Furthermore, assume the weight parameter of the central node v d the logical feature encoding after passing through K layers of network The initial coupling information encoding of the central node can be determined through formula (3-4): Finally, the final coupling information encoding h of the central node is determined through formula (3-5). d,d′,t .

[0150] Furthermore, for the escaping entity b 1 and the neighbor entity house 1 , the initial information encoding of sub-problem Q after passing through K layers of network can be determined through formula (4). Specifically, assume the first-layer weight matrix is the second-layer weight matrix is σ represents the activation function. Assume σ(x) = max(0, 1). The attention weight vector calculated in the first layer is the attention weight vector calculated in the second layer is Assume the neighbor entity house 1 (e f ) the logical feature encoding at the initial layer is the neighbor entity house 1 the final coupling information encoding at the first layer is the neighbor entity house 1 the semantic information encoding of the corresponding relationship type ι3 at the first layer is

[0151] The initial information encoding of sub-problem Q after passing through 1 layer of network is determined through formula (4): Based on the same method, the initial information encoding of sub-problem Q after passing through 2 layers of network can be determined

[0152] Step 203-5. To effectively learn the sequentiality between entities, a gated recurrent unit GRU is introduced. The initial information encoding of sub-problem Q after passing through L layers of network can obtain the final information encoding of sub-problem Q through the gate control unit.

[0153] In this embodiment, assume the GRU unit has been trained. For the initial information encoding of sub-problem Q after passing through 2 layers of network Assume According to formula (5), the final information encoding for sub-problem Q can be obtained

[0154] In this embodiment, the GRU further mines the temporal dependencies in the sequence features through the processing of to provide a more timely and relevant feature representation for subsequent scoring.

[0155] Step 203-6, assume that the escaping entity b 1 has two predicted positions (34, 39, 3) and (33, 38, 2.8). Let these two positions be e1 and e2 respectively, and let the weight matrix For the first position e1 of the escaping entity b 1 assume that another entity e 1 associated with the first position e1 of the escaping entity b j has a logical feature encoding after passing through the network of L-1 layers in the sub-question Q as Then, through formula (6), the score of the first position e1 of the escaping entity b 1 in the sub-question Q can be obtained, that is, For the position e2, assume that another entity e 1 associated with the second position e2 of the escaping entity b j has a logical feature encoding after passing through the network of L-1 layers in the sub-question Q Through formula (6), the score of the second position e2 of the escaping entity b 1 in the sub-question Q can also be obtained, that is,

[0156] In this embodiment, since the score of the first position e1 of the escaping entity b 1 in the sub-question Q is 0.5, and the score of the second position e2 of the escaping entity b 1 in the sub-question Q is 0.47, the first position e1 (34, 39, 3) is selected as the maximum probability position coordinate of the escaping entity b 1 at the next moment. At the same time, the second embedding feature of the escaping entity b 1 [encoding b 1 , t = 3, (34, 39, 3)] is output. Through the scoring process, various factors affecting the position are comprehensively considered, providing a quantitative basis for the decision-making of the escaping entity b 1 to determine the most likely position at the next moment in the current scenario.

[0157] Step 204, assume that the pursuing entity r 1 is at the position (15, 20, 5) at time step t = 3, and the pursuing entity r 2The position at time step t = 3 is (18, 22, 4), and the fleeing entity is b 1 The position at time step t = 3 is (30, 35, 2). Assume the surrounding circle θ = 10, the surrounding error δ1 = 1, the minimum safety distance ε = 3, and the constant ρ = 0.6.

[0158] Assume the pursuing entity r 1 and the pursuing entity r 2 Both selected the acceleration action in step 203. Determine the reward coefficient matrix D according to the action table and relevant rules. For the acceleration action, let the corresponding reward coefficient be [0.8, 0, 0.2, 0] (the coefficients here are determined according to the pre-set association rules between actions and rewards. For example, the acceleration action pays more attention to quickly approaching the fleeing entity, so the guiding reward coefficient is higher).

[0159] Calculate the guiding reward U1 for the pursuing entity r 1 , For the pursuing entity r 2 ,

[0160]

[0161] Calculate the task success reward U2, calculate the distance between the pursuing entity and the fleeing entity and the surrounding circle Related conditions Since the distances between the two pursuing entities and the fleeing entity are both greater than So Both are equal to zero.

[0162] Calculate the anti-collision penalty U3. Assume the distance between the two pursuing entities is: So Both are equal to zero.

[0163] Calculate the cooperation reward U4 for the pursuing entity r 1 , For the pursuing entity r 2 , Here the cooperation reward is negative because the current distance between the pursuing entity and the fleeing entity is far and the cooperation effect is not good.

[0164] Furthermore, calculate the reward = D × [d1, d2, d3, d4]′ for the pursuing entity r 1 : reward, r 1 = 0.0368; corresponding to the pursuing entity r 2 : reward, r 2 = 0.0448.

[0165] Compare the pursuing entity r 1and the pursuing entity r 2 's gain, because the pursuing entity r 2 's gain of 0.0448 is greater than that of the pursuing entity r 1 's gain of 0.0368. Therefore, in the discrete action space, although the gains of the currently selected acceleration actions are not high, relatively speaking, the strategy of the pursuing entity r 2 is slightly better.

[0166] Step 204-1. Since the pursuing entity r 1 and the pursuing entity r 2 's strategy determined in step 203 is to continue to accelerate and approach the escaping entity b 1 . Assume that the speed update rule of the pursuing entity is to increase the speed by 50% (because the acceleration action is selected). The original speed of the pursuing entity r 1 was The updated speed The pursuing entity r 2 The original speed was The updated speed

[0167] According to the speed and the current position, the target position of the pursuing entity r 1 is (15 + 7.5×1, 20 + 7.5×1.5, 5), that is, (22.5, 27.5, 5). The target position of the pursuing entity r 2 is (27, 31, 4).

[0168] Use sensors (simulate sensor data acquisition through programs in the simulation environment) to sample the information of the escaping entity b 1 ) and map elements in real time. For example, it is found that the speed of the escaping entity b 1 becomes 4 m / s and the acceleration becomes 0.8 m / s 2 . In terms of map elements, a new obstacle appears on the road road 2 , with the position at (25, 30, 0), the size of length 5, width 3, and height 2.

[0169] Step 204-2. Judge whether the task is completed, calculate the distance between the pursuing entity and the escaping entity and the encirclement condition. For the pursuing entity r 1 , For the pursuing entity r 2 ,

[0170] Because the encirclement condition of the pursuing entity r 1 and the pursuing entity r 2The surrounding conditions are all greater than the surrounding error, so it does not meet the requirements. Therefore, none of the above pursuit entities have completed the surrounding.

[0171] In the embodiment of the present invention, if the pursuit entity has not completed the surrounding, then return to step 201 to continue execution, update the entities, attribute features and relationships between entities again, and perform a series of operations such as motion prediction and action selection again until the surrounding condition task is completed.

[0172] The method provided by the embodiment of the present invention has three core innovation points. First, a temporal knowledge graph is constructed. By introducing temporal features, the implicit information contained in the time series is deeply mined, effectively reducing the dependence on other environmental data. This enables the prediction of future relationships between entities under the condition of a small sample size, and can accurately infer the answer to questions such as "the maximum probability that the escaping entity [number] moves from the current position [coordinate] to [target coordinate] at t = [moment]". Second, a key subgraph and logical encoding strategy are proposed. Based on weight screening, key subgraphs are extracted, and a relational logic graph system is constructed to depict the relationships between entities. By converting the relational logic graph into a mathematical expression, the relationship scores between different entities are quantitatively calculated, and for specific sub-questions, the entity with the highest relationship score is selected as the answer to fill in, significantly improving the accuracy of predicting future relationships between entities; finally, an autonomous reward function reasoning mechanism is designed. With the prediction ability of the temporal knowledge graph, analyze "the possibility of task success when the pursuit entity [number] selects [action] at t = [moment]", and each pursuit entity can autonomously select the benefit function and action target that best suit the current situation according to the prediction result, thereby greatly enhancing the autonomy and flexibility of the agent's decision-making.

[0173] Based on the same inventive concept, the embodiment of the present invention provides a cluster cooperative pursuit and escape decision optimization device. Since the principle of the device for solving technical problems is similar to that of the cluster cooperative pursuit and escape decision optimization method, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.

[0174] As Figure 7 shown, the device includes an update unit 701, a prediction unit 702, and a pursuit unit 703.

[0175] The update unit 701 is used to obtain the initial attribute features of multiple entities in the pursuit and escape environment according to the task instruction within the set time and obtain the initial knowledge graph and the initial relationship matrix; at the current moment, update the initial knowledge graph and the initial relationship matrix according to the current attribute features of the multiple entities;

[0176] A prediction unit 702 is configured to determine a first escaping entity as a central node in a pursuit and escape environment, construct a sub-problem regarding the first escaping entity, sequentially determine relationship weights related to the first escaping entity and a logical feature encoding of the first escaping entity based on a key sub-graph and a relationship logic graph including the first escaping entity; obtain a final information encoding of the sub-problem based on the logical feature encoding of the first escaping entity, a final time information encoding, and a final coupling information encoding; and predict a maximum probability position of the first escaping entity at the next moment according to the final information encoding of the sub-problem.

[0177] A pursuit unit 703 is configured to determine a revenue strategy of a first pursuing entity and a second pursuing entity according to a revenue coefficient matrix and a reward and punishment condition of the first pursuing entity and the second pursuing entity, and update target positions of the first pursuing entity and the second pursuing entity with the highest revenue strategy; and determine that the pursuit of the first escaping entity is completed if it is determined that the first pursuing entity and the second pursuing entity complete the encirclement of the first escaping entity according to the updated position of the first pursuing entity, the updated position of the second pursuing entity, and an encirclement condition.

[0178] It should be understood that the units included in the above cluster collaborative pursuit and escape decision optimization device are only logical divisions according to the functions implemented by the device. In practical applications, the above units can be superimposed or split. Moreover, the functions implemented by the cluster collaborative pursuit and escape decision optimization device provided in this embodiment correspond one by one to those of the cluster collaborative pursuit and escape decision optimization method provided in the above embodiment. For the more detailed processing flow implemented by this device, it has been described in detail in the first method embodiment above, and will not be described in detail here.

[0179] Another embodiment of the present invention further provides a computer device, which includes: a processor and a scenario database; the scenario database is configured to store computer program code, and the computer program code includes computer instructions; when the processor executes the computer instructions, the electronic device executes each step of the cluster collaborative pursuit and escape decision optimization method shown in the above method embodiment.

[0180] Another embodiment of the present invention further provides a computer-readable storage medium, in which computer instructions are stored. When the computer instructions run on a computer device, the computer device is caused to execute each step of the cluster collaborative pursuit and escape decision optimization method shown in the above method embodiment.

[0181] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the present invention.

[0182] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention also intends to include these modifications and variations.

Claims

1. Cluster collaborative pursuit and escape decision optimization method, characterized in that Including: Obtain the initial attribute features of multiple entities in the pursuit and escape environment within a set time according to the task instruction, and obtain the initial knowledge graph and the initial relationship matrix; At the current moment, update the initial knowledge graph and the initial relationship matrix according to the current attribute features of the multiple entities; Determine the first escaping entity as the central node in the pursuit and escape environment, construct a sub-problem regarding the first escaping entity, and sequentially determine the relationship weights related to the first escaping entity and the logical feature encoding of the first escaping entity based on the key sub-graph and the relationship logic graph including the first escaping entity; Obtain the final information encoding of the sub-problem based on the logical feature encoding of the first escaping entity, the final time information encoding, and the final coupling information encoding; Predict the most probable position of the first escaping entity at the next moment according to the final information encoding of the sub-problem; Determine the revenue strategies of the first pursuing entity and the second pursuing entity according to the revenue coefficient matrix and the reward and punishment conditions of the first pursuing entity and the second pursuing entity, and update the target positions of the first pursuing entity and the second pursuing entity with the highest revenue strategy; If it is determined that the first pursuing entity and the second pursuing entity complete the encirclement of the first escaping entity according to the updated position of the first pursuing entity, the updated position of the second pursuing entity, and the encirclement condition, it is determined that the pursuit of the first escaping entity is completed.

2. The method according to claim 1, characterized in that, Determine the relationship weights related to the first escaping entity through the following formula: Among them, represents the central node of the key subgraph and the neighbor nodes The edge between The weight of represents the entity The central node of the key subgraph of represents the entity The timestamp of represents the number of the represents the neighbor nodes of the middle node represents the entity The timestamp of represents the number of the represents the entity and the entity The edge between represents the key subgraph of the entity represents the entity All neighbor nodes in the key subgraph of connected to the central node represents the entity The relationship between represents the current query time represents the relationship weight.​ 3. The method according to claim 1, characterized in that, Determine the logical feature encoding of the first escaping entity through the following formula: Among them, represents the central node after the logical feature encoding after passing through the layers of the network, represents the non-linear activation function, represents the weight matrix of the th convolutional neural network layer, represents the set of all neighbor nodes of the central node represents the set the number of elements in, represents the neighbor node after the logical feature encoding after passing through the layers of the network, represents the semantic information encoding of the edge between the central node and the neighbor node represents the central node after the logical feature encoding after passing through the layers of the network, represents the edge between the central node and the neighbor node.

4. The method according to claim 1, wherein The final time information encoding is represented by the following formula: The final coupling information encoding is represented by the following formula: Among them, represents the final time information encoding of the central node, represents the activation function, represents the weight matrix of the network that encodes time, represents the weighting parameter, represents time in the short - time encoding of the encoding, represents time in the periodic encoding of the encoding; represents the final coupling information encoding of the central node, represents the initial coupling information encoding of the central node, represents the weight parameter, represents the central node and the neighbor node the semantic information encoding of the edge between them, represents the central node after passing through the logical feature encoding after the layer network, represents vector concatenation, represents the transpose operation of a matrix or vector.

5. The method according to claim 1, characterized in that, The final information encoding of the sub-problem is represented by the following formula: Among them, represents the final information encoding of sub-problem after layers of the network, represents the initial information encoding of sub-problem after layers of the network, represents a gated recurrent unit, represents an activation function, represents the weight matrix of the network for encoding sub-problem , represents the coupling weight between the central node and neighbor nodes after layers of the network, represents entity after passing through layers of the network, the logical feature encoding of sub-problem , represents the final coupling information encoding of the central node, represents the neighbor nodes of the central node , represents the set of all neighbor nodes of the central node , represents the coupling weight for and for , represents the final coupling information encoding of the central node after layers of the network, represents the semantic information encoding of the edge between the central node and neighbor node after layers of the network, represents vector concatenation.

6. The method according to claim 1, wherein Before predicting the most probable position of the first escaping entity at the next moment according to the final information encoding of the sub-problem, it also includes determining the score of the first escaping entity in the sub-problem through the following formula; Among them, represents an entity 's score in the sub-question . represents the transpose of the weight matrix . represents the th entity number. represents the logical feature encoding of the entity after passing through layers of the network for the sub-question .

7. The method according to claim 1, wherein The reward and punishment conditions include guiding reward, task success reward, anti-collision punishment, and cooperation reward; The revenue strategies of the first pursuing entity and the second pursuing entity are determined through the following formula: The encirclement condition is: Among them, represents the revenue function of the pursuit entity, represents the guiding reward, represents the task success reward, represents the anti-collision penalty, represents the cooperation reward, represents the revenue coefficient matrix, represents the radius of the surrounding circle, represents the surrounding error, represents the pursuit entity at moment's position, represents the escaping entity at moment's position, represents the number of escaping entities, represents the number of pursuit entities, represents transpose of, in the inner layer represents the modulus of the vector, and the outer layer represents taking the absolute value.

8. Cluster collaborative pursuit and escape decision optimization device, characterized in that Including: An update unit, configured to obtain the initial attribute features of multiple entities in the pursuit and escape environment within a set time according to the task instruction, and obtain the initial knowledge graph and the initial relationship matrix; At the current moment, update the initial knowledge graph and the initial relationship matrix according to the current attribute features of the multiple entities; A prediction unit, configured to determine the first escaping entity as the central node in the pursuit and escape environment, construct a sub-problem regarding the first escaping entity, and sequentially determine the relationship weights related to the first escaping entity and the logical feature encoding of the first escaping entity based on the key sub-graph and the relationship logic graph including the first escaping entity; Obtain the final information encoding of the sub-problem based on the logical feature encoding of the first escaping entity, the final time information encoding, and the final coupling information encoding; Predict the most probable position of the first escaping entity at the next moment according to the final information encoding of the sub-problem; A pursuit unit, configured to determine the revenue strategies of the first pursuit entity and the second pursuit entity according to the revenue coefficient matrix and the reward and punishment conditions of the first pursuit entity and the second pursuit entity, and update the target positions of the first pursuit entity and the second pursuit entity with the highest revenue strategy; If it is determined that the first pursuit entity and the second pursuit entity have completed the encirclement of the first escape entity according to the updated position of the first pursuit entity, the updated position of the second pursuit entity, and the encirclement condition, it is determined that the pursuit of the first escape entity has been completed.

9. A computer device, characterized in that, The computer device includes a memory and a processor. When the computer program stored in the memory is executed by the processor, the processor is caused to execute the cluster cooperative pursuit and escape decision optimization method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, A computer program is stored. When the computer program is executed by a processor, the processor is caused to execute the cluster cooperative pursuit and escape decision optimization method according to any one of claims 1-7.