A node variable trapping control method and system based on deep learning

By introducing escapee prediction models and decision network models, and combining local observations and prior motion trends, a variable node encirclement control system was designed, which solves the problem of low success rate of encirclement in complex environments in existing technologies and achieves a more efficient and stable encirclement effect.

CN119047547BActive Publication Date: 2026-07-31SHANGHAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANGHAI UNIV
Filing Date
2024-08-22
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing capture methods struggle to handle transient behaviors such as communication blockages and individual disconnection in complex and high-risk environments, resulting in low capture success rates. Furthermore, existing deep learning solutions ignore human experience and prior knowledge, leading to unstable training and difficulty in handling complex tasks.

Method used

A node-variable capture control method based on deep learning is adopted. By combining a fugitive prediction model, a decision network model, and a location tracking network model with local observation and communication constraints, the fugitive's position is predicted using dual models. The decision is made based on the node's azimuth information and prior motion trend. A node-variable capture system is designed.

Benefits of technology

It improves the accuracy and stability of the encirclement and capture, can dynamically handle changes in the number of nodes and the entry and exit behavior of obstacles, reduces the dependence on network instability features, and enhances the reliability and success rate of the encirclement and capture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119047547B_ABST
    Figure CN119047547B_ABST
Patent Text Reader

Abstract

This invention relates to a deep learning-based method and system for variable node encirclement control. The method includes the following steps: acquiring the current state information of the escapee, its prior movement trend, and the current state information of each node; predicting the escapee's predicted position at the next moment based on the current state information of each node and the escapee's state information; obtaining the ideal position of each node at the next moment through a decision network model based on the predicted position, the state information of each node and its neighboring nodes, and the prior movement trend; and calculating the optimal encirclement path for each node through an orientation tracking network based on the current state information of each node, the escapee's predicted position at the next moment, and the ideal position of each node at the next moment. Compared with existing technologies, this invention has advantages such as independent decision-making by each node, high stability, and support for dynamic tracking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of dynamic search technology, and in particular to a method and system for variable encirclement control of nodes based on deep learning. Background Technology

[0002] When carrying out various encirclement and capture missions, especially those involving wide-ranging, complex, and high-risk environments, the individuals involved often face transient behaviors such as communication blockage or loss, forced mission interruption, or individual separation. They may also face irreversible behaviors such as device damage or serious malfunctions, causing one or more individuals to temporarily or permanently withdraw from the mission, resulting in a significant reduction in the success rate of the encirclement and capture. However, existing algorithms are mostly designed for ideal conditions and are difficult to effectively handle such scenarios.

[0003] Existing capture methods often rely on sensors or idealized methods to capture the transient behavior of escapees, failing to capture their long-term behavior. In real-world scenarios, escapees employ escape strategies and don't simply wait to be captured, posing challenges to efficient capture. Existing deep learning solutions for capture are overly focused on end-to-end processing, directly processing data from raw data to the target result, with all behaviors captured and represented by the network in a single step. However, real-world scenarios are highly complex, and this approach, requiring consideration of all features, leads to unstable training and necessitates large amounts of training data for tasks with multiple repetitions and cross-contamination features. Furthermore, deep learning models for complex tasks rely entirely on the network for feature extraction. While this approach is advantageous for handling complex tasks with features difficult for humans to understand and define, it neglects human experience and prior knowledge. This results in training difficulties and dependence on the network's inherent instability, often limiting model performance. Summary of the Invention

[0004] The purpose of this invention is to overcome the shortcomings of the existing technology and provide a node variable capture control method and system based on deep learning.

[0005] The objective of this invention can be achieved through the following technical solutions:

[0006] A deep learning-based variable capture control method for nodes, including escapees and multiple nodes, further includes the following steps:

[0007] S1: Obtain the current state information of the escapee, the prior movement trend, and the current state information of each node;

[0008] S2: Input the current state information of each node and the escapee's state information into the first prediction model corresponding to each node, and input the current state information of all nodes into the second prediction model corresponding to each node. Through the first prediction model and the second prediction model, predict the escapee's predicted position at the next moment.

[0009] S3: Based on the predicted position, the state information of each node and its neighboring nodes, and the prior movement trend, the ideal position of each node in the next moment is obtained through the decision network model.

[0010] S4: Based on the current state information of each node, the predicted position of the escapee in the next moment, and the ideal position of each node in the next moment, the optimal capture path for each node is calculated through the orientation tracking network.

[0011] Furthermore, in step S2, the predicted location of the escapee at the next moment is predicted as follows: First, the first prediction model is used to make a prediction, while the second prediction model is running. Both the first and second prediction models are updated online. After a certain time step, the prediction model reward of the first prediction model and the prediction model reward of the second prediction model are compared, and the prediction model with the higher prediction model reward is taken as the prediction model for the next time step.

[0012] The expression for online update is:

[0013]

[0014]

[0015] In the formula, W t+1 p1 W represents the network parameters of the first prediction model at the next time step. t p1 Let α be the network parameters of the first prediction model at the current time. p1 R is the learning rate of the first prediction model. p1 The reward for the first predictive model; W t+1 p2 W represents the network parameters of the second prediction model for the next time step. t p2 Let α be the network parameters of the second prediction model at the current time. p2 R is the learning rate for the second prediction model. p2 Rewards for the second prediction model;

[0016] The expression for the reward of the prediction model is:

[0017]

[0018]

[0019] In the formula, For the escapee's actual location at the next moment, The first prediction model predicts the escapee's location at the next moment. This is the predicted location of the escapee at the next moment by the second prediction model.

[0020] Furthermore, the expression for the input state space of the first prediction model is:

[0021]

[0022] In the formula, Let be the input state space of the first prediction model. Information on the current status of the escapee. This provides the current state information of the current node. The x-coordinate represents the current state of the escapee. The vertical axis represents the current state of the escapee. The x-coordinate of the current node's current state. The y-coordinate represents the current state of the current node. The x-axis component represents the relative velocity between the current node and the escapee. The vertical component represents the relative velocity between the current node and the escapee.

[0023] The expression for the output of the first prediction model is:

[0024]

[0025] In the formula, These represent the predictions of the escapee's x and y coordinates at time t+1, respectively.

[0026] Furthermore, it also includes an abnormal interval centered on the escapee, which includes an ideal enclosing circle. When the current node observes that a neighboring node is within the abnormal interval, the state information of the neighboring node is adjusted by projecting the position information of the neighboring node onto the ideal enclosing circle while keeping the velocity information of the neighboring node unchanged.

[0027] Furthermore, each node has an observation range. When nodes are evenly distributed on the ideal enclosing circle, the observation range is greater than the distance between adjacent nodes. When an adjacent node on one side is outside the abnormal interval, the position information of the adjacent node on that side is projected onto the intersection of the observation range of the current node and the ideal enclosing circle on the same side.

[0028] Furthermore, when the second prediction model cannot obtain the current state information of all nodes, for missing nodes, the second prediction model uses the estimated state information of the missing nodes instead of the current state information of the missing nodes. The expression for calculating the estimated state information of the missing nodes is as follows:

[0029]

[0030]

[0031] In the formula, This is the estimated state information of the current node for the k-th missing node at time t. The x-coordinate of the missing node. y = the ordinate of the missing node. The x-axis component represents the relative velocity between the missing node and the escapee. The vertical axis represents the relative velocity between the missing node and the escapee, where n is the total number of nodes, and θ is the vertical component of the velocity. i Let r be the azimuth angle of the current node i, and r be the radius of the ideal enclosing circle.

[0032] Furthermore, when an obstacle enters the observation range, the obstacle is treated as a new node until the obstacle leaves the observation range.

[0033] Furthermore, the expression for calculating the prior tendency of motion is:

[0034]

[0035] In the formula, Let represent the prior motion trend of the current node i at time t, where l1 and l2 represent the distances of the current node i to its left and right neighbors, respectively, and r is the radius of the ideal enclosing circle. The expression for calculating the ideal position of the current node at the next time step is:

[0036]

[0037]

[0038]

[0039]

[0040] In the formula, The x-coordinate of the ideal position. The ordinate of the ideal position. Let be the azimuth angle of the current node i at time t. As the overall trend, The derivative of the azimuth angle of the current node i, , where is the confidence level.

[0041] Furthermore, the state space expression of the orientation tracking network is:

[0042]

[0043] In the formula, For the state space of the orientation tracking network, Let x be the x-coordinate of the current node at the current moment. Let be the y-coordinate of the current node at the current moment. The vertical component represents the relative velocity between the current node and the escapee. The vertical component represents the relative velocity between the current node and the escapee. The predicted location of the escapee in the next moment.

[0044] In a second aspect, a deep learning-based variable encirclement control system for nodes includes an escapee and multiple nodes. The variable encirclement control system is used to execute any of the deep learning-based variable encirclement control methods described above. Each of the multiple nodes includes an escapee prediction model, a decision network model, and a location tracking network model. The escapee prediction model comprises a first prediction model and a second prediction model. The inputs of the first prediction model are the current node's state information and the escapee's state information. The inputs of the second prediction model are the state information of all nodes. The output of the prediction model is the predicted position of the escapee at the next moment, obtained through the first and second prediction models. The decision network model obtains the ideal position of all nodes at the next moment based on the predicted position, the state information of all nodes, and the prior movement trend. The location tracking network model calculates the optimal encirclement control result based on the state information of all nodes and the ideal positions of all nodes at the next moment.

[0045] Compared with the prior art, the present invention has the following beneficial effects:

[0046] 1) This invention introduces an escapee prediction model to capture the behavior of escapees, thereby providing reliable information support for all individuals in the pursuit. The escapee prediction model adopts a dual model and selects the appropriate model to obtain the prediction result through real-time utility. Considering the observation range and communication constraints, the decision network model adopts independent decision-making by individuals, which does not rely on the pursuit individuals' observation of the whole situation, and is more in line with the needs of real-world scenarios. By utilizing the prediction information of escapee behavior, the pursuit individuals can better plan their next actions and make more accurate and effective decisions.

[0047] 2) This invention uses azimuth information and angular offset as motion control information for nodes. While conforming to node observation and capture cooperation rules, it expresses the node entry and exit processes and their characteristics in the network state space design using estimated values ​​through rule models and virtual behaviors. This allows for handling a variable number of capture decision-makers without changing the network's input and output dimensions. Thus, it considers the impact of transient and permanent entry and exit behaviors of nodes and obstacles on the capture process without degrading network performance, and supports dynamic changes in the number of capture nodes.

[0048] 3) This invention introduces prior motion trends, providing intuitive decisions when using the network to estimate the motion trends of escapees, reducing the network's dependence on implicit unstable features, obtaining motion trends that better conform to the motion laws during escape, and improving the stability of the network model and the solution.

[0049] 4) In this invention, the escapee prediction model, decision network model and orientation tracking network model are trained separately, and there will be no mutual influence that would lead to a decrease in performance. Attached Figure Description

[0050] Figure 1 This is a schematic diagram of the structure of the present invention.

[0051] Figure 2 A schematic diagram illustrating the method for estimating the state of the remaining pursuers.

[0052] Figure 3 A schematic diagram of the encirclement strategy is provided.

[0053] Figure 4 A diagram illustrating the introduction of intuitive prior knowledge.

[0054] Figure 5 This is a diagram illustrating a normal capture operation.

[0055] Figure 6 This is a diagram illustrating the exit points in the encirclement and capture process.

[0056] Figure 7 This is a schematic diagram of the entry points in the encirclement and capture process.

[0057] Figure 8 This is a diagram illustrating obstacles encountered during the encirclement and capture process. Detailed Implementation

[0058] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. These embodiments are based on the technical solution of the present invention and provide detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments.

[0059] Example 1

[0060] This invention relates to a method and system for variable node containment control based on deep learning. The method includes the following steps:

[0061] S1: Obtain the current state information of the escapee, the prior movement trend, and the current state information of each node;

[0062] S2: Input the current state information of each node and the escapee's state information into the first prediction model corresponding to each node, and input the current state information of all nodes into the second prediction model corresponding to each node. Through the first prediction model and the second prediction model, and through the decision network model, predict the escapee's predicted position at the next moment.

[0063] S3: Based on the predicted position, the state information of each node and its adjacent nodes, and the prior movement trend, obtain the ideal position of each node in the next moment;

[0064] S4: Based on the current state information of each node, the predicted position of the escapee in the next moment, and the ideal position of each node in the next moment, the optimal capture path for each node is calculated through the orientation tracking network.

[0065] like Figure 1 As shown, the system includes multiple pursuers as nodes. Each node includes an escapee prediction model, a decision network model, and a location tracking network model. The escapee prediction model comprises a first prediction model and a second prediction model. The input of the first prediction model is the current node's state information and the escapee's state information. The input of the second prediction model is the state information of all nodes. The output of the prediction model is the predicted position of the escapee at the next moment, obtained through the first and second prediction models. The decision network model obtains the ideal position of all nodes at the next moment based on the predicted position, the state information of all nodes, and the prior movement trend. The location tracking network model calculates the optimal capture control result based on the state information of all nodes and the ideal position of all nodes at the next moment.

[0066] I. Design of the Escapee Prediction Model:

[0067] The escapee prediction model includes a first prediction model and a second prediction model. The escapee's movement pattern should not be limited to accommodate escapees with various escape tendencies. The observation space of the escapee that the capturer's node can acquire at time t is defined as follows: in These represent the positions of the escapee in the x and y directions, respectively.

[0068] Preferably, by equipping more sensors, the escapee's speed and even acceleration information can be obtained. In this case, the escapee's state should be defined as... in Let x and y represent the escapee's velocity and acceleration in the x and y directions, respectively. By utilizing a state space that contains more complete information about the escapee, the prediction model will be able to predict the escapee's behavior more accurately.

[0069] This invention does not limit the state space of the escapee, but for the sake of generality, the escapee's state space will be used below. For example.

[0070] The prediction network runs on all hunter nodes and provides assistance to them; therefore, the prediction model needs to consider the communication constraints and local observation constraints faced by the hunters. Given this, the hunters can only obtain observation information within a certain range. Since the hunters have complete observability of their own state, we define the observation of hunter i at time t as follows: in Let x and y represent the positions of the pursuer i in the x and y directions at time t, respectively. Since the escapee may react accordingly based on the state and behavior of the pursuer, the pursuer's own state information is included in the state space of the prediction network. Furthermore, the prediction network is not limited to a specific model form and can be implemented using n linear layers.

[0071] For the first prediction model, its input state space is designed as follows:

[0072]

[0073] Since the prediction model is mounted on the encircling party, the state information it obtains comes from the encircling party. However, the encircling party is constrained by communication and local observation constraints. Therefore, the first model, based on the actually available information, considers its state space as the information of the encircler i itself and the state information of the escapee. The output of the first prediction model is the encircler i's prediction of the escapee's state at time t for the next time step:

[0074]

[0075] In the formula, Let represent the predicted positions of the escapee in the x and y directions at time t+1, respectively. Let x and y represent the true positions of the escapee in the x and y directions at time t+1, and let y be the true state that the escapee transitions to in the next time step. The reward function is:

[0076]

[0077] In the formula, For the escapee's actual location at the next moment, This represents the predicted location of the escapee at the next moment, as determined by the first prediction model.

[0078] Let the network parameters of the first prediction model be W. p1 Let the learning rate be α. p1 Then, a network update can be represented as:

[0079]

[0080] In the formula, W t+1 p1 W represents the network parameters of the first prediction model at the next time step. t p1 Let α be the network parameters of the first prediction model at the current time. p1 R is the learning rate of the first prediction model. p1 The first prediction model will be rewarded.

[0081] From the perspective of the first prediction network, only the behavioral decisions made by the escapee under the influence of its own pursuer i are considered. This is useful in many aggressive or tense scenarios, but it is insufficient to cover the entire pursuit process.

[0082] Therefore, for the second prediction model, its input state space includes the state information of all the pursuers. That is, the design of the input state space of the second prediction network also considers the state information of other cooperating pursuers. This is because the escapee's behavior will change according to the state of all the pursuers, not just a single pursuer. Since the state information of other cooperating pursuers cannot be obtained during the prediction network usage phase, this part of the design is further divided into a training phase and a usage phase.

[0083] During the training phase, the input state space of the second prediction model considers the state information of all encircling parties. Among the n encircling parties participating in the encirclement task, the input state space of the second prediction model fitted to any encircling party i is designed as follows: in This represents the state information of the nth pursuer, including its position in the x and y directions. and speed The reward function has the same form as the first prediction model, and is expressed as:

[0084]

[0085] In the formula, This is the predicted location of the escapee at the next moment by the second prediction model.

[0086] During the usage phase, the second prediction model cannot obtain the state information of all surrounding parties. Therefore, its input state space is designed as follows: in This represents an estimate of the state information of other pursuers, including estimates of their positions in the x and y directions. And the estimation of speed. Various estimation methods can be used to obtain this state, and different state spaces can be selected according to the actual situation. For example... Figure 2 As shown, the estimation method here assumes that the remaining pursuers are in an ideal encirclement state, and the dispersion angle between the pursuers is estimated to be 2π / n, where n is the number of pursuers in the encirclement task.

[0087] For the second prediction model mounted on the i-th predator, assuming the radius of the ideal encirclement is r, its state estimate for the k-th (k≠i) predator is: in:

[0088]

[0089] Only when Support is only considered when it is acquired. In the state of [condition], otherwise its speed estimate is set to 0. In this way, the second prediction model maintains the same state space during the usage phase as during the training phase. Let the network parameters of the second prediction model be W. p2 Let the learning rate be α. p2 The network update applicable to both the training and usage phases can be represented as:

[0090]

[0091] In the formula, W t+1 p2 W represents the network parameters of the second prediction model for the next time step. t p2 Let α be the network parameters of the second prediction model at the current time. p2 R is the learning rate for the second prediction model. p2 Rewards for the second prediction model;

[0092] During the encirclement task, both the first and second prediction models are used, following certain rules to dynamically select the prediction model based on the operational status. The rules are as follows: Initially, the first prediction model is used preferentially, but the second prediction model also runs simultaneously and performs online updates to adapt to different scenarios and environmental changes, exhibiting online learning characteristics. After a certain time step K, the rewards for both the first and continuously updated second prediction models are evaluated. The higher value is selected as the prediction model for the next time step K. After this step, the model is evaluated again to select a suitable prediction model.

[0093] II. Policy Network Model Design:

[0094] Traditional strategies that directly control node inputs struggle to achieve effective cooperation and encirclement, making it difficult to guarantee a high success rate. Considering that the encirclement task is cooperative, and each participant can only obtain information about others through local observations, this invention's strategy network does not directly provide the final control strategy, but instead provides the derivative of an azimuth angle. Furthermore, since the number of participants may change during the encirclement process, the spatial distribution characteristics among them are difficult to model and represent by the network. Therefore, this invention designs an encirclement strategy that allows nodes to enter and exit and possesses high stability, starting from the node's azimuth offset angle.

[0095] For a bounding square containing n nodes (4 nodes in this example), such as Figure 3 As shown, the observation range of any node is d, and assuming that all nodes are uniformly distributed on the ideal enclosing circle O0, any node can observe at least its neighboring nodes. Taking node i as an example, when node i can simultaneously observe its left and right neighboring nodes, and all of them are on the ideal enclosing circle O0, its state space is designed as follows: in These represent the observations of node i at time t of its left node i-1 and right node i+1 in the x and y directions, respectively. Here, the state prediction of the escapee at the next time step obtained from the prediction model in step one is presented. The state space of node i is also considered, with the aim of enabling the policy network model to account for the states that the escapee might reach in the future, thus giving the decision a certain predictive nature and making the decision more accurate. Furthermore, since the prediction model is designed and trained separately, it will not interfere with the decision network and cause performance degradation.

[0096] Similarly, when the velocity information of the observed node To enable richer state representations, support is provided when data is acquired (e.g., through estimation algorithms or sensor support). It can be represented as:

[0097]

[0098] When an observed neighbor node is not on circle O0, but lies between the upper bound O1 and the lower bound O2 of the anomalous region, it is considered a normal neighbor node, but with a certain degree of anomalousness. When anomaly exists, projection is used to project the node onto the ideal enclosing circle O0. The projection point is represented by the intersection of circle O3 (with the radius from node i to the observed neighbor node) and the ideal enclosing circle O0. This process... Figure 3 This can be represented as p1 1 or p1 2Project onto P1. Before and after a node is projected, if velocity information exists, that velocity information is not projected; only its position is projected. The more a neighboring node deviates from the ideal enclosing circle O0, the higher its degree of anomalousness, and the greater the offset caused by the observed node's projection onto the ideal enclosing circle O0 relative to its own node i. By using projection, the cost of a neighboring node returning to the ideal enclosing circle can be represented in node i's observation state of its neighboring nodes. This method not only allows for early planning of node exits or anomalous nodes but also provides early obstacle avoidance functionality.

[0099] When an observed neighbor node is outside O1 or inside O2, it is considered a node exiting the state. To ensure the approach behavior from node i to the point of node loss and the integrity of the state space, a position projection is also performed. At this point, regardless of the location of the neighbor node or whether it is observed, it is uniformly projected onto the intersection of the observation boundary of node i and the ideal enclosing circle O0. This point best simulates the characteristics of a lost node state. This process... Figure 3 The middle can be represented as or Project onto P2.

[0100] Specifically, consider a scenario with four capture nodes and one escape node. Ideally, the four capture nodes are evenly distributed around the escapee, maintaining a uniform capture radius and equal angular spacing between them, such as... Figure 5 As shown in the diagram, when a node exits abnormally, its left and right neighboring nodes will tend to move closer to the exit point to ensure that the remaining nodes are distributed as evenly as possible. Figure 6 As shown. Similarly, when a new node enters, the left and right nodes at the entry point will tend to disperse to maintain a uniform distribution among them, as shown. Figure 7 As shown. Furthermore, entering obstacles can also be considered nodes because their behavior is determined by observation and interaction between nodes. When an obstacle enters the observation range, it is treated as a new node until it leaves the observation range. There is no strict distinction between obstacles and nodes, thus providing obstacle avoidance functionality, such as... Figure 8 As shown.

[0101] The projected position information of neighbor nodes i-1 and i+1 can be represented as follows: and The state space after projection onto the nodes can then be represented as: in:

[0102]

[0103] or

[0104]

[0105] At this point, the input to the policy network model of a node can be represented as a single-dimensional... or Its output is the derivative of the offset angle of the hunter i at time t, denoted as . And a confidence level

[0106] III. Calculation of Final Orientation in the Strategy Model

[0107] To account for node entry and exit scenarios, the policy model does not directly output control over the target. Instead, it outputs the derivative of the offset angle and the confidence level, determining the ideal state of the node in the next time step through angle offset. However, since the encirclement task also needs to consider cooperation between nodes, and under the constraints of communication, local observation, and individual decision-making, the policy model becomes dependent on implicit instability features. Therefore, it is difficult for the policy network to independently complete the angle offset task of individual nodes and achieve the optimal control objective. Therefore, this patent introduces prior human knowledge to assist the network learning. This helps provide the policy network with some intuitive perspectives, reduces dependence on implicit instability features, and thus achieves better performance.

[0108] like Figure 4 As shown, intuitive information introduced artificially can be represented as a trend. Where l1 and l2 represent the distances from the left and right neighbors of the target node i to itself, respectively. When the nodes are unevenly distributed, intuitively, this manifests as a difference in the distances from node i to its left and right neighbors. Node i should tend to move towards more distant neighbors to eliminate this difference. The greater this difference, the stronger the tendency for the target node to move towards the more distant neighbor. Larger. Output of the integrated network and confidence level and intuitive trends The overall trend is:

[0109]

[0110] Therefore, the azimuth of the ideal position of the hunter i at time t+1 can be calculated. in:

[0111]

[0112]

[0113] In the formula, The x-coordinate of the ideal position. The ordinate of the ideal position. Let be the azimuth angle of the current node i at time t. As the overall trend, The derivative of the azimuth angle of the current node i, , where is the confidence level.

[0114] Because the prediction network provides a prediction of the escapee's location in the next moment. There is a slight error, therefore the ideal azimuth point here is... The current location information of the escapee is not used in the calculation; more accurate location information of the escapee at the current moment is needed here. Similarly, when the escapee's velocity and acceleration information can be obtained, the ideal azimuth point can be extended. For example, the escapee's current velocity and acceleration can be included as the velocity and acceleration that the pursuer i expects to reach at the next moment, so as to keep in sync with the escapee and improve the pursuit performance. The orientation tracking network trained separately in step three will drive the pursuer i to reach the ideal point, but considering network error, the pursuer i will reach the actual point at time t+1. This policy network model is trained using a reinforcement learning algorithm, and the reward can be defined as R. π =r dist +r dire The expression for the reward in the policy network model is:

[0115] R π =r dist +r dire

[0116]

[0117] In the formula, R π For the reward of the policy network model, r dire For azimuth bonus, r dist As a distance reward, Let be the azimuth angle of the left neighbor node at the next time step. Let be the azimuth angle of the current node at the next moment. Let be the azimuth angle of the right neighbor node at the next time step, and r be the radius of the ideal enclosing circle. This represents the location information of the pursuer i at time t+1. This represents the location information of the escapee at time t+1.

[0118] IV. Design of the Orientation Tracking Network Model:

[0119] For the ideal azimuth point obtained in step two It cannot be reached directly; the control input of the apprehending party is needed to guide it to the target location. Considering the numerous uncertainties and modeling errors in the real environment, and also considering versatility, this patent uses a location tracking network to provide optimal control for the apprehending party. The state space design of the location tracking network model is as follows:

[0120]

[0121] In the formula, For the state space of the orientation tracking network, Let x be the x-coordinate of the current node at the current moment. Let be the y-coordinate of the current node at the current moment. The vertical component represents the relative velocity between the current node and the escapee. The vertical component represents the relative velocity between the current node and the escapee. The predicted location of the escapee in the next moment.

[0122] The network update method is consistent with the escaper prediction model, and a loss function can also be designed and solved iteratively using adaptive dynamic programming. Here, the reward function is designed as follows:

[0123]

[0124] In a homogeneous surround scenario, all networks can be configured as a general network shared by all surroundrs. Based on the characteristics of the surround task, this invention designs three network models. Considering communication constraints and local observation constraints, it proposes an adaptive, multi-model surround task solution with nodes that can enter and exit. The escapee prediction model includes a first prediction model with a local perspective and a second prediction model with online updates and a global perspective. These two models, from different perspectives, are selected optimally during the surround task. The surroundr strategy network model designs its state space using inner and outer rings to facilitate node entry and exit. Simultaneously, it introduces intuitive human priors to guide the network learning process and improve network performance. The orientation tracking network model is responsible for ensuring nodes accurately reach their given desired locations, driving node actions to complete the surround task.

[0125] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.

Claims

1. A deep learning-based variable encirclement control method for nodes, comprising escapees and multiple nodes, characterized in that, It also includes the following steps: S1: Obtain the current state information of the escapee, the prior movement trend, and the current state information of each node; S2: Input the current state information of each node and the escapee's state information into the first prediction model corresponding to each node, and input the current state information of all nodes into the second prediction model corresponding to each node. Through the first prediction model and the second prediction model, predict the escapee's predicted position at the next moment. S3: Based on the predicted position, the state information of each node and its neighboring nodes, and the prior movement trend, the ideal position of each node in the next moment is obtained through the decision network model. S4: Based on the current state information of each node, the predicted position of the escapee in the next moment, and the ideal position of each node in the next moment, the optimal capture path for each node is calculated through the orientation tracking network. In step S2, the prediction of the escapee's predicted location at the next moment is specifically as follows: First, a first prediction model is used for prediction, while a second prediction model is running. Both the first and second prediction models are updated online. After a certain time step, the prediction model reward of the first prediction model and the prediction model reward of the second prediction model are compared, and the prediction model with the higher prediction model reward is taken as the prediction model for the next time step. The expression for the online update is: In the formula, The network parameters of the first prediction model for the next time step. These are the network parameters of the first prediction model at the current moment. The learning rate of the first prediction model. The first prediction model will be rewarded; The network parameters for the second prediction model at the next time step. The network parameters of the second prediction model at the current time. The learning rate for the second prediction model. Rewards for the second prediction model; The expression for the reward of the prediction model is: In the formula, The actual location of the escapee at the next moment. The first prediction model predicts the escapee's location at the next moment. The second prediction model predicts the escapee's location at the next moment; The expression for the input state space of the first prediction model is: In the formula, Let be the input state space of the first prediction model. Information on the current status of the escapee. This provides the current state information of the current node. The x-coordinate represents the current state of the escapee. The vertical axis represents the current state of the escapee. The x-coordinate of the current node's current state. The y-coordinate represents the current state of the current node. The x-axis component represents the relative velocity between the current node and the escapee. The vertical component represents the relative velocity between the current node and the escapee. The expression for the output of the first prediction model is: In the formula, They represent in Predicting the horizontal and vertical coordinates of the escapee at all times; The calculation expression for the prior motion trend is: In the formula, For the current node In t The a priori trend of motion at any given moment. Representing the current node Distances to the left and right neighbor nodes. Let the radius of the ideal enclosing circle be denoted by the expression for calculating the ideal position of the current node at the next moment: In the formula, The x-coordinate of the ideal position. The ordinate of the ideal position. For the current node In Azimuth at time, As the overall trend, Current node The derivative of the azimuth angle, , where is the confidence level.

2. The node variable capture control method based on deep learning according to claim 1, characterized in that, It also includes an abnormal interval centered on the escapee, which includes an ideal enclosing circle. When the current node observes that a neighboring node is within the abnormal interval, the state information of the neighboring node is adjusted by projecting the position information of the neighboring node onto the ideal enclosing circle while keeping the velocity information of the neighboring node unchanged.

3. The node variable encirclement control method based on deep learning according to claim 2, characterized in that, The node has an observation range. When the nodes are evenly distributed on the ideal enclosing circle, the observation range is greater than the distance between adjacent nodes. When an adjacent node on one side is outside the abnormal interval, the position information of the adjacent node on that side is projected onto the intersection of the observation range of the current node and the ideal enclosing circle on the same side, as an estimate of the state information of the missing node.

4. The deep learning-based variable encirclement control method for nodes according to claim 3, characterized in that, When the second prediction model cannot obtain the current state information of all nodes, for missing nodes, the second prediction model uses the estimated state information value of the missing node instead of the current state information of the missing node. The calculation expression for the estimated state information value of the missing node is as follows: In the formula, For the current node to the first k The missing node is t State information estimate at time step The x-coordinate of the missing node. y = ... The x-axis component represents the relative velocity between the missing node and the escapee. The vertical axis component represents the relative velocity between the missing node and the escapee. The total number of all nodes. For the current node azimuth angle, The radius of the ideal enclosing circle.

5. The node variable capture control method based on deep learning according to claim 3, characterized in that, When an obstacle enters the observation range, it is treated as a new node until the obstacle leaves the observation range.

6. The node variable capture control method based on deep learning according to claim 1, characterized in that, The state space expression of the orientation tracking network is: In the formula, For the state space of the orientation tracking network, Let x be the x-coordinate of the current node at the current moment. Let be the y-coordinate of the current node at the current moment. The vertical component represents the relative velocity between the current node and the escapee. The vertical component represents the relative velocity between the current node and the escapee. The predicted location of the escapee in the next moment.

7. A node-based variable containment control system based on deep learning, comprising an escapee and multiple nodes, characterized in that, The variable encirclement control system is used to execute a deep learning-based variable encirclement control method as described in any one of claims 1-6. Each of the plurality of nodes includes an escapee prediction model, a decision network model, and a location tracking network model. The escapee prediction model comprises a first prediction model and a second prediction model. The inputs of the first prediction model are the current node's state information and the escapee's state information. The inputs of the second prediction model are the state information of all nodes. The output of the prediction model is the predicted position of the escapee at the next moment, obtained through the first and second prediction models. The decision network model obtains the ideal position of all nodes at the next moment based on the predicted position, the state information of all nodes, and the prior movement trend. The location tracking network model calculates the optimal encirclement control result based on the state information of all nodes and the ideal position of all nodes at the next moment.