Vehicle chasing path planning method in urban scene

Through the rasterized processing and design of ATT-D3RQN network model, combined with long and short-term memory network and self-attention module, DQN algorithm is improved for training, which solves the real-time and resource efficiency problems of hunting path planning in urban market scenarios, and achieves more efficient path planning and hunting success rate.

CN120278358APending Publication Date: 2025-07-08XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510299281.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing pursuit path planning algorithms have limitations in real-time, complex environment adaptability and resource efficiency in urban market scenarios, especially the low efficiency of A-star algorithms, high complexity of game theory methods, and high demand for computing resources of reinforcement learning methods.

Method used

The city tracking scenario is processed in a grid, the ATT-D3RQN network model is designed, combined with long and short-term memory networks and self-attention modules, and the DQN algorithm is improved for training, and the path planning in a dynamic environment is realized through discrete data processing and weighted extraction.

Benefits of technology

It improves the real-time and resource efficiency of path planning algorithms in complex environments, enhances dynamic adaptability, and improves the pursuit success rate and path planning performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278358A_ABST
    Figure CN120278358A_ABST
Patent Text Reader

Abstract

The invention discloses a vehicle chasing path planning method in an urban scene. The method comprises the following steps: modeling an urban tracking scene; an ATT-D3RQN network model suitable for path planning is designed; based on a modeling result, training the designed ATT-D3RQN network model by using an improved DQN algorithm to obtain a trained ATT-D3RQN network model; inputting the state information of two pursuing parties to be subjected to path planning into the trained ATT-D3RQN network model, calculating the Q value of each action in the action space, selecting the action corresponding to the maximum Q value as the execution action of the pursuing vehicle, obtaining the state information of the two new pursuing parties through the selected execution action of the pursuing vehicle, and performing path planning on the two new pursuing parties according to the state information of the two new pursuing parties. And inputting the state information of the two new pursuing parties into the trained ATT-D3RQN network model until the pursuing vehicle successfully pursues the pursued vehicle. According to the invention, the vehicle chasing success rate is improved, and the vehicle chasing time is shortened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of vehicle path planning, and particularly relates to a method for planning a vehicle pursuit path in an urban scenario. Background Art

[0002] In the fields of control science and machine learning, path planning is one of the main research contents of motion planning, which involves the process of a computer generating an optimal path under given environments and constraints. This technology has wide applications in many fields, such as the autonomous collision-free movement of robots, the obstacle avoidance and penetration flight of unmanned aerial vehicles, etc. In the control field, the pursuit problem of mobile intelligent agents is one of the classical problems. Such problems mainly consider a group of pursuers capturing another group of evaders by planning reasonable paths. In such scenarios, path planning technology can help pursuers dynamically plan the optimal pursuit path according to environmental information and target states. Therefore, the pursuit path planning technology is particularly important in this context because it directly relates to the pursuit efficiency and success rate of pursuers. In modern society, pursuit scenarios mainly occur in cities. Therefore, the method for planning a vehicle pursuit path for urban road scenarios is very important and has broad application prospects.

[0003] Currently, there are few pursuit path planning methods specifically for urban scenarios, and other path planning algorithms are used for substitution. The existing technical solutions related to pursuit path planning are as follows:

[0004] 1. A* algorithm. This algorithm is a heuristic search algorithm widely used in path planning. It combines the advantages of the breadth-first search of Dijkstra's algorithm and heuristic search, and can efficiently find the optimal path from the starting point to the ending point in a complex environment. The core of the A* algorithm is to evaluate the priority of each node through a heuristic function, so as to select the most promising node for expansion. Its goal is to find the shortest path from the starting point to the ending point while minimizing the search space. In an urban scenario, the urban road network can be abstracted as a topological graph, and then the A* algorithm can be used on the topological graph to obtain the path between two static points. The advantages of the A* algorithm are high operating efficiency and the ability to adapt to various scenarios. The disadvantages of the A* algorithm are: it has very strict requirements for the selection of the heuristic function, and improper selection will lead to low efficiency; in large-scale problems, this algorithm may occupy a large amount of memory; the efficiency is very low when performing dynamic path planning.

[0005] 2. Algorithms based on game theory. The pursuit problem is essentially a differential game problem. The Hamilton-Jacobi-Isaacs equation can be used to establish a model for the whole process, and then the optimal solution of the pursuit process can be obtained by solving the differential equation. However, the complexity of this calculation method is relatively high, and it is only applicable to scenarios with relatively simple maps. For complex scenarios such as urban roads, the equations obtained by this method will be extremely complex and difficult to solve.

[0006] 3. Path planning algorithms based on reinforcement learning. Such algorithms are a method of learning the optimal path through the interaction between an agent and the environment. In recent years, with the development of deep reinforcement learning, such algorithms have shown significant advantages in path planning tasks in dynamic and uncertain environments. The core of reinforcement learning path planning is to learn the optimal policy through the interaction between the agent and the environment. The agent selects actions based on the current state and adjusts the policy according to the reward signal feedback from the environment to maximize the cumulative reward. Although the reinforcement learning method has the advantage of adapting to dynamic environments, it faces the technical bottleneck of high computational resource requirements.

[0007] Therefore, the existing pursuit path planning algorithms still have certain limitations in urban scenarios, mainly reflected in real-time performance, complex environment adaptability, and resource efficiency. Summary of the Invention

[0008] To solve the above problems existing in the prior art, the present invention provides a vehicle pursuit path planning method in an urban scenario. The technical problems to be solved by the present invention are realized through the following technical solutions:

[0009] An embodiment of the present invention provides a vehicle pursuit path planning method in an urban scenario, including:

[0010] Modeling the urban tracking scenario; the modeling includes rasterizing the urban tracking scenario, defining state information and action space, and designing a total reward function;

[0011] Designing an ATT-D3RQN network model suitable for path planning; wherein, the ATT-D3RQN network model is composed of a long short-term memory network, a self-attention module, and a fully connected network;

[0012] Based on the modeling results, training the designed ATT-D3RQN network model using an improved DQN algorithm to obtain a trained ATT-D3RQN network model; wherein, the improved DQN algorithm is based on the original DQN algorithm. For each complete set of state information: the complete set of state information is segmented according to a preset segmentation length and stored in an experience buffer pool. Each time training is performed, several segmented state information fragments are extracted from the experience buffer pool, and the extraction is performed in a weighted extraction manner;

[0013] Input the state information of the pursuer and the pursued to be path - planned into the trained ATT - D3RQN network model. In the trained ATT - D3RQN network model, calculate the Q - value of each action in the action space, and select the action corresponding to the maximum Q - value as the execution action of the pursuit vehicle, so as to obtain the new state information of the pursuer and the pursued through the selected execution action of the pursuit vehicle. Input the new state information of the pursuer and the pursued into the trained ATT - D3RQN network model until the pursuit vehicle successfully captures the pursued vehicle.

[0014] In an embodiment of the present invention, the defined state information includes the position coordinates of the pursuit vehicle and the pursued vehicle.

[0015] In an embodiment of the present invention, the defined action space includes five actions: moving in the four directions of up, down, left, and right and staying still.

[0016] In an embodiment of the present invention, the designed total reward function is expressed by the formula:

[0017] R = R1 + R2 + R3 - n;

[0018] Wherein, R represents the total reward, R1 represents the reward generated by successful pursuit, R2 represents the negative reward due to distance, R3 represents the negative reward generated by the pursuit vehicle due to collision or attempting to rush out of the boundary, and n represents the number of time steps experienced in the pursuit process;

[0019] The reward generated by successful pursuit is expressed by the formula:

[0020]

[0021] The negative reward due to distance is expressed by the formula:

[0022]

[0023] Wherein, (x p , y p ) represents the position coordinates of the pursuit vehicle, and (x e , y e ) represents the position coordinates of the pursued vehicle;

[0024] The negative reward generated by the pursuit vehicle due to collision or attempting to rush out of the boundary is expressed by the formula:

[0025]

[0026] In an embodiment of the present invention, the designed ATT - D3RQN network model includes a first long - short - term memory network, a second long - short - term memory network, a self - attention module, a first fully - connected layer network, and a second fully - connected layer network connected in sequence; wherein,

[0027] Both the first long short-term memory network and the second long short-term memory network are LSMT or Bi-LSMT;

[0028] The self-attention module is a self-attention module based on a single-head attention mechanism;

[0029] The first fully connected layer network includes a first fully connected layer and a second fully connected layer connected in sequence;

[0030] The second fully connected layer network includes a third fully connected layer and a fourth fully connected layer respectively connected to the first fully connected layer. The third fully connected layer is connected to a fifth fully connected layer, and the fourth fully connected layer is connected to a sixth fully connected layer; among them, the third fully connected layer and the fifth fully connected layer constitute a value branch network, and the fourth fully connected layer and the sixth fully connected layer constitute an advantage branch network; a value is obtained according to the value branch network, an advantage value is obtained according to the advantage branch network, and a Q value is calculated according to the value and the advantage value.

[0031] In an embodiment of the present invention, the process of training the designed ATT-D3RQN network model using an improved DQN algorithm based on the modeling result includes:

[0032] Initialize the number of iterations, the experience buffer pool, and the network parameters of the ATT-D3RQN network model;

[0033] Initialize the urban tracking scenario;

[0034] According to the initialized urban tracking scenario, obtain multiple groups of complete state information sets. For each group of complete state information sets: divide the complete state information set into several state information segments according to a preset segmentation length;

[0035] For each state information segment s t , the steps to be executed include: S3041. Select the action a of the pursuit vehicle using the ε-greedy strategy t ; S3042. Execute the action a t , observe the corresponding total reward r t and the next state information segment s t+1 , and put (s t , a t , r t , s t+1 ) into the experience buffer pool and record the index corresponding to (s t , a t , r t , s t+1 ); S3043. Count in the state information segment s t and s t+1Among them, the number of times the moving direction of the pursued vehicle changes and the number of times when the Manhattan distance between the pursued vehicle and the pursuing vehicle is less than or equal to 2 are used. The sum of these two numbers is incremented by 1 to obtain the empirical sample weight, and according to the said index, the empirical sample weight is stored in the corresponding position in the weight table; S3044. Randomly extract several groups of state information segments from the empirical buffer pool. The extraction probability of each group of state information segments is determined by the empirical sample weight; calculate the target Q value and the actual Q value corresponding to each group of state information segments, calculate the total loss based on all the target Q values and actual Q values, and use the gradient descent method to update the network parameters of the ATT-D3RQN network model;

[0036] Determine whether the pursuing vehicle has successfully pursued the pursued vehicle. If so, return to the step of initializing the urban tracking scenario. If not, return to the execution step for each state information segment s t until the iteration stop condition is met, and output the trained ATT-D3RQN network model.

[0037] In an embodiment of the present invention, the extraction probability of each group of state information segments is equal to the empirical sample weight of the state information segment divided by the sum of the empirical sample weights of all the extracted groups of state information segments.

[0038] In an embodiment of the present invention, calculating the target Q value corresponding to each group of state information segments is expressed by the formula:

[0039]

[0040] where s t represents the state information segment at time t, y t represents the target Q value corresponding to the state information segment s t r represents the total reward observed under the state information segment s t represents the total reward observed under the state information segment s t γ represents the discount factor, s t+1 represents the state information segment at the next moment observed under the state information segment s t θ - , θ both represent the network parameters of the ATT-D3RQN network model, Q() represents the function for calculating the Q value through the ATT-D3RQN network model, represents inputting the state information segment s t+1 and all actions in the action space into the ATT-D3RQN network model to obtain the Q values corresponding to all actions, and selecting the action a corresponding to the maximum Q value.

[0041] In an embodiment of the present invention, the method further includes:

[0042] Divide the urban tracking scenario into regions, execute S301 - S305 for each region to obtain the trained ATT-D3RQN network model for the corresponding region, and store the network parameters corresponding to the trained ATT-D3RQN network model of this region in the network parameter table.

[0043] In an embodiment of the present invention, if the pursuer and the pursued to be path-planned are in the urban tracking scenario of a certain region after division, read the network parameters corresponding to this region from the network parameter table, load the ATT-D3RQN network model using the read network parameters, calculate the Q value of each action in the action space in the loaded ATT-D3RQN network model, and select the action corresponding to the maximum Q value as the execution action of the pursuit vehicle, so as to obtain the new state information of the pursuer and the pursued through the selected execution action of the pursuit vehicle, and input the new state information of the pursuer and the pursued into the loaded ATT-D3RQN network model until the pursuit vehicle successfully captures the pursued vehicle.

[0044] The beneficial effects of the present invention:

[0045] The vehicle pursuit path planning method in the urban scenario proposed by the present invention simplifies the urban pursuit scenario through rasterization processing, discretizes the originally continuous urban pursuit scenario and the actions of the vehicle, which can reduce the complexity of the deep reinforcement learning model, thereby reducing the requirements of the algorithm for computing resources and making the algorithm easier to deploy; designs an ATT-D3RQN network model suitable for path planning. This ATT-D3RQN network model combines a long short-term memory network and a self-attention module, integrates the long short-term memory network into the network model. Through the long short-term memory network, the position of the target vehicle can be predicted, and advance planning can be carried out based on the predicted position, thereby improving the success rate of pursuit. On the basis of the long short-term memory network, a self-attention mechanism is introduced, which can effectively improve the learning ability of the recurrent neural network for long sequences, and thus improve the performance of the path planning algorithm; adopts a path planning algorithm based on the deep reinforcement learning DQN algorithm. This algorithm can adjust the strategy in real time according to the dynamic environment, improving the operation efficiency of the algorithm. And the present invention improves the DQN algorithm, slices the input data of the algorithm, so that when the agent extracts data from the experience buffer pool, it can obtain scattered time series, which better meets the requirement of the DQN algorithm for data to be independently and identically distributed, and the extraction adopts a weighted extraction method. This experience replay with priority can let the agent learn more key information, thereby accelerating the convergence of the strategy. Through experimental verification, the method proposed by the present invention is superior to the existing algorithms in terms of convergence speed, training stability, pursuit success rate, and pursuit steps. Generally speaking, the present invention proposes a new vehicle pursuit path planning method. This method enhances the dynamic adaptability through reinforcement learning, enhances the adaptability to complex environments through the designed ATT-D3RQN network model, and improves the execution efficiency by discretizing the road environment, thereby breaking through the constraints of the existing methods in terms of real-time performance, adaptability to complex environments, and resource efficiency, and providing a systematic solution for urban pursuit scenarios that takes into account path optimality, dynamic response speed, and computational economy.

[0046] The following will further elaborate on the present invention in conjunction with the drawings and embodiments. Brief Description of the Drawings

[0047] Figure 1 is a schematic flow chart of a vehicle pursuit path planning method in an urban scenario provided by an embodiment of the present invention;

[0048] Figure 2 is a schematic structural diagram of the ATT-D3RQN network model provided by an embodiment of the present invention;

[0049] Figure 3 is a schematic diagram of the training process of the ATT-D3RQN network model provided by an embodiment of the present invention;

[0050] Figure 4 It is a schematic diagram of the traditional free - space pursuit path planning;

[0051] Figure 5 It is a schematic diagram of the average reward comparison of the traditional DQN algorithm, DRQN algorithm, and the method proposed in the present invention under different numbers of iteration rounds;

[0052] Figure 6 It is a schematic diagram of the capture success rate comparison of the traditional BFS algorithm, traditional DQN algorithm, and the method proposed in the present invention under different map sizes;

[0053] Figure 7 It is a schematic diagram of the capture step comparison of the traditional BFS algorithm, traditional DQN algorithm, and the method proposed in the present invention under different map sizes. Detailed implementation manners

[0054] The following further describes the present invention in detail with specific embodiments, but the implementation manners of the present invention are not limited thereto.

[0055] Please refer to Figure 1 , an embodiment of the present invention provides a vehicle pursuit path planning method in an urban scenario, which specifically includes the following steps:

[0056] S10. Model the urban pursuit scenario; this modeling includes rasterizing the urban pursuit scenario, defining state information and action space, and designing a total reward function.

[0057] In the embodiment of the present invention, the map corresponding to the urban pursuit scenario is first rasterized, that is, the road is defined as a walkable grid, and other parts are defined as non - walkable grids. Each of the pursuer and the pursued occupies a grid. In the rasterization process, each grid corresponds to a square with a side length of 5 meters on the actual map. Then, the time is discretized, that is, the time is divided into uniform time steps, and each of the pursuer and the pursued takes one step of action at each time step. Here, it is stipulated that only one action can be selected at each time step. After the urban pursuit scenario is defined, the relevant content of deep reinforcement learning needs to be determined:

[0058] Define state information: This state information is used to represent the state of the pursuer and the pursued at each time step, and is also the input data for the subsequent DQN algorithm. In the embodiment of the present invention, the state information is defined as the position information of the pursuer and the pursued at each time step, expressed as: (x p , y p , x e , y e ). Among them, (x p , y p ) represents the position coordinates of the pursuing vehicle, and (x e , y e ) represents the position coordinates of the pursued vehicle.

[0059] Define the action space: The action space is used to represent the actions that both the pursuer and the pursued can take at each time step. In the embodiments of the present invention, the action space is defined to include five actions: moving in the four directions of up, down, left, and right, and staying still.

[0060] Design the total reward function: Reward is the core mechanism guiding the learning of the agent. The reward provides feedback to the agent on the quality of its behavior, thereby helping the agent learn how to make optimal decisions in the environment. The total reward function designed in the embodiments of the present invention is expressed by the formula:

[0061] R = R1 + R2 + R3 - n (1);

[0062] Wherein, R represents the total reward, R1 represents the reward generated by successful pursuit, R2 represents the negative reward due to distance, R3 represents the negative reward generated by the pursuit vehicle due to collision or attempting to break out of the boundary, and n represents the number of time steps experienced in the pursuit process;

[0063] The reward generated by successful pursuit is expressed by the formula:

[0064]

[0065] The negative reward due to distance is expressed by the formula:

[0066]

[0067] Wherein, (x p , y p ) represents the position coordinates of the pursuit vehicle, and (x e , y e ) represents the position coordinates of the pursued vehicle;

[0068] The negative reward generated by the pursuit vehicle due to collision or attempting to break out of the boundary is expressed by the formula:

[0069]

[0070] S20. Design an ATT-D3RQN network model suitable for path planning; wherein, the ATT-D3RQN network model is composed of a long short-term memory network, a self-attention module, and a fully connected network.

[0071] The ATT-D3RQN network model designed in the embodiments of the present invention is as Figure 2As shown, it includes a first long short-term memory network L1, a second long short-term memory network L2, a self-attention module ATT, a first fully connected layer network, and a second fully connected layer network connected in sequence; wherein the first long short-term memory network and the second long short-term memory network are both LSMT or Bi-LSMT; the self-attention module is a self-attention module based on a single-head attention mechanism; the first fully connected layer network includes a first fully connected layer FC1 and a second fully connected layer FC2 connected in sequence; the second fully connected layer network includes a third fully connected layer val_hidden and a fourth fully connected layer adv_hidden respectively connected to the second fully connected layer FC2, the third fully connected layer val_hidden is connected to the fifth fully connected layer val, and the fourth fully connected layer adv_hidden is connected to the sixth fully connected layer adv; wherein the third fully connected layer val_hidden and the fifth fully connected layer val constitute a value branch network, and the fourth fully connected layer adv_hidden and the sixth fully connected layer adv constitute an advantage branch network; the value is obtained according to the value branch network, the advantage value is obtained according to the advantage branch network, and the Q value is calculated according to the value and the advantage value. More specifically:

[0072] The deep Q network model is the core of deep reinforcement learning. The current state information is input into the deep Q network model. The Q value of each action can be obtained through the calculation of the network model. The intelligent agent selects the action with the highest Q value to complete a complete decision. The deep Q network model of the embodiment of the present invention is the designed ATT-D3RQN network model, which is a mixture of a long short-term memory network, a multi-head attention module and a fully connected network. When the state information is input into the ATT-D3RQN network model, it will first be calculated in turn by the first long short-term memory network L1 and the second long short-term memory network L2, and then pass through the self-attention module ATT to obtain the prediction result, and then pass through the first fully connected network to further mine the hidden information, and finally pass through the second fully connected network to output the Q value of each action. The formula for calculating the Q value of the ATT-D3RQN network model is expressed as: Q(S t ,a t ;θ), where S t represents the state information of both parties at time t, a t represents the action of chasing the vehicle, and θ represents the network parameters.

[0073] The specific information of each layer in the ATT-D3RQN network model is shown in Table 1. The N in the input size and output size represents the number of batches, that is, the number of data processed by the ATT-D3RQN network model each time. The input of the entire network is an N*64*4 matrix, where N represents the batch size. In this method, the batch size is 256; 64 represents that the length of the entire time series is 64, that is, the sequence consists of the data of the first 64 time steps; 4 represents the data size of each time step, and these four numbers represent the abscissa of the pursuing vehicle, the ordinate of the pursuing vehicle, the abscissa of the pursued vehicle, and the ordinate of the pursued vehicle respectively. The input data will first pass through two long short-term memory networks. After passing through the first long short-term memory network, without passing through an activation function, it is directly sent into the second long short-term memory network. Before entering the first long short-term memory network, the width of the input data is 4. After passing through the second long short-term memory network, the width of the data expands to 128, and the length remains 64.

[0074] Table 1 Specific information of each layer in the ATT-D3RQN network model

[0075]

[0076] Subsequently, it is processed by the self-attention module ATT, and the self-attention module ATT can be a self-attention module based on the single-head self-attention mechanism. In the ATT-D3RQN network model, the calculation method of the self-attention module ATT is as follows:

[0077] (1) For the input sequence x, multiply it by the parameter matrices W Q 、W K and W V respectively to obtain the query vector Q, the key vector K, and the value vector V.

[0078] Q = x·W Q (5);

[0079] K = x·W K (6);

[0080] V = x·W V (7);

[0081] (2) Calculate the attention scores through the following formula:

[0082]

[0083] where d k represents the number of columns of the key vector K. In the embodiments of the present invention, d k can take the value of 128.

[0084] (3) Multiply the scores by the value vector to obtain the output:

[0085] context = score·V (9);

[0086] (4) Calculate the average value of the second dimension of the output sequence context to obtain an output vector out of N * 128.

[0087] The sizes of each matrix in the above calculation process are shown in Table 2:

[0088] Table 2 Sizes of each matrix in the calculation process of the self-attention module

[0089] Matrix Name Matrix Size x (N,64,128) <![CDATA[W Q > (128,128) <![CDATA[W K > (128,128) <![CDATA[W V > (128,128) Q (N,64,128) K (N,64,128) V (N,64,128) score (N,64,64) context (N,64,128) out (N,128)

[0090] It should be noted that in the operation process of the self-attention module ATT, whenever a matrix multiplication operation involving a three-dimensional tensor and a two-dimensional tensor is involved, batch matrix multiplication is used. The operation method of batch matrix multiplication is as follows. Assume there are a three-dimensional tensor X and a two-dimensional tensor Y, then the calculation method of their matrix product Z is:

[0091] Z i = X i ·Y (10);

[0092] where Z i represents the matrix product of the i-th batch, and X i represents the two-dimensional tensor corresponding to the i-th batch in X.

[0093] After being processed by the self-attention module ATT, the data originally sized N*64*128 becomes N*128. Then this data passes through the first fully connected layer FC1 in the first fully connected layer network, and the data size changes from N*128 to N*256. Then this data passes through the second fully connected layer FC2 in the first fully connected layer network, and the data size changes from N*256 to N*512. Then this data is sent to the second fully connected layer network for calculation. The second fully connected layer network includes a parallel advantage branch network and a value branch network. In the value branch network, the data first passes through the third fully connected layer val_hidden, and the data size changes from N*512 to N*256. Then it passes through the fifth fully connected layer val, and the data size changes from N*256 to N*1. This data is called the value V. In the advantage branch network, the data first passes through the fourth fully connected layer adv_hidden, and the data size changes from N*512 to N*256. Then it passes through the sixth fully connected layer adv, and the data size changes from N*256 to N*5. This data is called the advantage value A. Here, the data width 5 represents 5 actions: move up, move down, move left, move right, and stay still. After obtaining the value V and the advantage value A, first calculate the average of the advantage values with a width of 5, and this number is called the average advantage value. Then the Q value of each action can be calculated based on the value V, the advantage value A, and the average advantage value. The calculation formula is:

[0094] Q = V + A - mean(A) (11);

[0095] where mean(A) represents taking the average of the advantage values.

[0096] S30. Based on the modeling results, use the improved DQN algorithm to train the designed ATT-D3RQN network model to obtain the trained ATT-D3RQN network model. Among them, the improved DQN algorithm is based on the original DQN algorithm. For each complete set of state information: divide the complete set of state information according to the preset segmentation length and store it in the experience buffer. Each time during training, extract several segmented state information fragments from the experience buffer, and use a weighted extraction method when extracting.

[0097] Existing path planning technologies for pursuit scenarios are often relatively simple in map setting, usually a two-dimensional continuous free space environment. And the existing multi-agent cooperative pursuit algorithm based on Multi-Agent Deep Deterministic Policy Gradient (MADDPG) is mainly applicable to this kind of continuous free space scenario, such as Figure 3As shown in the figure. There are fewer obstacles in this scenario, the environment has fewer restrictions on the agent, and the movement of the agent is relatively free. This scenario is suitable for the pursuit of drones and unmanned ships. However, in road scenarios, the movement rules of vehicles are quite different from those of drones and unmanned ships. At the same time, vehicles can only drive on existing roads and are subject to many restrictions when driving. Therefore, the existing MADDPG-based agent collaborative hunting algorithm has the following shortcomings: (1) Since MADDPG adopts a deterministic strategy to directly output continuous action values, it lacks a random exploration mechanism for the action space. If external noise (such as OU noise) is used to enhance the exploration, this will increase the complexity of the algorithm, so the training efficiency of the algorithm is low. (2) The policy network of MADDPG selects actions by maximizing the Q value. If the critical network has errors in estimating the Q value, it is easy to cause policy deviation and overestimation of the Q value. This problem is more significant in the continuous action space. (3) The MADDPG network has high coupling. MADDPG needs to optimize the action network and the critical network at the same time. The updates of the two are interdependent, which can easily lead to unstable training or difficult convergence. (4) The convergence of the MADDPG algorithm has high requirements on hyperparameters, and the convergence of the algorithm depends on the careful selection of parameters. Therefore, it is relatively difficult to apply the existing MADDPG-based pursuit algorithm to vehicles.

[0098] Based on the above problems of the MADDPG algorithm, the present invention selects the DQN algorithm, which is more stable in training and consumes less resources, as the basic algorithm of this method. Since the DQN algorithm can only process scenes in discrete action spaces, the embodiment of the present invention performs raster modeling of urban road scenes to adapt to the discrete action space of the DQN algorithm. This approach can reduce the complexity of the algorithm while simulating the real scene as much as possible and improve the reliability of the algorithm.

[0099] The embodiment of the present invention uses the improved DQN algorithm to train the designed ATT-D3RQN network model based on the modeling results. Figure 4 As shown, including:

[0100] S301, initializing the number of iterations, the experience buffer pool, and the network parameters of the ATT-D3RQN network model;

[0101] S302, initializing the city tracking scene;

[0102] S303, according to the initialized city tracking scene, multiple groups of complete status information sets are obtained, and for each group of complete status information sets: the complete status information set is segmented according to a preset segmentation length to obtain a plurality of status information segments;

[0103] S304: for each state information segment s t, the execution steps include: S3041. Select the action a of the pursuit vehicle using the ε-greedy strategy t ; S3042. Execute the action a t , observe the corresponding total reward r t and the next state information segment s t+1 , and put (s t , a t , r t , s t+1 ) into the experience buffer pool and record the index corresponding to (s t , a t , r t , s t+1 ); S3043. Count the number of times the action direction of the pursued vehicle changes and the number of times when the Manhattan distance between the pursued vehicle and the pursuit vehicle is less than or equal to 2 in the state information segments s t and s t+1 . Take the sum of the two numbers plus 1 as the experience sample weight, and store the experience sample weight at the corresponding position in the weight table according to the index; S3044. Randomly extract several groups of state information segments from the experience buffer pool, and the extraction probability of each group of state information segments is determined by the experience sample weight; calculate the target Q value and the actual Q value corresponding to each group of state information segments, calculate the total loss based on all the target Q values and actual Q values, and update the network parameters of the ATT-D3RQN network model using the gradient descent method;

[0104] S305. Determine whether the pursuit vehicle has successfully pursued the pursued vehicle. If so, return to the step of initializing the urban tracking scenario. If not, return to the execution steps for each state information segment s t , until the iteration stop condition is met, and output the trained ATT-D3RQN network model.

[0105] Train the ATT-D3RQN network model. Through the interaction between the agent and the environment, let the agent learn an optimal strategy as soon as possible. For S301~S305, more specifically:

[0106] S301. Let the iteration number ep = 0, initialize the experience buffer pool D, the network parameters θ, θ - of the ATT-D3RQN network model; where ep is the iteration number; during the training process, the size of the experience buffer pool D is set to 50000, and it is set to store cyclically, that is: when the experience buffer pool D is full, the subsequent data continues to be written from the first position of the buffer.

[0107] S302. Initialize the urban tracking scenario.

[0108] S303. Obtain multiple sets of complete state information sets according to the initialized urban tracking scenario. Each set of complete state information sets contains the state information of all time steps. In the embodiments of the present invention, each set of complete state information sets is segmented according to a preset segmentation length to obtain several state information segments. For example, if the preset segmentation length is c, and c can take the value of 64, then the state information segment of the pursuer and the pursued at time t (the t-th time step) after segmentation is denoted as s t =(loc t-c+1 ;…;loc t-1 ;loc t ), that is, this state information segment s t contains the state information at time t and the state information of the previous c - 1 time steps before time t. loc t represents the position vector of the pursuer and the pursued at time t, and loc t =(x p ,y p ,x e ,y e ).

[0109] For each state information segment s t , the steps to be executed include:

[0110] S3041. Use the ε-greedy strategy to select the action a t of the pursuit vehicle: Select a random action with probability ε, and select the action with the highest expected reward with probability 1 - ε. The calculation formula of ε is: where ep max represents the maximum number of iterations. For example, ep max can take the value of 40000.

[0111] S3042. Execute the action a t and observe the corresponding total reward r t and the next state information segment s t+1 . Put (s t ,a t ,r t ,s t+1 ) into the experience buffer D and record the index of (s t ,a t ,r t ,s t+1 ).

[0112] S3043. Count the number of times the action direction of the pursued vehicle changes and the number of times the Manhattan distance between the pursued vehicle and the pursuit vehicle is less than or equal to 2 in the state information segments s t and s t+1 . Add 1 to the sum of the two numbers as the weight of the experience sample, and according to (s t ,at , r t , s t+1 ) The corresponding index places the empirical sample weights into the corresponding positions in the weight table.

[0113] S3044. Randomly extract k groups of state information segments from the experience buffer D. Among them, the extraction probability of each group of state information segments is determined by the empirical sample weights. For example, the extraction probability of each group of state information segments is equal to the empirical sample weight of this group of state information segments divided by the sum of the empirical sample weights of the k groups of state information segments extracted. For example, the value of k is 256. For each group of state information segments, calculate its target Q value and the actual Q value Q(s t , a t ; θ). Calculate the loss Loss(θ) corresponding to this group of state information segments according to the target Q value and the actual Q value. Loss(θ) = (y t - Q(s t , a t ; θ)) 2 . Then add up all the losses and take the average to get the total loss. Then use the gradient descent method to update the network parameter θ, and update the network parameter θ every certain number of training rounds - . For example, update the network parameter θ after 100 training times - .

[0114] In the embodiment of the present invention, the target Q value corresponding to each group of state information segments is calculated. The formula is expressed as:

[0115]

[0116] Among them, s t represents the state information segment at time t, y t represents the target Q value corresponding to the state information segment s t , r t represents the total reward observed under the state information segment s t , which is calculated through formula (1). γ represents the discount factor, which is used to balance the current and future benefits. Here, for example, the value is 0.95. s t+1 represents the state information segment at the next moment observed under the state information segment s t . θ - and θ both represent the network parameters of the ATT-D3RQN network model. Q() represents the function for calculating the Q value through the ATT-D3RQN network model. represents inputting the state information segment s t+1 and all actions in the action space into the ATT-D3RQN network model to obtain the Q values corresponding to all actions, and selecting the action a corresponding to the maximum Q value.

[0117] Note that during the training process, the network parameters θ are optimized according to the loss function in each training. However, when calculating the target Q value in the loss function, the network parameters θ are required. If the same network parameters are used, it will cause dynamic drift of the target Q value itself, resulting in unstable training. Therefore, when calculating the target Q value, the network parameters θ - , the network parameters θ - will be synchronized with the network parameters θ after a certain number of iterations. This delayed update maintains the stability of the training. Note that although the network parameters used to calculate the target Q value are θ - , but in the calculation the network parameters θ are used again. Such a design is to prevent the problem of overestimation of the target Q value caused by the accumulation of forward errors.

[0118] S305. Determine whether the pursuing vehicle successfully pursues the pursued vehicle. If so, return to S302. If not, if ep ≤ ep max , then ep = ep + 1, return to S304 until the iteration stop condition is met, such as reaching the maximum number of iterations, and output the trained ATT-D3RQN network model.

[0119] Since the ATT-D3RQN network model designed in the embodiment of the present invention embeds a long short-term memory network, it needs to be trained using time series data. This requires that continuous time series information must be stored in the experience buffer D. When the existing DQN algorithm for embedding time series networks stores experience data, it often arranges the complete state information of a whole game in all time orders and then regards it as a whole and puts it into the experience buffer D. Each complete data of a game is regarded as a sample. During training, the agent will also randomly select k complete data for training. However, such a design violates the setting that the training data are independent and identically distributed. Because each complete data of a game itself is not independent and identically distributed. To solve this problem, the embodiment of the present invention divides a complete data according to a preset segmentation length and then stores it in the experience buffer D. In this way, when the agent extracts data from the experience buffer D, the obtained are scattered time series sequences, which meets the requirement that the training data are independent and identically distributed.

[0120] Secondly, when the agent extracts data for training, it does not use traditional equal-probability extraction, but weighted extraction. Each state information segment in the experience buffer D has a corresponding experience sample weight, which is equal to 1 by default. If a key game event occurs in this state information segment (the Manhattan distance between the pursuer and the pursued is less than or equal to 2, or the strategy of the pursued vehicle, i.e., the action method, changes), then each time a key game event occurs, the experience sample weight will be incremented by 1. When extracting data, the probability of each piece of data being selected is equal to its experience sample weight divided by the sum of the experience sample weights of all the data extracted. This experience replay with priorities allows the agent to learn more key information, thereby accelerating the convergence of the strategy.

[0121] S40. Input the state information of the pursuer and the pursued to be path-planned into the trained ATT-D3RQN network model. In the trained ATT-D3RQN network model, calculate the Q value of each action in the action space, and select the action corresponding to the maximum Q value as the execution action of the pursuing vehicle, so as to obtain the new state information of the pursuer and the pursued through the selected execution action of the pursuing vehicle. Input the new state information of the pursuer and the pursued into the trained ATT-D3RQN network model until the pursuing vehicle successfully catches the pursued vehicle.

[0122] After obtaining the trained ATT-D3RQN network model through S30 in the embodiment of the present invention, the state information of the pursuer and the pursued to be path-planned can be obtained according to the current urban tracking scenario, and this state information is input into the trained ATT-D3RQN network model to calculate the Q values of all actions in the action space. Here, it is defined that the action space includes 5 actions, so the Q values corresponding to the 5 actions are calculated, and the action corresponding to the maximum Q value among these Q values is selected as the execution action of the pursuing vehicle. After the pursuing vehicle executes this action, it is judged whether the pursuing vehicle has successfully caught the pursued vehicle. If so, the process ends and the pursuit path planning task is completed. Otherwise, the state information of the pursuer and the pursued is obtained again, and the new state information is input into the trained ATT-D3RQN network model, and the process of calculating the Q value and selecting the execution action is repeated until the pursuing vehicle successfully catches the pursued vehicle.

[0123] For practical use, the method proposed in the present invention also includes: dividing the urban tracking scene into regions, executing S301 to S305 for each region, obtaining the ATT-D3RQN network model trained in the corresponding region, and storing the network parameters corresponding to the ATT-D3RQN network model trained in the region in the network parameter table. If the two parties to be pursued for path planning are located in an urban tracking scene in a certain area after the division, the network parameters corresponding to the area are read from the network parameter table, and the ATT-D3RQN network model is loaded using the read network parameters, so that the Q value of each action in the action space is calculated in the loaded ATT-D3RQN network model, and the action corresponding to the maximum Q value is selected as the execution action of the pursuit vehicle, so as to obtain new state information of the two parties through the selected execution action of the pursuit vehicle, and the new state information of the two parties is input into the loaded ATT-D3RQN network model until the pursuit vehicle successfully pursues the pursued vehicle. More specifically:

[0124] In actual use, the method proposed in the present invention can be deployed in the vehicle command center. During deployment, the urban tracking scene can be divided into regions in advance, and then for each region, simulated pre-training is first performed according to S10~S30, and the network parameters of each trained ATT-D3RQN network model are saved in the network parameter table to realize discrete reinforcement learning. When a pursuit task is generated, the vehicle being pursued initiates a corresponding request. After receiving the request, the vehicle command center first retrieves the network parameters of the corresponding area to load the ATT-D3RQN network model, and then obtains the position information sequence of the vehicles of both parties in the pursuit in real time according to the drone and road monitoring means, and then inputs the position information sequence of both parties in the pursuit into the ATT-D3RQN network model, calculates the Q value of each action in the action space, and selects the action corresponding to the maximum Q value as the execution action of the pursuit vehicle, and then transmits the real-time planning data to the vehicle computer of the pursuit vehicle through the Internet of Vehicles, and finally realizes the vehicle pursuit path planning.

[0125] In order to verify the effectiveness of the vehicle pursuit path planning method in the urban scene provided by the embodiment of the present invention, the following verification is performed:

[0126] Figure 5 The average reward change curves of the traditional DQN algorithm, DRQN ​​algorithm and the method proposed in this invention during the training process are shown. The training scenario is a square map with a side length of 150 meters. Figure 5It can be seen that: the traditional ATT-D3RQN network model (Deep Q Network, DQN) algorithm without any optimization performs the worst and has the lowest reward; while the Deep Recurrent Q Network (DRQN) algorithm with the introduction of the long short-term memory network performs between the DQN algorithm and the method proposed in the present invention (denoted as the patent algorithm in the figure). However, from Figure 5 It can also be seen that: the DRQN algorithm has poor stability, and the reward value is prone to large fluctuations during the training process. And from Figure 5 It can be seen that: the method proposed in the present invention has the best effect, whether it is the convergence speed, algorithm stability or reward value, it is better than the other two algorithms.

[0127] Figure 6 shows a comparison chart of the capture success rates of the traditional Breadth-First Search (BFS) algorithm, the traditional DQN algorithm, and the method proposed in the present invention under different map sizes, Figure 7 shows the comparison of the capture steps of the traditional BFS algorithm, the traditional DQN algorithm, and the method proposed in the present invention under different map sizes. In the test, maps of different sizes were selected for testing. The maps are all square, and the side lengths of the maps are 150 meters, 300 meters, 450 meters, and 600 meters respectively. In the test, the embodiments of the present invention selected two algorithms as the comparison algorithms for the method proposed in the present invention: the BFS algorithm means that in each time step, the shortest path to the pursued vehicle is found by applying breadth-first search according to the positions of the two vehicles; while the DQN algorithm is a traditional reinforcement learning algorithm without any improvement. From Figure 6 and Figure 7 It can be seen that: on various map sizes, whether it is the success rate of the pursuit or the number of steps consumed to complete the pursuit, the method proposed in the present invention is significantly better than the other two traditional algorithms.

[0128] In summary, the vehicle pursuit path planning method in the urban scenario proposed in the embodiments of the present invention simplifies the urban pursuit scenario through rasterization processing, discretizes the originally continuous urban pursuit scenario and the actions of the vehicle, which can reduce the complexity of the deep reinforcement learning model, thereby reducing the requirements of the algorithm for computing resources and making the algorithm easier to deploy; designs an ATT-D3RQN network model suitable for path planning. This ATT-D3RQN network model combines a long short-term memory network and a self-attention module, integrates the long short-term memory network into the network model. Through the long short-term memory network, the position of the target vehicle can be predicted, and advance planning can be carried out based on the predicted position, thereby improving the success rate of pursuit. On the basis of the long short-term memory network, a self-attention mechanism is introduced, which can effectively improve the learning ability of the recurrent neural network for long sequences, thereby improving the performance of the path planning algorithm; adopts a path planning algorithm based on the deep reinforcement learning DQN algorithm. This algorithm can adjust the strategy in real time according to the dynamic environment, improving the operation efficiency of the algorithm. And the present invention improves the DQN algorithm, slices the input data of the algorithm, so that when the agent extracts data from the experience buffer pool, it can obtain scattered time series, which better meets the requirement of the DQN algorithm for data to be independently and identically distributed, and the extraction adopts a weighted extraction method. This experience replay with priority can let the agent learn more key information, thereby accelerating the convergence of the strategy. Through experimental verification, the method proposed in the present invention is superior to the existing algorithms in terms of convergence speed, training stability, pursuit success rate, and pursuit steps. Generally speaking, the present invention proposes a new type of vehicle pursuit path planning method. This method enhances the dynamic adaptability through reinforcement learning, enhances the adaptability to complex environments through the designed ATT-D3RQN network model, and improves the execution efficiency by discretizing the road environment, thereby breaking through the constraints of the existing methods in terms of real-time performance, adaptability to complex environments, and resource efficiency, and providing a systematic solution for urban pursuit scenarios that takes into account path optimality, dynamic response speed, and computational economy.

[0129] In the description of the present invention, it should be understood that the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality" means two or more unless otherwise specifically defined.

[0130] Although the present invention has been described in connection with various embodiments, those skilled in the art will understand and realize other variations of the disclosed embodiments by referring to the specification and its accompanying drawings during the implementation of the claimed invention. In the specification, the term "comprising" does not exclude other components or steps, and the singular forms "a" or "an" do not exclude a plurality. Certain measures are described in different embodiments, but this does not mean that these measures cannot be combined to achieve good results.

[0131] The above content is a further detailed description of the present invention in connection with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited only to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as falling within the protection scope of the present invention.

Claims

1. A vehicle pursuit path planning method in an urban scene, characterized in that, The method includes: Modeling the urban pursuit scenario; the modeling includes rasterizing the urban pursuit scenario, defining state information and action space, and designing a total reward function; Designing an ATT-D3RQN network model suitable for path planning; wherein, the ATT-D3RQN network model is composed of a long short-term memory network, a self-attention module, and a fully connected network; Based on the modeling results, training the designed ATT-D3RQN network model using an improved DQN algorithm to obtain a trained ATT-D3RQN network model; wherein, the improved DQN algorithm is based on the original DQN algorithm. For each complete set of state information: the complete set of state information is sliced according to a preset slicing length and stored in an experience buffer. Each time during training, several sliced state information segments are sampled from the experience buffer, and the sampling is performed in a weighted sampling manner; Inputting the state information of the pursuer and the pursued to be path-planned into the trained ATT-D3RQN network model. In the trained ATT-D3RQN network model, calculate the Q value of each action in the action space, and select the action corresponding to the maximum Q value as the execution action of the pursuit vehicle, so as to obtain new state information of the pursuer and the pursued through the selected execution action of the pursuit vehicle. Input the new state information of the pursuer and the pursued into the trained ATT-D3RQN network model until the pursuit vehicle successfully catches the pursued vehicle.

2. The vehicle pursuit path planning method in the urban scenario according to claim 1, wherein, The defined state information includes the position coordinates of the pursuit vehicle and the pursued vehicle.

3. The vehicle pursuit path planning method in the urban scene according to claim 1, wherein, The defined action space includes five actions: moving in the four directions of up, down, left, and right and staying still.

4. The vehicle pursuit path planning method in an urban scenario according to claim 1, wherein The designed total reward function is expressed by the formula: R = R1 + R2 + R3 - n; Wherein, R represents the total reward, R1 represents the reward generated by successful pursuit, R2 represents the negative reward due to distance, R3 represents the negative reward generated by the pursuit vehicle due to collision or attempting to break out of the boundary, and n represents the number of time steps experienced during the pursuit process; The reward generated by successful pursuit is expressed by the formula: The negative reward due to distance is expressed by the formula: Wherein, (xp, yp) represents the position coordinates of the pursuit vehicle, and (xe, ye) represents the position coordinates of the pursued vehicle; The negative reward generated by the pursuit vehicle due to collision or attempting to break out of the boundary is expressed by the formula:

5. The vehicle pursuit path planning method in an urban scenario according to claim 1, characterized in that The designed ATT-D3RQN network model includes a first long short-term memory network, a second long short-term memory network, a self-attention module, a first fully connected layer network, and a second fully connected layer network connected in sequence; wherein, The first long short-term memory network and the second long short-term memory network are both LSMT or Bi-LSMT; The self-attention module is a self-attention module based on a single-head attention mechanism; The first fully connected layer network includes a first fully connected layer and a second fully connected layer connected in sequence; The second fully-connected layer network includes a third fully-connected layer and a fourth fully-connected layer respectively connected to the first fully-connected layer. The third fully-connected layer is connected to a fifth fully-connected layer, and the fourth fully-connected layer is connected to a sixth fully-connected layer. Among them, the third fully-connected layer and the fifth fully-connected layer form a value branch network, and the fourth fully-connected layer and the sixth fully-connected layer form an advantage branch network. A value is obtained according to the value branch network, an advantage value is obtained according to the advantage branch network, and a Q value is calculated according to the value and the advantage value.

6. The vehicle pursuit path planning method in the urban scenario according to claim 1, wherein The process of training the designed ATT-D3RQN network model using the improved DQN algorithm based on the modeling result includes: Initializing the number of iterations, the experience buffer pool, and the network parameters of the ATT-D3RQN network model; Initializing the urban tracking scenario; According to the initialized urban tracking scenario, obtaining multiple groups of complete state information sets. For each group of complete state information sets: splitting the complete state information set into several state information segments according to a preset splitting length; For each fragment s of state information t , the execution steps include: S3041. Use the ε-greedy strategy to select the action a of the pursuing vehicle t ; S3042. Execute the action a t , observe the corresponding total reward r t and the next fragment s of state information t+1 , and put (s t , a t , r t , s t+1 ) into the experience buffer pool and record the index corresponding to (s t , a t , r t , s t+1 ); S3043. Count the number of times the action direction of the pursued vehicle changes and the number of times when the Manhattan distance between the pursued vehicle and the pursuing vehicle is less than or equal to 2 in the fragments s t and s t+1 . Take the sum of the two numbers plus 1 as the weight of the experience sample, and store the weight of the experience sample at the corresponding position in the weight table according to the index; S3044. Randomly extract several groups of fragments of state information from the experience buffer pool, and the extraction probability of each group of fragments of state information is determined by the weight of the experience sample; calculate the target Q value and the actual Q value corresponding to each group of fragments of state information, calculate the total loss according to all the target Q values and actual Q values, and update the network parameters of the ATT-D3RQN network model using the gradient descent method; Determine whether the pursuit vehicle has successfully pursued the pursued vehicle. If so, return to the step of initializing the urban tracking scenario. If not, return to the execution step for each state information segment s t until the iteration stop condition is met, and output the trained ATT-D3RQN network model.

7. The vehicle pursuit path planning method in the urban scenario according to claim 6, wherein, The extraction probability of each group of state information segments is equal to the empirical sample weight of the state information segment divided by the sum of the empirical sample weights of all the extracted groups of state information segments.

8. The vehicle pursuit path planning method in the urban scene according to claim 6, wherein Calculating the target Q value corresponding to each group of state information segments, which is expressed by the formula: Among them, s t represents a fragment of state information at time t, y t represents the target Q value corresponding to the state information fragment s t r represents the total reward observed under the state information fragment s t γ represents the discount factor, s t represents the state information fragment at the next moment observed under the state information fragment s t+1 θ t both represent the network parameters of the ATT-D3RQN network model, and Q() represents the function for calculating the Q value through the ATT-D3RQN network model. - represents inputting the state information fragment s t+1 and all actions in the action space into the ATT-D3RQN network model to obtain the Q values corresponding to all actions, and selecting the action a corresponding to the maximum Q value.

9. The vehicle pursuit path planning method in an urban scenario according to claim 6, wherein The method further includes: Dividing the urban tracking scenario into regions, and executing S301 - S305 for each region to obtain a trained ATT-D3RQN network model for the corresponding region, and storing the network parameters corresponding to the trained ATT-D3RQN network model for the region in the network parameter table.

10. The vehicle pursuit path planning method in an urban scenario according to claim 9, wherein If the pursuer and the pursued to be path-planned are in the urban tracking scenario of a certain region after division, read the network parameters corresponding to the region from the network parameter table, load the ATT-D3RQN network model using the read network parameters, calculate the Q value of each action in the action space in the loaded ATT-D3RQN network model, and select the action corresponding to the maximum Q value as the execution action of the pursuit vehicle, so as to obtain the new state information of the pursuer and the pursued through the selected execution action of the pursuit vehicle, and input the new state information of the pursuer and the pursued into the loaded ATT-D3RQN network model until the pursuit vehicle successfully captures the pursued vehicle.