Intelligent agent path planning method in dynamic perception limited environment
By combining the A* path search algorithm and an improved deep Q network in a dynamic perception-constrained environment, extracting local and global information and designing a multi-reward mechanism, the problems of low efficiency and inaccurate navigation in the agent's path planning are solved, and an efficient, safe and smooth path is generated.
Patent Information
- Application Number
- CN202510822678.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-09-23
AI Technical Summary
Existing technologies have problems with low efficiency and inaccurate navigation in intelligent agent path planning in dynamic perception-constrained environments, especially in the early stages of training, where learning efficiency is low, action selection strategies are single, reward functions are irrationally designed, and classic path guidance strategies lack diversity.
A neural network based on an improved dual deep Q network initialized by the A* path search algorithm is used to extract local spatial features and global auxiliary information through the neural network. Combined with the ε-greedy strategy and a custom safe action selection mechanism, a structured reward function is designed, including goal achievement rewards, collision penalties, proximity rewards and turning penalties, to generate high-quality path planning.
It significantly improves the navigation efficiency and accuracy of intelligent agents in perception-constrained environments, generates shorter, smoother and safer paths, shortens training convergence time, and improves the rationality and stability of path planning.
Smart Images

Figure CN120686824A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of navigation technology, and in particular to a method for intelligent agent path planning in a dynamic perception-constrained environment. Background Art
[0002] In applications such as robot navigation, autonomous driving systems, and drone path control, path planning is a core task in enabling autonomous decision-making by intelligent agents. Its fundamental goal is to plan the shortest, feasible, and safe path from a starting point to a destination in an environment with obstacles. However, existing technologies suffer from low efficiency and inaccurate navigation. Summary of the Invention
[0003] In view of this, an embodiment of the present application provides an intelligent agent path planning method in a dynamic perception-constrained environment to improve the efficiency and accuracy of the intelligent agent's navigation in the perception-constrained environment.
[0004] An aspect of an embodiment of the present application provides a method for intelligent agent path planning in a dynamic perception-constrained environment, the method comprising the following steps:
[0005] Obtain the multi-channel matrix of the local environment and the scalar auxiliary information of the global environment;
[0006] Extracting local spatial features of the multi-channel matrix using a main network; wherein the main network is a neural network initialized based on an A* path search algorithm and combined with an improved dual deep Q network;
[0007] Converting the scalar auxiliary information into auxiliary vector features using the main network;
[0008] Using the main network to splice the local spatial features and the auxiliary vector features to obtain a spliced vector;
[0009] Processing the concatenated vector using two fully connected layers in the main network to obtain a Q-value vector with the same dimension as the action space;
[0010] The action of the agent is controlled according to the Q-value vector.
[0011] In some embodiments, the step of obtaining the multi-channel matrix of the local environment and the scalar auxiliary information of the global environment includes the following steps:
[0012] Obtaining a local obstacle position matrix, a current target point matrix of the agent, an optional path prediction matrix, and a direction information matrix as the multi-channel matrix;
[0013] The Euclidean distance between the agent and the target point, the current position of the agent and the direction of the target point are obtained as the scalar auxiliary information.
[0014] In some embodiments, extracting the local spatial features of the multi-channel matrix using the main network comprises the following steps:
[0015] The local spatial features of the multi-channel matrix are extracted using two convolutional layers activated by ReLU in the main network, and then the local spatial features are converted into vector form through a Flatten layer.
[0016] In some embodiments, the converting the scalar auxiliary information into auxiliary vector features using the main network comprises the following steps:
[0017] The scalar auxiliary information is converted into auxiliary vector features using a fully connected layer in the main network.
[0018] In some embodiments, processing the concatenated vector using two fully connected layers in the main network to obtain a Q-value vector with the same dimension as the action space comprises the following steps:
[0019] The concatenated vector is processed using the two fully connected layers in the main network to obtain the four Q-value vectors with the same dimension as the action space, which respectively correspond to the action sets of the agent in the up, down, left, and right directions at the current position.
[0020] In some embodiments, the method further comprises the step of training the main network, wherein the step of training the main network comprises the following steps:
[0021] Initializing the main network and the target network; wherein the target network is a neural network obtained by initializing based on the A* path search algorithm and combining it with an improved dual deep Q network;
[0022] Before each training round, a reference path is generated using the A* path search algorithm based on the current position of the agent and the position of the target point;
[0023] utilizing the agent to perform actions according to the reference path, the ε-greedy strategy, and a predefined safe action selection mechanism;
[0024] Determine the environmental feedback reward after the agent performs an action, and record the experience of performing the action into the experience replay pool;
[0025] Randomly sampling small batches of samples from the experience replay pool to update the parameters of the main network;
[0026] Synchronizing the parameters of the master network to the target network at fixed step intervals;
[0027] Calculating a loss function and then calculating a loss gradient using the Q value vector output by the main network and the Q value vector output by the target network;
[0028] Update the parameters of the main network according to the loss gradient.
[0029] In some embodiments, determining the environmental feedback reward after the agent performs an action includes the following steps:
[0030] The reaching reward, collision penalty, approach reward and turning penalty after the agent performs the action are determined as the environment feedback reward.
[0031] Another aspect of the present application further provides an intelligent agent path planning device in a dynamic perception-constrained environment, the device comprising:
[0032] The environment perception unit is used to obtain the multi-channel matrix of the local environment and the scalar auxiliary information of the global environment;
[0033] A feature extraction unit, configured to extract local spatial features of the multi-channel matrix using a main network; wherein the main network is a neural network initialized based on an A* path search algorithm and combined with an improved dual deep Q network;
[0034] an information conversion unit, configured to convert the scalar auxiliary information into auxiliary vector features using the main network;
[0035] A feature splicing unit, configured to splice the local spatial features and the auxiliary vector features using the main network to obtain a spliced vector;
[0036] an action prediction unit, configured to process the concatenated vector using the two fully connected layers in the main network to obtain a Q-value vector having the same dimension as the action space;
[0037] An action execution unit is used to control the action of the agent according to the Q value vector.
[0038] Another aspect of the embodiments of the present application further provides an electronic device, including a processor and a memory;
[0039] The memory is used to store programs;
[0040] The processor executes the program to implement any of the above methods.
[0041] Another aspect of the embodiments of the present application further provides a computer-readable storage medium, wherein the storage medium stores a program, and the program is executed by a processor to implement any of the above methods.
[0042] This application has at least the following beneficial effects:
[0043] This application can obtain the multi-channel matrix of the local environment and scalar auxiliary information of the global environment; use the main network to extract the local spatial features of the multi-channel matrix; wherein the main network is a neural network initialized based on the A* path search algorithm and combined with an improved dual-depth Q network; use the main network to convert the scalar auxiliary information into auxiliary vector features; use the main network to splice the local spatial features and the auxiliary vector features to obtain a spliced vector; use two fully connected layers in the main network to process the spliced vector to obtain a Q-value vector with the same dimension as the action space; and control the action of the intelligent agent based on the Q-value vector. This application initializes the main network through the A* path search algorithm, allowing the main network to refer to existing experience for path planning, effectively alleviating the problem of insufficient experience and reliance on random exploration in the early stages of reinforcement learning, and can improve the efficiency and accuracy of the intelligent agent's navigation in a perceptually constrained environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0045] Figure 1 A flow chart of a method for intelligent agent path planning in a dynamic perception-constrained environment provided by an embodiment of the present application;
[0046] Figure 2 This is an example flow chart of a method for intelligent agent path planning in a dynamic perception-constrained environment provided by an embodiment of the present application;
[0047] Figure 3 A training flow chart of a neural network provided in an embodiment of the present application;
[0048] Figure 4 A flowchart of another neural network training method provided in an embodiment of the present application;
[0049] Figure 5 This is a structural block diagram of the intelligent agent path planning device in a dynamic perception restricted environment provided in an embodiment of the present application. DETAILED DESCRIPTION
[0050] In order to make the purpose, technical solutions and advantages of this application more clearly understood, the present application is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0051] Before describing the embodiments of the present application in detail, some of the related technologies involved in the embodiments of the present application are first described as follows:
[0052] A*: A heuristic path search algorithm, also known as the A-star algorithm, combines the shortest path search of the Dijkstra algorithm with heuristic estimation to find the path with the minimum cost in a known map.
[0053] DDQN: Double Deep Q-Network, a reinforcement learning method used to alleviate the problem of Q-value overestimation in Q-learning. This method uses two neural networks for action selection and action evaluation respectively.
[0054] A*-IDDQN: The name of the algorithm proposed in this paper, which is based on the A* path initialization and improved DDQN path planning method, where "I" stands for Improved.
[0055] Collision: During the path planning process, the agent comes into contact with an obstacle, which is a sign that the path is not feasible.
[0056] Path turns: refers to the number of changes in direction in the path. Too many turns may mean that the path is not smooth enough.
[0057] In intelligent agent navigation applications such as robot navigation, autonomous driving systems, and drone path control, path planning is one of the core tasks that enable autonomous decision-making. Its fundamental goal is to plan the shortest, feasible, and safest path for the agent from its starting point to its destination in an environment with obstacles.
[0058] Related path planning technologies include:
[0059] 1) Classic path planning algorithm:
[0060] Dijkstra's algorithm: Used for searching the shortest path in a graph. It is suitable for static, known environments. However, it is computationally inefficient in high-dimensional, complex spaces.
[0061] A* algorithm: Based on Dijkstra, it introduces a heuristic function to improve efficiency and is currently one of the most widely used path search methods. However, it has problems such as high memory consumption, excessive path reentry, and strong dependence on the heuristic function.
[0062] RRT algorithm (Rapidly-exploring Random Tree): It is suitable for high-dimensional unknown environments and has strong exploration capabilities. However, the paths it generates are often long and uneven, and they are prone to being too close to obstacles.
[0063] 2) Reinforcement learning path planning algorithm:
[0064] DQN (Deep Q-Network): It estimates Q values through neural networks, is suitable for high-dimensional state spaces, and is widely used in path planning.
[0065] Improved DQN (such as Double DQN, Dueling DQN, etc.): used to alleviate Q-value overestimation and improve the accuracy of policy evaluation. Some methods also combine mechanisms such as prioritized experience replay to accelerate learning efficiency.
[0066] 3) Attempts to integrate classical algorithms with reinforcement learning:
[0067] For example, the A* path is used as a heuristic guide to overcome the problem of DQN's lack of initial experience. Although such methods have made some improvements, they lack the depth of integration and lack comprehensive control over the path quality (such as smoothness and safety).
[0068] Disadvantages of existing technology:
[0069] Although the above methods have their own advantages, there are still several key issues in practical applications, including:
[0070] 1) Low learning efficiency in the early stages of training:
[0071] Reinforcement learning agents rely heavily on random exploration in the early stages of training, resulting in poor experience quality and long training time. Especially in complex environments, they are prone to invalid attempts and difficult to converge quickly.
[0072] 2) Single action selection strategy:
[0073] Most reinforcement learning algorithms use a fixed ϵ-greedy strategy (not selecting the current optimal solution with a certain probability). Their balance between exploration and exploitation relies on static parameter decay and lacks flexible adaptation to different states and environmental changes, which may lead to unreasonable or repetitive path selection.
[0074] 3) The reward function is not designed properly or is too simple:
[0075] Existing methods generally only use target point rewards and collision penalties, ignoring the overall quality factors of the path, such as path smoothness, number of turns, safety distance, etc., which makes it easy for the intelligent agent to fall into local optimality or generate low-quality paths.
[0076] 4) Classic path guidance strategies lack diversity:
[0077] Some methods use classic algorithms to generate paths (such as A* or RRT) to guide training, but they often only provide a single path or direct intervention strategy, lacking path diversity and flexibility, which is not conducive to the generalization and stable convergence of the strategy.
[0078] Reference Figure 1 The embodiment of the present application provides a method for intelligent agent path planning in a dynamic perception-constrained environment, specifically comprising the following steps S100 to S150:
[0079] S100: Acquire the multi-channel matrix of the local environment and the scalar auxiliary information of the global environment;
[0080] S110: extracting local spatial features of the multi-channel matrix using a main network; wherein the main network is a neural network initialized based on an A* path search algorithm and combined with an improved dual deep Q network;
[0081] S120: using the main network to convert the scalar auxiliary information into auxiliary vector features;
[0082] S130: Using the main network to splice the local spatial features and the auxiliary vector features to obtain a spliced vector;
[0083] S140: Processing the concatenated vector using two fully connected layers in the main network to obtain a Q-value vector with the same dimension as the action space;
[0084] S150: Controlling the action of the agent according to the Q-value vector.
[0085] Optionally, the acquiring of the multi-channel matrix of the local environment and the scalar auxiliary information of the global environment comprises the following steps:
[0086] Obtaining a local obstacle position matrix, a current target point matrix of the agent, an optional path prediction matrix, and a direction information matrix as the multi-channel matrix;
[0087] The Euclidean distance between the agent and the target point, the current position of the agent and the direction of the target point are obtained as the scalar auxiliary information.
[0088] Optionally, the extracting the local spatial features of the multi-channel matrix using the main network comprises the following steps:
[0089] The local spatial features of the multi-channel matrix are extracted using two convolutional layers activated by ReLU in the main network, and then the local spatial features are converted into vector form through a Flatten layer.
[0090] Optionally, the converting the scalar auxiliary information into auxiliary vector features using the main network comprises the following steps:
[0091] The scalar auxiliary information is converted into auxiliary vector features using a fully connected layer in the main network.
[0092] Optionally, the step of processing the concatenated vector using two fully connected layers in the main network to obtain a Q-value vector having the same dimension as the action space comprises the following steps:
[0093] The concatenated vector is processed using the two fully connected layers in the main network to obtain the four Q-value vectors with the same dimension as the action space, which respectively correspond to the action sets of the agent in the up, down, left, and right directions at the current position.
[0094] Optionally, the method further comprises a step of training the main network, wherein the step of training the main network comprises the following steps:
[0095] Initializing the main network and the target network; wherein the target network is a neural network obtained by initializing based on the A* path search algorithm and combining it with an improved dual deep Q network;
[0096] Before each training round, a reference path is generated using the A* path search algorithm based on the current position of the agent and the position of the target point;
[0097] utilizing the agent to perform actions according to the reference path, the ε-greedy strategy, and a predefined safe action selection mechanism;
[0098] Determine the environmental feedback reward after the agent performs an action, and record the experience of performing the action into the experience replay pool;
[0099] Randomly sampling small batches of samples from the experience replay pool to update the parameters of the main network;
[0100] Synchronizing the parameters of the master network to the target network at fixed step intervals;
[0101] Calculating a loss function and then calculating a loss gradient using the Q value vector output by the main network and the Q value vector output by the target network;
[0102] Update the parameters of the main network according to the loss gradient.
[0103] Optionally, determining the environmental feedback reward after the agent performs an action comprises the following steps:
[0104] The reaching reward, collision penalty, approach reward and turning penalty after the agent performs the action are determined as the environment feedback reward.
[0105] Next, the solution of the embodiment of the present application will be introduced and explained in detail with reference to specific application examples.
[0106] Specifically, this embodiment may include the following technical solutions:
[0107] 1) State modeling and information input structure.
[0108] This embodiment designs a multi-channel state input structure by structurally modeling environmental information, which consists of the following two types of input:
[0109] i. Matrix Input: Spatial perception data used to describe the agent's local environment, including:
[0110] Local obstacle position matrix;
[0111] Current agent target point matrix;
[0112] Optional auxiliary channels such as path prediction and direction information.
[0113] Among them, each matrix channel has the same size, is constructed in a fixed window sliding manner, and is input into the convolutional neural network for feature extraction.
[0114] ii. Scalar Input: Auxiliary information used to describe the global state, including:
[0115] Euclidean distance to the target point;
[0116] Unit vector in the direction of the target.
[0117] 2) Neural network structure and Q-value output.
[0118] This embodiment uses a deep Q network with a fusion structure to process state information and output the Q value of each action:
[0119] i. Convolution processing part:
[0120] The matrix input first passes through two convolutional layers (Conv1 and Conv2) activated by ReLU to extract local spatial features;
[0121] The convolution output is converted into vector form through the Flatten layer.
[0122] ii. Scalar preprocessing part:
[0123] The scalar input is processed through a layer of full connection (or normalization) and converted into a high-dimensional representation;
[0124] Splice with the convolution extraction result.
[0125] iii. Decision-making part:
[0126] The concatenated feature vector enters two fully connected layers (FC1 and FC2), and finally outputs a Q value vector with the same dimension as the action space.
[0127] 3) A* path experience initialization mechanism:
[0128] To address the problems of low initial training efficiency and blind exploration in reinforcement learning, this embodiment introduces the A* path experience initialization mechanism:
[0129] In the early stage of training or the initialization stage of each round, use the A* algorithm to generate the shortest path for the current environment;
[0130] Encapsulate all state-action-reward information in the path into experience entries and inject them directly into the experience replay pool;
[0131] As a guidance sample in the early stage of reinforcement learning, it improves sample quality and shortens training convergence time.
[0132] 4) Improved action selection strategy.
[0133] Traditional ε-greedy strategies suffer from low exploration efficiency and frequent repeated paths in complex environments. This embodiment designs an action selection strategy that combines customized path safety judgment:
[0134] Based on ε-greedy, an obstacle prediction and target direction judgment mechanism is introduced;
[0135] Among multiple actions with similar maximum Q values, the action that is closer to the target and less likely to cause collision is preferred;
[0136] If the optimal action leads to collision risk, the suboptimal action is rolled back or resampled to improve path quality and safety.
[0137] This strategy takes into account both exploration and exploitation, avoiding serious path duplication or collision during training or deployment.
[0138] 5) Reward function and path quality control mechanism.
[0139] To ensure that the path is not only feasible but also efficient and explainable, this embodiment constructs a reward function consisting of multiple components:
[0140] Goal achievement reward: Give high positive incentives when the agent reaches the goal;
[0141] Collision penalty: a large negative incentive is applied when colliding with an obstacle;
[0142] Approach reward: The agent receives a small reward when it approaches the target point;
[0143] Path turning penalty: Penalties are imposed when the path makes sharp turns, encouraging path smoothing.
[0144] This mechanism encourages the agent to make a trade-off between the shortest path, low collision, and high efficiency, thereby achieving the compatibility of task orientation and path quality.
[0145] More specifically, the present application can be implemented through the following embodiments.
[0146] Example 1: Implementation of the A*-IDDQN system in a single-agent path planning task.
[0147] (1) Application scenarios
[0148] Consider a 20×20 2D grid map environment with several static obstacles and a single target point. The mobile agent has an independent target point and its perception range is limited to a 5×5 local window. It cannot observe the entire environment and relies solely on local information for path planning.
[0149] (2) State representation structure design.
[0150] In this embodiment, the status is composed of the following information:
[0151] Local environment matrix (multi-channel): contains the obstacle distribution, the relative position of the target point, etc., with a size of 5*5*N, where N represents the number of channels.
[0152] Auxiliary scalar information:
[0153] Euclidean distance to the target point;
[0154] The agent's current location;
[0155] The direction of the target point.
[0156] (3) Network structure and action output.
[0157] Using an improved deep Q network structure:
[0158] The matrix input is processed through two layers of convolution to extract spatial features;
[0159] After independent preprocessing, the scalar input is concatenated with the convolutional features;
[0160] The splicing result is processed by two layers of fully connected neurons and outputs 4 Q values, corresponding to the action set {up, down, left, right}.
[0161] (4) Experience initialization and training process.
[0162] At the beginning of training, the A* algorithm is used to generate a shortest path for the current environment. The state-action-reward sequence in this path is used as high-quality experience and injected into the experience replay pool to avoid the inefficiency caused by random exploration in the early stages of reinforcement learning.
[0163] (5) Description of the training process.
[0164] Reinforcement learning training includes the following process:
[0165] 1. Initialize the Q network and the target network;
[0166] 2. Before each round, A* is used to generate a path as a reference based on the current position and the target position;
[0167] 3. The agent executes actions based on the ε-greedy strategy combined with a custom safe action selection mechanism;
[0168] 4. Environmental feedback rewards, record experience to the experience replay pool;
[0169] 5. Randomly sample small batches of samples from the experience pool to update the Q network;
[0170] 6. Synchronize the target network parameters every fixed number of steps.
[0171] (6) Reward function design.
[0172] This example uses a structured multi-reward design to simultaneously optimize the path's accessibility, safety, efficiency, and smoothness. The total reward function consists of the following four parts:
[0173] 1. Arrival Rewards:
[0174] When the agent reaches the target point, it is given a one-time significant positive reward to reinforce the behavior of achieving the target.
[0175] 2. Collision Penalty:
[0176] When the agent collides with an obstacle, a strong negative reward is given to inhibit dangerous paths.
[0177] 3. Proximity Rewards:
[0178] Each time the agent moves, the reward is calculated based on the change in its distance from the target point. A positive reward is given when it is close to the target, and a negative reward is given when it is far away, guiding it to continue approaching the target.
[0179] 4. Turning Penalty:
[0180] If the current action direction is different from the previous moment (i.e., a turn occurs), a slight penalty is given to encourage smooth paths and reduce sharp turns.
[0181] The final reward value is obtained by weighted summation of the above four parts:
[0182] Total reward = arrival reward + collision penalty + approach reward + turning penalty.
[0183] This structured reward function avoids the policy deviation problem that may be caused by a single reward signal, and effectively improves the rationality and execution stability of the path.
[0184] By way of example, the solution of this embodiment is described below with reference to the accompanying drawings.
[0185] Figure 2 The state modeling process and neural network structure of a single agent in the path planning process in the A*-IDDQN scheme adopted in this embodiment are demonstrated.
[0186] Figure 2 The upper part is the process of environmental perception and state modeling:
[0187] The agent constructs a matrix of several channels based on local perception;
[0188] At the same time, global auxiliary information is extracted, such as target direction vector, distance and other scalar values;
[0189] All the information together constitutes the input state of reinforcement learning.
[0190] Figure 2 The lower part of the figure shows the structure and processing of the neural network:
[0191] The matrix information first passes through a two-layer convolutional network to extract spatial features;
[0192] After preprocessing, the scalar information is concatenated with the convolution output vector;
[0193] The concatenated results are input into the fully connected layer, and the final output is the Q value vector representing different actions for path decision making.
[0194] Figure 2 It intuitively demonstrates the entire process from state modeling, information fusion, convolution extraction, and action output, which is an important structural basis that distinguishes this embodiment from the traditional DQN method.
[0195] Figure 3 This is a training flow chart of this embodiment. Figure 3This paper demonstrates how the reinforcement learning system combines the A* algorithm, experience replay mechanism, main network (Q-Network) and target network (Target-Network) to complete the optimization process of the path planning strategy during the training phase of this embodiment.
[0196] Figure 3 The key components are described as follows:
[0197] Environment: Indicates the map information of the current training environment, including obstacle distribution, starting point, and target point.
[0198] A* algorithm: Calculates a feasible path based on environmental information, and writes this path into the experience replay pool as the initial high-quality experience.
[0199] Experience replay pool: stores state-action-reward-next-state data recorded in the form of four-tuples; contains A* initialization experience and real experience collected during training; supports random sampling of training data to break time correlation.
[0200] Q network: receives the current state input; outputs the action Q value; is used to select the current action and receives loss function feedback for update.
[0201] Target network: Periodically synchronizes parameters from the Q network; only used to generate target Q values and stabilize the training process.
[0202] Action selection and update mechanism: The agent selects and executes actions based on the Q value; collects environmental feedback, calculates the loss function and backpropagates updates; the action execution results are simultaneously fed back to the Replay Buffer to form new training data.
[0203] Figure 3 The key components and processes of the reinforcement learning training phase are clearly presented, reflecting the core innovative idea of this embodiment of combining A* to improve sample efficiency and using the target network to improve stability.
[0204] Figure 4 For a more specific training flow chart. Figure 4 The complete training process of single-agent reinforcement learning path planning in this embodiment is demonstrated, reflecting the temporal relationship and interaction logic between A* initialization, action selection strategy, reward calculation and network update.
[0205] Figure 4 The solution is described as follows:
[0206] Training start: Start the entire training system.
[0207] A* path initialization: Before each training round, the A* algorithm is called to generate a feasible path from the starting point to the target point for experience pool initialization or reference.
[0208] Episode start: The agent initializes its state and prepares for a round of training.
[0209] Select action strategy, including two types of strategies:
[0210] ε-greedy: randomly select an action with a certain probability, and select the action with the largest Q value;
[0211] Customized strategy: Introducing target direction and safety judgment to optimize exploration effect.
[0212] The two together form the action selection module, which calls the execution module to perform the selected action.
[0213] Execute actions: Update the agent state based on the selected action and obtain feedback from the environment.
[0214] Reward calculation: Based on the action results and environment status, the structured reward function is called to comprehensively calculate the four sub-reward values (goal achievement, approach, collision, and turning).
[0215] Update the main network: Sample from the experience pool and update the main Q network through gradient descent.
[0216] Periodically update the target network: every certain number of steps, synchronize the main network parameters to the target network to improve training stability.
[0217] Determine the termination condition:
[0218] If the goal or maximum number of steps is reached, the evaluation process ends;
[0219] If the maximum number of rounds is reached, the training ends.
[0220] End of Round Evaluation:
[0221] Count the training indicators of the current round (such as path length, number of turns, number of collisions, etc.);
[0222] Fine-tune and optimize reward parameters based on progress.
[0223] The beneficial effects of this embodiment include:
[0224] 1. Significantly improve training efficiency:
[0225] This example introduces the A* algorithm path initialization experience mechanism, effectively alleviating the problem of insufficient experience and reliance on random exploration in the early stages of reinforcement learning. By pre-injecting high-quality path experience, training convergence time is significantly shortened and training efficiency is improved.
[0226] 2. Optimize action selection strategy:
[0227] Based on the traditional ε-greedy strategy, this embodiment designs a custom action selection strategy that combines path safety and target direction judgment. It can dynamically adjust the action selection method according to the environmental state, avoid invalid exploration and path duplication, and improve the rationality and safety of path planning.
[0228] 3. The reward function design is more targeted:
[0229] This embodiment constructs a structured multi-reward function that organically integrates goal achievement rewards, collision penalties, proximity rewards, and path turning penalties. It fully considers the efficiency, smoothness, and safety of path planning, avoids the problem of overly simplistic reward design in traditional methods, and can effectively guide the intelligent agent to learn high-quality paths.
[0230] 4. Improve path quality and explainability:
[0231] By combining local perception information with global auxiliary scalar information, the state input structure designed in this embodiment realizes the joint modeling of the local environment and the overall goal, enhances the model's adaptability to environmental changes, and the ultimately generated path has excellent characteristics such as shorter path length, lower collision risk, and fewer turns.
[0232] 5. Good scalability and practical value:
[0233] The technical solution of this embodiment can be widely used in scenarios such as robot navigation, drone path planning, and automatic driving systems. It has good scalability and is suitable for environments of different scales and complexities, and has high application promotion value.
[0234] Reference Figure 5 The embodiment of the present application provides an intelligent agent path planning device in a dynamic perception-constrained environment, comprising:
[0235] The environmental perception unit is used to obtain the multi-channel matrix of the local environment and the scalar auxiliary information of the global environment;
[0236] A feature extraction unit, configured to extract local spatial features of the multi-channel matrix using a main network; wherein the main network is a neural network initialized based on an A* path search algorithm and combined with an improved dual deep Q network;
[0237] an information conversion unit, configured to convert the scalar auxiliary information into auxiliary vector features using the main network;
[0238] A feature splicing unit, configured to splice the local spatial features and the auxiliary vector features using the main network to obtain a spliced vector;
[0239] an action prediction unit, configured to process the concatenated vector using the two fully connected layers in the main network to obtain a Q-value vector having the same dimension as the action space;
[0240] An action execution unit is used to control the action of the agent according to the Q value vector.
[0241] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0242] In some optional embodiments, the function / operation mentioned in the block diagram may not occur in the order mentioned in the operation diagram. For example, depending on the function / operation involved, the two boxes shown in succession can actually be executed substantially simultaneously or the boxes can sometimes be executed in reverse order. In addition, the embodiments presented and described in the flow chart of the present application are provided in an exemplary manner for the purpose of providing a more comprehensive understanding of the technology. The disclosed method is not limited to the operations and logical flows presented herein. Optional embodiments are contemplated in which the order of the various operations is changed and the sub-operations described as a part of a larger operation are performed independently.
[0243] In addition, although the present application is described in the context of functional modules, it should be understood that, unless otherwise stated, one or more of the functions and / or features described may be integrated into a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It is also understood that a detailed discussion of the actual implementation of each module is not necessary for understanding the present application. More specifically, given the properties, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the module will be understood within the routine skills of an engineer. Therefore, a person skilled in the art can implement the present application as set forth in the claims using ordinary techniques without undue experimentation. It is also understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present application, which is determined by the full scope of the appended claims and their equivalents.
[0244] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program code.
[0245] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0246] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, and then editing, interpreting, or processing in another suitable manner as necessary, and then storing it in a computer memory.
[0247] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having logic gate circuits for implementing logic functions on data signals, an application-specific integrated circuit having suitable combinational logic gate circuits, a programmable gate array (PGA), a field-programmable gate array (FPGA), etc.
[0248] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present application. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0249] Although the embodiments of the present application have been shown and described, those skilled in the art will appreciate that various changes, modifications, substitutions, and variations may be made to the embodiments without departing from the principles and intent of the present application, and that the scope of the present application is defined by the claims and their equivalents.
[0250] The above is a specific description of the preferred implementation of the present application, but the present application is not limited to the embodiments. Those skilled in the art may make various equivalent modifications or substitutions without violating the spirit of the present application, and these equivalent modifications or substitutions are all included in the scope defined by the claims of the present application.
Claims
1. An intelligent agent path planning method in a dynamic perception constrained environment, characterized by: The method comprises the following steps: Obtain the multi-channel matrix of the local environment and the scalar auxiliary information of the global environment; Extracting local spatial features of the multi-channel matrix using a main network; wherein the main network is a neural network initialized based on an A* path search algorithm and combined with an improved dual deep Q network; Converting the scalar auxiliary information into auxiliary vector features using the main network; Using the main network to splice the local spatial features and the auxiliary vector features to obtain a spliced vector; Processing the concatenated vector using two fully connected layers in the main network to obtain a Q-value vector with the same dimension as the action space; The action of the agent is controlled according to the Q-value vector.
2. The method for intelligent agent path planning in a dynamic perception-constrained environment according to claim 1, characterized in that: The method of obtaining the multi-channel matrix of the local environment and the scalar auxiliary information of the global environment includes the following steps: Obtaining a local obstacle position matrix, a current target point matrix of the agent, an optional path prediction matrix, and a direction information matrix as the multi-channel matrix; The Euclidean distance between the agent and the target point, the current position of the agent and the direction of the target point are obtained as the scalar auxiliary information.
3. The method for intelligent agent path planning in a dynamic perception-constrained environment according to claim 1, characterized in that: The method of extracting the local spatial features of the multi-channel matrix using the main network comprises the following steps: The local spatial features of the multi-channel matrix are extracted using two convolutional layers activated by ReLU in the main network, and then the local spatial features are converted into vector form through a Flatten layer.
4. The method for intelligent agent path planning in a dynamic perception-constrained environment according to claim 1, characterized in that: The method of converting the scalar auxiliary information into auxiliary vector features using the main network comprises the following steps: The scalar auxiliary information is converted into auxiliary vector features using a fully connected layer in the main network.
5. The method for intelligent agent path planning in a dynamic perception-constrained environment according to claim 1, characterized in that: The method of processing the concatenated vector using two fully connected layers in the main network to obtain a Q-value vector with the same dimension as the action space includes the following steps: The concatenated vector is processed using the two fully connected layers in the main network to obtain the four Q-value vectors with the same dimension as the action space, which respectively correspond to the action sets of the agent in the up, down, left, and right directions at the current position.
6. The method for intelligent agent path planning in a dynamic perception-constrained environment according to any one of claims 1 to 5, characterized in that: The method further comprises the step of training the main network, wherein the step of training the main network comprises the following steps: Initializing the main network and the target network; wherein the target network is a neural network obtained by initializing based on the A* path search algorithm and combining it with an improved dual deep Q network; Before each training round, a reference path is generated using the A* path search algorithm based on the current position of the agent and the position of the target point; utilizing the agent to perform actions according to the reference path, the ε-greedy strategy, and a predefined safe action selection mechanism; Determine the environmental feedback reward after the agent performs an action, and record the experience of performing the action into the experience replay pool; Randomly sampling small batches of samples from the experience replay pool to update the parameters of the main network; Synchronizing the parameters of the master network to the target network at fixed step intervals; Calculating a loss function and then calculating a loss gradient using the Q value vector output by the main network and the Q value vector output by the target network; Update the parameters of the main network according to the loss gradient.
7. The method for intelligent agent path planning in a dynamic perception-constrained environment according to claim 6, characterized in that: Determining the environmental feedback reward after the agent performs an action includes the following steps: The reaching reward, collision penalty, approach reward and turning penalty after the agent performs the action are determined as the environment feedback reward.
8. An intelligent agent path planning device in a dynamic perception-constrained environment, characterized in that: The device comprises: The environmental perception unit is used to obtain the multi-channel matrix of the local environment and the scalar auxiliary information of the global environment; A feature extraction unit, configured to extract local spatial features of the multi-channel matrix using a main network; wherein the main network is a neural network initialized based on an A* path search algorithm and combined with an improved dual deep Q network; an information conversion unit, configured to convert the scalar auxiliary information into auxiliary vector features using the main network; A feature splicing unit, configured to splice the local spatial features and the auxiliary vector features using the main network to obtain a spliced vector; an action prediction unit, configured to process the concatenated vector using the two fully connected layers in the main network to obtain a Q-value vector having the same dimension as the action space; An action execution unit is used to control the action of the agent according to the Q value vector.
9. An electronic device, characterized in that: The electronic device includes a processor and a memory; The memory is used to store programs; The processor executes the program to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The storage medium stores a program, and the program is executed by a processor to implement the method according to any one of claims 1 to 7.