AGV path planning method based on hybrid exploration strategy

By combining the A* algorithm and the spiking neural network SNN-DQN hybrid exploration strategy, the path planning problem of multiple AGV systems in dynamic environments was solved, a low-energy and efficient path planning method was implemented, and training efficiency and real-time performance were improved.

CN120630989APending Publication Date: 2025-09-12ZHENGZHOU UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510759851.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing technologies cannot adapt to dynamic obstacles when multiple AGVs collaborate, have poor real-time performance and scalability, and reinforcement learning algorithms consume high computing energy on resource-constrained AGV devices. Pulse neural network training is unstable and has difficulty converging.

Method used

A hybrid exploration strategy is adopted, combining the global path prior of the A* algorithm with the dynamic adaptability of the spiking neural network SNN-DQN, and through deep reinforcement learning algorithm and gradient surrogate function back propagation, energy consumption is reduced and training efficiency is improved.

Benefits of technology

It significantly reduces the energy consumption of AGV path planning, improves the adaptability in dynamic and complex scenarios, reduces the risk of collision, and improves training efficiency and convergence speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120630989A_ABST
    Figure CN120630989A_ABST
Patent Text Reader

Abstract

The invention discloses an AGV path planning method based on a hybrid exploration strategy, and the method comprises the steps: designing the hybrid exploration strategy, randomly selecting an AGV state as the current input, and fusing a deep reinforcement learning algorithm and an A * algorithm to output an optimal action; a deep Q network DQN algorithm is adopted in a deep reinforcement learning algorithm, and SNN is used as a strategy network and a target network of the DQN, so that energy consumption is remarkably reduced; the global path prior of the A * algorithm and the dynamic adaptive capacity of the SNN-DQN are combined, the adaptive capacity of the algorithm in a dynamic complex scene is improved, the collision risk is reduced, the convergence speed is increased by adopting a hybrid exploration strategy, and the training efficiency is improved; the error function gradient is calculated by adopting gradient substitution function back propagation, non-micro points are bypassed, end-to-end gradient back propagation is realized, and the problem of pulse neural network training is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of automatic guided vehicle (AGV) path planning, and in particular to an AGV path planning method based on a hybrid exploration strategy, which is suitable for autonomous navigation of AGVs in dynamic environments such as warehousing logistics and intelligent manufacturing. Background Art

[0002] With the rapid development of the social economy and the logistics industry, the scale of smart warehousing continues to expand, and the level of automation continues to improve. To cope with a large number of complex transportation tasks, automated guided vehicles (AGVs) are widely used in smart warehousing systems across the logistics industry. Therefore, accurate and efficient AGV path planning has become a key focus in the industry.

[0003] Currently, traditional path planning algorithms, such as the Dijkstra algorithm and the A* algorithm, are commonly used in AGV path planning. These algorithms have achieved good results in single AGV path planning, but they are overly dependent on global environmental information. When faced with multi-AGV collaboration, they cannot adapt to dynamic obstacles, have poor real-time performance, poor scalability, and low training efficiency.

[0004] Reinforcement learning (RL), particularly multi-agent reinforcement learning (MARL), provides an effective approach to solving the path planning problem for automated guided vehicles (AGVs) in dynamic environments. While RL can handle dynamic environments, it consumes high computational energy, making it difficult to deploy on resource-constrained AGVs.

[0005] In recent years, spiking neural networks (SNNs), by mimicking biological neural signal transmission, have demonstrated the potential to handle complex tasks with low energy consumption. However, direct use for path planning suffers from training instability and convergence difficulties.

[0006] Therefore, how to propose a real-time path planning method suitable for multi-AGV systems and complex application environments is an urgent problem to be solved. Summary of the Invention

[0007] The purpose of the present invention is to overcome the shortcomings of the existing technology and propose an AGV path planning method based on a hybrid exploration strategy. It uses SNN as the policy network and target network of DQN, significantly reduces energy consumption, and combines the global path prior of the A* algorithm with the dynamic adaptability of SNN-DQN to improve training efficiency.

[0008] The present invention provides an AGV path planning method based on a hybrid exploration strategy, which specifically includes the following steps:

[0009] S1, collects warehouse layout information and AGV operation parameters, and constructs the AGV state space, action space and reward function;

[0010] S2. Design a hybrid exploration strategy, randomly select an AGV state as the current input, and output the optimal action by fusing the deep reinforcement learning algorithm and the A* algorithm. The deep reinforcement learning algorithm uses the deep Q network DQN algorithm. The policy network and target network of the DQN algorithm use the spiking neural network (SNN) to calculate the Q value of each action.

[0011] S3. Calculate the next state Q value of the AGV action through the target network, and calculate the current state Q value of the AGV action and the difference between the current state Q value and the next state Q value through the policy network, and store the AGV current state, current state Q value, reward and next state Q value in the experience pool;

[0012] S4. Calculate the loss function value based on the difference between the current Q value and the next state Q value, use the gradient substitution function backpropagation to calculate the error function gradient, update the pulse neural network parameters in the policy network, and assign the pulse neural network parameters in the policy network to the pulse neural network in the target network when the number of pulse neural network parameter updates in the policy network reaches a preset threshold;

[0013] S5. After the training meets the requirements, the post-training deep reinforcement learning algorithm is deployed to track environmental changes in real time to complete AGV dynamic planning.

[0014] Furthermore, the reward function includes dense rewards and sparse rewards. The sparse rewards are used to give positive rewards when the AGV completes the current stage task, and negative rewards when the AGV runs out of the boundary area or collides. The dense rewards are given according to the distance between the current position and the target point position.

[0015] Dense rewards:

[0016] reward1=-k×(|x c -x t |+|y c -y t |)

[0017] reward2=10

[0018] reward 稠密 =reward1+reward2

[0019] Sparse rewards:

[0020]

[0021] Where k is the scaling factor, x c 、y c is the horizontal and vertical coordinates of the current position; x t 、y t are the horizontal and vertical coordinates of the target point.

[0022] Furthermore, in step S2, the optimal action is output by fusing the deep reinforcement learning algorithm with the A* algorithm, specifically including:

[0023] In the early stages of training, the A* algorithm is mainly used to guide the search process. The A* algorithm provides prior knowledge of the global optimal path and guides the discovery of the shortest path by regularly outputting actions, avoiding the blindness of AGV's random exploration. As training progresses, when the cumulative number of training rounds exceeds a threshold, the frequency of A* algorithm calls is automatically reduced, and the algorithm gradually transitions to the DQN algorithm's policy network to avoid over-reliance on prior knowledge.

[0024] Furthermore, the DQN algorithm's policy network uses a spiking neural network to calculate the Q value of each action, specifically including:

[0025] The current state is input into a pulse neural network based on the LIF model. The input state is converted into a pulse sequence through pulse coding, and then information processing is performed through the LIF model. The current state input is subjected to a convolution operation to obtain the membrane potential. Then, pulse input and membrane potential reset are performed based on the membrane potential. Finally, by outputting the membrane potential value of each action, the action corresponding to the optimal membrane potential is selected, and the current state Q value is obtained.

[0026] The LIF mathematical model is:

[0027] Where V is the membrane potential; V th is a given excitation threshold; V reset is the membrane potential after reset; τ is the membrane time constant; R is the resistance; and I is the output current.

[0028] Furthermore, the target network of the DQN algorithm uses a spiking neural network to calculate the Q value of each action, specifically including:

[0029] Extract the current state input state matrix, which is a local view in the AGV field of view, with the height and width set to 1. In this case, the input is a four-dimensional feature map B×C×H×W; expand the feature map to a five-dimensional feature map T×B×C×H×W, where T is the set number of time steps, B is the sampling batch size, C is the total number of channels, and H and W are the height and width of the feature map respectively;

[0030] The five-dimensional feature map is input into the convolution layer to extract spatial features, and the continuous convolution output is converted into a pulse signal through the LIF model; the feature map resolution is reduced through the pooling layer to reduce the amount of calculation while retaining significant features; the spatial features are mapped to high-level decision features through the fully connected layer; the pulse signal is matched with the action space dimension through the fully connected layer, and the next state Q value corresponding to each action is output.

[0031] Furthermore, the AGV current state, current state Q value, reward and next state Q value are stored in the experience pool, specifically including:

[0032] The difference between the current state Q value and the next state Q value of the AGV action is calculated through the policy network. According to the absolute value of the difference, a priority coefficient is added to the stored samples. When sampling, the sampling probability is adjusted according to the weight of the sample additional priority coefficient. After each training, the sample additional priority coefficient is updated according to the new difference.

[0033] Furthermore, the difference between the current state Q value and the next state Q value of the AGV action is calculated by the policy network, specifically including:

[0034] Q(t)=Q policy (s t ,a t θ policy )

[0035] Q(t+1)=r t +γQ policy (s t+1 ,a t+1 θ policy )

[0036] TD=Q(t+1)-Q(t)

[0037] Among them, Q(t) is the Q value of the current state at time t, Q(t+1) is the Q value of the state at time t+1, that is, the Q value of the next state, r t is the reward at time t, γ is the discount factor, Q policy is the policy network, θ policy is the weight parameter of the policy network, s t is the state at time t, a t is the action at time t, s t+1 is the state at time t+1, a t+1 is the action at time t+1, and TD is the difference.

[0038] Furthermore, the loss function value L(θ) is calculated based on the difference between the current Q value and the next state Q value, specifically including:

[0039] TD error =[rt +γQ target (s t+1 ,argmaxQ policy (s t+1 ,a t θ policy );θ target )]-Q policy (s t ,a;θ policy )

[0040]

[0041] Among them, θ target is the weight parameter of the target network, Q target For the target network.

[0042] Furthermore, the back propagation of the gradient substitution function to calculate the error function gradient specifically includes:

[0043] The LIF model is used when the membrane voltage V reaches the set voltage threshold V th A jump occurs at V, which is essentially a step function, and its derivative is V th Due to the non-differentiable problem of mutation, back propagation cannot be directly applied. In back propagation, the derivative of the inverse tangent function or the sigmoid function is used to approximate the gradient of the pulse trigger to bypass the non-differentiable point.

[0044] The derivative of the inverse tangent function is:

[0045] The derivative of the sigmoid function is

[0046] Where η is the net input of the spiking neural network and e is a natural constant.

[0047] The present invention also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor executes the steps of the above-mentioned AGV path planning method based on the hybrid exploration strategy.

[0048] The beneficial effects of the present invention are:

[0049] 1. The deep Q network DQN algorithm is adopted in the deep reinforcement learning algorithm, and SNN is used as the policy network and target network of DQN, which significantly reduces energy consumption;

[0050] 2. Combining the global path prior of the A* algorithm with the dynamic adaptability of SNN-DQN improves the adaptability of the algorithm in dynamic and complex scenarios, reduces the risk of collision, and adopts a hybrid exploration strategy to accelerate convergence and improve training efficiency.

[0051] 3. The gradient substitution function is used for back propagation to calculate the gradient of the error function, bypassing the non-differentiable point and realizing end-to-end gradient back propagation, thus solving the problem of pulse neural network training. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 It is a flow chart of the AGV path planning method based on spiking neural network implemented by the present invention;

[0053] Figure 2 It is a structural diagram of the AGV path planning model implemented in the present invention;

[0054] Figure 3 It is a schematic diagram of the structure of the pulse neural network implemented by the present invention. DETAILED DESCRIPTION

[0055] In order to enable those skilled in the art to better understand the technical solution of the present invention, the present invention is further described in detail below with reference to the accompanying drawings.

[0056] like Figure 1 As shown, the present invention provides an AGV path planning method based on a hybrid exploration strategy, which specifically includes the following steps:

[0057] S1, collects warehouse layout information and AGV operation parameters, and constructs the AGV state space, action space and reward function;

[0058] S2. Design a hybrid exploration strategy, randomly select an AGV state as the current input, and output the optimal action by fusing the deep reinforcement learning algorithm and the A* algorithm. The deep reinforcement learning algorithm uses the deep Q network DQN algorithm. The policy network and target network of the DQN algorithm use the spiking neural network (SNN) to calculate the Q value of each action.

[0059] S3. Calculate the next state Q value of the AGV action through the target network, and calculate the current state Q value of the AGV action and the difference between the current state Q value and the next state Q value through the policy network, and store the AGV current state, current state Q value, reward and next state Q value in the experience pool;

[0060] S4. Calculate the loss function value based on the difference between the current Q value and the next state Q value, use the gradient substitution function backpropagation to calculate the error function gradient, update the pulse neural network parameters in the policy network, and assign the pulse neural network parameters in the policy network to the pulse neural network in the target network when the number of pulse neural network parameter updates in the policy network reaches a preset threshold;

[0061] S5. After the training meets the requirements, the post-training deep reinforcement learning algorithm is deployed to track environmental changes in real time to complete AGV dynamic planning.

[0062] Furthermore, the step S1 includes the following specific steps:

[0063] S11. Collect warehouse layout information and construct a two-dimensional raster map, including the location and number of storage stations and picking stations, the operating areas, and the operating directions of the aisles;

[0064] S12: Collect AGV operating parameters and construct the AGV state space, including AGV name and number, standby position, current position, task phase, and load status, and determine whether the AGV can perform the task normally. Then, use the random function to generate random tasks, each of which includes a storage station (starting point) and a picking station (target location);

[0065] S13. Construct the AGV action space, including five actions: up, down, left, right, and stop. At the same time, map the action strings to numbers so that each action has a corresponding numerical value.

[0066] S14. Construct a reward function. The reward function is set with sparse rewards and dense rewards.

[0067] Furthermore, the reward function includes dense rewards and sparse rewards. The sparse rewards are used to give positive rewards when the AGV completes the current stage task, and negative rewards when the AGV runs out of the boundary area or collides. The dense rewards are given according to the distance between the current position and the target point position.

[0068] Dense rewards:

[0069] reward1=-k×(|x c -x t |+|y c -y t |)

[0070] reward2=10

[0071] dense reward=reward1+reward2

[0072] Sparse rewards:

[0073]

[0074] Where k is the scaling factor, which can be 0.01, x c 、y c is the horizontal and vertical coordinates of the current position; x t 、y t are the horizontal and vertical coordinates of the target point.

[0075] When collecting warehouse layout information, the following parts need to be clarified: 1) The relative position of the picking area and the storage area. That is, the storage area is located on the same side or the opposite side of the picking area; 2) The location and number of storage stations and picking stations in the warehouse; 3) The passable area; 4) The running direction of the channel, whether the channel is a one-way channel or a two-way channel. Based on the collected layout information, a two-dimensional raster map of the warehouse environment can be created to facilitate the definition of the subsequent state matrix. When collecting AGV operating parameters, the following information needs to be clarified: the number of AGVs that can be operated in the warehouse, the name of each AGV, the standby position, the current position, the task stage, the load condition, and whether the task is performed normally.

[0076] State Space: Based on the known warehouse layout information and AGV operating parameters, the state space is defined. The state space primarily consists of the following matrices: a traversable position matrix, a current position matrix, and a target position matrix for the AGV. Depending on the position requirements of each state space matrix, it is converted into a binary matrix, with the number of rows and columns corresponding to the number of grid cells occupied by the grid width and height. When defining the traversable position matrix, the locations of obstacles within the grid must be clearly defined. Obstacles include picking stations, other AGVs, and storage stations where pods are located. A binary matrix for traversable positions is generated, with 1 representing traversable positions and 0 representing obstacle positions. Furthermore, a binary matrix for obstacles is generated, with 1 representing obstacles and 0 representing all other positions. The obstacle matrix is ​​then compared with the traversable matrix to determine if they are complementary. The accuracy of the traversable matrix is ​​verified using the obstacle matrix. When defining the current position matrix, all positions are 0 except the AGV's current position, which is represented by 1. When defining the target position matrix, all positions are 0 except the AGV's current target position, which is represented by 1. The AGV passable position matrix, current position matrix and AGV target position matrix are tensor-concatenated to obtain the initial environment state tensor.

[0077] Furthermore, in step S2, the optimal action is output by fusing the deep reinforcement learning algorithm with the A* algorithm, specifically including:

[0078] In the early stages of training, the A* algorithm is mainly used to guide the search process. The A* algorithm provides prior knowledge of the global optimal path and guides the discovery of the shortest path by regularly outputting actions, avoiding the blindness of AGV's random exploration. As training progresses, when the cumulative number of training rounds exceeds a threshold, the frequency of A* algorithm calls is automatically reduced, and the algorithm gradually transitions to the DQN algorithm's policy network to avoid over-reliance on prior knowledge.

[0079] Furthermore, if Figure 2 As shown, the DQN algorithm's policy network uses a pulse neural network to calculate the Q value of each action, specifically including:

[0080] The current state is input into a pulse neural network based on the LIF model. The input state is converted into a pulse sequence through pulse coding, and then information processing is performed through the LIF model. The current state input is subjected to a convolution operation to obtain the membrane potential. Then, pulse input and membrane potential reset are performed based on the membrane potential. Finally, by outputting the membrane potential value of each action, the action corresponding to the optimal membrane potential is selected, and the current state Q value is obtained.

[0081] The LIF mathematical model is:

[0082] Where V is the membrane potential; V th is a given excitation threshold; V reset is the membrane potential after reset; τ is the membrane time constant; R is the resistance; and I is the output current.

[0083] Furthermore, the target network of the DQN algorithm uses a spiking neural network to calculate the Q value of each action, specifically including:

[0084] Extract the current state input state matrix, which is a local view in the AGV field of view, with the height and width set to 1. In this case, the input is a four-dimensional feature map B×C×H×W; expand the feature map to a five-dimensional feature map T×B×C×H×W, where T is the set number of time steps, B is the sampling batch size, C is the total number of channels, and H and W are the height and width of the feature map respectively;

[0085] The five-dimensional feature map is input into the convolution layer to extract spatial features, and the continuous convolution output is converted into a pulse signal through the LIF model; the feature map resolution is reduced through the pooling layer to reduce the amount of calculation while retaining significant features; the spatial features are mapped to high-level decision features through the fully connected layer; the pulse signal is matched with the action space dimension through the fully connected layer, and the next state Q value corresponding to each action is output.

[0086] Among them, the purpose of the policy network is action selection. By calculating the Q value of different actions taken by the AGV in the current state, the action output corresponding to the maximum Q value is selected; the purpose of the target network is to evaluate the Q value of the selected action and reduce the overestimation of the Q value by a single network. The network structure diagram is as follows Figure 3 shown.

[0087] Input encoding layer: Convert the AGV coordinates (x, y) into a two-dimensional pulse density matrix. The encoding formula is:

[0088]

[0089] Where v is the pulse propagation rate parameter; x and y are the two-dimensional coordinates of the AGV in the environment; t∈[0, T] is a discrete time variable; T is the total duration of the entire time calculation interval; and δ(·) is the unit pulse function.

[0090] Convolutional layers: conv1, 1 input channel, 32 output channels; conv2, 32 input channels, 64 output channels; conv3, 64 input channels, 128 output channels. The convolution kernel size is 3×3, and the padding value is set to 1 to ensure that the feature map size remains unchanged and to avoid spatial distortion of path information.

[0091] Pooling layer: The pooling window size for pool1 and pool2 is 2×2, the sliding step is set to 2, and the padding is set to 0 to ensure no border padding. The two-level pooling layer is bound to the feature map size, which reduces the feature map size. While ensuring path planning accuracy, it also improves the real-time performance and energy efficiency of the SNN.

[0092] In addition, the pooling result feedback can adjust the LIF neuron threshold. The specific formula is:

[0093]

[0094] Where V new th is the new threshold, V oldt h is the threshold at the previous moment, pool_output is the final output of the pooling layer, and α is a hyperparameter that is adjusted according to the actual model training and application scenarios.

[0095] The fully connected layer: fc1, with an input feature count of 128 and an output feature count of 256, processes the flattened feature dimensions of the convolutional layer output, maps them to a higher-dimensional feature space, and activates pre-pulse triggering preparations. fc2, with an input feature count of 256 and an output dimension equal to the number of actions, maps the hidden layer features to the Q-value of each possible action, ensuring that the output Q-value can be directly used for subsequent action execution, reward calculation, and backpropagation.

[0096] LIF neurons: Contains a total of 4 neurons. Three neurons (lif1, lif2, and lif3) are placed after each output feature of the convolutional layer, converting the output 32 / 64 / 128-channel feature maps into spatiotemporal pulse trains. One neuron (lif_fc) is placed after the fully connected layer and inputs the 256-dimensional abstract features output by fc1, converting them into a new time-encoded pulse model, which directly selects the final Q value of the final action.

[0097] Furthermore, the difference between the current state Q value and the next state Q value of the AGV action is calculated by the policy network, specifically including:

[0098] Q(t)=Q policy (s t ,at θ policy )

[0099] Q(t+1)=r t +γQ policy (s t+1 ,a t+1 θ policy )

[0100] TD=Q(t+1)-Q(t)

[0101] Among them, Q(t) is the Q value of the current state at time t, Q(t+1) is the Q value of the state at time t+1, that is, the Q value of the next state, r t is the reward at time t, γ is the discount factor, Q policy is the policy network, θ policy is the weight parameter of the policy network, s t is the state at time t, a t is the action at time t, s t+1 is the state at time t+1, a t+1 is the action at time t+1, and TD is the difference.

[0102] Furthermore, the AGV current state, current state Q value, reward and next state Q value are stored in the experience pool, specifically including:

[0103] The policy network calculates the difference between the Q-value of the AGV's current state and the Q-value of the next state. Based on the absolute value of this difference, a priority coefficient is added to the stored samples. During sampling, the sampling probability is adjusted based on the weight of the sample-added priority coefficient. After each training session, the sample-added priority coefficient is updated based on the new difference. Based on the difference, the experience samples are stored in the experience replay pool. During each network update, the agent randomly draws a batch of experiences from the experience pool for training. This random sampling helps reduce correlation between data and avoid overfitting.

[0104] Furthermore, the loss function value L(θ) is calculated based on the difference between the current Q value and the next state Q value, specifically including:

[0105] TD error =[r t +γQ target (s t+1 ,argmaxQ policy (s t+1 ,a t θ policy );θ target )]-Q policy (s t ,a;θ policy )

[0106]

[0107] Among them, θ target is the weight parameter of the target network, Q target For the target network.

[0108] Furthermore, the back propagation of the gradient substitution function to calculate the error function gradient specifically includes:

[0109] The LIF model is used when the membrane voltage V reaches the set voltage threshold V th A jump occurs at V, which is essentially a step function, and its derivative is V th Due to the non-differentiable problem of mutation, back propagation cannot be directly applied. In back propagation, the derivative of the inverse tangent function or the sigmoid function is used to approximate the gradient of the pulse trigger to bypass the non-differentiable point.

[0110] The derivative of the inverse tangent function is:

[0111] The derivative of the sigmoid function is

[0112] Where η is the net input of the spiking neural network and e is a natural constant.

[0113] The present invention also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor executes the steps of the above-mentioned AGV path planning method based on the hybrid exploration strategy.

[0114] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware through a computer program. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a ROM, a RAM, or the like.

[0115] The above disclosure is merely a preferred embodiment of the present invention and certainly cannot be used to limit the scope of the present invention. Therefore, equivalent changes made according to the claims of the present invention are still within the scope of the present invention.

Claims

1. An AGV path planning method based on a hybrid exploration strategy, characterized in that: S1, collects warehouse layout information and AGV operation parameters, and constructs the AGV state space, action space and reward function; S2. Design a hybrid exploration strategy, randomly select an AGV state as the current input, and output the optimal action by fusing the deep reinforcement learning algorithm and the A* algorithm. The deep reinforcement learning algorithm uses the deep Q network DQN algorithm. The policy network and target network of the DQN algorithm use the spiking neural network (SNN) to calculate the Q value of each action. S3. Calculate the next state Q value of the AGV action through the target network, and calculate the current state Q value of the AGV action and the difference between the current state Q value and the next state Q value through the policy network, and store the AGV current state, current state Q value, reward and next state Q value in the experience pool; S4. Calculate the loss function value based on the difference between the current Q value and the next state Q value, use the gradient substitution function backpropagation to calculate the error function gradient, update the pulse neural network parameters in the policy network, and when the number of pulse neural network parameter updates in the policy network reaches a preset threshold, assign the pulse neural network parameters in the policy network to the pulse neural network in the target network; S5. After the training meets the requirements, the post-training deep reinforcement learning algorithm is deployed to track environmental changes in real time to complete AGV dynamic planning.

2. The AGV path planning method based on a hybrid exploration strategy according to claim 1, characterized in that: The reward function includes dense rewards and sparse rewards. The sparse rewards are used to give positive rewards when the AGV completes the current stage task, and negative rewards when the AGV runs out of the boundary area or collides. The dense rewards are given according to the distance between the current position and the target point position. Dense rewards: reward1=-k×(|x c -x t |+|y c -y t |) reward2=10 reward 稠密 =reward1+reward2 Sparse rewards: Where k is the scaling factor, x c 、y c is the horizontal and vertical coordinates of the current position; x t 、y t are the horizontal and vertical coordinates of the target point.

3. The AGV path planning method based on a hybrid exploration strategy according to claim 1, characterized in that: In step S2, the optimal action is output by fusing the deep reinforcement learning algorithm with the A* algorithm, specifically including: In the early stages of training, the A* algorithm is mainly used to guide the search process. The A* algorithm provides prior knowledge of the global optimal path and guides the discovery of the shortest path by regularly outputting actions, avoiding the blindness of AGV's random exploration. As training progresses, when the cumulative number of training rounds exceeds a threshold, the frequency of A* algorithm calls is automatically reduced, and the algorithm gradually transitions to the DQN algorithm's policy network to avoid over-reliance on prior knowledge.

4. The AGV path planning method based on a hybrid exploration strategy according to claim 1, characterized in that: The DQN algorithm's policy network uses a pulse neural network to calculate the Q value of each action, specifically including: The current state is input into a pulse neural network based on the LIF model. The input state is converted into a pulse sequence through pulse coding, and then information processing is performed through the LIF model. The current state input is subjected to a convolution operation to obtain the membrane potential. Then, pulse input and membrane potential reset are performed based on the membrane potential. Finally, by outputting the membrane potential value of each action, the action corresponding to the optimal membrane potential is selected, and the current state Q value is obtained. The LIF mathematical model is: Where V is the membrane potential; V th is a given excitation threshold; V reset is the membrane potential after reset; τ is the membrane time constant; R is the resistance; and I is the output current.

5. The AGV path planning method based on a hybrid exploration strategy according to claim 1, characterized in that: The target network of the DQN algorithm uses a spiking neural network to calculate the Q value of each action, specifically including: Extract the current state input state matrix, which is a local view in the AGV field of view, with the height and width set to 1. In this case, the input is a four-dimensional feature map B×C×H×W; expand the feature map to a five-dimensional feature map T×B×C×H×W, where T is the set number of time steps, B is the sampling batch size, C is the total number of channels, and H and W are the height and width of the feature map respectively; The five-dimensional feature map is input into the convolution layer to extract spatial features, and the continuous convolution output is converted into a pulse signal through the LIF model; the feature map resolution is reduced through the pooling layer to reduce the amount of calculation while retaining significant features; the spatial features are mapped to high-level decision features through the fully connected layer; the pulse signal is matched with the action space dimension through the fully connected layer, and the next state Q value corresponding to each action is output.

6. The AGV path planning method based on a hybrid exploration strategy according to claim 1, characterized in that: The AGV's current state, current state Q value, reward, and next state Q value are stored in the experience pool, specifically including: The difference between the current state Q value and the next state Q value of the AGV action is calculated through the policy network. According to the absolute value of the difference, a priority coefficient is added to the stored samples. When sampling, the sampling probability is adjusted according to the weight of the sample additional priority coefficient. After each training, the sample additional priority coefficient is updated according to the new difference.

7. The AGV path planning method based on a hybrid exploration strategy according to claim 1, characterized in that: The calculation of the difference between the current state Q value and the next state Q value of the AGV action through the policy network specifically includes: Q(t)=Q policy (s t ,a t ;θ policy ) Q(t+1)=r t +γQ policy (s t+1 ,a t+1 ;θ policy ) TD=Q(t+1)-Q(t) Among them, Q(t) is the Q value of the current state at time t, Q(t+1) is the Q value of the state at time t+1, that is, the Q value of the next state, r t is the reward at time t, γ is the discount factor, Q policy is the policy network, θ policy is the weight parameter of the policy network, s t is the state at time t, a t is the action at time t, s t+1 is the state at time t+1, a t+1 is the action at time t+1, and TD is the difference.

8. The AGV path planning method based on a hybrid exploration strategy according to claim 7, characterized in that: The loss function value L(θ) is calculated based on the difference between the current Q value and the next state Q value, specifically including: TD error J[r t +γQ target (s t+1 ,argmaxQ policy (s t+1 ,a t θ policy )9θ target )]-Q policy (s t ,a6θ policy ) Among them, θ target is the weight parameter of the target network, Q target For the target network.

9. The AGV path planning method based on a hybrid exploration strategy according to claim 8, characterized in that: The method of calculating the error function gradient by back propagation using the gradient substitution function specifically includes: The LIF model is used when the membrane voltage V reaches the set voltage threshold V th A jump occurs at V, which is essentially a step function, and its derivative is V th Due to the non-differentiable problem of mutation, back propagation cannot be directly applied. In back propagation, the derivative of the inverse tangent function or the sigmoid function is used to approximate the gradient of the pulse trigger to bypass the non-differentiable point. The derivative of the inverse tangent function is: The derivative of the sigmoid function is Where η is the net input of the spiking neural network and e is a natural constant.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the processor is caused to perform the steps of the method according to any one of claims 1 to 9.