AGV path planning method based on improved SAC model
By combining the A* algorithm with the improved SAC model and spiking neural network, the real-time performance and energy efficiency issues of AGV path planning in complex environments are solved, achieving efficient and low-energy dynamic path planning.
Patent Information
- Application Number
- CN202511172085.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-21
- Publication Date
- 2025-12-05
AI Technical Summary
Existing AGV path planning models struggle to meet real-time and energy efficiency requirements when facing dynamic changes and complex environments. Traditional models have slow response speeds, biomimetic intelligent planning is prone to getting trapped in local optima, and reinforcement learning models have high computational complexity and high energy consumption.
By combining the A* algorithm with the improved SAC model, a spiking neural network (SNN) is used for path planning. The initialization is accelerated by pre-training with the A* algorithm. Combined with dense and sparse reward mechanisms, the spiking neural network is used for low-energy information transmission to achieve dynamic path planning.
It improves the efficiency and safety of AGV path planning in complex environments, reduces energy consumption, and is suitable for the real-time path planning needs of embedded devices.
Smart Images

Figure CN121072913A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of AGV path planning, in particular to an AGV path planning method based on an improved SAC model, aiming to improve the path planning performance and intelligent level in application scenarios such as warehouse and logistics automation systems, intelligent manufacturing and workshop automation. BACKGROUND
[0002] With the continuous modernization and intelligentization of the supply chain, the demand for efficient and automated transportation equipment is increasing. As the core logistics carrier, AGV(Automated Guided Vehicle) not only needs to meet the transportation efficiency requirements, but also must ensure the safety of the path, effectively avoid collision and congestion problems, and thus improve the reliability of the overall operation. At present, the mainstream path planning model can be mainly summarized into three categories: traditional planning model, planning method with bionic intelligent characteristics, and model based on reinforcement learning.
[0003] Traditional path planning models such as Dijkstra model and A* algorithm have the advantages of mature theory and high calculation efficiency, but when dealing with dynamic changes and complex environments, the response speed is slow and the flexibility is limited. Bionic intelligent planning methods based on particle swarm optimization, ant colony model, etc. exhibit strong global search ability, but when dealing with multi-modal and complex problems, they are prone to fall into local optimum, which limits their optimization effect. Reinforcement learning has the ability to learn optimal strategies in dynamic and unknown environments, can realize the self-adaptation of the environment, and exhibits strong generalization ability. The above three types of models in the face of multi-agent planning, the traditional model will be difficult to meet the real-time requirements of large-scale, multi-dynamic environment due to the exponential growth of state space; intelligent bionic planning can provide better global search ability, but its calculation complexity is high, which limits its practical application; in contrast, deep reinforcement learning shows better adaptability and effect in the task of path planning in dynamic environment, and can more effectively meet the needs of actual multi-agent, multi-environment scenarios.
[0004] The reinforcement learning model can autonomously learn to achieve the optimal strategy through continuous interaction with the environment. The SAC (Soft Actor-Critic) model can help achieve effective path planning in a complex and unknown environment by increasing the maximization of policy entropy and giving the control process higher flexibility. However, the reinforcement learning model usually relies on a large number of matrix operations, resulting in high hardware energy consumption and implementation complexity, which limits its efficient application on the hardware end. The SNN (Spiking Neural Networks) simulates the discharge mechanism of the biological nervous system, and the communication between neurons is sparse, which effectively reduces the energy consumption in the information transmission and processing process. Combined with neuromorphic chips, SNN can achieve very high energy efficiency and real-time performance, and is suitable for energy-limited and real-time performance required scenarios. However, in complex tasks, it is relatively difficult to directly use the spiking neural network for training, and the result stability needs to be improved.
[0005] Therefore, the combination of the spiking neural network and the reinforcement learning can not only reduce the energy consumption while ensuring the real-time performance and efficiency of the path planning process, but also improve the safety and adaptability of the path planning, so as to balance the performance and energy efficiency, and provide a new technical path for intelligent path planning. SUMMARY
[0006] To solve the above technical problems, the present application provides an AGV path planning method based on an improved SAC model, aiming to balance the performance indicators such as running efficiency, safety and energy consumption.
[0007] The present application provides an AGV path planning method based on an improved SAC model, which specifically includes the following steps:
[0008] S1, collecting warehouse layout data and AGV running state, calculating the state matrix corresponding to the current state, and taking it as an input variable;
[0009] S2, pre-training by using A* algorithm, global path planning by using A* algorithm, using the experience obtained from the path information, including s t , a t , r t , s t+1 , calculating the loss value required by the network of the SAC model, initializing the network parameters; at the same time, in the training process, using A* algorithm to guide action selection with a certain probability, so as to speed up the training speed, wherein s t is the current state, a t is the action selection, r t is the reward value, and s t+1 is the next state;
[0010] S3, the Actor network of the SAC model calculates the action and policy entropy according to the current state, and realizes low-energy operation in combination with the pulse neural network, and stores the action experience into an experience replay pool;
[0011] S4, small batches of data are sampled from the experience replay pool, the current Q value and the target Q value are calculated by using the Critic network and the target Critic network, and the timing difference error is calculated to update the Critic network parameters;
[0012] S5, according to the Q value output by the Critic network, the action generated by the Actor network and the policy entropy, a loss function is calculated, an error function gradient is calculated by using a replacement function back propagation, and the Actor network parameters are updated;
[0013] S6, it is judged whether the training reaches the set performance index, if yes, the training is stopped, the improved SAC model is adopted, the environment change is monitored in real time, and the dynamic path planning of the AGV is realized.
[0014] As a further improvement of the application, the step S1 is specifically: constructing an environment model and extracting current state information, combining the grid map of the warehouse layout and the running state information of all AGVs, extracting the passable matrix, the target position matrix and the current position matrix, and splicing these matrices to form the state input of the current AGV and as an input variable.
[0015] As a further improvement of the application, S2, pre-training is performed by using the A* algorithm, global path planning is performed by using the A* algorithm, experience including s t 、a t 、r t 、s t+1 is obtained from the path information, the loss value required by the network of the SAC model is calculated, and the network parameters are initialized; meanwhile, in the training process, the A* algorithm is used to guide the action selection with a certain probability to speed up the training speed, specifically: according to the input state, global path planning is performed by using the A* algorithm, the first time step action is executed, and the execution is continued until the target position is reached, thereby obtaining a complete trajectory data, the trajectory data includes the current state, the action selection, the next state, the reward value and the flag of whether the task is completed, wherein s t is the current state, a t is the action selection, r t is the reward value, s t+1 is the next state, and s tAs part of the input variables in step S1, the experience obtained by using the path information of the trajectory data is stored in the experience replay pool for the initialization training of the network parameters of the improved SAC model; thus, the trajectory data collected for a set number of training rounds is used to initialize the network parameters, and in the training process, the A* algorithm is used to guide the action selection with a certain probability to accelerate the training speed; the SAC model involves five networks, namely, the Actor network, the Critic network 1, the Critic network 2, the target Critic network 1, and the target Critic network 2.
[0016] As a further improvement of the present application, the A* algorithm is used to guide the action selection with a certain probability ε, specifically: ε ∈ [0, 1], and in this paper, the initial ε is set to 0.2, the action selection depends on the optimal action and policy entropy calculated by the Actor network with a probability of 1-ε, but under the probability ε, the random exploration action is not used, but the action recommended by the A* algorithm is used, and the probability ε decreases with the increase of the training round and decreases to 0 at the 200th round. In this way, while balancing exploration and utilization, the frequency of random exploration is reduced, over-exploration is avoided, and the training effect is improved.
[0017] As a further improvement of the present application, the Actor network is implemented by using a spiking neural network (SNN) and includes three parts: a pulse encoder, a pulse neural network, and a pulse decoder. The pulse encoder is composed of a convolutional layer conv and LIF (Leaky Integrate-and-Fire) neurons and is used to convert static state information into a corresponding pulse sequence. The pulse neural network includes two layers, each of which is composed of a synapse layer and a LIF neuron layer, extracts and converts spatially local pulse information into pulse representations with abstract features, wherein the synapse layer of the first layer is a convolutional layer conv, and the synapse layer of the second time is a fully connected layer FC. The pulse decoder is composed of a fully connected layer FC and converts the pulse sequence into membrane potential output for each action, which has more rich numerical information and more stable and smooth reflects the action probability. The Actor network is used to calculate the action and policy entropy under the current state, and the training process includes two parts: one is to calculate the corresponding action to realize the action selection by using the current state as the input; and the other is to sample a small batch of experience data from the experience replay pool, calculate the action and policy entropy of the corresponding state, and compare them with the Q value to obtain the loss function, and then update the network parameters through back propagation.
[0018] The mathematical model of the LIF neuron is as follows:
[0019] Charging U t = k v (V t-1 -V r )+ V r + WS't
[0020] Discharge
[0021] Reset
[0022] In the formula, U t is the membrane potential after synaptic stimulation; V t is the pulse trigger of the neuron at time t; V r is the membrane potential after reset; k v is the voltage decay coefficient; W is the weight of the pulse stimulation of other neurons; S t is the output pulse of the neuron at time t, S' t is the output pulse of other neurons at time t; V th is the given excitation threshold
[0023] As a further improvement of the application, the Actor network is used to calculate the action and policy entropy in the current state, and the optimal policy with respect to the policy entropy is defined as follows:
[0024]
[0025] In the formula, s t is the state at the current time t; a t is the action at the current time t; π is the policy adopted by the agent; τ is the state and action corresponding to the policy; γ is the discount factor, γ ∈ [0, 1), reflecting the time value of future rewards; R(s t ,a t ) is the reward obtained by the current state s t through the action a t ; α is the entropy temperature coefficient, α > 0, balancing the reward value and the randomness of the action; is the action distribution probability P corresponding to the current state s t ; φ is the Actor network parameter; H(P) function calculates the entropy of the probability distribution, and the entropy of the discrete action space is defined as follows:
[0026]
[0027] In the formula, p i (x) is the probability corresponding to the variable x, and ∑ i p i (x) = 1.
[0028] As a further improvement of the application, the current state corresponding action is obtained by A* algorithm or Actor network, the corresponding reward value is calculated by the reward function set by the environment, the next state information and task completion flag after execution are obtained, and the state position, action selection, reward value, next state and task completion flag are stored in the experience replay pool.
[0029] As a further improvement of the application, the reward function is set as follows: the reward function reward adopts a hybrid mechanism combining dense rewards and sparse rewards, and the specific design is as follows:
[0030]
[0031] reward = reward1 + reward2
[0032] In the formula, reward1 is a sparse reward setting. If the AGV enters an area outside the warehouse area or collides during operation, it is regarded as a violation of operation, and a negative reward is given. If the AGV returns to the position of the last step, a small negative reward is given to avoid repeated motion of the AGV on the path. When the AGV completes the current task and successfully reaches the target position, a positive reward is given to encourage the AGV to move towards the target position. Reward2 is a dense reward. The dense reward adopts a two-dimensional Gaussian function as part of the reward function to form a peak reward area around the target position, gradually guiding the AGV towards the target. r0 is the negative reward of single-step motion, r max is the maximum reward of reaching the target position, x c , y c is the coordinate of the current position, x t , y t is the coordinate of the target position, and sigma is the expansion range in the coordinate axis direction.
[0033] As a further improvement of the application, in step S4, a small batch of data is sampled from the experience replay pool, the current Q value and the target Q value are calculated by using the Critic network and the target Critic network, and the time difference error is calculated to update the Critic network parameters, and specifically, the definition of the Critic network about the optimal Q function under the policy entropy is as follows:
[0034]
[0035] In the formula, Q * (s t ,a t) is the action value corresponding to the current state-action pair, the calculation results of Critic network 1 and Critic network 2 are q1 and q2 respectively as the current Q value, the calculation results of target Critic network 1 and target Critic network 2 are q'1 and q'2 respectively, D is an experience replay pool, and R(s t t ) is the current state s t obtained by action a t .
[0036] The minimum value of the calculation results of the two target Critic networks is taken as the target Q value, the TD error is obtained by comparing the current Q value calculated by the Critic network, the loss function is minimized, the Critic network parameters are updated, and the target Critic network copies the parameters from the Critic network through soft update;
[0037] The Critic network loss function and soft update are defined as follows:
[0038]
[0039] θ i =τθ′ i +(1-τ)θ i i=1,2
[0040] In the formula, θ' i is the parameter of the target Critic network i; and τ is a soft update coefficient.
[0041] As a further improvement of the application, S5, according to the Q value output by the Critic network, the action generated by the Actor network and the policy entropy, the loss function is calculated, the error function gradient is calculated by using the replacement function back propagation, and the Actor network parameters are updated, specifically: the Q value corresponding to the state is calculated by using the Critic network to calculate the loss value, the loss function is minimized to update the Actor network parameters by back propagation, and the loss function is defined as follows:
[0042]
[0043] In the formula, is the Q value corresponding to the state calculated by using the Critic network i (i = 1, 2), is the minimum value of the calculation results of Critic network 1 and Critic network 2;
[0044] Considering the non-differentiable problem of LIF neurons in back propagation, the sign function is replaced by the replacement function to realize the minimization of loss and the effective update of network parameters, and the inverse tangent function is selected as the replacement function.
[0045] Compared with the prior art, the application has the following advantages and technical effects:
[0046] The application provides an AGV path planning method based on an improved SAC model, which has the advantages that the SAC initialization process is effectively accelerated by pre-training of the A* algorithm, the training efficiency is improved, the efficient exploration strategy of the SAC and the ability of the SAC to process uncertainty help the AGV to realize autonomous learning and path adjustment in an unknown or dynamic environment, the pulse neural network (SNN) structure is adopted, and information transmission is performed by using pulse signals, so that the energy efficiency of the system is significantly improved. This design makes the application particularly suitable for embedded devices, and effectively meets the requirements of low energy consumption and high real-time path planning. BRIEF DESCRIPTION OF DRAWINGS
[0047] Figure 1 FIG. 1 is a general framework diagram of the AGV path planning method based on the improved SAC model of the application;
[0048] Figure 2 FIG. 2 is a specific flowchart of the AGV path planning method based on the improved SAC model;
[0049] Figure 3 FIG. 3 is a framework diagram of the improved SAC model;
[0050] Figure 4 FIG. 4 is an architecture diagram of the Actor network and the Critic network of the improved SAC model. DETAILED DESCRIPTION
[0051] In order to enable personnel in the technical field to better understand the technical solutions of the application, the application will be further described in detail below with reference to the drawings.
[0052] As shown in FIG. 1, the application provides an AGV path planning method based on an improved SAC model, which specifically includes the following steps: Figure 1
[0053] S1, collecting warehouse layout data and AGV running state, calculating a state matrix corresponding to the current state, and taking the state matrix as an input variable;
[0054] S2, pre-training by using the A* algorithm, performing global path planning by using the A* algorithm, using the experience obtained from the path information, including s t , a t , r t , s t+1 , calculating the loss value required by the network of the SAC model, initializing the network parameters; meanwhile, in the training process, the A* algorithm is used to guide the action selection at a certain probability, so as to accelerate the training speed, wherein s t is the current state, a t is the action selection, and rt is a reward value, s t+1 is a next state.
[0055] S3, an Actor network of the SAC model calculates an action and a policy entropy according to a current state, and realizes low-energy operation in combination with a spiking neural network, and stores an action experience into an experience replay pool.
[0056] S4, small batch data is sampled from the experience replay pool, a current Q value and a target Q value are calculated by using a Critic network and a target Critic network, and a time difference error is calculated to update a Critic network parameter.
[0057] S5, a loss function is calculated according to a Q value output by the Critic network, an action and a policy entropy generated by the Actor network, an error function gradient is calculated by using a substitute function back propagation, and an Actor network parameter is updated.
[0058] S6, whether the training reaches a set performance index is judged, if yes, the training is stopped, an improved SAC model after the training is adopted, environmental changes are monitored in real time, and dynamic path planning of the AGV is realized.
[0059] According to Figure 1 , 2 , step S1, warehouse layout data and AGV running states are collected, a state matrix corresponding to a current state is calculated, and is used as an input variable, and specifically includes: an environment model is constructed and current state information is extracted.
[0060] The environment of the application is based on a grid map of a logistics warehouse layout, and the layout is divided into vertical layout, horizontal layout, flying-V layout and fishbone layout, and the application adopts horizontal layout. The grid map can be divided into picking, storage, waiting for picking and AGV storage areas, and only when the AGV has no load, the storage area is passable. Since the AGV does not rotate the direction when it is loaded, in order to facilitate picking operation, when the AGV moves to the waiting for picking area, the pod needs to be rotated towards the picking station. At the same time, since picking needs time, the AGV first runs to the waiting for picking area to wait for picking. The task completion process of the AGV is as follows: first, the task information is obtained, and the storage station and the picking station positions of the task pod are determined; then, the AGV moves from the current position to the storage station, obtains the pod, and at this time, the passing area below the pod is allowed to pass; then, the pod is moved to the waiting for picking area near the target picking station, and the pod is rotated to face the picking station; after the picking task is completed, the pod is returned to the storage station, and the original orientation is restored.
[0061] The grid map of the warehouse layout and the running state information of all AGVs are combined to extract the passable matrix, target position matrix and current position matrix, and these matrices are spliced to form the state input of the current AGV. The original layout size is a 13x13 grid map. The passable area and the current position are cropped into a 7x7 matrix with the current position as the center. According to the relative position of the target position and the current position, the target position is mapped in the matrix or the point closest to the target position in the matrix to form a 7x7 target position matrix. The obtained passable matrix, target position matrix and current position matrix are spliced as the current state input to the model to obtain the output action. The action space includes five discrete actions: up, down, left, right and stop to meet different operation requirements. The environment will execute the action according to the feedback strategy, and judge whether the action violates the rules through the constraint condition, and calculate the corresponding reward value by using the set reward function. The specific reward function will be described in the following part.
[0062] Step S2, pre-training by A* algorithm, global path planning is performed by A* algorithm, and the experience obtained from the path information, including s t , a t , r t , s t+1 , is used to calculate the loss value required by the network of the SAC model, and the network parameters are initialized; at the same time, during the training process, A* algorithm is used to guide the action selection with a certain probability, so as to speed up the training speed, wherein s t is the current state, a t is the action selection, r t is the reward value, and s t+1 is the next state, which specifically includes:
[0063] A* algorithm finds the shortest path from the starting point to the ending point by combining the actual path cost and heuristic estimation. Since the logistics warehouse environment is a combination of dynamic and static environments, the dynamic factors of moving obstacles need to be considered, so the global path planned by A* only takes the current time step, and constantly replans the path until the AGV reaches the target position. According to the input state, global path planning is performed by A* algorithm, the first time step action is executed, and the execution is continued until the target position is reached, thereby obtaining a complete trajectory data, which includes each current state, action selection, next state, reward value and flag indicating whether the task is completed. The experience obtained from the path information of the trajectory data is stored in the experience replay pool for network parameter initialization training of the improved SAC model. As shown in Figure 2 A* algorithm is used for pre-training, and the trajectory of the set number of training rounds is collected to initialize the network parameters. During the training process, A* algorithm is used to guide the action selection with a certain probability, which speeds up the training speed. Another part of the combination of A* algorithm and reinforcement learning will be described in the subsequent part.
[0064] The A* algorithm is used to guide action selection with a certain probability, which accelerates the training speed, specifically: the SAC model is improved to select actions, unlike the traditional epsilon-greedy strategy, the action selection of the method mainly depends on the optimal action calculated by the Actor network and the policy entropy, but under a certain proportion of random probability, instead of random exploration, the action recommended by the A* algorithm is used. This strategy design balances exploration and utilization while effectively reducing the frequency of random exploration, avoiding over-exploration and improving training effect. The initial setting is 20%, the A* algorithm action selection is performed, and the probability is continuously reduced as the training increases, and is reduced to 0 when the training reaches 200 rounds.
[0065] According to Figure 3 、 Figure 4 , the improved SAC model involves 5 networks, namely Actor network, Critic network 1, Critic network 2, target Critic network 1 and target Critic network 2, wherein the Actor network is implemented by using a spiking neural network (SNN) to calculate the action and policy entropy under the current state. The training process of the Actor network includes two parts: one is to use the current state as input to calculate the corresponding action to realize action selection; the other is to sample a small batch of experience data from the experience replay pool, calculate the action and policy entropy of the corresponding state, and compare with the Q value to obtain the loss function, and then update the network parameters through back propagation.
[0066] The Actor network action selection is realized by continuously interacting with the environment to obtain the current state information, outputting the corresponding action and its probability through the Actor network, and executing the action according to the selected strategy. The environment then feeds back the reward value, the next state and the state information of the task completion, providing a data basis for the next round of learning.
[0067] The Actor network generates action strategy, maximizes the target containing entropy regularization term to adjust its parameters, so that the strategy can explore the environment and improve the reward. The Actor network structure includes three parts: pulse encoder, spiking neural network and pulse decoder. The pulse encoder is composed of a convolution layer conv and LIF (Leaky Integrate-and-Fire) neurons, which is used to convert static state information into corresponding pulse sequence. The spiking neural network includes two layers, each layer is composed of a synaptic layer and a LIF neuron layer, which extracts and converts the spatial local pulse information into pulse representation with abstract features, wherein the synaptic layer of the first layer is a convolution layer conv, and the synaptic layer of the second time is a fully connected layer FC. The pulse decoder is composed of a fully connected layer FC, which converts the pulse sequence into membrane potential output of each action, has more rich numerical information, and more stably and smoothly reflects the action probability.
[0068] The mathematical model of LIF neuron is as follows:
[0069] Charging U t = k v (V t-1 -V r )+ V r + WS' t
[0070] Discharge
[0071] Reset
[0072] In the formula, U t is the membrane potential after synaptic stimulation; V t is the pulse trigger of the neuron at time t; V r is the membrane potential after reset; k v is the voltage decay coefficient; W is the weight of the pulse stimulation of other neurons; S t is the output pulse of the neuron at time t, S' t is the output pulse of other neurons at time t; V th is the given excitation threshold.
[0073] The calculation process of the spiking neural network is as follows: the input state is encoded into a pulse sequence by a pulse encoder; the feature information is extracted through the synaptic layer and the neuron layer; then the action corresponding to the current state and its value evaluation are output through the pulse decoder; finally, the action probability and the policy entropy of each action are calculated.
[0074] In the action selection stage, the current state is input, the Actor network outputs the corresponding action and action probability, and the action is executed according to the selected policy. The reward value, the next state and the task completion of the environment feedback are used as the basis for the next round of learning. In the training process, a small batch of state data is sampled from the experience replay pool, and the corresponding action and policy entropy are obtained by using the Actor network output. The optimal policy of the Actor network with respect to the policy entropy is defined as follows:
[0075]
[0076] In the formula, s t is the state at time t; a t is the action at time t; π is the policy adopted by the agent; τ is the state and action corresponding to the policy; γ is the discount factor, γ ∈ [0, 1), which reflects the time value of future rewards; R(s t ,a t ) is the reward value of the current state s t through the action a tThe reward obtained is calculated by a reward function reward; a is an entropy temperature coefficient, a > 0, and a trade-off between the reward value and the action randomness is made; is the current state s t The corresponding action distribution probability P, and φ is the Actor network parameter; the H(P) function calculates the entropy of the probability distribution, and the definition of the entropy of the discrete action space is as follows:
[0077]
[0078] In the formula, p i (x) is the probability corresponding to the variable x, and i i (x) = 1.
[0079] The experience information obtained by performing the action is stored in the experience replay pool.
[0080] The action corresponding to the current state is obtained by the A* algorithm or the Actor network, the corresponding reward value is calculated by the reward function set by the environment, the next state information and the task completion flag after the execution are obtained, and the state position, action selection, reward value, next state and task completion flag are stored in the experience replay pool.
[0081] In order to realize the movement of the AGV towards the target position and reduce the number of movement steps, the following reward function is designed to guide the path optimization and efficiency improvement. The reward function is as follows:
[0082] The reward function reward adopts a hybrid mechanism combining dense rewards and sparse rewards, and the specific design is as follows:
[0083]
[0084] reward = reward1 + reward2
[0085] In the formula, reward1 is the sparse reward setting, if the AGV enters the area outside the warehouse area or collides during operation, it is regarded as a violation operation, and a negative reward is given; if the AGV returns to the previous position, a small negative reward is given to avoid the repeated motion of the AGV on the path; when the AGV completes the current task and successfully reaches the target position, a positive reward is given to encourage the AGV to move towards the target position. Reward2 is the dense reward, the dense reward adopts a two-dimensional Gaussian function as part of the reward function, forming a peak reward area around the target position, gradually guiding the AGV towards the target, r0 is the negative reward of single-step motion, r max is the maximum reward of reaching the target position, x c , y c is the coordinate of the current position, x t , yt is the coordinate of the target position, and s is the extension range of the coordinate axis direction. In view of the action space of the present application being five discrete actions, the Manhattan distance is selected as the distance measure of the Gaussian kernel, and if the discrete action space is eight directions and a stop action, the Euclidean distance is selected as the Gaussian kernel. In this way, the clear incentive to the target achieved by sparse rewards is utilized, and the continuous guidance provided by dense rewards is combined, thereby enhancing the guidance effect in the training process. The advantages of the two can be combined to effectively improve the efficiency and effect of learning. In a complex task or environment, this mechanism helps to better balance the stability and exploration efficiency of learning, and achieve a more optimal path planning. If no violation operation occurs or the target is reached, the task completion flag is set to True, the current task is ended, otherwise the flag is set to False, and the state is updated to continue moving.
[0086] As shown in Figure 3 , S4, a small batch of data is sampled from the experience replay pool, the current Q value and the target Q value are calculated by using the Critic network and the target Critic network, and the time difference error is calculated to update the Critic network parameters, including: the Critic network estimates the value of the action in the current state, guiding the update of the Actor policy, the Critic network includes Critic network 1 and Critic network 2, which can improve the stability of training and avoid overestimation. The target Critic network copies the parameters from the Critic network through soft update, avoiding oscillation or divergence in the Critic network training, making the update target smoother and improving the training stability.
[0087] The definition of the Critic network with respect to the optimal Q function under the policy entropy is as follows:
[0088]
[0089] In the formula, Q * (s t ,a t ) is the action value corresponding to the current state-action pair, the calculation results of the Critic network 1 and the Critic network 2 are q1 and q2 respectively, as the current Q value, the calculation results of the target Critic network 1 and the target Critic network 2 are q'1 and q'2 respectively, D is the experience replay pool, and R(s t ,a t ) is the reward obtained by the action a t in the current state s t .
[0090] The Q value corresponding to the state is calculated by using the Critic network to calculate the loss value, and the loss function is minimized to update the Actor network parameters by back propagation. The loss function is defined as follows:
[0091]
[0092] In the formula, is the Q value corresponding to the state calculated by the Critic network i (i = 1, 2), is the minimum value of the calculation results of the Critic network 1 and the Critic network 2. Considering the non-differentiable problem of the LIF neuron in the back propagation, a substitute function is used to replace the sign function to realize the minimization of the loss and the effective update of the network parameters. The present application selects the inverse tangent function as the substitute function.
[0093] The minimum value of the calculation results of the two target Critic networks is taken as the target Q value, compared with the current Q value calculated by the Critic network to obtain the TD error, minimize the loss function, and update the Critic network parameters. The target Critic network copies the parameters from the Critic network through soft update.
[0094] The Critic network loss function and soft update are defined as follows:
[0095]
[0096] θ i =τθ' i +(1-τ)θ i i = 1, 2
[0097] In the formula, θ' i is the target Critic network i parameter; τ is the soft update coefficient.
[0098] After the model converges through continuous training, the parameters in the network are saved, and the optimized parameters are loaded into the chip. Through the collection of the current environment state, real-time environment dynamic tracking and path planning are realized.
[0099] Finally, it should be noted that the above enumeration is only a few specific embodiments of the present application. Obviously, the present application is not limited to the above embodiments, and there can be many variations. All variations that can be directly derived or inferred from the content disclosed by the person skilled in the art should be considered as the protection scope of the present application.
Claims
1. An AGV path planning method based on an improved SAC model, characterized in that, The method comprises the following steps: S1, collecting warehouse layout data and AGV running state, calculating the state matrix corresponding to the current state, and taking it as an input variable; S2, pre-training using A* algorithm, global path planning is carried out through A* algorithm, and experience obtained by using path information, including s t , a t , r t , s t+1 , the loss value required for calculating the network of SAC model is initialized, and the network parameters are initialized; during the training process, A* algorithm is used to guide action selection with a certain probability, so as to accelerate the training speed, wherein s t is the current state, a t is the action selection, r t is the reward value, and s t+1 is the next state; S3, the Actor network of the SAC model calculates the action and policy entropy according to the current state, and realizes the low-energy operation in combination with the pulse neural network, and stores the action experience into the experience replay pool; S4, sampling small batches of data from the experience replay pool, calculating the current Q value and the target Q value by using the Critic network and the target Critic network, and then calculating the time difference error to update the Critic network parameters; S5, calculating the loss function according to the Q value output by the Critic network, the action generated by the Actor network and the policy entropy, calculating the error function gradient by using the substitute function back propagation, and updating the Actor network parameters; S6, judging whether the training reaches the set performance index, if yes, stopping the training, adopting the improved SAC model after the training is completed, and realizing the dynamic path planning of the AGV by monitoring the environmental changes in real time. 2.The AGV path planning method based on the improved SAC model of claim 1, wherein, The step S1 is specifically: constructing an environment model and extracting current state information, combining the grid map of the warehouse layout and the running state information of all AGVs, extracting the passable matrix, the target position matrix and the current position matrix, splicing these matrices to form the state input of the current AGV, and taking it as an input variable.
3. The AGV path planning method based on the improved SAC model according to claim 2, characterized in that, S2. Pre-train using the A* algorithm, perform global path planning using the A* algorithm, and utilize the experience gained from path information, including s t a t r t s t+1 The required loss value for the SAC model network is calculated, and the network parameters are initialized. Simultaneously, during training, the A* algorithm is used with a certain probability to guide action selection, thereby accelerating the training speed. Specifically, based on the input state, global path planning is performed using the A* algorithm, selecting the action to execute at the first time step, and continuing execution until the target position is reached. This yields a complete trajectory data set, which includes the current state, action selection, next state, reward value, and a flag indicating whether the task is completed. t For the current state, a t For action selection, r t For reward value, s t+1 For the next state, the experience gained from the path information of the trajectory data is stored in the experience replay pool for the initial training of the network parameters of the improved SAC model. In this way, the network parameters are initialized by collecting trajectory data for a set number of training rounds. During the training process, the A* algorithm is used to guide the action selection with a certain probability to speed up the training. The SAC model involves a total of 5 networks, namely Actor network, Critic network 1, Critic network 2, target Critic network 1, and target Critic network 2.
4. The AGV path planning method based on the improved SAC model according to claim 3, characterized in that, The A* algorithm is used to guide the action selection with a certain probability, specifically, the action selection depends on the optimal action and the policy entropy calculated by the Actor network, but under a certain proportion of random probability, the random exploration action is not used, but the action recommended by the A* algorithm is used, which balances exploration and utilization, reduces the frequency of random exploration, avoids excessive exploration and improves the training effect.
5. The AGV path planning method based on the improved SAC model according to claim 4, characterized in that, The Actor network is realized by using a pulse neural network (SNN) and comprises three parts: a pulse encoder, a pulse neural network and a pulse decoder. The pulse encoder is composed of a convolution layer conv and a LIF (Leaky Integrate-and-Fire) neuron and is used for converting static state information into a corresponding pulse sequence. The pulse neural network comprises two layers, each of which is composed of a synapse layer and a LIF neuron layer. The pulse information in a local space is extracted and converted into a pulse representation with abstract features. The synapse layer of the first layer is a convolution layer conv, and the synapse layer of the second time is a fully connected layer FC. The pulse decoder is composed of a fully connected layer FC and is used for converting the pulse sequence into a membrane potential output of each action, which has more rich numerical information and more stable and smooth action probability. The Actor network is used for calculating the action and the policy entropy under the current state. The training process comprises two parts: one is to calculate the corresponding action by using the current state as the input to realize the action selection; and the other is to sample small batches of experience data from the experience replay pool, calculate the action and the policy entropy of the corresponding state, compare them with the Q value, obtain the loss function, and then update the network parameters by using the back propagation. The mathematical model of the LIF neuron is as follows: Charge U t = k v (V t-1 -V r )+ V r + WS' t Discharge reset where U t is the membrane potential after the synaptic stimulus; V t is the spike trigger at time t for the neuron; V r is the membrane potential after the reset; k v is the voltage decay coefficient; W is the weight of other neuron spike stimuli; S t is the output spike of neuron t, S' t is the output spike of other neuron t; V th is the given firing threshold.
6. The AGV path planning method based on the improved SAC model according to claim 5, characterized in that, The Actor network is used to calculate the action and the optimal policy in the policy entropy in the current state, and the formula of the optimal policy with respect to the policy entropy is defined as follows: where s t is the state at the current time t; a t is the action at the current time t; π is the policy adopted by the agent; τ is the state and action corresponding to the policy; γ is the discount factor, γ ∈ [0, 1), which reflects the time value of future rewards; R(s t ,a t ) is the reward obtained by the action a t in the current state s t ; α is the entropy temperature coefficient, α > 0, which balances the reward value and the randomness of the action; is the action distribution probability P corresponding to the current state s t ; φ is the Actor network parameter; and H(P) is a function for calculating the entropy of the probability distribution, and the definition of the entropy of the discrete action space is as follows: where p i (x) is the probability that the variable x corresponds to, and ∑ i p i (x) = 1.
7. The AGV path planning method based on the improved SAC model according to claim 6, characterized in that, The current state corresponding action is obtained by A* algorithm or Actor network, the corresponding reward value is calculated by the reward function set by the environment, and the next state information and task completion flag after execution are obtained, and the state position, action selection, reward value, next state and task completion flag are stored in the experience replay pool.
8. The AGV path planning method based on the improved SAC model according to claim 7, characterized in that, The reward function is set as follows: the reward function reward adopts a hybrid mechanism combining dense reward and sparse reward, and the specific design is as follows: reward = reward1 + reward2 wherein reward1 is a sparse reward setting, if the AGV enters an area outside the warehouse area or collision occurs during operation, it is considered as a violation of operation, and a negative reward is given; if the AGV returns to the position of the previous step, a smaller negative reward is given to avoid repeated movement of the AGV on the path; when the AGV completes the current task and successfully reaches the target position, a positive reward is given to encourage the AGV to move towards the target position; Reward2 is a dense reward, the dense reward adopts a two-dimensional Gaussian function as part of the reward function, forming a peak reward area around the target position, gradually guiding the AGV towards the target, r0 is the negative reward of single-step movement, r max is the maximum reward of reaching the target position, x c , y c is the coordinate of the current position, x t , y t is the coordinate of the target position, and σ is the expansion range for the coordinate axis direction.
9. The AGV path planning method based on the improved SAC model according to claim 8, characterized in that, In step S4, a small batch of data is sampled from the experience replay pool, the current Q value and the target Q value are calculated by using the Critic network and the target Critic network, and the time difference error is calculated to update the Critic network parameters, and the specific process is as follows: the definition of the Critic network with respect to the optimal Q function under the policy entropy is as follows: In the formula, Q * (s t ,a t ) is the action value corresponding to the current state-action pair, the calculation results of Critic network 1 and Critic network 2 are q1 and q2 respectively as the current Q value, the calculation results of target Critic network 1 and target Critic network 2 are q'1 and q'2 respectively, D is an experience replay pool, and R(s t ,a t ) is the reward obtained by the current state s t through the action a t . The minimum value of the calculation results of the two target Critic networks is taken as the target Q value, which is compared with the current Q value calculated by the Critic network to obtain the TD error, the loss function is minimized, and the Critic network parameters are updated; the target Critic network copies the parameters from the Critic network through soft update; The Critic network loss function and soft update are defined as follows: θ i = τθ i + (1 - τ)θ i i = 1, 2 In the formula, θ i is the target Critic network i parameter; τ is a soft update coefficient.
10. The AGV path planning method based on the improved SAC model according to claim 9, characterized in that, In step S5, according to the Q value output by the Critic network, the action generated by the Actor network and the policy entropy, the loss function is calculated, the error function gradient is calculated by using the replacement function back propagation, and the Actor network parameters are updated, and the specific process is as follows: the Q value corresponding to the state is calculated by using the Critic network, the loss value is calculated, the loss function is minimized to update the Actor network parameters, and the loss function is defined as follows: In the formula, is the Q value corresponding to the state calculated by Critic network i (i = 1, 2), is the minimum value of the calculation results of Critic network 1 and Critic network 2. Considering the non-differentiable problem of LIF neurons in back propagation, the replacement function is used to replace the sign function to realize the minimization of loss and the effective update of network parameters, and the inverse tangent function is selected as the replacement function.
Citation Information
Cited By
Value function training method based on model prediction extension
CN121257643A