Improved intelligent vehicle path planning method

By improving the SAC algorithm and using a hybrid reward function, combined with the priority experience replay technique, the problem of slow training speed of robot path planning in unknown or dynamic environments is solved, and efficient path planning results are achieved.

CN120949768APending Publication Date: 2025-11-14CHINA THREE GORGES UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511070325.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing robot path planning methods rely on known data, which cannot meet the needs of rapid path planning in unknown or dynamic environments. In particular, the lack of data support in scenarios such as disaster relief leads to a lack of flexibility and real-time performance in path planning.

Method used

An improved SAC algorithm is used to construct a policy network model. The intelligent vehicle path planning is trained in an unknown environment through reinforcement learning. The hybrid reward function and priority experience replay technique are combined to improve training efficiency and sample utilization efficiency.

Benefits of technology

It accelerates the training process of reinforcement learning algorithms, shortens training time, improves sample utilization efficiency, and achieves efficient path planning in unknown or dynamic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120949768A_ABST
    Figure CN120949768A_ABST
Patent Text Reader

Abstract

The invention discloses an improved intelligent vehicle path planning method, and the algorithm comprises the steps: firstly collecting the parameters of a data training strategy network model, inputting the state information of laser radar data, the position of an intelligent vehicle, the position of a target and the like to a network model for reasoning after the training process is completed, outputting the motion of the intelligent vehicle, and repeating the process until a target point is reached; in order to solve the problems of long training time, low training speed and low sample use efficiency of a reinforcement learning algorithm, the invention designs a mixed reward function for a path planning task, and provides better environment feedback. Meanwhile, an improved priority experience playback technology is added, different weights are given to data, high-quality samples are selected from an experience pool, a probability distribution sampling method is used for replacing a sampling process of a traditional Sum-Tree structure, the sampling speed is higher, the reinforcement learning algorithm training process can be effectively accelerated, the training time is shortened, and the sample utilization efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of robot path planning technology, and in particular relates to an improved intelligent vehicle path planning method. Background Technology

[0002] Robot path planning is a core area of ​​robotics, involving determining a feasible path from a starting point to a target point for a robot in complex and often dynamic environments. In numerous applications such as industrial production, logistics, search and rescue, and healthcare, robots need efficient and safe path planning capabilities to autonomously complete tasks. For example, in logistics warehouses, Automated Guided Vehicles (AGVs) need to navigate through numerous obstacles such as shelves, forklifts, and workers to transport goods from storage areas to delivery areas; in hospital environments, service robots need to deliver medicines, meals, and other items to patients in narrow and high-traffic spaces such as wards, corridors, and elevators. An excellent path planning algorithm enables robots to efficiently find the optimal path in these complex scenarios, avoid obstacles, and complete tasks on time, thereby greatly improving production efficiency and service quality.

[0003] Traditional robot path planning methods primarily rely on manually designed algorithms and models, such as the A* algorithm, Dijkstra's algorithm, and artificial potential field methods. These methods depend on accurate environmental modeling and known information, such as map data and the shape and location of obstacles. However, in practical applications, robots often need to operate in unknown or partially unknown environments. For example, in disaster relief sites, robots need to quickly enter unfamiliar ruins and building interiors to search for people and detect materials. Traditional methods often fail to work effectively in such situations due to a lack of prior knowledge of the environment. Furthermore, traditional methods struggle to adapt to dynamic environmental changes, such as the sudden appearance of obstacles or changes in target location, resulting in a lack of flexibility and real-time performance in robot path planning.

[0004] Reinforcement learning is a machine learning method that learns optimal behavioral strategies through interaction between an agent and its environment, offering a novel approach to robot path planning. In reinforcement learning, an agent can learn, through trial and error in an unknown environment, how to take optimal actions in different states to maximize long-term cumulative rewards. In robot path planning, this means that the robot can use successful arrival at the target point, obstacle avoidance, and energy conservation as reward signals, learning the optimal path planning strategy through interaction with the environment. Reinforcement learning enables robots to adapt to various complex, unknown, and dynamically changing environments without relying on precise environmental models and extensive prior knowledge. This gives robots greater adaptability and autonomy when facing complex tasks; therefore, an improved intelligent vehicle path planning method is needed to address these issues. Summary of the Invention

[0005] The technical problem to be solved by this invention is to provide an improved intelligent vehicle path planning method, which aims to solve the problem that the existing robot path planning methods rely too much on known data for modeling and cannot meet the needs of application scenarios such as rescue and disaster relief where data support is lacking and rapid deployment is required. It has the characteristics of improving the training convergence speed in path planning tasks and shortening the training time.

[0006] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows: An improved intelligent vehicle route planning method includes the following steps: S1, based on the SAC algorithm intelligent agent, constructs a policy network model for the process of the intelligent vehicle moving from any position to the target point; constructs updatable data samples, trains the policy network model, and completes the parameter calibration of the policy network model; S2: By acquiring the parameters of the policy network model, a policy network based on the path planning task is built, and the reinforcement learning model inference process is executed to complete the path planning task.

[0007] Preferably, step S1 specifically includes: S101, Problem Modeling: Model the process of an intelligent vehicle moving from an arbitrary location to a target point as a multi-step Markov decision process; including state representation, reward function, and action design. S101, Network Setup: Build the SAC algorithm intelligent agent, construct the policy network, value network and target value network, and initialize the sample experience pool; S103, Initialize Environment: Initialize the location of the intelligent vehicle, randomly generate target points, and start the round; S104, State Acquisition: At time t, acquire the LiDAR data of the intelligent vehicle. This data covers 36 directions, with adjacent directions spaced 10° apart. This data is recorded in an array. Simultaneously, acquire the coordinate information of the intelligent vehicle and the target point, and integrate the LiDAR data and coordinate information to construct a complete state representation S at time t. t ; S105, Action Acquisition: The constructed state representation S at time t... t The input is fed into the policy network, which outputs the linear velocity and angular velocity at that the intelligent vehicle should take at time t based on the current state St, thereby controlling the motion of the intelligent vehicle. S106, State Update: Obtain the state representation S of the intelligent vehicle at time t+1. t+1 Round end marker "done" t and taking action a at time t. t The reward r tProvide data support for subsequent network update steps; S107, Sample Storage: Construct a data sample, containing (S t ,a t ,r t ,S t+1 ,P t ,done t In this process, the sample priority P is set to 1, indicating that new samples have equal importance in the initial stage. Subsequently, the samples are stored in the experience pool, and S is updated. t For S t+1 ; S108, Sample extraction: A fixed number of sample data are extracted from the experience pool according to the sample priority to form a batch; S109, Network Update: The neural network is updated using the extracted batch samples, and for each data in the batch, its sample priority Pi (i=1,2,3……batch_Size) is re-evaluated and updated to optimize the targeting and effectiveness of subsequent sample extraction. S110, Round End Judgment: Determine if the round has ended. If the round has ended, perform a training end judgment. If the round has not ended, repeat steps S105-S109. S111, Training End Judgment: If the maximum number of training rounds is reached, the training ends; otherwise, repeat steps S103-S109. S112, Save Model Parameters: Save the policy network model parameters and end the entire algorithm training process.

[0008] Preferably, in step S2, the specific steps of the reinforcement learning model inference part are as follows: S201, Begin the path planning task and build the policy network; S202, Read the parameters of the trained network model; S203, State Acquisition: At time t, acquire the LiDAR data of the intelligent vehicle. This data covers 36 directions, with adjacent directions spaced at 10° intervals. This data is recorded in an array. Acquire the coordinate information of the intelligent vehicle and the target point, and integrate the LiDAR data and coordinate information to construct a complete state representation S at time t. t ; S204, Action Acquisition: The constructed state representation S at time t. t The input is fed into the policy network, which then adjusts the input based on the current state S. t Output the linear velocity and angular velocity a that the intelligent vehicle should take at time t. t This enables the control of the movement of intelligent vehicles; S205, State Update: Obtain the state representation S of the intelligent vehicle at time t+1. t+1 Update St For S t+1 ; S206, Arrival Judgment: Based on S t+1 Determine whether the target point has been reached. If not, repeat process S202-S205; if the target point has been reached, end the path planning task.

[0009] Preferably, the state representation in step S101 has 40 dimensions, with the first 36 dimensions being the LiDAR data and the last 4 dimensions being the horizontal and vertical coordinates of the intelligent vehicle and the horizontal and vertical coordinates of the target point; the reward function includes distance reward, arrival reward, collision penalty, and direction reward, and the action dimension is 2, including angular velocity and linear velocity.

[0010] Preferably, the SAC algorithm framework in step S102 contains a total of 3 neural networks, namely 1 policy network and 2 value networks.

[0011] Preferably, after acquiring the lidar data in steps S103 and S105, the data in the direction whose value exceeds a certain threshold is set to a fixed value to simplify the state representation.

[0012] Preferably, the reward in step S105 includes a primary reward and a secondary reward, and the calculation of the primary reward is shown in formula (1): (1) The auxiliary rewards include angle rewards and distance rewards.

[0013] Preferably, the angle reward design is as shown in formula (2): (2) In the formula, angle Let be the radian value of the angle between the robot's orientation and the target, ranging from . ; The angle reward design is shown in formula (3): (3) In the formula, distance Let Euclidean distance be the distance between the robot and the target. This represents the straight-line distance between the robot's starting position and the target point. This represents the straight-line distance between the robot and the target point. Adjust the coefficient for rewards. When the angle If the value equals 0, the maximum reward is obtained.

[0014] Preferably, the total reward is as shown in formula (4); (4) When the robot neither reaches the target nor collides with it, the reward is set to a distance-based reward, meaning the closer the robot is to the target, the smaller the absolute value of the reward.

[0015] Preferably, during data sampling in step S107, a probability distribution is generated based on the sample priority, data is sampled from the probability distribution, and the priority is updated after the network update; the priority calculation is as shown in formula (5): (5) It is a small positive number used to ensure that all experiences are sampled with a non-zero probability; Priority normalization is used to construct a probability distribution, and sampling is performed based on this distribution. The normalization process is shown in formula (6): (6) in It is the sampling probability. The sample priority is calculated by formula (5). It is an adjustable parameter, ranging from [0,1], that controls the degree of influence of priority. If, then all samples are sampled with equal probability; if If so, sampling will be performed entirely according to priority.

[0016] The beneficial effects of this invention are as follows: This invention addresses the problems of long training time, slow training speed, and low sample utilization efficiency in reinforcement learning algorithms. It designs a hybrid reward function for path planning tasks, providing better environmental feedback. Simultaneously, it incorporates an improved priority experience replay technique, assigning different weights to data to select high-quality samples from the experience pool. Furthermore, it replaces the traditional Sum-Tree sampling process with a probability distribution sampling method, achieving a more efficient sampling speed. This invention effectively accelerates the training process of reinforcement learning algorithms, shortens training time, and improves sample utilization efficiency. Attached Figure Description

[0017] Figure 1 This is a network structure diagram of the strategy network and value network of the present invention; Figure 2 This is a flowchart of the improved SAC algorithm training process in an embodiment of the present invention; Figure 3 This is a flowchart of the improved SAC algorithm path planning in an embodiment of the present invention; Figure 4 This is a model diagram of the simulated robot turtlebot3 burger in this embodiment of the invention; Figure 5 These are four simulation environment diagrams constructed in this embodiment of the invention; Figure 6 This is an experimental comparison diagram of scenario a in the embodiments of the present invention; Figure 7 This is an experimental comparison diagram of scenario b in the embodiments of the present invention; Figure 8 This is an experimental comparison diagram of scenario c in the embodiments of the present invention; Figure 9 This is an experimental comparison diagram of scenario d in the embodiments of the present invention. Detailed Implementation

[0018] Example 1: like Figure 1 As shown, an improved intelligent vehicle path planning method includes the following steps: S1, based on the SAC algorithm intelligent agent, constructs a policy network model for the process of the intelligent vehicle moving from any position to the target point; constructs updatable data samples, trains the policy network model, and completes the parameter calibration of the policy network model; S2: By acquiring the parameters of the policy network model, a policy network based on the path planning task is built, and the reinforcement learning model inference process is executed to complete the path planning task.

[0019] Preferably, step S1 specifically includes: S101, Problem Modeling: Model the process of an intelligent vehicle moving from an arbitrary location to a target point as a multi-step Markov decision process; including state representation, reward function, and action design. S101, Network Setup: Build the SAC algorithm intelligent agent, construct the policy network, value network and target value network, and initialize the sample experience pool; S103, Initialize Environment: Initialize the location of the intelligent vehicle, randomly generate target points, and start the round; S104, State Acquisition: At time t, acquire the LiDAR data of the intelligent vehicle. This data covers 36 directions, with adjacent directions spaced 10° apart. This data is recorded in an array. Simultaneously, acquire the coordinate information of the intelligent vehicle and the target point, and integrate the LiDAR data and coordinate information to construct a complete state representation S at time t. t ; S105, Action Acquisition: The constructed state representation S at time t... t The input is fed into the policy network, which outputs the linear velocity and angular velocity at that the intelligent vehicle should take at time t based on the current state St, thereby controlling the motion of the intelligent vehicle. S106, State Update: Obtain the state representation S of the intelligent vehicle at time t+1. t+1 Round end marker "done" t and taking action a at time t.t The reward r t Provide data support for subsequent network update steps; S107, Sample Storage: Construct a data sample, containing (S t ,a t ,r t ,S t+1 ,P t ,done t In this process, the sample priority P is set to 1, indicating that new samples have equal importance in the initial stage. Subsequently, the samples are stored in the experience pool, and S is updated. t For S t+1 ; S108, Sample extraction: A fixed number of sample data are extracted from the experience pool according to the sample priority to form a batch; S109, Network Update: The neural network is updated using the extracted batch samples, and for each data in the batch, its sample priority Pi (i=1,2,3……batch_Size) is re-evaluated and updated to optimize the targeting and effectiveness of subsequent sample extraction. S110, Round End Judgment: Determine if the round has ended. If the round has ended, perform a training end judgment. If the round has not ended, repeat steps S105-S109. S111, Training End Judgment: If the maximum number of training rounds is reached, the training ends; otherwise, repeat steps S103-S109. S112, Save Model Parameters: Save the policy network model parameters and end the entire algorithm training process.

[0020] Preferably, in step S2, the specific steps of the reinforcement learning model inference part are as follows: S201, Begin the path planning task and build the policy network; S202, Read the parameters of the trained network model; S203, State Acquisition: At time t, acquire the LiDAR data of the intelligent vehicle. This data covers 36 directions, with adjacent directions spaced at 10° intervals. This data is recorded in an array. Acquire the coordinate information of the intelligent vehicle and the target point, and integrate the LiDAR data and coordinate information to construct a complete state representation S at time t. t ; S204, Action Acquisition: The constructed state representation S at time t. t The input is fed into the policy network, which then adjusts the input based on the current state S. t Output the linear velocity and angular velocity a that the intelligent vehicle should take at time t. t This enables the control of the movement of intelligent vehicles; S205, State Update: Obtain the state representation S of the intelligent vehicle at time t+1. t+1 Update S t For S t+1 ; S206, Arrival Judgment: Based on S t+1 Determine whether the target point has been reached. If not, repeat process S202-S205; if the target point has been reached, end the path planning task.

[0021] Preferably, the state representation in step S101 has 40 dimensions, with the first 36 dimensions being the LiDAR data and the last 4 dimensions being the horizontal and vertical coordinates of the intelligent vehicle and the horizontal and vertical coordinates of the target point; the reward function includes distance reward, arrival reward, collision penalty, and direction reward, and the action dimension is 2, including angular velocity and linear velocity.

[0022] Preferably, the SAC algorithm framework in step S102 contains a total of 3 neural networks, namely 1 policy network and 2 value networks.

[0023] Preferably, after acquiring the lidar data in steps S103 and S105, the data in the direction whose value exceeds a certain threshold is set to a fixed value to simplify the state representation.

[0024] Preferably, the reward in step S105 includes a primary reward and a secondary reward, and the calculation of the primary reward is shown in formula (1): (1) The auxiliary rewards include angle rewards and distance rewards.

[0025] Preferably, the angle reward design is as shown in formula (2): (2) In the formula, angle Let be the radian value of the angle between the robot's orientation and the target, ranging from . ; The angle reward design is shown in formula (3): (3) In the formula, distance Let Euclidean distance be the distance between the robot and the target. This represents the straight-line distance between the robot's starting position and the target point. This represents the straight-line distance between the robot and the target point. Adjust the coefficient for rewards. When the angle If the value equals 0, the maximum reward is obtained.

[0026] Preferably, the total reward is as shown in formula (4); (4) When the robot neither reaches the target nor collides with it, the reward is set to a distance-based reward, meaning the closer the robot is to the target, the smaller the absolute value of the reward.

[0027] Preferably, during data sampling in step S107, a probability distribution is generated based on the sample priority, data is sampled from the probability distribution, and the priority is updated after the network update; the priority calculation is as shown in formula (5): (5) It is a small positive number used to ensure that all experiences are sampled with a non-zero probability; Priority normalization is used to construct a probability distribution, and sampling is performed based on this distribution. The normalization process is shown in formula (6): (6) in It is the sampling probability. The sample priority is calculated by formula (5). It is an adjustable parameter, ranging from [0,1], that controls the degree of influence of priority. If, then all samples are sampled with equal probability; if If so, sampling will be performed entirely according to priority.

[0028] Furthermore, the target value network has the same structure as the value network.

[0029] Example 2: like Figures 2-3 As shown, the specific technical solution of the present invention is as follows, and the training steps are as follows: S1.1. Problem Modeling: The process of the intelligent vehicle moving from any position to the target point is modeled as a multi-step Markov decision process (MDP); the process has several key elements, namely state space design (40-dimensional), reward function, and action space design (2-dimensional).

[0030] State-space design (40-dimensional): The first 36 dimensions: LiDAR ranging data (covering 360°, 10° intervals), truncated by a threshold; readings outside the detection range are set to a fixed value (MAX_RANGE). The last 4 dimensions: the current coordinates of the intelligent vehicle (x_robot, y_robot) + the coordinates of the target point (x_goal, y_goal) Motion space design (2D): linear velocity ; angular velocity ; S1.2. Network Construction: Construct a SAC framework containing three neural networks: a policy network, a value network, a target value network, and an experience pool (Memory).

[0031] Policy network: Input state S t Output action a t .

[0032] Value network: Output action value Q(s,a).

[0033] Target Value Network: Dual Critic structure (to prevent Q-value overestimation), outputs target action value target_Q(s,a).

[0034] Experience pool: Initial capacity is memory_size, storing samples (S t ,a t ,r t ,S t+1 ,P t ,done t ).

[0035] S1.3 Initialize the environment: This includes robot position, target point position, robot speed, and target point position.

[0036] S1.4. State Acquisition: Real-time acquisition of LiDAR data (36-dimensional) and coordinate information (4-dimensional), stitched together to form state S. t .

[0037] S1.5. Action Acquisition: S t Input policy network generates action a t .

[0038] S1.6. After the intelligent vehicle performs an action in the simulation environment, it obtains the state S at time t+1. t+1 .

[0039] Reward Acquisition: Enter S t+1 Calculate the reward r based on the reward function. t Set the round end flag "done" based on the status. t If a collision occurs, the target is reached, or the timeout occurs, "done" will be displayed. t Set to true, otherwise set to false.

[0040] The reward function is divided into primary reward and secondary reward, and the formula is as follows: The main rewards are shown in formula (1): (1) The auxiliary rewards are shown in formulas (2)-(3): angle : The radian value of the angle between the robot's orientation and the target, ranging from .

[0041] distance : The Euclidean distance between the robot and the target.

[0042] The angle reward design is shown in Formula (2), and the distance reward design based on the potential energy function is shown in Formula (3).

[0043] (2) (3) in, This represents the straight-line distance between the robot's starting position and the target point. This represents the straight-line distance between the robot and the target point. Adjust the coefficient for rewards. When the angle If the result equals 0, the maximum reward is obtained. The total reward is shown in formula (4).

[0044] (4) S1.7. Sample storage: Store sample data in the experience pool.

[0045] Initialization priority P t=1 This ensures that new samples have the same priority.

[0046] Construct a data sample containing (S) t ,a t ,r t ,S t+1 ,P t ,done t ).

[0047] Store the sample in the experience pool (memory), increment the total number of samples in the experience pool (memory_number) by 1, and update S. t For S t+1 .

[0048] S1.8. Sample extraction: Extract sample data.

[0049] The total number of samples in the experience pool (memory_number) must be greater than the specified batch size.

[0050] Construct a probability distribution based on the priority of data in the experience pool.

[0051] A fixed number (batch_size) of samples are repeatedly drawn from a probability distribution to form a batch.

[0052] S1.9. Network update: Perform a network update using batch sample data. Update the strategy network, value network, and target value network using batch data.

[0053] The priority of each sample in this batch is recalculated to optimize the targeting and effectiveness of subsequent sample extraction. The calculation formula is shown in formula (5): (5) It is a small positive number used to ensure that all experiences are sampled with a non-zero probability. .

[0054] S1.10. End of round judgment: If done=true, then proceed to the end of training judgment.

[0055] If done=false, then repeat steps S3-S7.

[0056] S1.11. Training End Judgment: If the maximum number of training rounds has not been reached, proceed to the next round and increment the round number by 1.

[0057] Training ends when the maximum number of training rounds is reached.

[0058] S1.12. Save model parameters: Save the strategy network model parameters as model.pth.

[0059] The entire algorithm training process is now complete.

[0060] The reasoning steps are as follows: S2.1. Network Setup Strategy

[0061] S2.2. Model Parameter Reading: Read the strategy network model parameters model.pth.

[0062] S2.3. State Acquisition: Acquire lidar data (36-dimensional) and coordinate information (4-dimensional), and concatenate them into state information S. t .

[0063] S2.4. Action Acquisition: Input the state information into the policy network P_net and obtain the network inference output a. t .

[0064] S2.5. State update, obtain the state S at the next moment. t+1 .

[0065] S2.6. Arrival Judgment: Task completion judgment, based on status St+1, determines whether the target point has been reached. If the destination is not reached, repeat steps S2.2-S2.5.

[0066] If the target point is reached, the path planning task ends. Example 3: This embodiment provides a specific experimental procedure for applying this method in a practical application, as follows: Experimental environment: The experiment is based on the Linux system, uses the Python language, and is conducted on the Ubuntu 18.04 system.

[0067] Experimental equipment: The training equipment used a 12900kf CPU and an RTX 3060 12G graphics card.

[0068] Simulation Scenario: The physical environment is simulated using ROS Melodic + Gazebo, with the TurtleBot3 burger smart car. The smart car is equipped with LiDAR, which can return real-time information about obstacles in the surrounding area; the smart car model is attached. Figure 4 As shown.

[0069] Algorithm Implementation: The original SAC algorithm and the improved SAC algorithm are implemented using Python and PyTorch.

[0070] The simulation scenario experimental design includes four simulation scenarios (a, b, c, d), as shown in the appendix. Figure 5 As shown.

[0071] The SAC algorithm and the improved SAC algorithm were used respectively for 1000 rounds of training in the simulation environment.

[0072] The maximum number of steps in each round is 500. Starting from the 100th round, the probability of the robot reaching the target, the probability of collision, and the average reward per round are calculated every 100 rounds.

[0073] After training, save the policy network model as policy.pth.

[0074] Comparison of algorithm performance in four scenarios Figures 6-9 As shown.

[0075] A three-line table is used to record data such as the algorithm's success rate, collision rate, and the number of convergence rounds. The three-line table is shown in Table 1 below: Table 1: Algorithm success rate, collision rate, and number of convergence rounds (three-line table);

[0076] The results show that the path planning algorithm based on reinforcement learning can complete the navigation task well in all four scenarios. With the same number of training rounds, the SAC algorithm with improved experience replay starts convergence earlier and requires less total data during training. Since the acquisition time for each data point is the same, less data means less total training time, faster algorithm convergence, and higher sample utilization efficiency.

Claims

1. An improved intelligent vehicle path planning method, characterized in that, Includes the following steps: S1, based on the SAC algorithm intelligent agent, constructs a policy network model for the process of the intelligent vehicle moving from any position to the target point; constructs updatable data samples, trains the policy network model, and completes the parameter calibration of the policy network model; S2: By acquiring the parameters of the policy network model, a policy network based on the path planning task is built, and the reinforcement learning model inference process is executed to complete the path planning task.

2. The improved intelligent vehicle path planning method according to claim 1, characterized in that, Step S1 specifically includes: S101, Problem Modeling: Model the process of an intelligent vehicle moving from an arbitrary location to a target point as a multi-step Markov decision process; including state representation, reward function, and action design. S101, Network Setup: Build the SAC algorithm intelligent agent, construct the policy network, value network and target value network, and initialize the sample experience pool; S103, Initialize Environment: Initialize the location of the intelligent vehicle, randomly generate target points, and start the round; S104, State Acquisition: At time t, acquire the LiDAR data of the intelligent vehicle, and simultaneously acquire the coordinate information of the intelligent vehicle and the target point. Integrate the LiDAR data and coordinate information to construct a complete state representation S at time t. t ; S105, Action Acquisition: The constructed state representation S at time t... t The input is fed into the policy network, which then adjusts the input based on the current state S. t Output the linear velocity and angular velocity a that the intelligent vehicle should take at time t. t ; S106, State Update: Obtain the state representation S of the intelligent vehicle at time t+1. t+1 Round end marker "done" t and taking action a at time t. t The reward r t Provide data support for subsequent network update steps; S107, Sample Storage: Construct a data sample, containing (S t ,a t ,r t ,S t+1 ,P t ,done t In this process, the sample priority P is set to 1, indicating that new samples have equal importance in the initial stage. Subsequently, the samples are stored in the experience pool, and S is updated. t For S t+1 ; S108, Sample extraction: A fixed number of sample data are extracted from the experience pool according to the sample priority to form a batch; S109, Network Update: Update the neural network using the extracted batch samples, and re-evaluate and update the sample priority P for each data point in the batch. i (i=1,2,3……batch_Size); S110, Round End Judgment: Determine if the round has ended. If the round has ended, perform a training end judgment. If the round has not ended, repeat steps S105-S109. S111, Training End Judgment: If the maximum number of training rounds is reached, the training ends; otherwise, repeat steps S103-S109. S112, Save Model Parameters: Save the policy network model parameters and end the entire algorithm training process.

3. The improved intelligent vehicle path planning method according to claim 1, characterized in that, In step S2, the specific steps of the reinforcement learning model inference are as follows: S201, Begin the path planning task and build the policy network; S202, Read the parameters of the trained network model; S203, State Acquisition: At time t, acquire the LiDAR data of the intelligent vehicle; acquire the coordinate information of the intelligent vehicle and the target point, and integrate the LiDAR data and coordinate information to construct a complete state representation S at time t. t ; S204, Action Acquisition: The constructed state representation S at time t. t The input is fed into the policy network, which then adjusts the input based on the current state S. t Output the linear velocity and angular velocity a that the intelligent vehicle should take at time t. t This enables the control of the movement of intelligent vehicles; S205, State Update: Obtain the state representation S of the intelligent vehicle at time t+1. t+1 Update S t For S t+1 ; S206, Arrival Judgment: Based on S t+1 Determine whether the target point has been reached. If not, repeat process S202-S205; if the target point has been reached, end the path planning task.

4. The improved intelligent vehicle path planning method according to claim 1, characterized in that, In step S101, the state representation has a total of 40 dimensions. The first 36 dimensions are the LiDAR data, and the last 4 dimensions are the horizontal and vertical coordinates of the intelligent vehicle and the horizontal and vertical coordinates of the target point. The reward function includes distance reward, arrival reward, collision penalty, and direction reward. The action dimension is 2, which includes angular velocity and linear velocity.

5. An improved intelligent vehicle path planning method according to claim 1, characterized in that, The SAC algorithm framework in step S102 contains a total of 3 neural networks: 1 policy network and 2 value networks.

6. The improved intelligent vehicle path planning method according to claim 1, characterized in that, After acquiring the lidar data in steps S103 and S105, the data in the direction whose value exceeds a certain threshold is set to a fixed value to simplify the state representation.

7. An improved intelligent vehicle path planning method according to claim 1, characterized in that, The rewards in step S105 include primary rewards and secondary rewards. The calculation of the primary rewards is shown in formula (1): ;(1) The auxiliary rewards include angle rewards and distance rewards.

8. An improved intelligent vehicle path planning method according to claim 7, characterized in that, The angle reward design is shown in formula (2): ; (2) In the formula, angle Let be the radian value of the angle between the robot's orientation and the target, ranging from . ; The angle reward design is shown in formula (3): ;(3) In the formula, distance Let Euclidean distance be the distance between the robot and the target. This represents the straight-line distance between the robot's starting position and the target point. This represents the straight-line distance between the robot and the target point. Adjust the coefficient for rewards; when the angle If the value equals 0, the maximum reward is obtained.

9. An improved intelligent vehicle path planning method according to claim 8, characterized in that, The total reward is shown in formula (4); ; (4) When the robot neither reaches the target nor collides with it, the reward is set to a distance-based reward, meaning the closer the robot is to the target, the smaller the absolute value of the reward.

10. An improved intelligent vehicle path planning method according to claim 1, characterized in that, In step S107, during data sampling, a probability distribution is generated based on the sample priority. Data is sampled from the probability distribution, and the priority is updated after the network update. The priority calculation is shown in formula (5): ;(5) It is a small positive number used to ensure that all experiences are sampled with a non-zero probability; Priority normalization is used to construct a probability distribution, and sampling is performed based on this distribution. The normalization process is shown in formula (6): ;(6) in It is the sampling probability. The sample priority is calculated by formula (5). It is an adjustable parameter, ranging from [0,1], that controls the degree of influence of priority. Then all samples are sampled with equal probability; if If so, sampling will be performed entirely according to priority.