Vehicle path planning method and device based on reinforcement learning

The vehicle path planning method designed with a dual network architecture and a multi-layer fully connected neural network solves the problems of overestimation of expected return values ​​and poor training stability in traditional reinforcement learning, achieves improved robustness and stability in high-dimensional state space, and is suitable for complex and changeable vehicle path planning tasks.

CN120800422AActive Publication Date: 2025-10-17SHENZHEN UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511285720.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-10
Publication Date
2025-10-17
Estimated Expiration
2045-09-10

AI Technical Summary

Technical Problem

Existing vehicle path planning methods based on traditional reinforcement learning have problems such as policy deviation and poor training stability caused by overestimation of expected return values. In particular, they are prone to policy fluctuations and convergence difficulties in high-dimensional state spaces, and lack robustness.

Method used

A dual-network architecture approach is adopted, with the decoupling design of the first model and the second model to perform action selection and expected reward value evaluation respectively. Combined with exploration rate control and simulated annealing mechanism, a multi-layer fully connected neural network is constructed and the Dropout mechanism is introduced. A variety of path planning operators are designed to adapt to complex and changeable vehicle path planning scenarios.

Benefits of technology

It effectively suppresses the overestimation of expected return values, improves the learning stability and convergence quality of the vehicle path planning strategy, ensures robustness in high-dimensional state space, and adapts to complex and diverse vehicle path planning tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120800422A_ABST
    Figure CN120800422A_ABST
Patent Text Reader

Abstract

The invention provides a vehicle path planning method and device based on reinforcement learning, and relates to the technical field of data processing, and the method comprises the steps: inputting a state vector of a t time step into a first model, and obtaining a first expected profit value corresponding to a first sample action; selecting a target first sample action, and updating to obtain a state vector of a t + 1 time step; inputting the state vector of the t + 1 time step into a second model, and obtaining a plurality of second sample actions and corresponding second expected profit values; determining a target expected profit value based on the second expected profit value, determining training loss based on the target expected profit value and a first expected profit value corresponding to the target first sample action, and updating parameters of the first model based on the training loss; performing soft updating on the parameters of the second model based on the parameters of the first model after a plurality of time steps; and obtaining a vehicle path planning result based on the output data of the trained first model. According to the invention, the robustness of vehicle path planning can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, in particular to a vehicle path planning method and device based on reinforcement learning. BACKGROUND

[0002] Reinforcement learning is a sequential decision-making method based on the interaction between an agent and an environment. The agent learns a strategy through continuous exploration and feedback to maximize its long-term cumulative reward. Reinforcement learning has been applied to solve vehicle path planning problems.

[0003] In actual application scenarios, vehicle path planning problems are often complex and changeable. However, in existing vehicle path planning methods based on traditional reinforcement learning, there is often a problem of strategy bias caused by overestimation of expected return value (Q value), and there is a lack of training stability guarantee mechanism in high-dimensional state space, which is prone to strategy fluctuations and convergence difficulties. The robustness of existing vehicle path planning methods is poor. SUMMARY

[0004] The present application provides a vehicle path planning method and device based on reinforcement learning to solve the problem of poor robustness of existing vehicle path planning methods and improve the robustness of vehicle path planning methods.

[0005] The present application provides a vehicle path planning method based on reinforcement learning, comprising: Inputting a state vector at time step t into a first model to obtain first expected return values corresponding to a plurality of first sample actions output by the first model, each first sample action corresponding to a call to a path planning operator, and the state vector including a sequence of operators that have been called; Selecting one of the plurality of first sample actions as a target first sample action, updating the state vector based on the target first sample action to obtain the state vector at time step t+1; Inputting the state vector at time step t+1 into a second model to obtain a plurality of second sample actions output by the second model and second expected return values corresponding to each second sample action; Determining a target expected return value based on the second expected return values, determining a training loss based on the target expected return value and the first expected return value corresponding to the target first sample action, and updating parameters of the first model based on the training loss; After a plurality of time steps, soft updating parameters of the second model based on the parameters of the first model; Obtaining a vehicle path planning result based on output data of the first model after training.

[0006] According to the vehicle path planning method based on reinforcement learning provided by the application, each path planning operator corresponds to a path adjustment operation, each path planning operator is associated with a parameter k, and the parameter k is used for controlling the size of the path adjustment operation corresponding to the path planning operator.

[0007] According to the vehicle path planning method based on reinforcement learning provided by the application, the selecting one action as a target first sample action from the plurality of first sample actions comprises: Based on a first formula, the exploration rate is determined, and the first formula is: Wherein, The exploration rate corresponding to the t time step is represented as The exploration rate reference value is represented as T, T is an adjustment parameter, and the size of T is negatively related to the size of the time step. The target first sample action is selected from the plurality of first sample actions based on the exploration rate.

[0008] According to the vehicle path planning method based on reinforcement learning provided by the application, the determining a target expected return value based on the second expected return value comprises: The second sample action corresponding to the maximum second expected return value is obtained as a target second sample action; Based on the effect value of the path planning scheme corresponding to the plurality of second sample actions, a reward value of a t time step is determined; The target expected return value is determined based on the reward value of the t time step.

[0009] According to the vehicle path planning method based on reinforcement learning provided by the application, the determining a reward value of a t time step based on the effect value of the path planning scheme corresponding to the plurality of second sample actions comprises: The maximum value and the average value of the effect improvement values of the path planning schemes corresponding to the plurality of second sample actions are obtained, An reward change trend value is obtained, and the reward change trend value reflects the change trend of the reward values of the latest N time steps; The reward value of the t time step is determined based on the maximum value, the average value and the reward change trend value.

[0010] According to the vehicle path planning method based on reinforcement learning provided by the application, the first model is a multi-layer fully connected neural network, and each hidden layer in the first model has a Dropout mechanism.

[0011] The application further provides a vehicle path planning device based on reinforcement learning, comprising: The first expected return obtaining module is configured to input a state vector at a t time step into a first model to obtain first expected return values corresponding to a plurality of first sample actions output by the first model, each of the first sample actions corresponding to calling a path planning operator, and the state vector including a sequence of operators that have been called; The action selection module is configured to select one of the plurality of first sample actions as a target first sample action, update the state vector based on the target first sample action to obtain the state vector at a t+1 time step; The second expected return obtaining module is configured to input the state vector at the t+1 time step into a second model to obtain a plurality of second sample actions output by the second model and second expected return values corresponding to the second sample actions; The target expected return determining module is configured to determine a target expected return value based on the second expected return values, determine a training loss based on the target expected return value and the first expected return value corresponding to the target first sample action, and update parameters of the first model based on the training loss; The soft updating module is configured to perform soft updating of parameters of the second model based on the parameters of the first model after a plurality of time steps; The result output module is configured to obtain a vehicle path planning result based on output data of the first model after training.

[0012] The application further provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the vehicle path planning method based on reinforcement learning according to any one of the above when executing the computer program.

[0013] The application further provides a non-transitory computer-readable storage medium having a computer program stored thereon, and the computer program implements the vehicle path planning method based on reinforcement learning according to any one of the above when executed by a processor.

[0014] The application further provides a computer program product including a computer program, and the computer program implements the vehicle path planning method based on reinforcement learning according to any one of the above when executed by a processor.

[0015] The vehicle path planning method and device based on reinforcement learning provided by the application can effectively inhibit the overestimation of expected return values in traditional reinforcement learning by decoupling action selection and expected return value evaluation through the introduction of a double-network architecture of a first model and a second model, thereby improving the stability and convergence quality of vehicle path planning strategy learning and ensuring robustness in a high-dimensional state space. BRIEF DESCRIPTION OF DRAWINGS

[0016] In order to make the technical solutions in the present application or prior art clearer, the drawings needed to be used in the embodiments or prior art description will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present application, and all other embodiments obtained by those of ordinary skill in the art without creative labor based on these drawings also belong to the protection scope of the present application.

[0017] Figure 1 is a flowchart of the vehicle path planning method based on reinforcement learning provided by the present application.

[0018] Figure 2 is a schematic diagram of path planning operator types and influences in the vehicle path planning method based on reinforcement learning provided by the present application.

[0019] Figure 3 is a block diagram of the model training process in the vehicle path planning method based on reinforcement learning provided by the present application.

[0020] Figure 4 is an effect diagram of the vehicle path planning method based on reinforcement learning provided by the present application. Figure 1 .

[0021] Figure 5 is an effect diagram of the vehicle path planning method based on reinforcement learning provided by the present application. Figure 2 .

[0022] Figure 6 is an effect diagram of the vehicle path planning method based on reinforcement learning provided by the present application. Figure 3 .

[0023] Figure 7 is a structural diagram of the vehicle path planning device based on reinforcement learning provided by the present application.

[0024] Figure 8 is a structural diagram of the electronic device provided by the present application. DETAILED DESCRIPTION

[0025] In order to make the technical solutions in the present application or prior art clearer, the drawings needed to be used in the embodiments or prior art description will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present application, and all other embodiments obtained by those of ordinary skill in the art without creative labor based on these drawings also belong to the protection scope of the present application.

[0026] It should be understood that the word "comprising" when used in the specification and claims herein, specifies the presence of stated features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0027] It should also be understood that the terminology used in the description of the application herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. As used in this description and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise.

[0028] It should further be understood that the term "and / or" as used in the specification and in the claims, means any one of the items, any combination of the items, and all possible combinations of the items in the related list.

[0029] As used in this description and the following claims, the term "if' can be interpreted as meaning "when," or "once," or "in response to determining," or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [a described condition or event] is detected" can be interpreted as meaning "once it is determined" or "in response to determining" or "once [the described condition or event] is detected" or "in response to detecting [the described condition or event]," depending on the context.

[0030] The following description is provided in connection with Figures 1-6 A reinforcement learning-based vehicle path planning method is described. As shown in Figure 1 The reinforcement learning-based vehicle path planning method includes the steps of: S110, inputting a state vector at a t time step into a first model to obtain first expected return values corresponding to a plurality of first sample actions output by the first model, each first sample action corresponding to calling a path planning operator, and the state vector including a sequence of operators that have been called; S120, selecting one of the plurality of first sample actions as a target first sample action, updating the state vector based on the target first sample action, and obtaining a state vector at a t+1 time step; S130, inputting the state vector at the t+1 time step into a second model to obtain a plurality of second sample actions output by the second model and second expected return values corresponding to each second sample action; S140, determining a target expected return value based on the second expected return values, determining a training loss based on the target expected return value and a first expected return value corresponding to the target first sample action, and updating parameters of the first model based on the training loss; S150, after a plurality of time steps, soft update parameters of the second model based on parameters of the first model; S160, obtaining a vehicle path planning result based on output data of the first model after training is completed.

[0031] The vehicle path planning method based on reinforcement learning provided in the application can effectively inhibit the overestimation of expected return value in traditional reinforcement learning by decoupling action selection and expected return value evaluation through the introduction of the dual network architecture of the first model and the second model, thereby improving the stability and convergence quality of vehicle path planning strategy learning and ensuring robustness in a high-dimensional state space.

[0032] In the vehicle path planning method based on reinforcement learning provided in the application, each path planning operator corresponds to a path adjustment operation, and different targeted operators are designed for vehicle path planning to realize learning of complex and diverse vehicle path planning strategies.

[0033] More specifically, the vehicle path planning method based on reinforcement learning provided in the application is for vehicle path planning for performing a delivery task, and in the vehicle path planning method for performing a delivery task, there are multiple roles: a customer, a distribution center, and a vehicle. The vehicle needs to deliver goods from the distribution center to the customer. Before training the first model and the second model, the vehicle path planning method based on reinforcement learning provided in the application first performs data preparation and environment state definition, and can use an existing standard VRPTW (vehicle routing problem with time windows) benchmark data set (for example, Solomon, Gehring-Homberger data set, etc.), which covers multiple customer distribution scenarios such as C, R, and RC. Each problem instance contains customer coordinates, service time length, demand volume, vehicle capacity, etc. For each instance, the current solution state can be saved, including: customer composition of each path, residual capacity of each vehicle, current path length, cost of each path, and other information.

[0034] In the method provided in the present application, a plurality of highly adaptive heuristic operators are constructed for the vehicle routing problem with time window (VRPTW), and three levels of design are respectively designed for intra-path, inter-path and overall structure optimization, which can effectively deal with key problems such as time window restriction, capacity constraint and path scheduling. In the method provided in the present application, the action space is composed of two path planning operators: operator operation in the distribution center and operator operation between distribution centers. Specifically, the operator operation in the distribution center refers to the vehicle path adjustment operation of the delivery task of the customers allocated to the same distribution center, and the operator operation between distribution centers refers to the vehicle path adjustment operation of the delivery task of the customers corresponding to different distribution centers. The path planning mode constructed for processing between distribution centers can be expanded from the operation level of the path to the level of the distribution center, and can be adapted to more complex multi-depot vehicle routing problem (MDVRP).

[0035] As shown in Figure 2 , the operator operation in the distribution center includes the following: 1. xchg_in: intra-path exchange of consecutive k customers; 2. ins_in: intra-path removal and reinsertion of consecutive k customers; 3. xchg_bw: inter-path exchange of consecutive k customers; 4. ins_bw: inter-path removal and reinsertion of consecutive k customers; 5. 2opt : merging and reconstructing two paths; 6. ruin_recreate: destroying and reconstructing path structure.

[0036] The operator operation between distribution centers includes the following: 1. reassign: reassigning k consecutive customers to different distribution centers; 2. xchg_out: inter-path exchange of consecutive k customers between different distribution centers; 3. ins_out: inter-path removal and reinsertion of consecutive k customers between different distribution centers; 4. mer_spl: merging two paths of different distribution centers and re-splitting; 5. no_op: no operation.

[0037] The specific path adjustment operation corresponding to each path planning operator described above is as follows.

[0038] xchg_in (in-path exchange): exchange consecutive k customers in a single path, mainly used to optimize local access order, compress service time span, and relieve time window conflicts.

[0039] ins_in (in-path insertion): remove consecutive k customers from a path and reinsert them at new positions, used to adjust customer order, eliminate local overload or latency, and improve path feasibility and compactness.

[0040] xchg_bw (between-path exchange): select consecutive k customers from each of two paths and exchange their positions, coordinate load and time window tension between paths, and enhance global balance.

[0041] ins_bw (between-path insertion): remove consecutive k customers from one path and insert them into another path, suitable for resource imbalance or single-path overload scenarios, and assist in transferring tasks between paths.

[0042] 2opt : merge two paths into one and then reassign customers and build paths, improve path structure quality through overall rearrangement, often used to escape local optimality.

[0043] ruin_recreate (ruin_recreate): destroy part of the path structure from the current solution and then rebuild it, suitable for global search phase, reshape low-quality paths, and improve solution space diversity.

[0044] reassign (cross-center reassignment): select consecutive k customers from one distribution center and reassign them to other distribution centers, then reconstruct the path according to the feasibility of the receiving center. This operation helps to break the regional distribution rigidity, relieve local resource pressure, and improve overall distribution flexibility and adaptability.

[0045] xchg_out (cross-center exchange): select consecutive k customers from two different distribution centers and exchange them, used to optimize cross-regional resource utilization efficiency, coordinate task load between centers, and improve global service quality.

[0046] ins_out (cross-center insertion): remove consecutive k customers from a path in a distribution center and insert them into a path in another center, suitable for scenarios where some distribution centers have resource bottlenecks or customer geographic distribution is uneven, effectively improving overall path feasibility and distribution efficiency.

[0047] mer_spl (merge-split): extract one path from each of the two distribution centers, merge them and then split them into two new paths, and assign them to the original distribution centers. This operation reconstructs the customer distribution structure by path reorganization, enhances local rearrangement ability, helps to break the disadvantages of the original path structure, and improves the quality of the overall solution.

[0048] no_op (null operation): keep the current solution unchanged and do not perform any path adjustment. It is used for strategy control to keep the current solution stable, and when the estimated income is low or the disturbance is not conducive to convergence, it avoids invalid operations to improve the robustness of the strategy.

[0049] Further, in the method provided by the application, for each path planning operator, there is an associated parameter k, which is used to control the size of the path adjustment operation corresponding to the path planning operator. Each operator introduces a parameter k to control the operation size (such as the number of exchanged customers), and by constructing a combination of [operator x k], a complete action space is obtained. For example, xchg_in (k=2) represents exchanging 2 consecutive customers in the path, and ins_bw (k=3) represents inserting 3 customers between paths. Such combinations enable the strategy to flexibly switch between large disturbances and small adjustments.

[0050] By controlling the size of the path adjustment operation corresponding to the operator, the path time window and vehicle capacity constraints are explicitly considered in the operator design, ensuring that the action has structural feasibility in state updating. Moreover, in the method provided by the application, the above-mentioned operators are not only covered by the multi-level operations inside and outside the distribution center and the path, but also have the k value parameter built-in to control the operation size, which can adjust the disturbance intensity according to the task state, and has good adjustability and adaptability. The application not only supports intelligent selection of operator categories, but also first introduces the internal parameters (such as k value) of the operator into the automatic adjustment system, combines the simulated annealing mechanism to realize dynamic control of search intensity, and achieves a better balance between global exploration and local optimization. By jointly encoding the path planning operator and the adjustable parameter k value, a structured action space is constructed, each action not only corresponds to an operation type, but also corresponds to a disturbance size, so that the agent can learn "what operation" and "how much operation" in the process of strategy learning.

[0051] In the method provided in the application, according to the instance size of the solved standard set, the fixed network length is designed, the first model and the second model are policy networks, which will select the next operator according to the selection of the historical operator and the reward brought by the operation. The first model is the main network, and the second model is the target network. The network structures of the first model and the second model are consistent. The main body is a three-layer fully connected neural network, the number of nodes in each layer is determined according to the size of the vehicle path planning problem, the activation function adopts ReLU, and the Dropout mechanism is added to each hidden layer to suppress overfitting. The output is consistent with the size of the action space, and the expected return value corresponding to each operator combination is output. The main network and the target network are decoupled for action selection and evaluation, which effectively alleviates the problem of overestimation of Q value and improves the stability of policy convergence. Through the end-to-end joint learning mechanism, the overall framework realizes the cooperative optimization of search strategy and operator parameters, significantly improves the adaptability and self-evolution ability of the algorithm, reduces the dependence on artificial experience, and is more suitable for complex and variable actual application scenarios.

[0052] In the method provided in the application, during the training and testing of the reinforcement learning model, the current path state is first represented as a vector form input. According to the size of the problem, a sufficient state space is defined for convergence to a local feasible solution. The state space mainly focuses on the selection of historical actions, and the selection strategy of the action is learned, while the remaining capacity, path distance, service time window and other conditions are uniformly processed in the operation of the operator. When training the first model and the second model based on reinforcement learning, after each action execution, the state transition is recorded and stored in the experience replay buffer.

[0053] Among them, (state at time t): the state at time t. It represents an observation of the environment, which in the application represents a sequence of operators that have been called.

[0054] (action at time t): the action selected by the agent under the state In the application scenario, it refers to the operator that can be selected.

[0055] (reward at time t): the feedback given by the environment to the agent after executing the action It is usually a scalar value that represents the "goodness" of the current action. In the application scenario, it refers to whether the optimization effect of the solution has improved.

[0056] (reward at next state): the feedback given by the environment to the agent after executing the action After, the environment moves to a new state. This embodies the system dynamics, i.e. actions cause changes in states.

[0057] During training, as shown in the figure, small batches of data are periodically sampled from the experience pool, the first model outputs the expected return value at the current time step, the second model outputs the target expected return value, and the model is updated based on the outputs of the two models to realize the training of the model. Figure 3

[0058] The method provided in the application constructs a neural network architecture that can learn historical operator sequence information, designs a three-layer fully connected network structure, and introduces a ReLU activation function and a Dropout suppression mechanism, thereby enhancing the stability of policy convergence, and significantly improving the expression ability of problem structure and the policy generalization performance.

[0059] Specifically, after the first model receives the input state vector, it outputs a Q value vector with the same dimension as the action space, and each dimension represents the expected return value of a certain perturbation operator (such as xchg_in(k=2)) at the current state. That is, the first model outputs multiple first sample actions and the first expected return values corresponding to each first sample action. Based on the output of the first model, an action is selected as the target first sample action, and the action is executed on the current path solution to generate a new path solution, and the current state vector is updated to obtain the state vector at time step t+1.

[0060] Specifically, the action is selected as the target first sample action from the multiple first sample actions, which can be selected by using a greedy strategy or an epsilon-greedy strategy. In one possible implementation, the action is selected as the target first sample action from the multiple first sample actions, including: Based on the first formula, the exploration rate is determined, and the first formula is: wherein, represents the exploration rate corresponding to time step t, represents an exploration rate reference value, and T is an adjustment parameter, and the size of T is negatively related to the size of the time step; The target first sample action is selected from the multiple first sample actions based on the exploration rate.

[0061] ​In this implementation, the epsilon-greedy strategy is used for initial action exploration, and the simulated annealing mechanism is introduced to adjust the exploration intensity. In the initial training stage, a higher temperature T is set to encourage large-scale perturbation (larger operator size k) to promote global search. As the training round progresses, the temperature gradually decreases, making the perturbation tend to be local fine-tuning (smaller k). The exploration rate epsilon and the temperature T are adjusted through the annealing function linkage mechanism, which combines the simulated annealing idea to maintain a high temperature (high perturbation) in the initial training stage to promote solution space exploration, and gradually reduce the perturbation amplitude in the later stage to guide the strategy to fine search convergence, realizing the linkage of search intensity and strategy evolution stage, while maintaining the stability of the strategy and maintaining a moderate exploration ability.

[0062] The method provided in the application embeds the operator category and the parameter (such as the exchange number k value) in the action space together, and designs a temperature control mechanism based on the simulated annealing idea, realizes the control strategy of gradually converging the parameter value (the perturbation scale of the problem) with the training stage, thereby enhancing the adaptability and search efficiency of the algorithm in different optimization stages.

[0063] After determining the target first sample action and obtaining the state vector of the t+1 time step based on the target first sample action, the state vector of the t+1 time step is input into the second model. The second model is the same as the first model, and outputs a Q value vector with the same dimension as the action space. Each dimension represents the expected return value of a certain perturbation operator (such as xchg_in(k=2)) under the current state (the state vector of the t+1 time step). The expected return value output by the second model is called the second expected return value.

[0064] The target expected return value is determined based on the second expected return value, specifically including: Obtaining the second sample action corresponding to the maximum second expected return value as the target second sample action; Determining the reward value of the t time step based on the effect value of the path planning scheme corresponding to the plurality of second sample actions; Determining the target expected return value based on the reward value of the t time step.

[0065] The calculation formula of the target expected return value can be expressed as: ; Wherein, is the target expected return value, r is the reward value of the t time step, is a discount factor (discount factor) between 0 and 1, indicating the importance of future rewards, is the state vector of the t+1 time step, represents the expected return value of the first model predicting the execution of the target second sample action a in the state , representing evaluating the Q value using the second model in state.

[0066] The reward value of the t time step is determined based on the effect value of the path planning scheme corresponding to the plurality of second sample actions, including: Obtaining the maximum value and the average value of the effect improvement value of the path planning scheme corresponding to the plurality of second sample actions, Obtaining a reward change trend value, the reward change trend value reflecting the change trend of the reward value of the latest N time steps; Determine the reward value of the t time step based on the maximum value, the average value and the reward change trend value.

[0067] The effect value of the path planning scheme reflects the performance of the path planning scheme on the preset indicators (such as path length, transportation cost, delivery time, etc.). Specifically, in order to balance short-term improvement and long-term convergence effect, the reward function is designed as follows: ; wherein, and respectively represent the maximum and average improvement values of all solution actions in the current step; T is a reward change trend value calculated based on the last 50 training; if the historical data is sufficient, w=0.3, otherwise 0. All reward components are standardized to the range of [-1, 1] to enhance the stability of training. This reward function encourages robust long-term convergence and avoids strategies falling into short-sighted optimization, especially suitable for the later stage of training.

[0068] Based on the difference between the first expected revenue value corresponding to the target first sample action and the target expected revenue value, a training loss is determined, and the parameters of the first model are updated based on the training loss. Specifically, the training loss can be represented by the formula: ; wherein L is the training loss, is the target expected revenue value, represents the first expected revenue value corresponding to the target first sample action output by the first model based on the state vector of the t time.

[0069] Based on the training loss, the parameters of the first model are updated by back propagation, and the parameters of the first model are soft updated to the second model every fixed step number, to maintain the stability of the second model.

[0070] The training process continues until a custom stopping condition is met, such as reaching a maximum number of runs or convergence of the solution quality. The final output includes the currently found optimal solution (the path with the best performance), the applied operator sequence, the corresponding reward trajectory, and the final trained parameters of the main network. The currently found optimal solution can be used as the path planning result. It can also be used as a reference for other path planning problems based on the operator sequence and corresponding reward trajectory applied during training. This method can achieve adaptive operator selection by learning from historical experience and generalize to solution sets that are difficult to achieve in complex optimization problems.

[0071] In order to verify the superiority of the vehicle path planning method based on reinforcement learning provided by this application, an experimental verification was carried out. The experiment was based on multiple groups of instances, all of which belong to the Solomon benchmark dataset. Each instance is represented by letters and numbers. The letter instance categories include C, R, and RC instances. Class C is an instance in which customer point locations are clustered, class R is an instance in which customer points are uniformly randomly distributed, and class RC is an instance in which some customer points are clustered and some are randomly distributed. The first digit "1 / 2" in the number represents the time window / capacity group. Series 1 indicates a tight time window and a small vehicle capacity, and series 2 indicates a wide time window and a large vehicle capacity. The other digits of the number represent the instance number, for example, R111 indicates that the instance is an R-class instance, series 1 (tight time window, small capacity) is numbered 11, C208 indicates that the instance is a C-class instance, and series 2 (wide time window, large capacity) is numbered 08.

[0072] Table 1 compares the solution (path planning result) found by the reinforcement learning-based vehicle path planning method proposed in this application on a standard VRPTW instance with the known optimal solution. The results include the target value, number of vehicles used, and total path length, and calculate the relative deviation between the target value and path length. Experimental results show that in concentrated instances (such as C102 and C109), the proposed method can achieve a near-optimal number of vehicles with a small deviation from the target value. In instances with more complex distributions (such as R111 and RC206), the deviation is relatively large, but in most cases the number of vehicles used remains consistent with the optimal solution, demonstrating the strong capabilities of the proposed method in vehicle scheduling.

[0073] Table 1

[0074] The results of comparing the method provided in the application with the policy results of the random selection operator are shown in Table 2. It can be seen from the results that the method provided in the application is significantly better than the random method in all examples, and in particular, in examples such as C109 and RC105, the path length is shortened by more than 20%, which reflects the good utilization ability of the method for the spatial structure. Even in the case of the same number of vehicles, the path quality of the random method is still significantly worse than the method provided in the application, which shows that the method provided in the application achieves a better balance between cost and the number of vehicles, and shows stronger robustness.

[0075] Table 2

[0076] As Figures 4 to 6 shown, the training convergence processes of the method provided in the application and other reinforcement learning methods (DQN, PPO, Q-learning) in four representative examples in C, R and RC are respectively shown. The yellow dotted line in the figure represents the method provided in the application, the blue solid line represents the Q-learning method, the green dashed line represents the DQN method, and the red dashed line represents the PPO method. The horizontal coordinate is the training batch, and the vertical coordinate is the action value.

[0077] It can be seen from Figures 4-6 : the C class (centralized distribution example): the method provided in the application converges quickly and has the best final result; the DQN performs well; the PPO and Q-Learning converge slowly and are unstable; the R and RC classes (random distribution and mixed distribution examples): although the examples are more complex, the method provided in the application still shows faster convergence speed and better solution quality, and other methods generally have large fluctuations and early convergence problems. These results further verify the stability and effectiveness of the method provided in the application in various scenarios.

[0078] The reinforcement learning-based vehicle path planning device provided in the application is described below. The reinforcement learning-based vehicle path planning device described below can be referred to in correspondence with the reinforcement learning-based vehicle path planning method described above. As Figure 7 shown, the reinforcement learning-based vehicle path planning device provided in the application comprises: A first expected return acquisition module 710 is configured to input a state vector at a t time step into a first model, and acquire first expected return values corresponding to a plurality of first sample actions output by the first model, each first sample action corresponding to calling one path planning operator, and the state vector comprising a sequence of operators that have been called. The action selection module 720 is configured to select one action as a target first sample action from the plurality of first sample actions, update the state vector based on the target first sample action to obtain a state vector at a t+1 time step; The second expected return acquisition module 730 is configured to input the state vector at the t+1 time step into the second model, and acquire a plurality of second sample actions output by the second model and a second expected return value corresponding to each second sample action. The target expected return determination module 740 is configured to determine a target expected return value based on the second expected return value, determine a training loss based on the target expected return value and a first expected return value corresponding to the target first sample action, and update parameters of the first model based on the training loss. The soft update module 750 is configured to perform soft update on parameters of the second model based on the parameters of the first model after a plurality of time steps. The result output module 760 is configured to obtain a vehicle path planning result based on output data of the first model after training.

[0079] Figure 8 An example of an entity structure diagram of an electronic device is shown in FIG. 8. Figure 8 As shown in FIG. 8, the electronic device can include a processor 810, a communications interface 820, a memory 830, and a communications bus 840, wherein the processor 810, the communications interface 820, and the memory 830 can communicate with each other through the communications bus 840. The processor 810 can invoke a logical instruction in the memory 830 to execute a vehicle path planning method based on reinforcement learning. The vehicle path planning method based on reinforcement learning includes: inputting a state vector at a t time step into a first model, acquiring a first expected return value corresponding to a plurality of first sample actions output by the first model, each first sample action corresponding to invoking a path planning operator, and the state vector including a sequence of operators that have been invoked; selecting one action as a target first sample action from the plurality of first sample actions, updating the state vector based on the target first sample action to obtain a state vector at a t+1 time step; inputting the state vector at the t+1 time step into a second model, acquiring a plurality of second sample actions output by the second model and a second expected return value corresponding to each second sample action; determining a target expected return value based on the second expected return value, determining a training loss based on the target expected return value and a first expected return value corresponding to the target first sample action, and updating parameters of the first model based on the training loss; performing soft update on parameters of the second model based on the parameters of the first model after a plurality of time steps; and obtaining a vehicle path planning result based on output data of the first model after training.

[0080] In addition, the logic instructions in the memory 830 described above can be implemented in the form of software functional units and sold or used as independent products, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the parts that contribute to the prior art or parts of the technical solutions can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the method described in various embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0081] In another aspect, the present application also provides a computer program product, which comprises a computer program, the computer program can be stored on a non-transitory computer readable storage medium, and the computer program can be executed by a processor, so that the computer can execute the vehicle path planning method based on reinforcement learning provided by the above-mentioned method. The vehicle path planning method based on reinforcement learning comprises: inputting a state vector at t time step into a first model, obtaining a first expected return value corresponding to a plurality of first sample actions output by the first model, each first sample action corresponds to calling a path planning operator, and the state vector includes a sequence of operators that have been called; selecting one action as a target first sample action from the plurality of first sample actions, updating the state vector based on the target first sample action to obtain a state vector at t+1 time step; inputting the state vector at t+1 time step into a second model, obtaining a plurality of second sample actions output by the second model and a second expected return value corresponding to each second sample action; determining a target expected return value based on the second expected return value, determining a training loss based on the target expected return value and a first expected return value corresponding to the target first sample action, updating parameters of the first model based on the training loss; after a plurality of time steps, soft updating parameters of the second model based on the parameters of the first model; obtaining a vehicle path planning result based on the output data of the first model after the training is completed.

[0082] In yet another aspect, the present application also provides a non-transitory computer-readable storage medium having stored thereon a computer program, which, when executed by a processor, implements the reinforcement learning-based vehicle path planning method provided by the above method. The reinforcement learning-based vehicle path planning method comprises: inputting a state vector at a t time step into a first model to obtain first expected return values corresponding to a plurality of first sample actions output by the first model, each first sample action corresponding to calling a path planning operator, and the state vector comprising a sequence of operators that have been called; selecting one action as a target first sample action from the plurality of first sample actions, updating the state vector based on the target first sample action to obtain a state vector at a t+1 time step; inputting the state vector at the t+1 time step into a second model to obtain a plurality of second sample actions output by the second model and second expected return values corresponding to each second sample action; determining a target expected return value based on the second expected return values, determining a training loss based on the target expected return value and a first expected return value corresponding to the target first sample action, and updating parameters of the first model based on the training loss; after a plurality of time steps, performing soft updating of parameters of the second model based on the parameters of the first model; and obtaining a vehicle path planning result based on output data of the first model after training is completed.

[0083] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e., can be located in one place or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiment scheme according to actual needs. Those skilled in the art can understand and implement without creative labor.

[0084] From the above description of the embodiments, those skilled in the art can clearly understand that the embodiments can be implemented by means of software plus necessary universal hardware platforms, and of course can also be implemented by hardware. Based on such understanding, the above technical solutions, essentially or in other words, the part that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.

[0085] It should be pointed out finally that the above embodiments are only used to illustrate the technical solutions of the present application, but not to limit the same; and although the present application has been described in detail with reference to the foregoing embodiments, it should be appreciated by those skilled in the art that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features thereof can be replaced equivalently; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A vehicle path planning method based on reinforcement learning, characterized in that: include: Inputting the state vector of time step t into the first model, obtaining first expected return values ​​corresponding to multiple first sample actions output by the first model, each of which corresponds to calling a path planning operator, and the state vector includes a sequence of called operators; Selecting an action from the multiple first sample actions as a target first sample action, and updating the state vector based on the target first sample action to obtain the state vector at time step t+1; Inputting the state vector at time step t+1 into the second model, obtaining a plurality of second sample actions output by the second model and a second expected return value corresponding to each second sample action; determining a target expected return value based on the second expected return value, determining a training loss based on the target expected return value and the first expected return value corresponding to the target first sample action, and updating parameters of the first model based on the training loss; After a plurality of time steps, soft-update the parameters of the second model based on the parameters of the first model; Based on the output data of the first model after training, a vehicle path planning result is obtained.

2. The vehicle path planning method based on reinforcement learning according to claim 1, characterized in that: Each of the path planning operators corresponds to a path adjustment operation, and each of the path planning operators is associated with a parameter k, where the parameter k is used to control the scale of the path adjustment operation corresponding to the path planning operator.

3. The vehicle path planning method based on reinforcement learning according to claim 1, characterized in that: The selecting one action from the plurality of first sample actions as a target first sample action includes: The exploration rate is determined based on the first formula, which is: ,in, represents the exploration rate corresponding to time step t, It represents the reference value of exploration rate, T is the adjustment parameter, and the size of T is negatively correlated with the size of the time step; The target first sample action is selected from the plurality of first sample actions based on the exploration rate.

4. The vehicle path planning method based on reinforcement learning according to claim 1, characterized in that: The determining of a target expected return value based on the second expected return value includes: Obtaining the second sample action corresponding to the maximum second expected return value as the target second sample action; Determining a reward value at time step t based on effect values ​​of the path planning solutions corresponding to the plurality of second sample actions; Based on the reward value at time step t, the target expected return value is determined.

5. The vehicle path planning method based on reinforcement learning according to claim 4, characterized in that: The determining of the reward value at time step t based on the effect values ​​of the path planning solutions corresponding to the plurality of second sample actions includes: Obtain the maximum and average values ​​of the effect improvement of the path planning solutions corresponding to the second sample actions, Obtaining a reward change trend value, where the reward change trend value reflects a change trend of the reward value over the latest N time steps; The reward value at time step t is determined based on the maximum value, the average value, and the reward change trend value.

6. The vehicle path planning method based on reinforcement learning according to claim 1, characterized in that: The first model is a multi-layer fully connected neural network, and each hidden layer in the first model has a Dropout mechanism.

7. A vehicle path planning device based on reinforcement learning, characterized in that: include: A first expected benefit acquisition module is configured to input a state vector at time step t into a first model and obtain first expected benefit values ​​corresponding to a plurality of first sample actions output by the first model, wherein each first sample action corresponds to calling a path planning operator, and the state vector includes a sequence of called operators; an action selection module, configured to select an action from the plurality of first sample actions as a target first sample action, and update the state vector based on the target first sample action to obtain the state vector at time step t+1; A second expected benefit acquisition module is configured to input the state vector at time step t+1 into a second model, and acquire a plurality of second sample actions output by the second model and a second expected benefit value corresponding to each second sample action; a target expected return determining module, configured to determine a target expected return value based on the second expected return value, determine a training loss based on the target expected return value and the first expected return value corresponding to the target first sample action, and update the parameters of the first model based on the training loss; a soft update module, configured to soft update the parameters of the second model based on the parameters of the first model after a plurality of time steps; The result output module is used to obtain the vehicle path planning result based on the output data of the first model after training.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the vehicle path planning method based on reinforcement learning as described in any one of claims 1 to 6 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the vehicle path planning method based on reinforcement learning as described in any one of claims 1 to 6 is implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the vehicle path planning method based on reinforcement learning as described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Spatial search method and device based on deep reinforcement learning

    CN112633591A

  • Multi-agent path planning method based on deep reinforcement learning

    CN114815840A

  • Electric vehicle path planning method based on evolutionary reinforcement learning in dual-network fusion scene

    CN120146335A

  • Automatic driving vehicle navigation control method and system

    WO2024087654A1