Vehicle dynamic path planning method based on Harris eagle fusion deep reinforcement learning

By combining Harris Eagle optimization and deep reinforcement learning, dynamically adjusting its participation ratio and performing information interaction, the problem of local optimal traps and insufficient adaptability in vehicle dynamic path planning is solved, and more efficient path planning is achieved.

CN120333489AActive Publication Date: 2025-07-18四川吉利学院
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510820554.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-07-18
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

The existing technology has problems of local optimal traps and insufficient adaptability in vehicle dynamic path planning. Traditional methods are difficult to quickly obtain global near-optimal solutions in high-dimensional, multimodal and nonlinear complex optimization problems. The Harris Hawk Optimization Method (HHO) and Deep Reinforcement Learning (DRL) each have problems of insufficient exploration or slow convergence, and lack global search capabilities.

Method used

Combining the Harris Eagle optimization method and deep reinforcement learning, the participation ratio of the two is dynamically adjusted through the dynamic weight fusion mechanism and the two-way feedback mechanism, and the global exploration of the Harris Eagle population optimization method and the local development of deep reinforcement learning are used to realize information interaction and collaborative optimization.

Benefits of technology

It improves the adaptability and collaborative optimization capabilities of vehicle dynamic path planning, improves the search efficiency and optimization accuracy in complex environments, and achieves more efficient path planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120333489A_ABST
    Figure CN120333489A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of vehicle path planning, and discloses a Harris eagle fusion deep reinforcement learning vehicle dynamic path planning method, which comprises the steps of initializing the position of a Harris eagle population optimization method and strategy network parameters of deep reinforcement learning; calculating an escape energy function to judge that global exploration is carried out through a Harris eagle optimization method or local development is carried out through deep reinforcement learning; based on a dynamic weight fusion mechanism, according to the fitness improvement rate and the population diversity, dynamically adjusting the participation proportion of the Harris eagle population optimization method and the deep reinforcement learning; based on a bidirectional feedback mechanism, injecting a global optimal solution of deep reinforcement learning into a population of the Harris eagle population optimization method, and adding a high-quality solution sequence of the Harris eagle population optimization method into a deep reinforcement learning experience pool; and outputting the optimal path of the vehicle after iterative optimization. According to the method, the adaptivity and collaborative optimization capability of vehicle dynamic path planning can be enhanced, and more efficient vehicle dynamic planning is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of vehicle path planning, and specifically relates to a vehicle dynamic path planning method that combines Harris Hawks with deep reinforcement learning. Background Art

[0002] Intelligent Transportation Systems (ITS) have developed rapidly in recent years, providing new technical means for urban traffic governance and vehicle scheduling. In this trend, real-time path planning for autonomous vehicles in dynamic, variable, and uncertain road environments has become a core challenge. Traditional deterministic path planning methods often struggle to quickly find a globally near-optimal solution when faced with optimization problems of high-dimensionality, multi-modality, and non-linear complexity. Meta-heuristic methods exhibit good robustness and adaptability in high-dimensional non-linear optimization and are widely used to solve complex decision-making problems. However, even meta-heuristic methods may still face local optimum traps and search imbalance problems in scenarios with significant multi-modal, dynamic, and stochastic characteristics. The Harris Hawks Optimization (HHO) method, as an emerging bionic method, achieves a certain balance between global and local searches by simulating the hunting behavior of hawk flocks. Research shows that HHO can outperform classical meta-heuristic methods (such as PSO, GWO, WOA) in multi-modal functions and high-dimensional problems, demonstrating strong performance in complex optimization tasks. However, when HHO faces highly dynamic vehicle path planning scenarios, there are still potential problems of falling into local optima and insufficient adaptability. Therefore, some studies have attempted to enhance the performance of HHO through adaptive parameters and hybrid strategies to improve the coordination of global and local explorations. However, in real-time path planning problems with strong time-varying characteristics, HHO still needs further optimization to improve the convergence speed and optimization accuracy.

[0003] In contrast, Deep Reinforcement Learning (DRL) shows great potential in the field of decision control. It continuously optimizes strategies through the state-action-reward loop of interacting with the environment. DRL has been successfully applied to intelligent decision-making, autonomous driving, and path planning, and can gradually discover high-value strategies in complex and uncertain environments. With the help of deep neural networks, DRL can extract features in high-dimensional state spaces and achieve efficient mapping from perception to decision-making. However, DRL usually has problems such as insufficient exploration and slow policy convergence in the initial learning stage, and it takes a long time to approach an optimal solution. In addition, DRL is prone to falling into sub-optimal strategies in high-dimensional continuous action spaces and lacks sufficient global search ability to ensure global exploration. Therefore, combining DRL with meta-heuristic methods becomes a natural choice to provide heuristic guidance and potential optimal region references for DRL. Existing research shows that the fusion of evolutionary methods or swarm intelligence methods with DRL can significantly improve the exploration efficiency and convergence performance of DRL in complex problems. However, there is still a lack of systematic research on deeply integrating HHO with DRL to adapt to the specific challenge of vehicle dynamic path planning, and there are obvious gaps in this field. Summary of the Invention

[0004] In view of the above deficiencies in the prior art, the present invention provides a vehicle dynamic path planning method that fuses Harris Hawks with deep reinforcement learning.

[0005] To achieve the above invention objective, the technical solution adopted by the present invention is as follows: A vehicle dynamic path planning method that fuses Harris Hawks with deep reinforcement learning, comprising the following steps: Initialize the positions of the Harris Hawks population optimization method and the policy network parameters of deep reinforcement learning; Calculate the escape energy function to determine whether to conduct global exploration through the Harris Hawks optimization method or local exploitation through deep reinforcement learning; Based on the dynamic weight fusion mechanism, dynamically adjust the participation ratios of the Harris Hawks population optimization method and deep reinforcement learning according to the fitness improvement rate and population diversity; Based on the two-way feedback mechanism, inject the global optimal solution of deep reinforcement learning into the population of the Harris Hawks population optimization method, and add the high-quality solution sequence of the Harris Hawks population optimization method to the deep reinforcement learning experience pool; Output the optimal vehicle path after iterative optimization.

[0006] Furthermore, the calculation formula for the dynamic weight factor in the dynamic weight fusion mechanism is:

[0007] Wherein, is the dynamic weight factor, is the reference value of population diversity, is the population diversity, is a preset positive number, is the reference value of fitness improvement rate, is the fitness improvement rate.

[0008] Furthermore, injecting the global optimal solution of deep reinforcement learning into the population of the Harris hawk population optimization method includes: When deep reinforcement learning feeds back to the Harris hawk population optimization method, using the global optimal solution of deep reinforcement learning as the guiding center to re-initialize a set number of individuals in the Harris hawk population, expressed as:

[0009] where, is the position of the i th individual in the Harris hawk population at t +1 moment, is the global optimal solution of deep reinforcement learning, is the hyperparameter for controlling the perturbation intensity, is a Gaussian random vector with a mean of 0 and a covariance of the identity matrix I.

[0010] Furthermore, adding the sequence of high-quality solutions of the Harris hawk population optimization method to the deep reinforcement learning experience pool includes: When the Harris hawk population optimization method feeds back to deep reinforcement learning, screening a set number of high-quality solution sets as vehicle state and control decision sequences and adding them to the deep reinforcement learning experience pool according to the weight factor, expressed as:

[0011] where, w k is the weight factor of the k th high-quality solution, exp is the natural exponential function, is the steepness parameter for controlling the weight distribution, is the t th high-quality solution obtained by the Harris hawk population optimization method in the k th generation, is the t th high-quality solution obtained by the Harris hawk population optimization method in the m th generation, K is the number of high-quality solutions.

[0012] Furthermore, after adding the sequence of high-quality solutions of the Harris hawk population optimization method to the deep reinforcement learning experience pool, deep reinforcement learning performs weighted sampling in the experience pool according to the weight factor, expressed as:

[0013] Among them, P is the sampling probability, sample is the sampling operation, is the individual state of the k-th high-quality solution corresponding to the j-th state-action trajectory, is the individual action of the k-th high-quality solution corresponding to the j-th state-action trajectory, w k is the k weight factor of the k-th high-quality solution, w m is the m weight factor of the k-th high-quality solution, K is the number of high-quality solutions.

[0014] Furthermore, the policy gradient update formula of deep reinforcement learning is:

[0015] Among them, is the gradient of the policy objective function with respect to the parameter , P is the sampling probability, sample is the sampling operation, is the state of the k-th high-quality solution corresponding to the j-th state-action trajectory, is the action of the k-th high-quality solution corresponding to the j-th state-action trajectory, is the gradient term of the logarithmic policy, is the value function estimation, K is the number of high-quality solutions, J k is the number of discrete states on the path.

[0016] Furthermore, the calculation formula of the fitness improvement rate is:

[0017] Among them, I t is the t fitness improvement rate of the g-th generation, is the fitness of the global optimal solution in the t- 1st generation, is the t fitness of the global optimal solution in the g-th generation, max is the maximum value function, is a preset positive number.

[0018] Furthermore, the calculation formula of the population diversity is:

[0019] Among them, Dt is the population diversity of the t generation, is the position vector of the i -th individual in the t generation, is the mean vector of the population positions in the t generation, N is the number of individuals.

[0020] Furthermore, the position update method of the Harris hawk optimization method is:

[0021] where is the position of the Harris hawk population in the t +1 generation, is the position vector of the hawk with the best fitness among all Harris hawks in the -th generation, is the arithmetic mean of all hawk positions in the

[0022] Furthermore, the objective function of the deep reinforcement learning is:

[0023] where is the objective function, is the policy parameter, is the expectation operator, is the discount factor in the t -th generation, is the instantaneous reward function for taking the action in the state

[0024] The present invention has the following beneficial effects: By combining the Harris hawk optimization method and deep reinforcement learning, the present invention can make full use of the advantages of both to achieve more efficient vehicle dynamic planning. At the same time, the present invention further introduces a dynamic weight fusion mechanism and a two-way feedback mechanism to enhance the self-adaptability and collaborative optimization ability of vehicle dynamic path planning. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 is a schematic flow chart of a vehicle dynamic path planning method integrating Harris hawks and deep reinforcement learning; Figure 2 is a comparison chart of different vehicle dynamic planning methods in a simple scenario; Figure 3 is a comparison chart of different vehicle dynamic planning methods in a complex scenario. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] ​The following describes the specific embodiments of the present invention to facilitate those skilled in the art of this technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of this technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.

[0027] The present invention proposes a vehicle path planning method that fuses HHO and DRL to give full play to the global search advantage of HHO and the strategy learning ability of DRL. The present invention goes further on the basis of simple superposition, introducing a dynamic weight fusion mechanism, and flexibly adjusting the participation degrees of HHO and DRL according to measurable indicators such as the fitness improvement rate and population diversity to achieve an adaptive balance between global exploration and local development. When the method falls into a local optimum or improves slowly, increasing the global exploration ratio of DRL can help the method jump out of the trap. And when the method makes obvious progress or maintains a high population diversity, enhancing the local search intensity of HHO can further refine the goodness of the solution. In addition, the present invention proposes a two-way feedback mechanism: after each round of iteration, the high-quality solution information discovered by DRL is fed back to HHO to guide the next round to focus on potential areas; at the same time, the excellent solution sequence screened by HHO is injected into the DRL experience pool in a weighted form to accelerate the DRL strategy update. Through this two-way closed-loop information interaction, HHO and DRL achieve co-evolution, making up for the deficiencies when they are used independently and showing stronger adaptability in a highly dynamic and complex path planning environment. This framework helps to achieve significant improvements in global-local search, adaptive regulation, and real-time decision-making, providing new ideas for the research on the fusion of HHO and DRL.

[0028] As Figure 1 shown, a vehicle dynamic path planning method that fuses Harris hawks and deep reinforcement learning provided by an embodiment of the present invention includes the following steps S1 to S5: S1. Initialize the positions of the Harris hawk population optimization method and the parameters of the policy network of deep reinforcement learning; In an optional embodiment of the present invention, step S1 initializes the positions and speeds of the Harris hawk population, which is expressed as:

[0029] Among them, is the position of the th individual in the Harris hawk population, is the population size.

[0030] Initialize the parameters of the policy network (actor) in deep reinforcement learning and the parameters of the value network (critic) , providing a starting point for subsequent policy updates and value estimations.

[0031] S2. Calculate the escape energy function to determine whether to conduct global exploration through the Harris hawk optimization method or local exploitation through deep reinforcement learning; In an alternative embodiment of the present invention, the Harris hawk optimization method simulates the behavior of catching rabbits. When the escape energy is in this state, this operation will be performed. By judging the surrounding mode and using a random number to confirm whether to dive. The escape energy function is as follows:

[0032] where is a random initial value in (-1, 1); is the current iteration number; is the total iteration number.

[0033] When the escape energy is in this state, a random walk is performed, and its update formula is:

[0034]

[0035] where is a random individual in the Harris hawk population; is the position of the best individual in the Harris hawk population, i.e., the rabbit position; is the average position of the population; is the population size; both follow a normal distribution parameters; and are the upper and lower bounds of the problem to be solved.

[0036] When the escape energy is in this state, local search is performed, and four update methods are selected according to the energy value.

[0037] (1) When the escape energy value and are in this state, the method performs a soft surround, and its position update formula is:

[0038]

[0039] where , representing the update step size; represents the gap between the position with the highest fitness value and the current position.

[0040] (2) When the escape energy value and the method performs a hard encirclement, and the rabbit will be directly killed. Its position update formula is:

[0041] (3) When the escape energy value and the rabbit cannot escape, and the Harris hawk adopts a smarter soft encirclement. The specific method is:

[0042]

[0043]

[0044]

[0045] Among them, is the dimension of the problem to be solved, is a random vector in the dimension, is the Lévy flight formula, is a random number in (0, 1),

[0046] (4) When the escape energy value and the rabbit may escape, and the Harris hawk adopts an asymptotic soft encirclement. The specific method is:

[0047]

[0048] This embodiment is based on the Harris hawk optimization method, and judges the global exploration and local exploitation behaviors in the current iteration process through the escape energy function :

[0049] Among them, is a random number, and its value range is (-1, 1), is the current iteration number, is the total iteration number.

[0050] When global exploration is performed, and the update formula is:

[0051] Among them, is the randomly selected individual position, is a random number.

[0052] In the policy-based deep reinforcement learning Actor-Critic method, in deterministic and stochastic scenarios, the policy represented by a neural network can be updated by gradient ascent. This usually requires an estimate of a value function. A commonly used architecture is actor-critic, where the actor is responsible for the policy and the critic is responsible for the estimation of the value function. In deep reinforcement learning, both can be represented by a non-linear neural network. The actor uses the gradient in the policy gradient theorem to adjust the policy parameters, while the critic estimates the approximate value function of the current policy.

[0053] In vehicle dynamic programming, the Markov decision process (MDP) model is used to describe the transition process of the vehicle in different states. It is defined as follows: (1) State : Describes the specific situation of the vehicle at a certain moment, such as position, speed, etc.

[0054] (2) Action : Operations that the vehicle can take, such as accelerating, decelerating, steering, etc.

[0055] (3) State transition probability : The probability of transitioning from state by taking action to state

[0056] (4) Reward function : The reward obtained by taking action in state

[0057] The state transition formula is:

[0058] Where is the state transition probability, is the state at time t+1, is the state at time t, is the action taken at time t.

[0059] In vehicle dynamic programming, the state at the next moment only depends on the current state and action.

[0060] In reinforcement learning, it is necessary to learn a policy (actor) and a value function (critic) .

[0061] (1) Value function Represents taking action in state ​ The expected total reward obtained.

[0062]

[0063] Among them, represents the expected value; represents the discount factor to the power, used to determine the present value of future rewards, the smaller the future reward weight is; represents the reward obtained at time obtained.

[0064] (2) In reinforcement learning, the policy 𝜋 is the policy used by the agent to determine which actions to take in different states. The policy 𝜋 is a probability distribution of taking different actions in a given state, representing the probability of taking action in state as follows:

[0065] Among them, represents the conditional probability of taking action when the state is at time obtained.

[0066] The policy gradient method optimizes the policy by maximizing the cumulative reward.

[0067] The gradient of the objective function is:

[0068] Among them, represents the gradient of the objective function with respect to the parameter , indicating how to adjust the parameter to maximize the objective function ; represents the expected value; represents the parameterized policy ; represents the action value function of taking action in the state under the policy , which is the total expected reward obtained by starting from the state and taking action obtained.

[0069] Update the policy parameter by gradient ascent to optimize the policy:

[0070] Among them, represents the parameter vector of the policy, which controls the behavior of the policy; represents the learning rate, which determines the step size of each update.

[0071] The Deep Q-Network (DQN) combines a deep neural network and Q-learning to estimate the Q-value function. The update formula for the Q-value function is:

[0072] Among them, the value of taking action in state represents the expected total reward given the state and action; represents the immediate reward obtained at time ; the maximum Q-value in state represents the Q-value of the best action in the next state.

[0073] DQN approximates the Q-value function through a deep neural network , where are the parameters of the neural network. The specific steps are as follows: Step 1: Neural network architecture Use a deep neural network with input state and output the Q-values corresponding to all possible actions.

[0074] Step 2: Calculation of the target Q-value Calculate the target Q-value using the Bellman equation:

[0075] Among them, represents the target Q-value; represents the immediate reward obtained at time t+; represents the maximum Q-value obtained from the target network parameters in state , that is, the Q-value of the best action in the next state.

[0076] Step 3: Loss function Define the loss function as the mean squared error between the current Q-value and the target Q-value:

[0077] Among them, is the loss function.

[0078] Step 4: Gradient descent Use the gradient descent method to update the network parameters , as follows:

[0079] Among them, represents the loss function with respect to the parameter gradient To solve the problem of sparse rewards, reward shaping is introduced. The new reward function is defined as:

[0080] Among them, is the target reward, is the additional reward used to guide the learning process.

[0081] The additional reward can be defined by the potential function which is a function mapping from the state to a real number and is used to measure the potential value of the state, as follows:

[0082] Among them, is the discount factor, represents the next state resulting from the action a. By this method, the additional reward can be added to the original reward so that the agent can obtain more abundant information during the learning process, thereby accelerating the learning speed.

[0083] In vehicle dynamic programming, the potential function needs to consider the following factors: 1. Target position: The target position that the vehicle needs to reach.

[0084] 2. Path smoothness: Ensure the smoothness and feasibility of the vehicle path.

[0085] 3. Obstacle avoidance: The vehicle needs to avoid obstacles.

[0086] 4. Continuity of state and action: Consider the physical characteristics and dynamic limitations of the vehicle.

[0087] Taking the above factors into comprehensive consideration, the present invention uses the Euclidean distance as the potential function because it can clearly guide the vehicle to move towards the target position and can also combine the requirements of path smoothness and obstacle avoidance.

[0088]

[0089] Among them, represents the potential function, represents the current state, Represents the target state, represents the Euclidean distance between the current state and the target state; PathSmoothness(s) represents the measure of path smoothness, the smoother the better; ObstacleProximity(s) represents the obstacle proximity, the farther away from the obstacle the better; and represent the weight factors used to balance the target position, path smoothness, and obstacle avoidance.

[0090] In deep reinforcement learning, the Markov decision process (MDP) model is used to describe the transition process of the vehicle in different states. In each state, the actor network selects an action according to the policy and the critic network estimates the value function of the current policy .

[0091] S3. Based on the dynamic weight fusion mechanism, dynamically adjust the participation ratio of the Harris hawk population optimization method and deep reinforcement learning according to the fitness improvement rate and population diversity; In an alternative embodiment of the present invention, in each iteration, the energy function was originally used to determine whether to use the global exploration of HHO or the local exploitation of DRL. However, to achieve more flexible adaptive regulation, this embodiment introduces a dynamic weight fusion mechanism and a two-way feedback mechanism: To better control the participation degrees of HHO and DRL at different time periods, this embodiment defines a dynamic weight factor to balance the contribution ratio of the two in the th iteration. To make the determination of objective, this study selects two quantitative indicators: "population diversity (Diversity)" and "fitness improvement rate (Improvement Ratio)".

[0092] Assume is the fitness of the global optimal solution (objective function value) in the th generation. Theoretically, the goal of the optimization problem is to make as small as possible (or as close to 0 as possible). Define the relative improvement rate between adjacent generations:

[0093] where, I t is the fitness improvement rate in the t th generation, is the fitness of the global optimal solution in the t- st generation, is the fitness of the global optimal solution in the t th generation, and max is the maximum value function. is a preset positive number, which is a very small positive number used to avoid a zero denominator. When is large, it indicates that the current iteration can significantly improve the optimal solution, showing that the method has good progress in this iteration; when is very small or even close to 0, it means that the method has limited progress in the most recent iteration and may be trapped in a local optimum or stagnation state.

[0094] To measure the distribution of the HHO population in the search space, the population diversity is defined as:

[0095] where D t is the population diversity of the t th generation, is the i th individual's position vector in the t th generation, is the mean vector of the population positions in the t th generation, N is the number of individuals. When is very small, it means that the population has gathered in a narrow area and is prone to being trapped in a local optimum; when is large, it means that the search still has high diversity and exploration space.

[0096] In order to increase the weight of DRL when the population diversity is low and the fitness improvement rate is small to break through the deadlock. When the fitness progress is good or the diversity is still high, it is inclined to let HHO perform more refined local exploitation. Considering comprehensively, its calculation method is:

[0097] where is the dynamic weight factor, is the reference value of population diversity, is the population diversity, is a preset positive number, is the reference value of fitness improvement rate, is the fitness improvement rate. When is small and is small, and will become larger, making approach 1, thereby increasing the global exploration proportion of DRL. On the contrary, when the improvement is obvious or the diversity is sufficient, approaches 0, making HHO dominant in this generation.

[0098] Based on the above process, in each generation of iteration, this embodiment passes through Combine DRL (global exploration) and HHO (local development) in a specific ratio and dynamically regulate them according to the data in the iterative process, rather than using a fixed threshold for judgment.

[0099] S4. Inject the global optimal solution of deep reinforcement learning into the population of the Harris hawk optimization method based on a two-way feedback mechanism, and add the high-quality solution sequence of the Harris hawk optimization method to the deep reinforcement learning experience pool; In an optional embodiment of the present invention, to enhance the in-depth cooperation between HHO and DRL during the search process, this study realizes two-way feedback through information exchange at the end stage of each round of iteration. The core goal of the two-way feedback is to transfer the excellent features obtained by DRL in global exploration to the HHO population, so that HHO can perform more targeted local search in the next round; at the same time, inject the high-quality solutions found by HHO in local refinement into the experience pool of DRL in the form of weighted samples to help DRL quickly focus on the high-value state-action space. Through this closed-loop interaction, the two parties continuously promote each other in subsequent iterations, improving the overall optimization performance.

[0100] After the end of this round of iteration, DRL obtains several relatively optimal solutions in global exploration, such as finding a route segment that is significantly better than the previous generation in path planning. Denote the global best solution obtained after the convergence of DRL in this round as . To make HHO focus more on this excellent area in the next generation of iteration, this solution can be used as the "guidance center" to re-initialize some individuals in the HHO population. Let be a scaling factor used to control the number of individuals to be re-initialized, such as individuals are reset, where is the population size, which can be expressed as:

[0101] where, is the position of the i th individual in the Harris hawk population at t +1 moment, is the global optimal solution of deep reinforcement learning, is a hyperparameter for controlling the perturbation intensity, is a Gaussian random vector with a mean of 0 and a covariance of the identity matrix I. , is the set of indices of individuals to be re-initialized. Through this formula, in the next round of iteration, some individuals in HHO will be concentrated near the high-potential area discovered by DRL, improving the pertinence and effectiveness of local search.

[0102] After HHO completes a round of iteration, a set of candidate solutions will be obtained, and the top % of the excellent solution set For the path planning problem, these high-quality solutions can correspond to several state-action-reward sequences, such as the vehicle state and control decision sequences sampled discretely along an excellent path. These sequences can be directly added to the experience replay buffer of DRL as training samples for subsequent policy updates.

[0103] To make DRL more inclined to utilize the excellent information transmitted by HHO during training, higher sampling weights can be assigned to these newly added samples. Let the objective function of the problem be , for each excellent solution of HHO corresponding sequence, define the weight factor as follows:

[0104] where, w k is the weight factor of the k th high-quality solution, exp is the natural exponential function, >0 is a parameter used to control the steepness of the weight distribution, is the t th generation of the Harris hawk population optimization method to obtain the k th high-quality solution, is the t th generation of the Harris hawk population optimization method to obtain the m th high-quality solution, K is the number of high-quality solutions. The larger it is, the higher the quality of the solution (the lower the objective function value), and thus it should be preferentially used in subsequent DRL training.

[0105] After injecting these weighted samples into the experience pool, when DRL draws samples from the buffer for training, it will sample the samples according to weighted sampling, that is:

[0106] where, P is the sampling probability, sample is the sampling operation, is the individual state of the th high-quality solution corresponding to the w k th state-action trajectory, k is the weight factor of the w m th high-quality solution, m is the weight factor of the K th high-quality solution,

[0107] Through weighted sampling, the policy gradient update formula of DRL can be expressed as:

[0108] where, is the policy objective function the gradient of the parameter is, P is the sampling probability, sample is the sampling operation, is the state of the j-th state-action trajectory corresponding to the k-th high-quality solution, is the action of the j-th state-action trajectory corresponding to the k-th high-quality solution, is the gradient term of the logarithmic policy, is the value function estimate, K is the number of high-quality solutions, J k is the number of discrete states on the path. The state-action pairs corresponding to high-quality solutions will participate in policy updates with a higher probability, strengthening DRL's preference for high-value regions, thereby accelerating policy convergence.

[0109] Through the above mathematical description of the two-way feedback mechanism, this embodiment clarifies the implementation method of population individual repositioning in the DRL → HHO direction, and clarifies the quantization index and sampling strategy for converting excellent solutions into weighted training samples in the HHO → DRL direction. This mechanism and the dynamic weight fusion mechanism work together: the former adjusts the proportion of HHO and DRL in each iteration, and the latter superimposes the advantages of the two through information feedback. This reasonable mathematical support helps to directly implement and verify the effectiveness of this two-way closed-loop optimization strategy in actual experiments.

[0110] S5. Output the optimal vehicle path after iterative optimization.

[0111] In an optional embodiment of the present invention, after the iteration ends, the optimized optimal vehicle path and the corresponding action strategy are output. The output optimal path at this time has incorporated the comprehensive advantages obtained by the collaborative optimization of HHO and DRL under dynamic weights and two-way feedback.

[0112] The position update method of the Harris hawk optimization method is:

[0113] where, is the position of the Harris hawk population in t the +1 generation, is the position vector of the hawk with the best fitness among all Harris hawks in the t-th generation, is the escape energy factor, is the arithmetic mean of all hawk positions in the t-th generation.

[0114] The objective function of deep reinforcement learning is as follows:

[0115] Among them, is the objective function, is the policy parameter, is the expectation operator, is the discount factor of the t th generation, is the instantaneous reward function for taking action in state .

[0116] The method of the present invention will be analyzed and described below in combination with simulation experiments.

[0117] The experimental environment for this experiment is: operating system: Windows 11 (64bit); processor: AMD Ryzen 7 5800H with Radeon Graphics, 3.20 GHz; running memory: 16G; simulation platform: Intellij IDEA.

[0118] To ensure the fairness of the experiment, each method is run thirty rounds, and its average value and standard deviation are statistically calculated.

[0119] This experiment uses standard test functions to verify the effectiveness of the following strategies: the separate Harris Hawks Optimization (HHO) method, the separate Deep Reinforcement Learning (DRL) method, and the method of fusing HHO and DRL.

[0120] The following 5 international standard test functions are selected as the standard test functions.

[0121] Sphere function, mainly testing the local search ability of the method. Due to its simple convex structure, the optimization method needs to perform fine local search in the entire search space to find the optimal solution.

[0122] Schwefel function, testing the global search ability of the method. Since this function has multiple local optimal solutions, the optimization method needs to have strong global search ability to avoid falling into local optima.

[0123] Rastrigin function, testing the global and local search abilities of the method. This function has multiple local optimal solutions, and at the same time, the global optimal solutions are distributed over a larger range, which is suitable for evaluating the performance of the method in dealing with complex search spaces.

[0124] The Ackley function is mainly used to test the global search ability of the method. This function has a deep global optimal basin, but there are many local optimal solutions in the periphery, and a strong global search ability is required to find the optimal solution.

[0125] The Griewank function is used to test the global and local search abilities of the method. This function has multiple global and local optimal solutions, and the optimization method needs to achieve a balance in global and local search abilities.

[0126] HHO (Harris Hawks Optimization) parameter settings: - Population Size: To ensure the diversity of the search space and avoid excessive computational effort, this value is set to 30.

[0127] - Max Iterations: To more comprehensively verify the performance of the algorithm, it is set to 1000 times.

[0128] DRL (Deep Reinforcement Learning) parameter settings: - State Space: In the optimization problem, the dimension of the state space is equal to the number of variables, which is set to 30.

[0129] - Action Space: It represents the direction and step size of movement in each dimension. Using a discrete action space, an action space of ±1 in each dimension is selected.

[0130] - Reward Function: It is set to the negative value of the objective function value, so that the minimization problem is transformed into maximizing the reward as: -objective_function_value.

[0131] - Neural Network Architecture: A three-layer fully connected neural network with 128 neurons in each layer.

[0132] - Learning Rate: It is set to 0.0001.

[0133] - Discount Factor: It is set to 0.95.

[0134] - Experience Replay: Experience replay is used to alleviate the problem of sample correlation, and the size of its experience pool is set to 10000.

[0135] - Target Network: The update frequency is set to 100.

[0136] Parameter Settings of the Hybrid Method (HHO-DRL) Comprehensively consider the parameter settings of HHO and DRL; the population size is set to 30, and the maximum number of iterations is 1000. At the same time, in each iteration, first use HHO to update the population, and then use DRL for fine-tuning of individuals.

[0137] The experimental results are shown in Table 1: Table 1 Comparison Table of Experiments with Different Intelligent Optimization Algorithms

[0138] As shown in Table 1, this experiment evaluated the performance of three different methods (HHO, DRL, HHO-DRL) for five different test functions (f1 to f5). The theoretical optimal solutions of the test functions are all 0, that is, the closer the optimization result is to 0, the better. The following is an in-depth analysis and discussion of the experimental results.

[0139] f1: Local search. In the local search function f1, the average value of the HHO method is 4.28e-152, and the standard deviation is 3.33e-149, showing extremely high stability and optimization performance. The average value of the DRL method is 2.59e-96, and the standard deviation is 3.22e-95. Although the optimization result is also relatively good, there is an obvious gap compared with the HHO method. The average value of the HHO-DRL method is 8.65e-97, and the standard deviation is 2.86e-97, which performs better than the single DRL method but worse than the HHO method. This indicates that in the local search scenario, the HHO method has significant advantages, probably because its local search strategy can more effectively avoid local optimal traps.

[0140] f2: Global search. In the global search function f2, the average value of the HHO method is 6.59e-49, and the standard deviation is 8.77e-55, still showing strong optimization ability. The average value of the DRL method is 8.29e-108, and the standard deviation is 2.85e-100, indicating that DRL performs poorly in global search. The average value of the HHO-DRL method is 6.85e-48, and the standard deviation is 8.24e-48, showing better performance than HHO. This shows that in global search, the method combining HHO and DRL can better balance exploration and exploitation and improve optimization performance.

[0141] f3: Comprehensive Search. In the comprehensive search function f3, the average value of the HHO method is 3.59e-79, and the standard deviation is 2.89e-76, showing good stability. The average value of the DRL method is 1.00e-80, and the standard deviation is 1.50e-79, and its performance is slightly better than that of the HHO method. The average value of the HHO-DRL method is 7.25e-132, and the standard deviation is 5.22e-125, and its optimization result is significantly better than that of the individual HHO and DRL methods. This indicates that in the comprehensive search scenario, the HHO-DRL method can effectively combine the advantages of both to achieve a better optimization effect.

[0142] f4: Global Search. For another global search function f4, the average value of the HHO method is 5.75e-50, and the standard deviation is 8.14e-49, continuing to show its strong global search ability. The average value of the DRL method is 2.22e-65, and the standard deviation is 6.55e-49, and its performance is far inferior to that of the HHO. The average value of the HHO-DRL method is 4.55e-50, and the standard deviation is 5.11e-48, and the optimization effect is close to but slightly worse than that of the HHO. This shows that in the global search, although the HHO-DRL performs well, the pure HHO method still has certain advantages.

[0143] f5: Comprehensive Search. In the comprehensive search function f5, the average values of both HHO and DRL are 1.26e+04, and the standard deviations are both 9.45e-01, showing almost the same performance. The average value of the HHO-DRL method is 5.00e+02, and the standard deviation is 5.00e-01, and its optimization result is significantly better than that of HHO and DRL. This indicates that in the f5 comprehensive search scenario, the HHO-DRL method significantly improves the optimization effect, probably due to its better balance between exploration and exploitation.

[0144] Analysis of influencing factors: The HHO method performs excellently in both local search and global search, especially in avoiding local optima. The DRL method performs poorly in global search but has certain competitiveness in comprehensive search. The HHO-DRL method performs well in multiple scenarios, especially in comprehensive search and some global search, showing the potential to combine the advantages of both.

[0145] Method characteristics: The HHO method is good at local search and avoiding local optima; the DRL method has adaptive learning ability and strong adaptability.

[0146] Problem characteristics: The characteristics of different test functions have a significant impact on the method performance, especially in terms of function complexity and multi-peak characteristics.

[0147] Method Fusion: The hybrid method of HHO-DRL combines the local search ability of HHO and the global exploration ability of DRL, achieving better optimization effects.

[0148] The HHO-DRL method has demonstrated its effectiveness and advantages in various search tasks, but its specific performance still needs to be selected according to the characteristics of specific problems.

[0149] The complex functions in CEC2014 have characteristics such as high dimensionality, multimodality, non-linearity, and irregularity, which pose higher requirements for optimization methods. Optimization problems in high-dimensional spaces are more complex, and methods need to have strong search and convergence capabilities. Multimodal functions have multiple local optimal solutions, requiring methods to be able to effectively jump out of local optima. Non-linear and irregular functions make the optimization path uncertain, and methods need to have strong adaptability and exploration capabilities. This study evaluates the method performance by selecting three different types of test functions. These test functions include unimodal functions (CEC01 and CEC06), multimodal functions (CEC14 and CEC15), and hybrid functions (CEC20 and CEC21). The experimental parameters are set as follows: population size 30, dimension 30, number of iterations 500 times, and 15 independent runs are carried out. After each run, the average value and standard deviation of the calculation results are calculated. This setting can comprehensively evaluate the performance of the method on different types of problems and ensure the stability and reliability of the experimental results, thus enhancing the persuasiveness of the research. The performance comparison of different intelligent optimization methods on the CEC2014 test functions is shown in Table 2.

[0150] Table 2 Performance Comparison Table of Different Intelligent Optimization Algorithms on CEC2014 Test Functions

[0151] As shown in Table 2, the research will be analyzed from two aspects: method characteristics and data performance.

[0152] In terms of method characteristics, different methods have different characteristics and optimization mechanisms, which directly affect their performance on complex functions. The PSO (Particle Swarm Optimization) method has strong global search ability, but may not be fine enough in local search and is prone to falling into local optima. The GWO (Grey Wolf Optimization) method simulates the hunting behavior of grey wolves. Although it balances the exploration and exploitation abilities, it performs poorly on high-dimensional complex functions. The WOA (Whale Optimization Algorithm) simulates the hunting behavior of humpback whales and has strong global search ability, but has low local search efficiency for complex functions. MHHO (Modified Harris Hawk Optimization) has certain advantages in high-dimensional spaces and complex functions, but the method complexity is relatively high. AHHO (Adaptive Harris Hawk Optimization) has flexibility in adjusting parameters and can perform well on complex functions. HHO-DRL (Harris Hawk Optimization combined with Deep Reinforcement Learning) is particularly prominent on complex functions due to its adaptive learning and optimization abilities.

[0153] In terms of data performance, the performance differences of different methods on different functions are significant. On the CEC01 function, HHO-DRL performs the best, indicating its significant advantage in dealing with such complex functions. On the CEC06 function, PSO and GWO perform better, showing their strong global search abilities. On the CEC14 function, AHHO and HHO-DRL are significantly better than other methods, indicating their obvious advantages in adaptive ability and learning ability for such functions. On the CEC15 function, HHO-DRL performs the best, showing its advantage in dealing with large-scale optimization problems. On the CEC20 function, PSO and WOA perform better, but the standard deviation is relatively high, indicating that they have certain ability in global search but poor stability. On the CEC21 function, AHHO and HHO-DRL perform excellently, indicating their significant advantages in dealing with complex non-linear functions.

[0154] By analyzing the performance of different methods on each function, the research can obtain the applicability of different methods. PSO is suitable for problems with high requirements for global search, but may perform poorly on complex multimodal functions. GWO is suitable for problems that require a balance between global and local search, but may need improvement on high-dimensional complex functions. WOA is suitable for problems that require strong global search ability, but needs to be enhanced in local search. MHHO is suitable for high-dimensional complex functions, but the method complexity is relatively high and may require optimizing the calculation efficiency. AHHO is suitable for complex non-linear functions and multimodal functions and has strong adaptability and flexibility. HHO-DRL is suitable for various complex functions, especially performs excellently on high-dimensional and non-linear problems, and its adaptive ability and optimization ability are particularly prominent.

[0155] The performance differences of different methods on complex functions mainly stem from their optimization mechanisms and adaptability. HHO-DRL and AHHO perform well on most functions, mainly due to their adaptive ability and learning optimization ability.

[0156] Combining the results of the previous international standard test functions and the results of the CEC2014 test functions, it can be seen that the AHHO method performs well on some standard test functions, but the performance of the HHO-DRL method is significantly improved in terms of the number of method iterations.

[0157] In this experiment, the HHO-DRL fusion method is applied to vehicle path planning to further verify that the improved method of this study has the ability to solve practical problems, and the optimization performance of the HHO-DRL method is highlighted by comparing it with other three methods. This study sets two scenarios. One is a simple scenario with a map size of 30×30 and an obstacle ratio of 15%. The other is a complex scenario with a map size of 30×30 and an obstacle ratio of 30%. In this subsection, the number of iterations of the experimental method T = 50, repeated ten times, and the average value and standard deviation of the results are taken. The comparison methods selected are GWO, MHHO, and AHHO (Grey Wolf Optimization, Modified Harris Hawk Optimization, Adaptive Harris Hawk Optimization).

[0158] The vehicle dynamic planning results of different methods in the simple scenario are as Figure 2 shown, where the horizontal and vertical coordinates are grid coordinates with the unit of dimensionless discrete index. As Figure 2 shown, the path planned by the HHO-DRL method is the shortest, with the starting point being the yellow point and the ending point being the blue point.

[0159] The vehicle dynamic planning results of different methods in the complex scenario are as Figure 3 shown, where the horizontal and vertical coordinates are grid coordinates with the unit of dimensionless discrete index. As Figure 3 shown, the path planned by the HHO-DRL method is still the shortest.

[0160] To better compare the effects of the method in this paper with other methods, the optimization rate is introduced to represent the percentage change of the effect of the method in this paper relative to the effects of other methods.

[0161]

[0162] Among them, represents the length of the path planned by the improved method, represents the length of the path of the method before improvement, represents a certain method. If is smaller than , then is a positive value, indicating that the optimization is effective.

[0163] At the same time, the study also introduces the efficiency ratio to measure the optimization effect.

[0164]

[0165] This formula represents how many times the improved effect is compared to the original effect. If , it means the improved method is more effective.

[0166] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0167] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0168] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0169] Specific embodiments are applied in the present invention to elaborate on the principles and implementation manners of the present invention. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, based on the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

[0170] Those of ordinary skill in the art will realize that the embodiments described herein are provided to assist the reader in understanding the principles of the present invention, and it should be understood that the scope of protection of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not depart from the essence of the present invention based on these technical revelations disclosed in the present invention, and these deformations and combinations are still within the scope of protection of the present invention.

Claims

1. A vehicle dynamic path planning method integrating Harris hawks and deep reinforcement learning, characterized in that, It includes the following steps: Initialize the positions of the Harris hawk population optimization method and the policy network parameters of deep reinforcement learning; Calculate the escape energy function to determine whether to conduct global exploration through the Harris hawk optimization method or local exploitation through deep reinforcement learning; Based on the dynamic weight fusion mechanism, dynamically adjust the participation ratios of the Harris hawk population optimization method and deep reinforcement learning according to the fitness improvement rate and population diversity; Based on the bidirectional feedback mechanism, inject the global optimal solution of deep reinforcement learning into the population of the Harris hawk population optimization method, and add the high-quality solution sequence of the Harris hawk population optimization method to the deep reinforcement learning experience pool; Output the optimal vehicle path after iterative optimization.

2. The vehicle dynamic path planning method integrating Harris hawk and deep reinforcement learning according to claim 1, wherein, The calculation formula for the dynamic weight factor in the dynamic weight fusion mechanism is: Among them, is the dynamic weight factor of the t +1-th generation, is the reference value of population diversity, is the t population diversity of the +1-th generation, is a preset positive number, is the reference value of fitness improvement rate, t is the fitness improvement rate of the +1-th generation.

3. A vehicle dynamic path planning method integrating Harris hawk and deep reinforcement learning according to claim 1, characterized in that Injecting the global optimal solution of deep reinforcement learning into the population of the Harris hawk population optimization method includes: When deep reinforcement learning feeds back to the Harris hawk population optimization method, use the global optimal solution of deep reinforcement learning as the guiding center to re-initialize a set number of individuals in the Harris hawk population, expressed as: Among them, is the position of the i -th individual in the Harris hawk population at the t +1-th generation, is the global optimal solution of deep reinforcement learning, is the hyperparameter for controlling the perturbation intensity, is a Gaussian random vector with mean 0 and covariance matrix being the identity matrix I.

4. A vehicle dynamic path planning method integrating Harris hawk and deep reinforcement learning according to claim 1, characterized in that, Adding the high-quality solution sequence of the Harris hawk population optimization method to the deep reinforcement learning experience pool includes: When the Harris hawk population optimization method feeds back to deep reinforcement learning, screen a set number of high-quality solution sets as the vehicle state and control decision sequence, and add them to the deep reinforcement learning experience pool according to the weight factor, expressed as: Among them, w k is the weight factor of the k th high-quality solution, exp is the natural exponential function, is the parameter controlling the steepness of the weight distribution, is the t th high-quality solution obtained by the Harris hawks optimization method in the k th generation, is the t th high-quality solution obtained by the Harris hawks optimization method in the m th generation, K is the number of high-quality solutions.

5. A vehicle dynamic path planning method integrating Harris hawk and deep reinforcement learning according to claim 1, characterized in that After adding the high-quality solution sequence of the Harris hawk population optimization method to the deep reinforcement learning experience pool, deep reinforcement learning performs weighted sampling in the experience pool according to the weight factor, expressed as: Among them, P is the sampling probability, sample is the sampling operation, is the individual state of the k-th high-quality solution corresponding to the j-th state-action trajectory, is the individual action of the k-th high-quality solution corresponding to the j-th state-action trajectory, w k is the weight factor of the k -th high-quality solution, w m is the weight factor of the m -th high-quality solution, K is the number of high-quality solutions.

6. The vehicle dynamic path planning method integrating Harris hawk and deep reinforcement learning according to claim 1, characterized in that The policy gradient update formula of deep reinforcement learning is: Among them, is the policy objective function of the gradient of the parameter . P is the sampling probability, sample is the sampling operation, is the state of the k-th high-quality solution corresponding to the j-th state-action trajectory, is the action of the k-th high-quality solution corresponding to the j-th state-action trajectory, is the gradient term of the logarithmic policy, is the value function estimation, K is the number of high-quality solutions, J k is the number of discrete states on the path.

7. A vehicle dynamic path planning method integrating Harris hawk and deep reinforcement learning according to claim 1, characterized in that, The calculation formula for the fitness improvement rate is: Among them, I t is the fitness improvement rate of the t th generation, is the fitness of the global optimal solution in the t- 1st generation, is the fitness of the global optimal solution in the t th generation, and max is the maximum value function, is a preset positive number.

8. A vehicle dynamic path planning method integrating Harris hawk and deep reinforcement learning according to claim 1, characterized in that, The calculation formula for population diversity is: Among them, D t is the population diversity of the t generation, is the position vector of the i th individual in the t generation, is the mean vector of the population positions in the t generation, N is the number of individuals.

9. A vehicle dynamic path planning method integrating Harris hawk with deep reinforcement learning according to claim 1, characterized in that, The position update method of the Harris hawk optimization method is: Among them, is the position of the Harris hawk population in t generation +1, is the position vector of the hawk with the optimal fitness among all Harris hawks in the t-th generation, is the escape energy factor, is the arithmetic mean of the positions of all hawks in the t-th generation.

10. A vehicle dynamic path planning method integrating Harris hawk and deep reinforcement learning according to claim 1, characterized in that, The objective function of deep reinforcement learning is: Among them, is the objective function, is the policy parameter, is the expectation operator, is the t discount factor of the th generation, is the instantaneous reward function for taking action in state

Citation Information

Patent Citations

  • Multi-intelligent agent reinforcement learning path planning method based on ant colony algorithm

    CN112286203A

  • Path planning method based on heuristic deep reinforcement learning

    CN112325897A

  • Vehicle path planning method, system and terminal based on circle search algorithm

    CN117029862A

  • Unmanned aerial vehicle path planning method and device for hanging falling protector

    CN119374595A

  • Robot path planning method based on multi-strategy improved Harris eagle algorithm

    CN119536245A