A vehicle dynamic path planning method integrating Harris Hawk with deep reinforcement learning

By combining Harris Eagle optimization and deep reinforcement learning, dynamically adjusting their participation ratios and conducting information interaction, the problems of local optimality and insufficient global search of traditional path planning methods in dynamic environments are solved, and more efficient vehicle path planning is achieved.

CN120333489BActive Publication Date: 2025-09-30四川吉利学院
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510820554.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-30
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

Traditional path planning methods have difficulty in quickly obtaining a global near-optimal solution in dynamic, changeable, and uncertain road environments. The Harris Eagle optimization method suffers from local optimality and insufficient adaptability in highly dynamic vehicle path planning scenarios. Deep reinforcement learning is prone to falling into sub-optimal strategies in high-dimensional continuous action spaces and lacks global search capabilities.

Method used

Combining the Harris Hawk optimization method and deep reinforcement learning, the participation ratio of the two is dynamically adjusted through a dynamic weight fusion mechanism and a two-way feedback mechanism. Information interaction is carried out through the high-quality solution sequence of the Harris Hawk population optimization method and the global optimal solution of deep reinforcement learning to achieve an adaptive balance between global exploration and local development.

Benefits of technology

It improves the adaptability and collaborative optimization capabilities of vehicle dynamic path planning, and improves the efficiency and accuracy of path planning in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120333489B_ABST
    Figure CN120333489B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of vehicle path planning, and discloses a vehicle dynamic path planning method that integrates Harris Hawk and deep reinforcement learning. The method comprises the following steps: initializing the position of a Harris Hawk population optimization method and the policy network parameters of deep reinforcement learning; calculating an escape energy function to determine whether to conduct global exploration through the Harris Hawk optimization method or to conduct local development using deep reinforcement learning; dynamically adjusting the participation ratio of the Harris Hawk population optimization method and deep reinforcement learning based on a dynamic weight fusion mechanism according to the fitness improvement rate and population diversity; injecting the global optimal solution of deep reinforcement learning into the population of the Harris Hawk population optimization method based on a two-way feedback mechanism, and adding the high-quality solution sequence of the Harris Hawk population optimization method to the deep reinforcement learning experience pool; and outputting the optimal path of the vehicle after iterative optimization. The present invention can enhance the adaptability and collaborative optimization capabilities of vehicle dynamic path planning and achieve more efficient vehicle dynamic planning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of vehicle path planning, and in particular to a vehicle dynamic path planning method integrating Harris Hawk with deep reinforcement learning. Background Art

[0002] Intelligent Transportation Systems (ITS) have developed rapidly in recent years, providing new technological solutions for urban traffic management and vehicle scheduling. Within this trend, real-time path planning for autonomous vehicles in dynamic, changing, and uncertain road environments has become a core challenge. Traditional deterministic path planning methods often struggle to quickly achieve near-optimal global solutions for optimization problems characterized by high dimensionality, multimodality, and nonlinear complexity. Metaheuristic methods demonstrate robustness and adaptability in high-dimensional, nonlinear optimization and are widely used to solve complex decision-making problems. However, even these metaheuristic methods can still encounter local optima and imbalanced search in scenarios characterized by multimodality, dynamics, and significant stochasticity. Harris Hawks Optimization (HHO), an emerging biomimetic method, achieves a balance between global and local search by simulating the hunting behavior of hawks. Research has shown that HHO can outperform classic metaheuristic methods (such as PSO, GWO, and WOA) for multimodal functions and high-dimensional problems, demonstrating strong performance in complex optimization tasks. However, when faced with highly dynamic vehicle path planning scenarios, HHO still faces the potential for falling into local optimality and insufficient adaptability. To this end, some research has attempted to enhance HHO's performance through adaptive parameters and hybrid strategies, improving the coordination between global and local exploration. However, in real-time path planning problems with strong time-varying characteristics, HHO still needs further optimization to improve convergence speed and optimization accuracy.

[0003] In contrast, deep reinforcement learning (DRL) demonstrates strong potential in decision-making and control, continuously optimizing policies through a state-action-reward cycle that interacts with the environment. DRL has been successfully applied to intelligent decision-making, autonomous driving, and path planning, enabling the gradual discovery of high-value policies in complex and uncertain environments. Leveraging deep neural networks, DRL can extract features from high-dimensional state spaces, enabling an efficient mapping from perception to decision-making. However, DRL often suffers from insufficient exploration and slow policy convergence in its initial learning phase, requiring a significant time to approach an optimal solution. Furthermore, DRL is prone to becoming trapped in suboptimal policies in high-dimensional continuous action spaces and lacks sufficient global search capabilities to ensure global exploration. Therefore, combining DRL with metaheuristic methods has become a natural choice, providing DRL with heuristic guidance and potential optimal regions. Previous studies have shown that integrating evolutionary or swarm intelligence methods with DRL can significantly improve DRL's exploration efficiency and convergence performance in complex problems. However, systematic research on the deep integration of HHO and DRL to address the specific challenge of dynamic vehicle path planning remains lacking, leaving a significant gap in this area. Summary of the Invention

[0004] In view of the above-mentioned deficiencies in the prior art, the present invention provides a vehicle dynamic path planning method integrating Harris Hawk with deep reinforcement learning.

[0005] In order to achieve the above-mentioned object of the invention, the technical solution adopted by the present invention is:

[0006] A vehicle dynamic path planning method integrating Harris Hawk with deep reinforcement learning includes the following steps:

[0007] Initialize the positions of the Harris Hawk population optimization method and the policy network parameters of deep reinforcement learning;

[0008] Calculate the escape energy function to determine whether to perform global exploration via Harris Hawk optimization or local exploitation using deep reinforcement learning;

[0009] Based on the dynamic weight fusion mechanism, the participation ratio of Harris Hawk population optimization method and deep reinforcement learning is dynamically adjusted according to the fitness improvement rate and population diversity;

[0010] Based on a two-way feedback mechanism, the global optimal solution of deep reinforcement learning is injected into the population of the Harris Hawk swarm optimization method, and the high-quality solution sequence of the Harris Hawk swarm optimization method is added to the deep reinforcement learning experience pool;

[0011] After iterative optimization, the optimal vehicle path is output.

[0012] Furthermore, the calculation formula of the dynamic weight factor in the dynamic weight fusion mechanism is:

[0013]

[0014] in, is the dynamic weight factor, is the reference value of population diversity, For population diversity, is a preset positive number, is the reference value of fitness improvement rate, is the fitness improvement rate.

[0015] Furthermore, the populations of Harris Hawk population optimization method that inject the global optimal solution of deep reinforcement learning into the population include:

[0016] When deep reinforcement learning is fed back to the Harris Hawk population optimization method, the global optimal solution of deep reinforcement learning is used as the guiding center to reinitialize a set number of individuals in the Harris Hawk population, which can be expressed as:

[0017]

[0018] in, Harris Hawk i Individuals in t +1 moment position, is the global optimal solution for deep reinforcement learning. is a hyperparameter that controls the perturbation intensity, is a Gaussian random vector with mean 0 and covariance equal to the identity matrix I.

[0019] Furthermore, the high-quality solution sequences of the Harris Hawk population optimization method are added to the deep reinforcement learning experience pool, including:

[0020] When the Harris Hawk population optimization method feeds back to the deep reinforcement learning, a set number of high-quality solution sets are selected as the vehicle state and control decision sequence, and added to the deep reinforcement learning experience pool according to the weight factor, which is expressed as:

[0021]

[0022] in, w k For the k The weight factor of the high-quality solution, exp is the natural exponential function, To control the steepness parameter of the weight distribution, Optimization methods for Harris hawk populations in t The first k A high-quality solution, Optimization methods for Harris hawk populations in t The first mA high-quality solution, K is the number of high-quality solutions.

[0023] Furthermore, after adding the high-quality solution sequence of the Harris Hawk population optimization method to the deep reinforcement learning experience pool, the deep reinforcement learning performs weighted sampling in the experience pool according to the weight factor, which is expressed as:

[0024]

[0025] in, P is the sampling probability, sample For sampling operation, is the individual state of the jth state-action trajectory corresponding to the kth high-quality solution, is the individual action of the jth state-action trajectory corresponding to the kth high-quality solution, w k For the k The weight factor of a good solution, w m For the m The weight factor of a good solution, K is the number of high-quality solutions.

[0026] Furthermore, the policy gradient update formula for deep reinforcement learning is:

[0027]

[0028] in, is the policy objective function Parameters The gradient, P is the sampling probability, sample For sampling operation, is the state of the jth state-action trajectory corresponding to the kth high-quality solution, is the action of the jth state-action trajectory corresponding to the kth high-quality solution, is the gradient term of the logarithmic strategy, is the value function estimate, K is the number of high-quality solutions, J k is the number of discrete states on the path.

[0029] Furthermore, the calculation formula for the fitness improvement rate is:

[0030]

[0031] in, I t For the t The fitness improvement rate of each generation, For the t- The global optimal solution fitness of generation 1, For the t The global optimal solution fitness of the generation, max is the maximum value function, The default positive number.

[0032] Furthermore, the calculation formula for population diversity is:

[0033]

[0034] in, D t For the t Generational population diversity, For the i Individuals in t The position vector of the generation, For the t The mean vector of the generation population positions, N The number of items.

[0035] Furthermore, the position update method of the Harris Hawk optimization method is:

[0036]

[0037] in, Harris hawk populations in t +1 generation position, is the hawk position vector with the best fitness among all Harris hawks of the tth generation, is the escape energy factor, is the arithmetic mean of all eagle positions in the tth generation.

[0038] Furthermore, the objective function of deep reinforcement learning is:

[0039]

[0040] in, is the objective function, is the strategy parameter, is the expectation operator, For the t The discount factor for generations, In state Take action The instantaneous reward function.

[0041] The present invention has the following beneficial effects:

[0042] By combining the Harris Eagle optimization method with deep reinforcement learning, this paper leverages the strengths of both to achieve more efficient vehicle dynamic planning. Furthermore, the paper introduces a dynamic weight fusion mechanism and a two-way feedback mechanism to enhance the adaptability and collaborative optimization capabilities of vehicle dynamic path planning. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 A flowchart of a vehicle dynamic path planning method integrating Harris Hawk with deep reinforcement learning;

[0044] Figure 2 A comparison chart of different vehicle dynamic planning methods in a simple scenario;

[0045] Figure 3 A comparison chart of different vehicle dynamic planning methods in complex scenarios. DETAILED DESCRIPTION

[0046] The specific embodiments of the present invention are described below to facilitate understanding of the present invention by those skilled in the art. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the appended claims, these changes are obvious, and all inventions and creations utilizing the concepts of the present invention are protected.

[0047] This paper proposes a vehicle path planning method that integrates HHO and DRL, leveraging the global search advantages of HHO and the policy learning capabilities of DRL. This method goes a step further than simple superposition by introducing a dynamic weighted fusion mechanism. This mechanism flexibly adjusts the participation of HHO and DRL based on measurable indicators such as fitness improvement rate and population diversity, achieving an adaptive balance between global exploration and local exploitation. When the method is stuck in a local optimum or experiencing slow improvement, increasing the global exploration ratio of DRL can help it escape. When the method achieves significant progress or maintains high population diversity, enhancing the local search intensity of HHO can further refine the solution quality. Furthermore, this paper proposes a bidirectional feedback mechanism: after each iteration, information on high-quality solutions discovered by DRL is fed back to HHO to guide the next iteration in focusing on promising areas. Simultaneously, the high-quality solution sequences selected by HHO are weighted and injected into the DRL experience pool, accelerating DRL policy updates. Through this bidirectional, closed-loop information exchange, HHO and DRL achieve co-evolution, overcoming the shortcomings of their independent use and demonstrating greater adaptability in highly dynamic and complex path planning environments. This framework helps to achieve significant improvements in global-local search, adaptive control and real-time decision-making, and provides new ideas for the fusion research of HHO and DRL.

[0048] like Figure 1As shown, an embodiment of the present invention provides a vehicle dynamic path planning method integrating Harris Hawk with deep reinforcement learning, comprising the following steps S1 to S5:

[0049] S1. Initialize the positions of Harris Hawk population optimization method and policy network parameters of deep reinforcement learning;

[0050] In an optional embodiment of the present invention, step S1 initializes the position and velocity of the Harris hawk population, which is expressed as:

[0051]

[0052] in, Harris Hawk individual positions, For the population size.

[0053] Initializing policy network (actor) parameters in deep reinforcement learning and the value network (critic) parameters , providing a starting point for subsequent strategy updates and value estimation.

[0054] S2, calculate the escape energy function to determine whether to perform global exploration through Harris Hawk optimization or local development using deep reinforcement learning;

[0055] In an optional embodiment of the present invention, the Harris Hawk optimization method simulates the behavior of catching a rabbit, when escaping energy This is done when Determine the encirclement mode and pass the random number Confirm whether to dive. The escape energy function is as follows:

[0056]

[0057] in, is a random initial value of (-1, 1); is the current iteration number; is the total number of iterations.

[0058] When escaping energy When , a random walk is performed, and the update formula is:

[0059]

[0060]

[0061] in, is a random individual in the Harris's hawk population; is the individual position of the best value for Harris's hawk, i.e., the rabbit position; is the population average position; is the population size; All obey normal distribution Parameters; and are the upper and lower bounds of the problem being solved.

[0062] When escaping energy When , a local search is performed and four update methods are selected according to the energy value.

[0063] (1) When escaping energy value and When , the method performs soft encirclement, and its position update formula is:

[0064]

[0065]

[0066] in, , represents the update step size; Indicates the distance between the position with the highest fitness value and the current position.

[0067] (2) When escaping energy value and When the method performs hard encirclement, the rabbit will be directly killed, and its position update formula is:

[0068]

[0069] (3) When escaping energy value and When the rabbit cannot escape, Harris's hawk adopts a smarter soft encirclement strategy, specifically:

[0070]

[0071]

[0072]

[0073]

[0074] in, is the dimension of the question being asked, yes A random vector in dimension, It is the Levi flight formula, is a random number from (0, 1), The value is 1.5.

[0075] (4) When escaping energy value and When the rabbit is likely to escape, Harris's hawk adopts a gradual soft encirclement, specifically:

[0076]

[0077]

[0078] This embodiment is based on the Harris Hawk optimization method by escaping the energy function Determine the global exploration and local development behaviors during the current iteration:

[0079]

[0080] in, is a random number with a value range of (-1, 1). is the current iteration number, is the total number of iterations.

[0081] when When , global exploration is performed and the update formula is:

[0082]

[0083] in, is a randomly selected individual position, is a random number.

[0084] In the actor-critic method for policy-based deep reinforcement learning, a policy represented by a neural network can be updated via gradient ascent in both deterministic and stochastic scenarios. This typically requires an estimate of a value function. A common architecture is the actor-critic approach, where the actor is responsible for the policy and the critic is responsible for estimating the value function. In deep reinforcement learning, both can be represented by nonlinear neural networks. The actor uses the gradients from the policy gradient theorem to adjust the policy parameters, while the critic estimates an approximate value function for the current policy.

[0085] In vehicle dynamics programming, the Markov decision process (MDP) model is used to describe the transition process of a vehicle in different states. It is defined as follows:

[0086] (1) Status : Describes the specific situation of the vehicle at a certain moment, such as position, speed, etc.

[0087] (2) Action : Actions that the vehicle can take, such as acceleration, deceleration, steering, etc.

[0088] (3) State transition probability : From the status Take action Transfer to state probability.

[0089] (4) Reward Function : In state Take action The rewards received.

[0090] The state transition formula is:

[0091]

[0092] in, is the state transition probability, is the state at time t+1, is the state at time t, is the action taken at time t.

[0093] In vehicle dynamics planning, the state at the next moment depends only on the current state and action.

[0094] In reinforcement learning, we need to learn a policy (actor) and a value function (critic) .

[0095] (1) Value function Indicates that the status Take action The expected total reward received.

[0096]

[0097] in, Indicates expected value; Represents the discount factor of Power, Used to determine the present value of future rewards, The smaller it is, the smaller the future reward weight will be; Indicates time The reward received.

[0098] (2) In reinforcement learning, policy 𝜋 is the strategy used by the agent to determine which actions to take in different states. Policy 𝜋 is a probability distribution of taking different actions in a given state, representing the probability of taking different actions in state Take action The probability is as follows:

[0099]

[0100] in, Indicates time Status is Take action The conditional probability of .

[0101] Policy gradient methods optimize policies by maximizing the cumulative reward.

[0102] Objective function Gradient:

[0103]

[0104] in, Represents the objective function Parameters The gradient of , which indicates how to adjust the parameters To maximize the objective function ; Indicates expected value; Represents a parameterized strategy ; Indicates that in the strategy Next, status Take action The action value function of Start, take action The total reward expected to be obtained.

[0105] Update the policy parameters by gradient ascent To optimize the strategy:

[0106]

[0107] in, A parameter vector representing the strategy, which controls the behavior of the strategy; Represents the learning rate, which determines the step size of each update.

[0108] The Deep Q-Network (DQN) combines deep neural networks and Q-learning to estimate the Q-value function. The update formula of the Q-value function is:

[0109]

[0110] in, In state Take action of value, which represents the expected total reward for a given state and action; Indicates time The immediate rewards you receive; In state The maximum Q value under the condition represents the Q value of the best action for the next state.

[0111] DQN approximates the Q-value function through a deep neural network ,in are the parameters of the neural network. The specific steps are as follows:

[0112] Step 1 Neural Network Architecture

[0113] Use a deep neural network whose input state , output the Q values ​​corresponding to all possible actions.

[0114] Step 2: Calculation of target Q value

[0115] The target Q value is calculated using the Bellman equation:

[0116]

[0117] in, represents the target Q value; represents the immediate reward obtained at time t+; Indicates status Next, the target network parameters The maximum Q value obtained is the Q value of the best action for the next state.

[0118] Step 3 Loss Function

[0119] The loss function is defined as the mean square error between the current Q value and the target Q value:

[0120]

[0121] in, is the loss function.

[0122] Step 4: Gradient Descent

[0123] Update network parameters using gradient descent , as follows:

[0124]

[0125] in, Represents the loss function About parameters Gradient

[0126] In order to solve the problem of reward sparsity, reward shaping is introduced. The new reward function Defined as:

[0127]

[0128] in, is the target reward, It is an additional reward to guide the learning process.

[0129] The additional reward can be obtained through the potential function To define, the function is a state The function mapped to real numbers is used to measure the potential value of the state, as shown below:

[0130]

[0131] in, is the discount factor, Indicates the next state caused by action a. In this way, the additional reward can be added to the original reward In this way, the intelligent agent can obtain richer information during the learning process, thereby accelerating the learning speed.

[0132] In vehicle dynamics programming, the potential function needs to take into account the following factors:

[0133] 1. Target location: The target location that the vehicle needs to reach.

[0134] 2. Path smoothness: Ensure the smoothness and feasibility of the vehicle path.

[0135] 3. Obstacle avoidance: The vehicle needs to avoid obstacles.

[0136] 4. Continuity of states and actions: Considering the physical characteristics and dynamic limitations of the vehicle.

[0137] Taking the above factors into consideration, the present invention uses Euclidean distance as the potential function because it can clearly guide the vehicle to move towards the target location while combining the requirements of path smoothness and obstacle avoidance.

[0138]

[0139] in, represents the potential function, Indicates the current state, represents the target state, Represents the Euclidean distance between the current state and the target state; PathSmoothness(s) represents the measure of path smoothness, the smoother the better; ObstacleProximity(s) represents the proximity to obstacles, the farther away from obstacles the better; and Represents a weight factor that balances target location, path smoothness, and obstacle avoidance.

[0140] In deep reinforcement learning, the Markov decision process (MDP) model is used to describe the vehicle's transition process in different states. In each state, the actor network follows the strategy. Select an action, and the critic network estimates the value function of the current policy .

[0141] S3. Based on the dynamic weight fusion mechanism, the participation ratio of Harris Hawk population optimization method and deep reinforcement learning is dynamically adjusted according to the fitness improvement rate and population diversity;

[0142] In an optional embodiment of the present invention, in each iteration, the energy function To decide whether to use HHO's global exploration or DRL's local development. However, to achieve more flexible adaptive control, this embodiment introduces a dynamic weight fusion mechanism and a two-way feedback mechanism:

[0143] In order to better control the participation of HHO and DRL in different time periods, this embodiment defines a dynamic weight factor To balance The contribution ratio of the two in the iteration. There is an objective basis for the determination of population diversity. This study uses two quantitative indicators: "population diversity" and "fitness improvement rate".

[0144] Assumptions For the The global optimal solution fitness (objective function value) of the generation, the goal of the optimization problem is to make The smaller the better (or the closer to 0 the better). Define the relative improvement rate between adjacent generations:

[0145]

[0146] in, I t For the t The fitness improvement rate of each generation, For the t- The fitness of the global optimal solution of generation 1, For the t The global optimal solution fitness of the generation, max is the maximum value function, is a preset positive number, which is a very small positive number used to avoid the denominator being 0. When it is larger, it means that the current iteration can significantly improve the optimal solution, indicating that the method has made good progress in this iteration; when When it is very small or even close to 0, it means that the method has made limited progress in the latest round of iterations and may fall into a local optimum or stagnation state.

[0147] In order to measure the distribution of HHO population in the search space, the population diversity is defined as:

[0148]

[0149] in, D t For the t Generational population diversity, For the i Individuals in t The position vector of the generation, For the t The mean vector of the generation population positions, N is the number of items. is very small, indicating that the population has gathered in a small area and is prone to fall into local optimality; when It is large, indicating that the search still has high diversity and exploration space.

[0150] In order to increase the weight of DRL when the population diversity is low and the fitness improvement rate is small, in order to break the dilemma, when the fitness progress is good or the diversity is still high, HHO is preferred for more refined local mining. Taking all factors into consideration, the calculation method is:

[0151]

[0152] in, is the dynamic weight factor, is the reference value of population diversity, For population diversity, is a preset positive number, is the reference value of fitness improvement rate, is the fitness improvement rate. Smaller and When smaller, and will become larger, Approaching 1, thus increasing the global exploration ratio of DRL. On the contrary, when the improvement is obvious or the diversity is sufficient, Approaching 0, HHO dominates the generation.

[0153] Based on the above process, in each iteration, this embodiment passes Combine DRL (global exploration) and HHO (local exploitation) in a specific ratio and dynamically adjust them based on data during the iteration process rather than using a fixed threshold.

[0154] S4. Based on the two-way feedback mechanism, the global optimal solution of deep reinforcement learning is injected into the population of Harris Hawk swarm optimization method, and the high-quality solution sequence of Harris Hawk swarm optimization method is added to the deep reinforcement learning experience pool;

[0155] In an optional embodiment of the present invention, to enhance the deep collaboration between HHO and DRL during the search process, this study implements bidirectional feedback through information exchange at the end of each iteration. The core goal of bidirectional feedback is to transfer the excellent features obtained by DRL in global exploration to the HHO population, so that HHO can achieve more targeted local search in the next round; at the same time, the high-quality solutions found by HHO in local refinement are injected into the experience pool of DRL in the form of weighted samples, helping DRL to quickly focus on high-value state-action space. Through this closed-loop interaction, the two parties continue to promote each other in subsequent iterations, improving the overall optimization performance.

[0156] After this round of iterations, DRL obtains several better solutions in global exploration, such as finding a route segment that is significantly better than the previous generation in path planning. The global best solution obtained after this round of DRL convergence is In order to make HHO focus more on this excellent area in the next iteration, this solution can be used as the "guidance center" to reinitialize some individuals in the HHO population. is a proportional factor used to control the number of reinitialized individuals, such as individuals are reset, of which is the population size, which can be expressed as:

[0157]

[0158] in, Harris Hawk i Individuals in t +1 moment position, is the global optimal solution for deep reinforcement learning. is a hyperparameter that controls the perturbation intensity, is a Gaussian random vector with mean 0 and covariance equal to the identity matrix I. , is the set of individual indexes that need to be reinitialized. Through this formula, HHO will concentrate some individuals near the high-potential areas discovered by DRL in the next iteration, improving the pertinence and effectiveness of local search.

[0159] After HHO completes a round of iteration, a set of candidate solutions will be obtained, and the top % of good solution sets For path planning problems, these high-quality solutions can correspond to several state-action-reward sequences, such as discretely sampled vehicle states and control decision sequences along good paths. These sequences can be directly added to the DRL experience replay buffer. , training samples for subsequent policy updates.

[0160] In order to make DRL more inclined to utilize the good information conveyed by HHO during training, a higher sampling weight can be assigned to these newly added samples. Suppose the objective function of the problem is , for each HHO good solution Corresponding sequence, define the weight factor as follows:

[0161]

[0162] in, w k For the k The weight factor of the high-quality solution, exp is the natural exponential function, >0 is a parameter used to control the steepness of the weight distribution. Optimization methods for Harris hawk populations in t The first k A high-quality solution, Optimization methods for Harris hawk populations in t The first m A high-quality solution, K is the number of high-quality solutions. The larger the value, the higher the quality of the solution (the lower the objective function value), and thus it should be used preferentially in subsequent DRL training.

[0163] After injecting these weighted samples into the experience pool, DRL extracts When extracting samples for training, the Perform weighted sampling on the samples, that is:

[0164]

[0165] in, P is the sampling probability, sample For sampling operation, is the individual state of the jth state-action trajectory corresponding to the kth high-quality solution, is the individual action of the jth state-action trajectory corresponding to the kth high-quality solution, w k For the k The weight factor of a good solution, w m For the m The weight factor of a good solution, K is the number of high-quality solutions.

[0166] Through weighted sampling, the policy gradient update formula of DRL can be expressed as:

[0167]

[0168] in, is the policy objective function Parameters The gradient, P is the sampling probability, sample For sampling operation, is the state of the jth state-action trajectory corresponding to the kth high-quality solution, is the action of the jth state-action trajectory corresponding to the kth high-quality solution, is the gradient term of the logarithmic strategy, is the value function estimate, K is the number of high-quality solutions, J k is the number of discrete states on the path. State-action pairs corresponding to high-quality solutions will participate in policy updates with a higher probability, strengthening DRL's preference for high-value areas and thus accelerating policy convergence.

[0169] Through the mathematical description of the bidirectional feedback mechanism described above, this embodiment defines the implementation method for relocalizing individuals in the population in the DRL → HHO direction, and the quantitative metrics and sampling strategy for converting excellent solutions into weighted training samples in the HHO → DRL direction. This mechanism works in conjunction with the dynamic weight fusion mechanism: the former adjusts the proportion of HHO and DRL in each iteration, while the latter leverages the advantages of both through information feedback. This well-founded mathematical support facilitates direct implementation and verification of the effectiveness of this bidirectional closed-loop optimization strategy in practical experiments.

[0170] S5. After iterative optimization, the optimal vehicle path is output.

[0171] In an optional embodiment of the present invention, after the iteration is completed, the optimized vehicle path and corresponding action strategy are output. The output optimal path at this time incorporates the comprehensive advantages of HHO and DRL coordinated optimization under dynamic weighting and bidirectional feedback.

[0172] The position update method of Harris Hawk optimization method is:

[0173]

[0174] in, Harris hawk populations in t +1 generation position, is the hawk position vector with the best fitness among all Harris hawks of the tth generation, is the escape energy factor, is the arithmetic mean of all eagle positions in the tth generation.

[0175] The objective function of deep reinforcement learning is:

[0176]

[0177] in, is the objective function, is the strategy parameter, is the expectation operator, For the t The discount factor for generations, In state Take action The instantaneous reward function.

[0178] The method of the present invention is analyzed and explained below in conjunction with simulation experiments.

[0179] The experimental environment is as follows: operating system: Windows 11 (64-bit); processor: AMD Ryzen 7 5800H with Radeon Graphics, 3.20 GHz; running memory: 16G; simulation platform: Intellij IDEA.

[0180] To ensure the fairness of the experiment, each method was run for thirty rounds, and its average value and standard deviation were calculated.

[0181] This experiment uses a standard test function to verify the effectiveness of the following strategies: Harris Hawk Optimization (HHO) alone, Deep Reinforcement Learning (DRL) alone, and a fusion of HHO and DRL.

[0182] Standard test function Select the following 5 international standard test functions.

[0183] The Sphere function mainly tests the local search capability of the method. Due to its simple convex structure, the optimization method needs to perform a fine local search in the entire search space to find the optimal solution.

[0184] The Schwefel function tests the global search capability of the method. Since this function has multiple local optimal solutions, the optimization method needs to have strong global search capabilities to avoid falling into local optimal solutions.

[0185] The Rastrigin function tests the global and local search capabilities of a method. This function has multiple local optima, while the global optimal solution is distributed over a larger range, making it suitable for evaluating the performance of a method in complex search spaces.

[0186] The Ackley function mainly tests the global search capability of the method. This function has a deep global optimal basin, but many local optimal solutions are located in the periphery, requiring a strong global search capability to find the optimal solution.

[0187] The Griewank function tests the global and local search capabilities of an optimization method. This function has multiple global and local optimal solutions, requiring the optimization method to achieve a balance between global and local search capabilities.

[0188] HHO (Harris Hawks Optimization) parameter settings:

[0189] -Population Size: To ensure the diversity of the search space and avoid excessive computation, this value is set to 30.

[0190] -Max Iterations: To more comprehensively verify the algorithm performance, it is set to 1000 times.

[0191] DRL (Deep Reinforcement Learning) parameter settings:

[0192] -State Space: In the optimization problem, the dimension of the state space is equal to the number of variables and is set to 30.

[0193] -Action Space: For the direction and step size of movement in each dimension, use a discrete action space representation, and choose an action space of ±1 in each dimension.

[0194] -Reward Function: Set to the negative of the objective function value so that the minimization problem is transformed into maximizing the reward: -objective_function_value.

[0195] -Deep Learning Model (Neural Network Architecture): 3-layer fully connected neural network with 128 neurons in each layer.

[0196] -Learning Rate: Set to 0.0001.

[0197] -Discount Factor: Set to 0.95.

[0198] -Experience Replay: Use experience replay to alleviate sample correlation issues, and the experience pool size is set to 10,000.

[0199] -Target Network: Update frequency is set to 100.

[0200] Fusion method (HHO-DRL) parameter settings

[0201] The parameter settings of HHO and DRL are comprehensively considered; the population size is set to 30 and the maximum number of iterations is 1000. At the same time, in each iteration, HHO is first used to update the population, and then DRL is used to fine-tune the individuals.

[0202] The experimental results are shown in Table 1:

[0203] Table 1 Experimental comparison of different intelligent optimization algorithms

[0204]

[0205] As shown in Table 1, this experiment evaluated the performance of three different methods (HHO, DRL, and HHO-DRL) on five different test functions (f1 to f5). The theoretical optimal solution for each test function was 0, meaning the closer the optimization result is to 0, the better. The following is an in-depth analysis and discussion of the experimental results.

[0206] f1: Local search. In the local search function f1, the HHO method has an average of 4.28e-152 and a standard deviation of 3.33e-149, demonstrating extremely high stability and optimization performance. The DRL method has an average of 2.59e-96 and a standard deviation of 3.22e-95. Although the optimization results are relatively good, they lag significantly behind the HHO method. The HHO-DRL method has an average of 8.65e-97 and a standard deviation of 2.86e-97, outperforming the DRL method alone but not as well as the HHO method. This shows that the HHO method has a significant advantage in local search scenarios, likely because its local search strategy more effectively avoids the local optimum trap.

[0207] f2: Global search. In the global search function f2, the HHO method achieved an average of 6.59e-49 and a standard deviation of 8.77e-55, demonstrating strong optimization capabilities. The DRL method achieved an average of 8.29e-108 and a standard deviation of 2.85e-100, indicating that DRL performed poorly in global search. The HHO-DRL method achieved an average of 6.85e-48 and a standard deviation of 8.24e-48, outperforming HHO. This demonstrates that combining HHO and DRL in global search can better balance exploration and exploitation, improving optimization performance.

[0208] f3: Comprehensive search. In the comprehensive search function f3, the HHO method achieved an average of 3.59e-79 and a standard deviation of 2.89e-76, demonstrating good stability. The DRL method achieved an average of 1.00e-80 and a standard deviation of 1.50e-79, slightly outperforming the HHO method. The HHO-DRL method achieved an average of 7.25e-132 and a standard deviation of 5.22e-125, significantly outperforming both the HHO and DRL methods individually. This demonstrates that in comprehensive search scenarios, the HHO-DRL method effectively combines the strengths of both methods to achieve superior optimization results.

[0209] f4: Global search. For another global search function, f4, the HHO method achieved an average of 5.75e-50 and a standard deviation of 8.14e-49, demonstrating its strong global search capabilities. The DRL method achieved an average of 2.22e-65 and a standard deviation of 6.55e-49, far inferior to HHO. The HHO-DRL method achieved an average of 4.55e-50 and a standard deviation of 5.11e-48, achieving optimization results close to but slightly inferior to HHO. This indicates that, despite the strong performance of HHO-DRL in global search, pure HHO still has some advantages.

[0210] f5: Comprehensive search. In the comprehensive search function f5, HHO and DRL both achieved an average of 1.26e+04 and a standard deviation of 9.45e-01, showing nearly identical performance. The HHO-DRL method achieved an average of 5.00e+02 and a standard deviation of 5.00e-01, significantly outperforming both HHO and DRL. This indicates that in the comprehensive search scenario of f5, the HHO-DRL method significantly improves optimization results, likely due to its better balance between exploration and exploitation.

[0211] Analysis of Influencing Factors: The HHO method excels in both local and global search, particularly in avoiding local optima. The DRL method performs poorly in global search but is competitive in comprehensive search. The HHO-DRL method performs well in multiple scenarios, particularly in comprehensive search and partial global search, demonstrating the potential of combining the strengths of both methods.

[0212] Method characteristics: The HHO method is good at local search and avoids local optimality; the DRL method has adaptive learning capabilities and strong adaptability.

[0213] Problem characteristics: The characteristics of different test functions have a significant impact on the performance of the method, especially in terms of the complexity and multimodal nature of the function.

[0214] Method fusion: The HHO-DRL hybrid method achieves better optimization results by combining the local search capability of HHO and the global exploration capability of DRL.

[0215] The HHO-DRL method has demonstrated its effectiveness and advantages in a variety of search tasks, but its specific performance still needs to be selected according to the characteristics of the specific problem.

[0216] The complex functions in CEC2014 exhibit high dimensionality, multimodality, nonlinearity, and irregularity, placing higher demands on optimization methods. Optimization problems in high-dimensional spaces are more complex, requiring methods to possess strong search and convergence capabilities. Multimodal functions have multiple local optima, requiring methods to effectively escape local optima. Nonlinearity and irregularity make the optimization path uncertain, requiring methods to possess strong adaptability and exploration capabilities. This study evaluated the performance of methods using three different types of test functions: unimodal functions (CEC01 and CEC06), multimodal functions (CEC14 and CEC15), and hybrid functions (CEC20 and CEC21). The experimental parameters were set as follows: population size 30, dimension 30, iterations 500, and 15 independent rounds of computation. After each computation, the mean and standard deviation of the results were calculated. This setup allows for a comprehensive evaluation of the performance of methods on different problem types and ensures the stability and reliability of the experimental results, thereby enhancing the persuasiveness of the research. Table 2 shows a comparison of the performance of different intelligent optimization methods on the CEC2014 test functions.

[0217] Table 2 Performance comparison of different intelligent optimization algorithms on CEC2014 test functions

[0218]

[0219] As shown in Table 2, the study will analyze the method characteristics and data performance.

[0220] Different methods have different characteristics and optimization mechanisms, which directly impact their performance on complex functions. Particle swarm optimization (PSO) has strong global search capabilities, but can be less precise in local search and easily fall into local optima. Gray wolf optimization (GWO) simulates the hunting behavior of gray wolves. While it balances exploration and exploitation, it performs poorly on high-dimensional, complex functions. Whale optimization (WOA) simulates the hunting behavior of humpback whales and has strong global search capabilities, but its local search efficiency is low for complex functions. MHHO (modified Harris Hawk optimization) has certain advantages in high-dimensional spaces and complex functions, but its method complexity is relatively high. AHHO (adaptive Harris Hawk optimization) offers flexibility in parameter adjustment and performs well on complex functions. HHO-DRL (Harris Hawk optimization combined with deep reinforcement learning) performs particularly well on complex functions due to its adaptive learning and optimization capabilities.

[0221] In terms of data performance, different methods performed significantly differently on different functions. On the CEC01 function, HHO-DRL performed best, indicating that it has significant advantages in processing such complex functions. On the CEC06 function, PSO and GWO performed well, demonstrating their strong global search capabilities. On the CEC14 function, AHHO and HHO-DRL significantly outperformed other methods, indicating that their adaptive and learning capabilities are significantly superior on such functions. On the CEC15 function, HHO-DRL performed best, demonstrating its advantages in processing large-scale optimization problems. On the CEC20 function, PSO and WOA performed well, but the standard deviation was high, indicating that they have certain capabilities in global search, but poor stability. On the CEC21 function, AHHO and HHO-DRL performed well, indicating that they have significant advantages in processing complex nonlinear functions.

[0222] By analyzing the performance of different methods on various functions, the study can conclude the applicability of different methods. PSO is suitable for problems with high global search requirements, but may perform poorly on complex multimodal functions. GWO is suitable for problems that require a balance between global and local search, but may need improvement on high-dimensional complex functions. WOA is suitable for problems that require strong global search capabilities, but needs to be enhanced in local search. MHHO is suitable for high-dimensional complex functions, but the method complexity is high, and computational efficiency may need to be optimized. AHHO is suitable for complex nonlinear functions and multimodal functions, and has strong adaptability and flexibility. HHO-DRL is suitable for various complex functions, especially for high-dimensional and nonlinear problems, and its adaptability and optimization capabilities are particularly outstanding.

[0223] The performance differences between different methods on complex functions mainly stem from their optimization mechanisms and adaptability. HHO-DRL and AHHO perform well on most functions, mainly due to their adaptability and learning optimization capabilities.

[0224] Combining the results of the previous international standard test functions and the results of the CEC2014 test function, it can be seen that the AHHO method performs well on some standard test functions, but the performance of the HHO-DRL method is significantly improved in terms of the number of method iterations.

[0225] This experiment applied the HHO-DRL fusion method to vehicle path planning to further validate the improved method's ability to solve practical problems. The optimization performance of the HHO-DRL method was also highlighted by comparison with three other methods. Two scenarios were used: a simple scenario with a 30×30 map size and a 15% obstacle ratio; and a complex scenario with a 30×30 map size and a 30% obstacle ratio. The experimental method in this section was repeated ten times with a T=50 iteration count, and the average and standard deviation of the results were calculated. The methods selected for comparison were GWO, MHHO, and AHHO (Grey Wolf Optimization, Improved Harris Hawk Optimization, and Adaptive Harris Hawk Optimization).

[0226] The results of vehicle dynamic planning using different methods in simple scenarios are as follows: Figure 2 As shown, the horizontal and vertical coordinates are grid coordinates, and the unit is dimensionless discrete index. Figure 2 As shown in Figure 3, the path planned by the HHO-DRL method is the shortest, with the starting point being the yellow point and the end point being the blue point.

[0227] The results of vehicle dynamic planning using different methods in complex scenarios are as follows: Figure 3 As shown, the horizontal and vertical coordinates are grid coordinates, and the unit is dimensionless discrete index. Figure 3 As shown in Figure 3, the path planned by the HHO-DRL method is still the shortest.

[0228] In order to better compare the effects of this method with other methods, the optimization rate is introduced to represent the percentage change of the effect of this method relative to the effects of other methods.

[0229]

[0230] in, Indicates the length of the planning path of the improved method, represents the path length of the method before improvement, Indicates a method. Compare Small, then A positive value indicates that the optimization is effective.

[0231] At the same time, the study also introduced the efficiency ratio to measure the optimization effect.

[0232]

[0233] This formula indicates how many times the improved effect is the original effect. , it means that the improved method is more effective.

[0234] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0235] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0236] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0237] Specific embodiments are used in the present invention to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. At the same time, for those skilled in the art, according to the ideas of the present invention, there may be changes in the specific implementation methods and application scopes. In summary, the contents of this specification should not be understood as limiting the present invention.

[0238] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and it should be understood that the scope of protection of the present invention is not limited to such specific descriptions and embodiments. Those skilled in the art can make various other specific variations and combinations based on the technical teachings disclosed in the present invention without departing from the essence of the present invention, and such variations and combinations are still within the scope of protection of the present invention.

Claims

1. A vehicle dynamic path planning method integrating Harris Hawk with deep reinforcement learning, characterized in that: The following steps are involved: Initialize the positions of the Harris Hawk population optimization method and the policy network parameters of deep reinforcement learning; Calculate the escape energy function to determine whether to perform global exploration via Harris Hawk optimization or local exploitation using deep reinforcement learning; Based on the dynamic weight fusion mechanism, the participation ratio of Harris Hawk population optimization method and deep reinforcement learning is dynamically adjusted according to the fitness improvement rate and population diversity; the calculation formula of the dynamic weight factor in the dynamic weight fusion mechanism is: in, For the t +1 generation of dynamic weighting factors, is the reference value of population diversity, For the t +1 generation of population diversity, is a preset positive number, is the reference value of fitness improvement rate, For the t +1 generation fitness improvement rate; The calculation formula for fitness improvement rate is: in, I t For the t The fitness improvement rate of each generation, For the t- The global optimal solution fitness of generation 1, For the t The global optimal solution fitness of the generation, max is the maximum value function, is a preset positive number; Based on a two-way feedback mechanism, the global optimal solution of deep reinforcement learning is injected into the population of the Harris Hawk swarm optimization method, and the high-quality solution sequence of the Harris Hawk swarm optimization method is added to the deep reinforcement learning experience pool; After iterative optimization, the optimal vehicle path is output.

2. The vehicle dynamic path planning method integrating Harris Hawk and deep reinforcement learning according to claim 1 is characterized in that: The population that injects the global optimal solution of deep reinforcement learning into the Harris Hawk population optimization method includes: When deep reinforcement learning is fed back to the Harris Hawk population optimization method, the global optimal solution of deep reinforcement learning is used as the guiding center to reinitialize a set number of individuals in the Harris Hawk population, which can be expressed as: in, Harris Hawk i Individuals in t +1 generation position, is the global optimal solution for deep reinforcement learning. is a hyperparameter that controls the perturbation intensity, is a Gaussian random vector with mean 0 and covariance equal to the identity matrix I.

3. The vehicle dynamic path planning method integrating Harris Hawk and deep reinforcement learning according to claim 1 is characterized in that: Adding high-quality solution sequences of the Harris Hawk population optimization method to the deep reinforcement learning experience pool includes: When the Harris Hawk population optimization method feeds back to the deep reinforcement learning, a set number of high-quality solution sets are selected as the vehicle state and control decision sequence, and added to the deep reinforcement learning experience pool according to the weight factor, which is expressed as: in, w k For the k The weight factor of the high-quality solution, exp is the natural exponential function, To control the steepness parameter of the weight distribution, Optimization methods for Harris hawk populations in t The first k A high-quality solution, Optimization methods for Harris hawk populations in t The first m A high-quality solution, K The number of high-quality solutions.

4. The vehicle dynamic path planning method integrating Harris Hawk and deep reinforcement learning according to claim 1, characterized in that: After adding the high-quality solution sequence of the Harris Hawk population optimization method to the deep reinforcement learning experience pool, deep reinforcement learning performs weighted sampling in the experience pool according to the weight factor, which is expressed as: in, P is the sampling probability, sample For sampling operation, is the individual state of the jth state-action trajectory corresponding to the kth high-quality solution, is the individual action of the jth state-action trajectory corresponding to the kth high-quality solution, w k For the k The weight factor of a good solution, w m For the m The weight factor of a good solution, K The number of high-quality solutions.

5. The vehicle dynamic path planning method integrating Harris Hawk and deep reinforcement learning according to claim 4 is characterized in that: The policy gradient update formula for deep reinforcement learning is: in, is the policy objective function Parameters The gradient, P is the sampling probability, sample For sampling operation, is the state of the jth state-action trajectory corresponding to the kth high-quality solution, is the action of the jth state-action trajectory corresponding to the kth high-quality solution, is the gradient term of the logarithmic strategy, is the value function estimate, K is the number of high-quality solutions, J k is the number of discrete states on the path.

6. The vehicle dynamic path planning method integrating Harris Hawk and deep reinforcement learning according to claim 1, characterized in that: The formula for calculating population diversity is: in, D t For the t Generational population diversity, For the i Individuals in t The position vector of the generation, For the t The mean vector of the generation population positions, N The number of items.

7. The vehicle dynamic path planning method integrating Harris Hawk and deep reinforcement learning according to claim 1, characterized in that: The position update method of Harris Hawk optimization method is: in, Harris hawk populations in t +1 generation position, is the hawk position vector with the best fitness among all Harris hawks of the tth generation, is the escape energy factor, is the arithmetic mean of all eagle positions in the tth generation.

8. The vehicle dynamic path planning method integrating Harris Hawk and deep reinforcement learning according to claim 1, characterized in that: The objective function of deep reinforcement learning is: in, is the objective function, is the strategy parameter, is the expectation operator, For the t The discount factor for generations, In state Take action The instantaneous reward function.

Citation Information

Patent Citations

  • Multi-intelligent agent reinforcement learning path planning method based on ant colony algorithm

    CN112286203A

  • Path planning method based on heuristic deep reinforcement learning

    CN112325897A