Rule-learning fusion-based personalized lane changing decision-making method for expressway automatic driving vehicle
By integrating rules and learning methods into autonomous vehicles, a hybrid decision-making model is constructed, which solves the problems of low scenario coverage and poor safety of traditional methods, and achieves higher lane-changing success rate and support for personalized driving styles.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-30
- Publication Date
- 2026-04-10
AI Technical Summary
Existing lane-changing decision-making methods for autonomous vehicles have shortcomings in terms of low scenario coverage, rigid decision-making, poor safety, and difficulty in achieving personalized driving styles. Traditional rule-based methods are difficult to cope with complex scenarios, while the training process of deep reinforcement learning is unstable and the decision-making logic is not transparent.
A rule-based learning fusion approach is adopted, which combines vehicle dynamics model, model predictive control and differentiated driving style to construct discrete state equation and rule-based lane-changing strategy model. Combined with the reward function of reinforcement learning and fuzzy rules, a hybrid decision model is formed to improve the scenario adaptability and safety of lane changing.
It significantly improves the success rate and safety of lane changes, can adapt to a variety of complex dynamic scenarios, meets personalized driving needs, and enhances the user driving experience and the initiative of intelligent vehicles.
Smart Images

Figure CN121822484A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of autonomous vehicle behavior decision-making, and specifically designs a personalized lane-changing decision-making method for autonomous vehicles on highways based on rule-learning fusion. Background Technology
[0002] With the deep integration of artificial intelligence and the automotive industry, autonomous driving technology has become the core driving force for innovation in the transportation field. Among the many sub-tasks of autonomous driving, lane changing decision-making is a key link that reflects the intelligence of vehicles and affects driving safety and efficiency.
[0003] Traditional decision-making methods are mostly based on rules or optimization theory. While they are stable in some aspects, they are difficult to cover all scenarios in real-world traffic environments. For example, rule-based methods rely excessively on expert knowledge and require programmers to pre-define explicit logical rules. Their advantages are clear logic, ease of verification, and predictable behavior; however, their disadvantages are extremely prominent: the rule base cannot exhaust all complex scenarios, resulting in low scenario coverage, rigid decision-making, inability to handle highly interactive game situations, and difficulty in achieving personalized driving styles.
[0004] Learning-based methods (such as Deep Reinforcement Learning) involve agents continuously interacting with their environment to learn the optimal strategy that maximizes accumulated rewards through trial and error. Its advantages include strong generalization ability, suitability for personalization, no need for precise explicit environment models, the ability to directly learn complex mapping relationships in high-dimensional states from data, and better adaptability to unseen new scenarios. Its disadvantages include poor security and difficulty in verification; the training process of DRL is unstable; the learned policy is a deep neural network with opaque decision-making logic, potentially leading to inexplicable and dangerous behavior; and its security is highly dependent on the coverage of the simulated training environment and the design of the reward function, posing a fatal risk in under-trained or rare scenarios. Summary of the Invention
[0005] To overcome the shortcomings of the prior art, this invention provides a personalized lane-changing decision-making method for autonomous vehicles on highways based on rule learning fusion. It aims to effectively improve the scenario adaptability, decision reliability, and driving style diversity of intelligent vehicles by combining vehicle dynamics models, model predictive control, near-end strategy optimization, and differentiated driving styles. This will improve the safety and success rate of lane changing and better meet the personalized lane-changing needs of intelligent vehicles.
[0006] To achieve the above objectives, the present invention adopts the following technical solution: The present invention provides a personalized lane-changing decision-making method for autonomous vehicles on highways based on rule-learning fusion, characterized by its application to lane-changing or overtaking scenarios on one-way two-lane highways. The one-way two lanes of the highway consist of a driving lane and an overtaking lane. A vehicle coordinate system is established with the vehicle Ego in the driving lane as the origin, the forward direction of vehicle Ego as the positive X-axis, and the left side facing the direction of vehicle Ego's travel as the positive Y-axis. Vehicles ahead of vehicle Ego in the driving lane are designated as potential risk vehicles A, and vehicles behind vehicle Ego in the overtaking lane are designated as potential risk vehicles B. The personalized lane-changing decision-making method for autonomous vehicles on highways includes the following steps: Step 1: Obtain the state variables of the vehicle's Ego under the i-th driving style at time k. The control variables of the vehicle's Ego at time k under the i-th driving style And the perturbation term of the vehicle's Ego at time k under the i-th driving style. Thus, the discrete state equation of the vehicle Ego under the i-th driving style at time k is constructed; where, It is the longitudinal relative distance error between the vehicle Ego with the i-th driving style and the potentially risky vehicle B at time k. It is the longitudinal relative velocity error between the vehicle Ego with the i-th driving style and the potentially risky vehicle B at time k. It is the acceleration of the vehicle Ego at time k for the i-th driving style. It is the expected acceleration of the vehicle Ego at time k for the i-th driving style. It is the expected yaw rate of the vehicle Ego at time k under the i-th driving style. It is the acceleration of the potentially hazardous vehicle B at time k; Step 2: Establish a prediction model for the regular lane-changing strategy model based on the vehicle kinematics model; Step 3: Based on the prediction model of the rule-based lane-changing strategy model, construct and solve the rule-based lane-changing strategy model of the vehicle Ego under the i-th driving style at time k, and obtain the optimal rule control variable of the vehicle Ego in the prediction time domain P under the i-th driving style at time k. ; Step 4: Construct the reward function for the vehicle's Ego at time k under the i-th driving style. ; Step 5: Construct a learning-based lane-changing strategy network, including a policy network and a value network, to... As the state at time k, Action at time k After execution by the vehicle Ego, the state at time k+1 is obtained. and the reward at time k Thus, a sample is obtained. The sample pool is then used to train the learning-based lane-changing strategy network, resulting in a learning-based lane-changing strategy model of the vehicle's Ego under the i-th driving style containing optimal parameters, and outputting the optimal learning control variable. ; Step 6: Calculate the collision avoidance acceleration of the vehicle's Ego at time k under the i-th driving style. The optimal control deviation values of the vehicle's Ego model under the i-th driving style at time k. And as fuzzy variables, they are used to construct a fuzzy rule standard graph, thereby determining the weight coefficients of the rule-based lane-changing strategy model. The weighting coefficients of the learning-based lane-changing strategy model are 1- This forms a lane-changing strategy model based on rule learning fusion, and outputs the optimal control variables for the two optimal rule learning fusions under the i-th driving style at time k. .
[0007] The personalized lane-changing decision-making method for autonomous vehicles on highways based on rule-learning fusion, as described in this invention, is also characterized in that step 1 includes: Step 1.1: Obtain using equation (1) The state variables in: (1) In equation (1), Let Ego be the relative distance between the vehicle with the i-th driving style and the potentially risky vehicle B at time k. Let Ego be the expected distance between the vehicle with the i-th driving style and the potentially risky vehicle B at time k. It is the instantaneous speed of the potentially hazardous vehicle B at time k. Let Ego be the instantaneous speed of the vehicle with the i-th driving style at time k. Let Ego be the expected speed of the vehicle with the i-th driving style at time k. Step 1.2: Construct the discrete state equation using equation (2): (2) In equation (2), Let Ego represent the predicted state component of the vehicle with the i-th driving style at time k+1 in the X-axis direction. Let E represent the predicted state component of the vehicle Ego with the i-th driving style at time k+1 along the Y-axis. C, D, G, and E are four coefficient matrices, and we have: (3) (4) (5) (6) In equations (3)-(6), This represents the time step of the rule-based lane-changing strategy model. This represents the time constant of a rule-based lane-changing strategy model. The time gain for a rule-based lane-changing strategy model.
[0008] Furthermore, in step 2, the prediction model of the rule-based lane-changing strategy model is constructed using equation (7): (7) In equation (7), It is the output variable of the vehicle's Ego in the prediction time domain p from time k to time k+p under the i-th driving style. This represents the time series of the output variables of the vehicle's Ego under the i-th driving style from time k to time k+p in the prediction time domain p. Let Ego be the state variable of the vehicle in the prediction time domain p from time k to time k+p under the i-th driving style. Let Ego represent the time series of the state variables of the vehicle from time k to time k+p within the prediction time domain p under the i-th driving style. Let Ego represent the control variables of the vehicle from time k to time k+p within the prediction time domain p under the i-th driving style. Let Ego represent the time series of the control variables of the vehicle under the i-th driving style from time k to time k+p in the prediction time domain p. Let Ego represent the perturbation variable of the vehicle from time k to time k+p within the prediction time domain p under the i-th driving style. Let Ego represent the time series of the perturbation variables of the vehicle from time k to time k+p within the prediction time domain p under the i-th driving style. It is the coefficient matrix of the state variables. It is the coefficient matrix of the control variables. Let be the coefficient matrix of the disturbance variable, and we have: (8) (9) (10).
[0009] Furthermore, step 3 includes: Step 3.1: Construct the constraints of the rule-based lane-changing strategy model using equation (11): (11) In equation (11), Let Ego represent the minimum safe distance that vehicle Ego, with the i-th driving style, should maintain during lane changing at time k. Let Ego represent the lower limit of acceleration for the vehicle with the i-th driving style during lane changing. Let Ego represent the upper limit of acceleration for the vehicle with the i-th driving style during lane changes. Let Ego be the maximum boundary value for the acceleration transformation rate of the vehicle with the i-th driving style. Let Ego be the minimum boundary condition for the rate of change of acceleration of the vehicle with the i-th driving style. This represents the rate of change of acceleration of the vehicle Ego for the i-th driving style. , , These are the distance relaxation factor, velocity relaxation factor, and acceleration relaxation factor, respectively. Let Ego be the minimum distance relaxation coefficient for the vehicle with the i-th driving style. Let Ego be the maximum speed relaxation coefficient for the vehicle with the i-th driving style. Let be the minimum relaxation coefficient of the vehicle's Ego acceleration transformation rate for the i-th driving style. Let Ego be the maximum relaxation coefficient for the vehicle's acceleration transformation rate under the i-th driving style. Let Ego be the minimum safe distance offset for the vehicle with the i-th driving style during lane changing; Step 3.2: Construct the objective function of the regular lane-changing strategy model for the i-th driving style at time k using equation (12). : (12) In equation (12), It is the relaxation factor of the objective function in the rule-based lane-changing strategy model. Let Ego represent the time series of the state variables of the vehicle with the i-th driving style from time k to time k+j. Let represent the weight matrix of the state variables under the i-th driving style at time j. Let Ego represent the time series of the control variables for the vehicle with the i-th driving style from time k to time k+j. Let represent the weight matrix of the control variables under the i-th driving style at time j. It is the weight coefficient matrix of the objective function of the rule-based lane-changing strategy model; T represents the transpose.
[0010] Furthermore, step 4 includes: Step 4.1: Construct the reward function for the i-th driving style at time k using equation (13). ; (13) In equation (13), It is the safety reward function of the vehicle's Ego under the i-th driving style at time k, and is obtained from equation (14). It is the personalized reward function of the vehicle Ego under the i-th driving style at time k, and is obtained by equation (17); (14) In equation (14), This is the speed limit of the vehicle's Ego at time k under the i-th driving style. Let Ego be the collision risk of the vehicle at time k under the i-th driving style, and we have: (15) (16) In equations (15) and (16), It is the minimum distance at time k such that the rear of vehicle Ego and the front of potential risk vehicle B do not collide under the i-th driving style; (17) In equation (17), Represents the mean function, It is the weight of the personalized reward under the i-th driving style.
[0011] Furthermore, step 6 includes: Step 6.1: Calculate the collision avoidance acceleration at time k for the i-th driving style using equation (14). : (14) In equation (14), Let Ego represent the first derivative of the instantaneous velocity of the vehicle with the i-th driving style at time k. Let Ego, representing the vehicle with the i-th driving style, be the first derivative of the longitudinal relative velocity error between the vehicle with the i-th driving style and the potentially risky vehicle B at time k. Let Ego, representing the vehicle with the i-th driving style, be the second derivative of the relative distance between it and the potentially risky vehicle B at time k. It is the second derivative of the minimum safe distance boundary; Step 6.2: Calculate the optimal control variable deviation between the two models under the i-th driving style at time k using equation (15). ; (15) Step 6.3: According to , With weighting coefficients Based on the relationships between them, construct a fuzzy rule standard graph; Step 6.4: Calculate the optimal control quantity of the fusion strategy using equation (16) to learn the rules. : (16) This leads to a hybrid lane-changing decision model for autonomous vehicles that incorporates the driver's driving style.
[0012] The present invention provides an electronic device, comprising a memory and a processor, characterized in that the memory is used to store a program supporting the processor in executing the personalized lane-changing decision method for autonomous vehicles on highways, and the processor is configured to execute the program stored in the memory. The present invention discloses a computer-readable storage medium storing a computer program, characterized in that the computer program, when executed by a processor, performs the steps of the personalized lane-changing decision method for autonomous vehicles on highways.
[0013] Compared with the prior art, the beneficial effects of the present invention are reflected in: 1. This invention deeply integrates the constraint optimization capability of traditional model predictive control methods with the perception and decision-making capability of proximal policy optimization methods of deep reinforcement learning, which significantly improves the intelligence and adaptability of the decision-making system. By using reinforcement learning, it overcomes the shortcomings of traditional rules in handling environmental uncertainty and interactive game, and can cope with a wider range of complex dynamic scenarios, thereby improving the success rate of lane changing while ensuring vehicle safety.
[0014] 2. This invention establishes a lane-changing model that can adapt to different driving styles by classifying driver styles. This allows the designed lane-changing model to adapt to users with different driving styles, meet diverse driving needs, and improve the user's driving experience. At the same time, the model combines game theory methods to further enhance the initiative of intelligent vehicles during lane-changing. Attached Figure Description
[0015] Figure 1 This is a flowchart of the personalized lane-changing decision-making method for autonomous vehicles on highways based on rule-learning fusion according to the present invention. Figure 2 This is a network training logic diagram of the personalized lane-changing decision-making method for autonomous vehicles on highways based on rule-learning fusion of the present invention. Figure 3 This is a schematic diagram of the minimum safe distance model for the rule-learning fusion-based personalized lane-changing decision-making method for autonomous vehicles on highways, as presented in this invention. Figure 4 This is a schematic diagram of a game model for the personalized lane-changing decision-making method for autonomous vehicles on highways based on rule-learning fusion, as described in this invention. Figure 5This is a schematic diagram of the fuzzy rule graph of the personalized lane-changing decision-making method for autonomous vehicles on highways based on rule-learning fusion of the present invention. Detailed Implementation
[0016] In this embodiment, a rule-based learning fusion-based personalized lane-changing decision-making method for autonomous vehicles on highways is applied to the field of intelligent vehicle behavior decision-making. It integrates rule-based and learning-based methods to construct a safe lane-changing decision-making model that considers driver style in various complex scenarios. This is a rational approach to solving current decision-making challenges. This method leverages the powerful representational learning capabilities of DRL to address environmental uncertainties and achieve personalization; it also uses rules to ensure that decision-making behavior meets safety constraints, thus achieving an optimal balance between openness and reliability. Specifically, for example... Figure 1 As shown, the method includes the following steps: Step 1: Establish a system based on the one-way two-lane highway scenario, such as... Figure 2 The vehicle coordinate system shown represents the driving lane and the overtaking lane. The origin of the coordinate system is the vehicle Ego in the driving lane, the positive X-axis is the direction of travel of vehicle Ego, and the positive Y-axis is the left side facing the direction of travel of vehicle Ego. The vehicle in front of vehicle Ego in the driving lane is denoted as potential risk vehicle A, and the vehicle behind vehicle Ego in the overtaking lane is denoted as potential risk vehicle B.
[0017] Step 2: Obtain the original driver data and filter out the safe lane-changing data according to traffic rules. Then, using the natural driver model as the standard, use the K-means clustering method to divide the safe lane-changing data into n classes, each representing the driving style of n drivers. Then, design corresponding lane-changing methods for different styles using the following methods.
[0018] Step 3: Based on the i-th driving style from Step 1 and the vehicle kinematics model, determine the system state variables under the i-th driving style. and the control variables of the system under the i-th driving style The system disturbance is Using the system's state variables as system inputs, the discrete state equations are derived, where... It is the longitudinal relative distance error between car Ego (driving style i) and car B at time k. The longitudinal relative speed error between car Ego (driving style i) and car B at time k is... It is the acceleration of the Ego car with the i-th driving style at time k. It is the expected acceleration of the Ego car with the i-th driving style. It is the expected yaw rate of the Ego car with the i-th driving style. It is the acceleration of car B at time k.
[0019] Step 3.1: Obtain using equation (1) The state variables in: (1) In equation (1), Let Ego be the relative distance between the vehicle with the i-th driving style and the potentially risky vehicle B at time k. Let Ego be the expected distance between the vehicle with the i-th driving style and the potentially risky vehicle B at time k. It is the instantaneous speed of the potentially hazardous vehicle B at time k. Let Ego be the instantaneous speed of the vehicle with the i-th driving style at time k. Let Ego be the expected speed of the vehicle with the i-th driving style at time k.
[0020] Step 3.2: Construct the discrete state equation using equation (2): (2) In equation (2), Let Ego represent the predicted state component of the vehicle with the i-th driving style at time k+1 in the X-axis direction. Let E represent the predicted state component of the vehicle Ego with the i-th driving style at time k+1 along the Y-axis. C, D, G, and E are four coefficient matrices, and we have: (3) (4) (5) (6) In equations (3)-(6), This represents the time step of the rule-based lane-changing strategy model. This represents the time constant of a rule-based lane-changing strategy model. The time gain for a rule-based lane-changing strategy model.
[0021] Step 4: Establish a prediction model for the regular lane-changing strategy model by combining the vehicle kinematics model, such as... Figure 3 As shown, a lateral and longitudinal kinematic model of the vehicle is established by analyzing the motion relationship between the vehicle and potential risk vehicles, and the following pre-set conditions are used to ensure the stability and reliability of the decision.
[0022] Step 4.1: Construct the prediction model for the regular lane-changing strategy model using equation (7): (7) In equation (7), It is the output variable of the vehicle's Ego in the prediction time domain p from time k to time k+p under the i-th driving style. This represents the time series of the output variables of the vehicle's Ego under the i-th driving style from time k to time k+p in the prediction time domain p. Let Ego be the state variable of the vehicle in the prediction time domain p from time k to time k+p under the i-th driving style. Let Ego represent the time series of the state variables of the vehicle from time k to time k+p within the prediction time domain p under the i-th driving style. Let Ego represent the control variables of the vehicle from time k to time k+p within the prediction time domain p under the i-th driving style. Let Ego represent the time series of the control variables of the vehicle under the i-th driving style from time k to time k+p in the prediction time domain p. Let Ego represent the perturbation variable of the vehicle from time k to time k+p within the prediction time domain p under the i-th driving style. Let Ego represent the time series of the perturbation variables of the vehicle from time k to time k+p within the prediction time domain p under the i-th driving style. It is the coefficient matrix of the time series of the output variables of the vehicle Ego under the i-th driving style in the prediction time domain p from time k to time k+p. It is the coefficient matrix of the time series of the control variables of the vehicle Ego under the i-th driving style, from time k to time k+p in the prediction time domain p. Let be the coefficient matrix of the time series of the perturbation variables of the vehicle's Ego under the i-th driving style, from time k to time k+p within the prediction time domain p, and we have: (8) (9) (10) Step 5: Based on the prediction model of the rule-based lane-changing strategy model, construct and solve the rule-based lane-changing strategy model of the vehicle Ego under the i-th driving style at time k, and obtain the optimal rule control variable of the vehicle Ego in the prediction time domain P under the i-th driving style at time k. .
[0023] Step 5.1: Construct the constraints of the rule-based lane-changing strategy model using equation (11): (11) In equation (11), Let Ego represent the minimum safe distance that vehicle Ego, with the i-th driving style, should maintain during lane changing at time k. Let Ego represent the lower limit of acceleration for the vehicle with the i-th driving style during lane changing. Let Ego represent the upper limit of acceleration for the vehicle with the i-th driving style during lane changes. Let Ego be the maximum boundary value for the acceleration transformation rate of the vehicle with the i-th driving style. Let Ego be the minimum boundary condition for the rate of change of acceleration of the vehicle with the i-th driving style. This represents the rate of change of acceleration of the vehicle Ego for the i-th driving style. , , These are the distance relaxation factor, velocity relaxation factor, and acceleration relaxation factor, respectively. Let Ego be the minimum distance relaxation coefficient for the vehicle with the i-th driving style. Let Ego be the maximum speed relaxation coefficient for the vehicle with the i-th driving style. Let be the minimum relaxation coefficient of the vehicle's Ego acceleration transformation rate for the i-th driving style. Let Ego be the maximum relaxation coefficient for the vehicle's acceleration transformation rate under the i-th driving style. Let Ego be the minimum safe distance offset for the vehicle with the i-th driving style during lane changing.
[0024] Step 5.2: Construct the objective function of the rule-based lane-changing strategy model using equation (12). : (12) In equation (12), It is the relaxation factor of the objective function in the rule-based lane-changing strategy model. Let Ego represent the weight matrix of the time series of the state variables of the vehicle with the i-th driving style from time k to time k+j. Let Ego represent the weight matrix of the control variables for the i-th driving style vehicle from time k to time k+j. It is the weight coefficient matrix of the objective function of the rule-based lane-changing strategy model; T represents the transpose.
[0025] Step 6: Construct the reward function for the vehicle's Ego at time k under the i-th driving style. ; Step 6.1: Reward function constructed using equation (13) ; (13) In equation (13), It is the safety reward function of the vehicle's Ego under the i-th driving style at time k, and is obtained from equation (14). It is the personalized reward function of the vehicle Ego under the i-th driving style at time k, and is obtained by equation (17).
[0026] (14) In equation (14), This is the speed limit of the vehicle's Ego at time k under the i-th driving style. Let Ego be the collision risk of the vehicle at time k under the i-th driving style, and we have: (15) (16) In equations (15) and (16), It is the minimum distance at time k such that the rear of the vehicle Ego with the i-th driving style does not collide with the front of the potentially risky vehicle B. (17) In equation (17), Represents the mean function, It is the weight of personalized rewards.
[0027] Step 7: Construct a learning-based lane-changing strategy network, including a policy network and a value network, to... As the state at time k, Action at time k After execution by the vehicle Ego, the state at time k+1 is obtained. and the reward at time k Thus, a sample is obtained. A sample pool is constructed to train the learning-based lane-changing strategy network, resulting in a learning-based lane-changing strategy model of the vehicle's Ego under the i-th driving style containing optimal parameters, and outputting the optimal learning control variable. The training process is as follows Figure 4 As shown.
[0028] Step 7.1: Initialize the parameters of the Actor policy network and the Critic value network respectively, and initialize the sample pool at the same time; Step 7.2: Use the state information of the vehicle Ego and the potentially at-risk vehicle B at time k. As the observation space for reinforcement learning, the control variables in the lane-changing decision of the vehicle Ego are used. This serves as the action space for the vehicle's Ego reinforcement learning.
[0029] Step 7.3: Determine the actions of the vehicle Ego at time k. The input is fed into the training environment, and after the autonomous vehicle Ego executes the action, the state at time k+1 is obtained. And calculate the reward at time k based on the above reward function, and apply the data experience. Store the data in the experience revisit buffer to build a sample pool; Step 7.4: Sample a batch of empirical data from the empirical replay buffer to train the Actor network and the Critic network.
[0030] Step 7.5: Based on the sampled empirical data, calculate the loss of the value network and update the network parameters of the Critic value network through backpropagation; Step 7.6: Calculate the loss of the policy network using the value assessment output of the Critic network, and backpropagate to update the policy network parameters of the Actor; Step 7.7: Continue training until the model converges.
[0031] Step 8: Calculate the collision avoidance acceleration of the vehicle's Ego at time k under the i-th driving style. The optimal control deviation values of the vehicle's Ego model under the i-th driving style at time k. And as fuzzy variables, they are used to design fuzzy rules, thereby determining the weight coefficients of the rule-based lane-changing strategy model. The weighting coefficients of the learning-based lane-changing strategy model are 1- This forms a lane-changing strategy model based on rule learning fusion, and outputs the optimal control variables for optimal rule learning fusion. .
[0032] Step 8.1: Calculate the collision avoidance acceleration using equation (14). get: (14) In equation (14), Let Ego represent the first derivative of the instantaneous velocity of the vehicle with the i-th driving style at time k. Let Ego, representing the vehicle with the i-th driving style, be the first derivative of the longitudinal relative velocity error between the vehicle with the i-th driving style and the potentially risky vehicle B at time k. Let Ego, representing the vehicle with the i-th driving style, be the second derivative of the relative distance between it and the potentially risky vehicle B at time k. It is the second derivative of the minimum safe distance boundary.
[0033] Step 8.2: Calculate the deviation of the optimal control variables for the two models using equation (15). (15) Step 8.3: According to , With weighting coefficients The relationships between them are used to construct a fuzzy rule standard graph, such as Figure 5 As shown, the k-th time can be found using this fuzzy rule standard graph. , Corresponding weight coefficients The optimal control quantity used in equation (16) to calculate the rule-learning fusion strategy .
[0034] Step 8.4: Calculate the optimal control quantity of the fusion strategy using Equation (16). : (16) This leads to a hybrid lane-changing decision model for autonomous vehicles that incorporates the driver's driving style (e.g., Figure 1 As shown in the figure, the model can guarantee that when the collision risk is not high, the hybrid decision-making strategy is in a fully learning decision-making mode; the higher the collision risk and the smaller the expected output difference between the two decision-making strategies, the more difficult it is for the learning-based decision-making strategy to guarantee the driving safety of the vehicle. At this time, according to the fuzzy rules, the weight coefficient is changed to increase the weight of the rule-based decision-making strategy, and the outputs of the two strategies are linearly superimposed through the weights to achieve hybrid decision-making.
[0035] In this embodiment, an electronic device includes a memory and a processor. The memory stores a program that supports the processor in executing the above-described method, and the processor is configured to execute the program stored in the memory.
[0036] In this embodiment, a computer-readable storage medium stores a computer program, which is executed by a processor to perform the steps of the above method.
Claims
1. A personalized lane-changing decision-making method for autonomous vehicles on highways based on rule-learning fusion, characterized in that, This method is applied to lane-changing or overtaking scenarios on one-way two-lane highways. The one-way two lanes of the highway consist of a driving lane and an overtaking lane. A vehicle coordinate system is established with the vehicle Ego in the driving lane as the origin, the forward direction of vehicle Ego as the positive X-axis, and the left side facing the direction of vehicle Ego's travel as the positive Y-axis. Vehicles ahead of vehicle Ego in the driving lane are designated as potential risk vehicles A, and vehicles behind vehicle Ego in the overtaking lane are designated as potential risk vehicles B. The personalized lane-changing decision-making method for autonomous vehicles on highways includes the following steps: Step 1: Obtain the state variables of the vehicle's Ego under the i-th driving style at time k. The control variables of the vehicle's Ego at time k under the i-th driving style And the perturbation term of the vehicle's Ego at time k under the i-th driving style. Thus, the discrete state equation of the vehicle Ego under the i-th driving style at time k is constructed; where, It is the longitudinal relative distance error between the vehicle Ego with the i-th driving style and the potentially risky vehicle B at time k. It is the longitudinal relative velocity error between the vehicle Ego with the i-th driving style and the potentially risky vehicle B at time k. It is the acceleration of the vehicle Ego at time k for the i-th driving style. It is the expected acceleration of the vehicle Ego at time k for the i-th driving style. It is the expected yaw rate of the vehicle Ego at time k under the i-th driving style. It is the acceleration of the potentially hazardous vehicle B at time k; Step 2: Establish a prediction model for the regular lane-changing strategy model based on the vehicle kinematics model; Step 3: Based on the prediction model of the rule-based lane-changing strategy model, construct and solve the rule-based lane-changing strategy model of the vehicle Ego under the i-th driving style at time k, and obtain the optimal rule control variable of the vehicle Ego in the prediction time domain P under the i-th driving style at time k. ; Step 4: Construct the reward function for the vehicle's Ego at time k under the i-th driving style. ; Step 5: Construct a learning-based lane-changing strategy network, including a policy network and a value network, to... As the state at time k, Action at time k After execution by the vehicle Ego, the state at time k+1 is obtained. and the reward at time k Thus, a sample is obtained. The sample pool is then used to train the learning-based lane-changing strategy network, resulting in a learning-based lane-changing strategy model of the vehicle's Ego under the i-th driving style containing optimal parameters, and outputting the optimal learning control variable. ; Step 6: Calculate the collision avoidance acceleration of the vehicle's Ego at time k under the i-th driving style. The optimal control deviation values of the vehicle's Ego model under the i-th driving style at time k. And as fuzzy variables, they are used to construct a fuzzy rule standard graph, thereby determining the weight coefficients of the rule-based lane-changing strategy model. The weighting coefficients of the learning-based lane-changing strategy model are 1- This forms a lane-changing strategy model based on rule learning fusion, and outputs the optimal control variables for the two optimal rule learning fusions under the i-th driving style at time k. .
2. The personalized lane-changing decision-making method for autonomous vehicles on highways based on rule-learning fusion as described in claim 1, characterized in that: Step 1 includes: Step 1.1: Obtain using equation (1) The state variables in: (1) In equation (1), Let Ego be the relative distance between the vehicle with the i-th driving style and the potentially risky vehicle B at time k. Let Ego be the expected distance between the vehicle with the i-th driving style and the potentially risky vehicle B at time k. It is the instantaneous speed of the potentially hazardous vehicle B at time k. Let Ego be the instantaneous speed of the vehicle with the i-th driving style at time k. Let Ego be the expected speed of the vehicle with the i-th driving style at time k. Step 1.2: Construct the discrete state equation using equation (2): (2) In equation (2), Let Ego represent the predicted state component of the vehicle with the i-th driving style at time k+1 in the X-axis direction. Let E represent the predicted state component of the vehicle Ego with the i-th driving style at time k+1 along the Y-axis. C, D, G, and E are four coefficient matrices, and we have: (3) (4) (5) (6) In equations (3)-(6), This represents the time step of the rule-based lane-changing strategy model. This represents the time constant of a rule-based lane-changing strategy model. The time gain for a rule-based lane-changing strategy model.
3. The personalized lane-changing decision-making method for autonomous vehicles on highways based on rule-learning fusion as described in claim 1, characterized in that, Step 2 involves using equation (7) to construct a prediction model for a rule-based lane-changing strategy: (7) In equation (7), It is the output variable of the vehicle's Ego in the prediction time domain p from time k to time k+p under the i-th driving style. This represents the time series of the output variables of the vehicle's Ego under the i-th driving style from time k to time k+p in the prediction time domain p. Let Ego be the state variable of the vehicle in the prediction time domain p from time k to time k+p under the i-th driving style. Let Ego represent the time series of the state variables of the vehicle from time k to time k+p within the prediction time domain p under the i-th driving style. Let Ego represent the control variables of the vehicle from time k to time k+p within the prediction time domain p under the i-th driving style. Let Ego represent the time series of the control variables of the vehicle under the i-th driving style from time k to time k+p in the prediction time domain p. Let Ego represent the perturbation variable of the vehicle from time k to time k+p within the prediction time domain p under the i-th driving style. Let Ego represent the time series of the perturbation variables of the vehicle from time k to time k+p within the prediction time domain p under the i-th driving style. It is the coefficient matrix of the state variables. It is the coefficient matrix of the control variables. Let be the coefficient matrix of the disturbance variable, and we have: (8) (9) (10)。 4. The personalized lane-changing decision-making method for autonomous vehicles on highways based on rule-learning fusion as described in claim 3, characterized in that, Step 3 includes: Step 3.1: Construct the constraints of the rule-based lane-changing strategy model using equation (11): (11) In equation (11), Let Ego represent the minimum safe distance that vehicle Ego, with the i-th driving style, should maintain during lane changing at time k. Let Ego represent the lower limit of acceleration for the vehicle with the i-th driving style during lane changing. Let Ego represent the upper limit of acceleration for the vehicle with the i-th driving style during lane changes. Let Ego be the maximum boundary value for the acceleration transformation rate of the vehicle with the i-th driving style. Let Ego be the minimum boundary condition for the rate of change of acceleration of the vehicle with the i-th driving style. This represents the rate of change of acceleration of the vehicle Ego for the i-th driving style. , , These are the distance relaxation factor, velocity relaxation factor, and acceleration relaxation factor, respectively. Let Ego be the minimum distance relaxation coefficient for the vehicle with the i-th driving style. Let Ego be the maximum speed relaxation coefficient for the vehicle with the i-th driving style. Let be the minimum relaxation coefficient of the vehicle's Ego acceleration transformation rate for the i-th driving style. Let Ego be the maximum relaxation coefficient for the vehicle's acceleration transformation rate under the i-th driving style. Let Ego be the minimum safe distance offset for the vehicle with the i-th driving style during lane changing; Step 3.2: Construct the objective function of the regular lane-changing strategy model for the i-th driving style at time k using equation (12). : (12) In equation (12), It is the relaxation factor of the objective function in the rule-based lane-changing strategy model. Let Ego represent the time series of the state variables of the vehicle with the i-th driving style from time k to time k+j. Let represent the weight matrix of the state variables under the i-th driving style at time j. Let Ego represent the time series of the control variables for the vehicle with the i-th driving style from time k to time k+j. Let represent the weight matrix of the control variables under the i-th driving style at time j. It is the weight coefficient matrix of the objective function of the rule-based lane-changing strategy model; T represents the transpose.
5. The personalized lane-changing decision-making method for autonomous vehicles on highways based on rule-learning fusion as described in claim 1, characterized in that, Step 4 includes: Step 4.1: Construct the reward function for the i-th driving style at time k using equation (13). ; (13) In equation (13), It is the safety reward function of the vehicle's Ego under the i-th driving style at time k, and is obtained from equation (14). It is the personalized reward function of the vehicle Ego under the i-th driving style at time k, and is obtained by equation (17); (14) In equation (14), This is the speed limit of the vehicle's Ego at time k under the i-th driving style. Let Ego be the collision risk of the vehicle at time k under the i-th driving style, and we have: (15) (16) In equations (15) and (16), It is the minimum distance at time k such that the rear of vehicle Ego and the front of potential risk vehicle B do not collide under the i-th driving style; (17) In equation (17), Represents the mean function, It is the weight of the personalized reward under the i-th driving style.
6. The personalized lane-changing decision-making method for autonomous vehicles on highways based on rule-learning fusion as described in claim 1, characterized in that, Step 6 includes: Step 6.1: Calculate the collision avoidance acceleration at time k for the i-th driving style using equation (14). : (14) In equation (14), Let Ego represent the first derivative of the instantaneous velocity of the vehicle with the i-th driving style at time k. Let Ego, representing the vehicle with the i-th driving style, be the first derivative of the longitudinal relative velocity error between the vehicle with the i-th driving style and the potentially risky vehicle B at time k. Let Ego, representing the vehicle with the i-th driving style, be the second derivative of the relative distance between it and the potentially risky vehicle B at time k. It is the second derivative of the minimum safe distance boundary; Step 6.2: Calculate the optimal control variable deviation between the two models under the i-th driving style at time k using equation (15). ; (15) Step 6.3: According to , With weighting coefficients Based on the relationships between them, construct a fuzzy rule standard graph; Step 6.4: Calculate the optimal control quantity of the fusion strategy using equation (16) to learn the rules. : (16) This leads to a hybrid lane-changing decision model for autonomous vehicles that incorporates the driver's driving style.
7. An electronic device comprising a memory and a processor, characterized in that, The memory is used to store programs that support the processor in executing the personalized lane-changing decision method for autonomous vehicles on highways as described in any one of claims 1-6, and the processor is configured to execute the programs stored in the memory.
8. A computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is run by the processor, it executes the steps of the personalized lane-changing decision method for autonomous vehicles on highways as described in any one of claims 1-6.