An intelligent vehicle decision-making method for signal-free intersections
By constructing a signal-free crossing intelligent vehicle decision model based on LSTM and PPO, the efficiency and safety balance problem in the signal-free crossing scenario in the existing technology is solved, and more efficient and safe vehicle traffic decisions are achieved.
Patent Information
- Application Number
- CN202510371231.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-27
AI Technical Summary
The prior art is difficult to balance efficiency and safety in the signal-free crossroad scenario, and reinforcement learning algorithms have problems such as high computational overhead and overconservative decision strategies in vehicle trajectory prediction and conflict point calculation.
A strategy network and value network composed of LSTM neural network and a fully connected layer are adopted, and a signal-free cross-border intelligent vehicle decision model is constructed in combination with the PPO algorithm. By obtaining vehicle status and traffic information in real time, the optimal accelerator, steering wheel angle and brake control signals are output.
It reduces the conservatism of the decision model, improves the vehicle traffic efficiency at signal-free intersections, enhances the exploration and risk tolerance of the decision system, and reduces the model calculation overhead.
Smart Images

Figure CN119889073B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent driving decision-making, and particularly relates to an intelligent vehicle decision-making method for unsignalized intersections. Background Art
[0002] In urban roads, unsignalized intersections are common traffic scenarios. Such intersections do not have traffic lights or other control devices, and rely on the mutual cooperation and real-time judgment among drivers to complete vehicle passage. However, there are significant challenges in vehicle passage at unsignalized intersections. Traditional rule-based vehicle decision-making methods are difficult to effectively cope with the highly complex traffic environment. Especially in the case of large traffic flow, unclear right-of-way, or significant differences in driving behaviors, it is easy to lead to low traffic efficiency, limited traffic capacity, and even safety accidents. In addition, due to the dynamic characteristics of intersections, the judgment and decision-making process of drivers within intersections is full of uncertainties, further exacerbating the complexity of traffic management. To improve the traffic efficiency and safety of unsignalized intersections, automated decision-making based on intelligent driving has become a research hotspot. In traditional technologies, rule algorithms or trajectory planning methods are often used for intersection decision-making, mainly including rule-based decision-making methods, trajectory optimization-based methods, and machine learning-based prediction models. Although these methods have certain effects in specific scenarios, they often show problems of being too conservative or too aggressive in the unsignalized intersection scenario, and it is difficult to achieve a balance between efficiency and safety. In addition, traditional methods usually rely on manually configured rules or limited data samples, and it is difficult to adapt to the dynamic and complex traffic characteristics of intersections.
[0003] In recent years, with the development of deep learning and reinforcement learning, more and more research has begun to explore data-driven methods. Through simulating a large number of real scenarios, reinforcement learning algorithms can gradually learn optimal decision-making strategies. However, when dealing with intersection decision-making problems, existing reinforcement learning algorithms usually need to calculate conflict points. Since the accuracy of conflict point calculation is usually based on the accuracy of the model's prediction of the vehicle trajectories of participants, existing reinforcement learning algorithms will lead to problems such as too simple vehicle trajectory prediction process, difficult algorithm convergence, relatively high computational overhead costs, and too conservative vehicle decision-making strategies resulting in reduced traffic efficiency. Summary of the Invention
[0004] The purpose of the present invention is to overcome the defects of the prior art and provide an intelligent vehicle decision-making method for unsignalized intersections, which can reduce the conservatism of the decision-making model during the decision-making process and improve the vehicle passing efficiency of unsignalized intersections on the premise of ensuring safety.
[0005] The technical solution provided by the present invention is as follows:
[0006] An intelligent vehicle decision-making method for unsignalized intersections, comprising:
[0007] Build an intelligent vehicle decision-making model for an intersection without signals. The decision-making model includes: a policy network and a value network;
[0008] Among them, both the policy network and the value network are composed of an LSTM neural network and a fully connected layer;
[0009] The vehicle information from t - n to t, the traffic participating vehicle information, the relative position information of the traffic intersection and the vehicle itself, and the relative position information of the target task point and the vehicle itself are combined to form state parameters as the input variables of the policy network, and the throttle control signal, the steering wheel angle control signal, and the brake control signal corresponding to the state parameters at time t are combined to form action parameters as the output variables of the policy network;
[0010] The state parameters and the action parameters are input into the value network, and the value network outputs the value corresponding to the state parameters and the action parameters;
[0011] Use the PPO algorithm to construct the loss functions of the policy network and the value network respectively, and train the decision-making model. Optimize the parameters of the policy network and the value network respectively with the minimum of the loss function as the optimization goal to obtain the optimal decision-making model;
[0012] When the vehicle travels to an intersection without signals, the state parameters are obtained in real time and input into the optimal decision-making model, and the optimal throttle control signal, steering wheel angle control signal, and brake control signal are output by the optimal decision-making model.
[0013] Preferably, the loss function of the policy network is:
[0014] ;
[0015] Among them, is the clipping function, is the importance sampling ratio, is the parameter of the policy network, is the policy advantage function, and are respectively the state and the corresponding action at time represents the clipping threshold; represents taking the expectation.
[0016] Preferably, the loss function of the value network is:
[0017] ;
[0018] Among them, is the state value function. For the input state , the output is the value of the current state; is the actual value at the moment.
[0019] Preferably, the calculation formula of the actual value is:
[0020] ;
[0021] where is the value advantage function;
[0022] ;
[0023] , , are respectively the moment, the moment, the moment, the TD error at the moment;
[0024] The calculation formula of the TD error at the moment is:
[0025] ;
[0026] In the formula, is the actual reward at the moment, is the value of the corresponding state output by the value network , is the attenuation factor.
[0027] Preferably, the actual reward is calculated through the following reward function:
[0028] ;
[0029] where is the task completion reward value, is the collision penalty value, is the progressive distance reward value, is the progressive collision risk penalty value, is the reasonable speed reward value, is the smoothness reward value.
[0030] Preferably, the calculation formula of the progressive reward is:
[0031] ;
[0032] where is the value of the current time step, is the normalized distance, , , is the weight coefficient; is the distance from the host vehicle to the target position at the initial moment, is the distance from the host vehicle to the target position at time
[0033] Preferably, the calculation method of the collision risk progressive penalty value includes the following steps:
[0034] Step 1: Calculate the relative position vector , ) according to the position of the host vehicle ( , ) and the position of the traffic participating vehicle ( ; Calculate the relative velocity vector , ) according to the speed of the host vehicle ( , ) and the speed of the traffic participating vehicle ( );
[0035] Step 2: Calculate the relative velocity of the vehicle along the relative position direction:
[0036]
[0037] Step 3: Calculate the predicted time to collision :
[0038]
[0039] Step 4: Determine the collision risk progressive penalty value according to the predicted time to collision:
[0040] ;
[0041] Among them, , , , are adjustment parameters, and satisfy .
[0042] Preferably, the number of layers of the LSTM neural network adopted by the policy network and the value network is 1 layer, and the number of hidden units is 128.
[0043] Preferably, the number of neurons in the fully connected layer adopted by the policy network and the value network is 64.
[0044] Preferably, the activation function adopted by the fully connected layer of the policy network is Tanh, and the activation function adopted by the fully connected layer of the value network is ReLU.
[0045] The beneficial effects of the present invention are as follows:
[0046] The intelligent vehicle decision-making method for unsignalized intersections provided by the present invention can reduce the conservativeness of the decision-making model during the decision-making process, and improve the vehicle passing efficiency of unsignalized intersections on the premise of ensuring safety. At the same time, by designing a reward function more suitable for the intersection scenario, the present invention can improve the understanding ability of the decision-making system for tasks, increase the model convergence speed, save storage and computing resources, reduce the model computing overhead, and further improve the exploration and risk tolerance of the decision-making system. Brief Description of the Drawings
[0047] Figure 1 It is a flowchart of the intelligent vehicle decision-making method for unsignalized intersections described in the present invention.
[0048] Figure 2 It is a schematic diagram of the intersection decision-making model based on Markov chain described in the present invention.
[0049] Figure 3 It is a schematic diagram of the relative speed calculation process of the vehicle along the relative position direction described in the present invention. Detailed Embodiment
[0050] The following further describes the present invention in detail with reference to the drawings, so that those skilled in the art can implement it according to the description in the specification.
[0051] As Figure 1 shown, the present invention provides an intelligent vehicle decision-making method for unsignalized intersections, including the following implementation processes.
[0052] I. Information Acquisition
[0053] The traffic information of the unsignalized intersection is obtained through the cooperation of on-vehicle sensors and communication devices, including: the vehicle obtains its own global state information such as position, speed, and orientation angle through GPS and inertial measurement unit (IMU); a three-dimensional perception map of the surrounding environment is constructed using lidar (LiDAR) and cameras to identify the intersection coordinates and surrounding vehicles; the relative distance and relative speed of other vehicles are accurately measured through millimeter-wave radar; combined with V2X (Vehicle-to-Everything) communication technology, the state information (such as position, speed, intention) of other vehicles is obtained to enhance the perception effect.
[0054] II. Constructing a Vehicle Decision-Making Model for Unsigned Intersections
[0055] As Figure 2 shown, in this embodiment, a vehicle decision-making model for unsignalized intersections is constructed based on Markov chain, and the specific process is as follows.
[0056] 2.1 Construction of the State Space
[0057] State Space The purpose is to express the state of the vehicle itself. The selection of the state space will directly affect the agent's understanding of the environment. For the environment of an intersection without signals, the present invention combines the PPO-LSTM algorithm to construct a state space that matches the characteristics of the algorithm, so as to give full play to the advantages of the algorithm in dealing with time series problems.
[0058] The state space of the present invention mainly consists of the following four parts:
[0059] (1) Information of the vehicle itself: the speed of the vehicle itself , acceleration , heading angle .
[0060] (2) Information of traffic participating vehicles: the relative distance from the traffic participating vehicle to the vehicle itself , angle , speed , acceleration and the heading angle of the traffic participating vehicle . Among them, is the serial number of the traffic participating vehicle.
[0061] Among them, with the center of the vehicle itself as the origin, the east-west direction of the intersection as the x-axis, and the north-south direction as the y-axis to establish a geodetic coordinate, the angle between the line connecting the center point of the vehicle itself and the center point of the traffic participating vehicle and the y-axis is the angle , and the angle between the speed direction of the traffic participating vehicle and the y-axis is the heading angle of the traffic participating vehicle .
[0062] (3) Information of the traffic intersection structure and task target points: the distance from the center point of the intersection to the vehicle itself , angle , the distance from the task target point to the vehicle itself , angle .
[0063] Among them, with the center of the intersection as the origin, the east-west direction as the x-axis, and the north-south direction as the y-axis to establish a geodetic coordinate, the angle between the line connecting the center point of the vehicle itself and the center point of the intersection and the y-axis is the angle , and the angle between the line connecting the center point of the vehicle itself and the task target point and the y-axis is the angle .
[0064] (4) Historical information: All state information except historical information in , that is, the information in (1), (2), and (3) corresponding to the time from t-1 to t-n. For the convenience of distinguishing from the state space, are respectively denoted as Here it is recommended (taking 5 here) is used to record the historical states of traffic participating vehicles, facilitating the agent to accurately predict the trajectories of traffic participating vehicles.
[0065] The finally determined state space is:
[0066] .
[0067] 2.2 Construction of the action space
[0068] Action space is the set of actions that the agent can select when in state . In the present invention, the action space of the agent is defined as the throttle, the steering angle of the steering wheel steer, and the brake brake, and their ranges are [0, 1], [-1, 1], and [0, 1] respectively. Each time the three action instructions of the agent can be arbitrarily selected within the above three ranges. In actual vehicle control, the actual output is the maximum output of the vehicle in this capability multiplied by the action value given by the model. For example: actual output throttle signal = maximum output value of vehicle throttle * action output by the agent. In a real vehicle, the throttle signal is converted into the torque output of the engine through the ECU (engine control unit), the steering signal is executed by the electric power steering system (EPS) or the steering gear, and the brake signal is controlled by the brake control unit (BCU) and acts on the hydraulic or electronic brake. The finally determined action space is:
[0069] .
[0070] 2.3 Construction of the reward function
[0071] Reward function Generally, there are two construction forms. The first is to define , and in this construction method, the reward at the current time step only depends on the state at the current moment, that is, a reward is given if the vehicle is in a safe interaction position, and a penalty is given if the vehicle is in a dangerous interaction position. The second reward construction method is to define , and in this construction method, the reward at the current time step not only depends on the state at the current moment, but also is related to the response actions made by the agent in the current state. Since the environment where the agent is located in the intersection with poor signals is complex and changeable, it is inevitable that the second definition method cannot cover all situations. Therefore, the first reward construction method is adopted in the present invention . The quality of the formulated reward function will directly affect the agent's judgment on whether the current action is correct. Therefore, formulating a reward function evaluation system that conforms to the scenario of the present invention is one of the key issues of the present invention. The reward function formulated in the present invention mainly includes the following parts:
[0072] 2.3.1 One-time Reward
[0073] Task Completion Reward: A one-time reward is given when the agent completes the task. This is to encourage the agent to complete the task.
[0074]
[0075] Among them, is an adjustment parameter, which is adjusted according to the actual task and is usually taken as a relatively large positive number. It is recommended that take 500 - 1000.
[0076] Collision Penalty: When the agent collides within the intersection, a one-time large penalty is given to prevent the agent from colliding.
[0077]
[0078] Among them, is an adjustment parameter, which is adjusted according to the actual task and is usually taken as a relatively large negative number. It is recommended that take -500 - -1000.
[0079] 2.3.2 Progressive Reward
[0080] (1) Set the progressive reward for the agent's distance from the target point. The specific calculation process is as follows. Assume that the position of the agent at time is , the position of the agent at time is , and the position of the task target point is , . First, calculate the distance from the agent to the target position at the initial time , and the distance from the agent to the target position at any time . The calculation formula is as follows:
[0081]
[0082]
[0083] Then normalize the distance at time, and denote the normalized distance as . The calculation formula is as follows:
[0084]
[0085] Finally, obtain the calculation formula for the progressive distance reward value:
[0086] ;
[0087] Among them, and are weight coefficients, which are determined by debugging according to the task scenario, the time required to complete the task, and the weights between various rewards. Their values may be different in different experimental scenarios. is the value at the current time step. is the normalized distance; in this formula, increasing the value of can encourage the agent to complete the task efficiently, but if the value of is too large, it may lead to aggressive vehicle behavior. When the value of exceeds a certain level, it will cause the progressive reward weight to exceed the collision reward weight, resulting in the agent ignoring the collision risk interaction. And increasing the value of will weaken this phenomenon. Therefore, it is crucial to balance the values between and . After experimental debugging, in a situation with a duration of 50 ms of , it is more appropriate that and satisfy .
[0088] (2) Progressive penalty for collision risk
[0089] In existing algorithms, the Time to Collision (TTC) is often used to measure the progressive collision risk. TTC refers to the time required for the ego vehicle to collide with the object in front (usually other vehicles, pedestrians, or obstacles) without taking any avoidance measures. Its calculation formula is:
[0090] ;
[0091] Among them, is the current distance between the ego vehicle and the target object. is the relative speed between the ego vehicle and the target object. Usually, it is the speed difference between the ego vehicle and the target object. In the intersection problem, usually, the collision point is predicted first. is the distance from the ego vehicle to the collision point. is the speed of the ego vehicle. However, since the prediction process of the collision point by this method is too simple, usually, the tangent intersection of the current speeds of each vehicle is used to predict the collision point, which may not be accurately predicted in the complex environmental interaction at intersections. Therefore, a new method is proposed in the present invention, that is, using the relative position and the relative position change rate of the vehicle for risk assessment. The specific method is as follows:
[0092] The first step: Calculate the relative position vector according to the position of the ego vehicle and the positions of traffic participating vehicles , then calculate the relative velocity vector , ), and the vehicle speeds of traffic participants ( , ), and calculate the relative velocity along the relative position direction using the relative velocity vector and the relative distance position vector. The calculation process is as shown, and the calculation formula is as follows: Figure 3 ;
[0093] ;
[0094] Finally, calculate the predicted time to collision (the time to collision is calculated by the following formula and represents the time from a certain moment when the vehicle maintains its current state unchanged until it collides with other vehicles). Note that here, to avoid the denominator being equal to 0, when the speed meets the condition, the speed is assigned the value ;
[0095] ;
[0096] And divide the vehicle position into low-risk area, medium-risk area, high-risk area and risk-free area according to the value of , and punish the agent by dividing the risk area. The specific division is as follows:
[0097] ;
[0098] The reward expression here is as follows:
[0099] ;
[0100] Here, , , , are adjustment parameters, which are adjusted according to the actual task, usually take negative numbers, and satisfy .
[0101] The recommended values here are: , , , .
[0102] 2.3.3 Reward for reasonable speed and smoothness:
[0103] This part of the reward content mainly aims to improve the comfort and efficiency of riding in intelligent driving vehicles. To encourage the vehicle to quickly pass through intersections, a small reward is given when the vehicle speed is between 3 - 30 km / h. At the same time, when the vehicle speed is greater than 40, a certain penalty is imposed. Also, to ensure riding comfort, when the absolute value of the jerk and the absolute value of the yaw rate are greater than 0.3, a certain penalty is given to the vehicle.
[0104] The reward expression here is:
[0105] When the vehicle speed is between 3 - 30 :
[0106] ;
[0107] When the vehicle speed exceeds 40 :
[0108] ;
[0109] Here, , are adjustment parameters, which are adjusted according to the actual task. is a positive number, is a negative number.
[0110] When the absolute value of the acceleration or the change rate of the yaw angle is greater than 0.3:
[0111] ;
[0112] Here, is an adjustment parameter, which is adjusted according to the actual task. is a negative number. The recommended values here are: j = 2, k = -10, l = -8.
[0113] Finally, the reward function expression at time is obtained:
[0114] ;
[0115] III. Construct a decision-making model and an evaluation model based on the LSTM neural network
[0116] To improve the decision-making level of the intelligent agent in the dynamic and changing signal-free intersection scenario and overcome the limitations of traditional MLP neural networks in dealing with complex time-series tasks, the present invention constructs an Actor-Critic architecture based on the LSTM neural network, where both the policy network (Actor network) and the value network (Critic network) introduce the LSTM structure to process the time-series information in the state space. This design enables the intelligent agent to more effectively capture the dependency relationship between the historical state and the current state, thus making more reasonable decisions in a dynamic environment.
[0117] 3.1 Data Input
[0118] Status Input: Input the status information of the previous step into the model
[0119] 3.2 LSTM Layer Design
[0120] Function of LSTM Layer: Extract the temporal features of the input status sequence and capture the dynamic relationships between states. Output the hidden state as the input for the subsequent network
[0121] LSTM Parameter Configuration:
[0122] Number of Hidden Units: 128
[0123] Number of Layers: 1
[0124] Dropout: 0.2, used to prevent overfitting
[0125] 3.3 Critic Network Construction
[0126] Network Structure:
[0127] The output of the LSTM layer is passed to the fully connected layer to generate the state value function
[0128] Loss Function of the State Value Function:
[0129]
[0130] Loss Function of the Action Value Function:
[0131] ;
[0132] Among them, is the state value function. For the input state , the output is the value of the current state, is the actual state value at time is the action value function. For the input state , taking the action , the output is the value obtained from the current state-action; is the actual action value at time. E represents taking the expectation
[0133] Output: : The evaluation value of the state
[0134] 3.4 Actor Network Construction
[0135] Network structure:
[0136] The output of the LSTM layer passes through a fully connected layer (FC) and is mapped to the action space, and the output is the probability distribution of the action or the specific action value.
[0137] Action selection: The output of the model is the mean and variance of the action, and the final action is obtained by sampling from a Gaussian distribution.
[0138] Among them, the specific network model parameters are shown in Table 1.
[0139] Table 1 Network model parameter table
[0140]
[0141] Here, the difference between batch size and n_step is explained. n_step represents the number of samples of the model. For example, if n_step is taken as 512, it means taking 512 time-step data from the environment. Batch size represents the number of samples used for each optimization, that is, the data used for parameter update is not all 512 data but the number of samples in the data, that is, the number of batch size, which is 128 here. n_epochs represents how many rounds of optimization are performed with the data sampled each time. Total_step represents the total number of training steps, which is 1*10^6 here.
[0142] IV. Parameter update based on the PPO algorithm
[0143] 4.1 Value function
[0144] In order to obtain a state value function and a policy value function with high evaluation accuracy, the present invention takes the state value function as an example to discuss the obtaining process of this function.
[0145] In the first step, the loss function of the state value function is defined. The loss function is a deviation function for measuring the deviation between the output value of the value network and the actual value. Our goal is to make the loss function take the minimum value, and finally achieve the situation where the value prediction is consistent with the actual value or the difference is within an acceptable range. The definition formula of the loss function is as follows:
[0146]
[0147] Among them, is the state value function. For the input state , the output is the value of the current state, is the actual value calculated by the Generalized Advantage Estimation (GAE) method. It can be understood as the actual value calculated by the GAE method when the model is at time. Note that the The calculation process can only be performed after the model completes a round of tasks. The calculation method is to record the action, state, reward value and other information of each time step, and then use the GAE method to calculate it. The specific calculation process is as follows:
[0148] First, calculate the TD error :
[0149]
[0150] in, yes The actual reward at the moment, in order to facilitate readers to understand the formula, we can Disassembly:
[0151]
[0152] in, It can be understood as the calculated value of the reward obtained by the model at the current moment. Type In the calculation formula, we can use It is understood as the difference between the model's calculated reward at the current moment and the actual reward.
[0153] Step 2: Calculate the value advantage function , and its calculation formula is as follows:
[0154] ;
[0155] , , They are time, time, time, Moment Error. Usually the value is the training step length n_step, which refers to the number of times the algorithm collects a set of states, actions, rewards, etc. from the environment and then uses these data to update the model parameters. For example, n_step=2048, which means that each training uses the data collected in 2048 time steps for multiple optimizations. The number of optimizations depends on n_epochs. If n_epochs is 10, it means that the 2048 data will be used for 10 training optimizations.
[0156] The PPO algorithm uses the Monte Carlo method to calculate this process. is the attenuation factor, parameter is the balance factor of GAE, usually taking values between 0 and 1 (e.g., 0.95), which controls the balance between bias and variance.
[0157] In the third step, calculate :
[0158]
[0159] After obtaining it is substituted into the loss function, and the value of the loss function can be calculated. According to the calculated loss function, the parameters of the value network are updated by the method of gradient descent θ . The specific steps are as follows:
[0160] (1) Take the partial derivative of all θ in the loss function to find the rate of change of the loss function with respect to each parameter .
[0161] (2) Use the obtained gradient for parameter update, and its update formula is:
[0162]
[0163] where is the learning rate, which controls the step size of each update. is the parameter of the current network. Thus, we have obtained the complete calculation process for the state value . The calculation method for the policy value function is exactly the same as the above process, and no further explanation will be given here.
[0164] 4.2 Policy Function
[0165] First, introduce the policy function. The PPO algorithm uses a probability-based policy, that is, given a state , the policy network calculates the probability of each possible action . Finally, the policy function will output the probability distribution of the agent taking each action in the state. For a state , the policy output of PPO is:
[0166]
[0167] To update the parameters of the policy function, first define the policy advantage function , which is used to measure how good the agent's action is in the state . Its calculation formula is as follows:
[0168]
[0169] Next, define the importance sampling ratio , which is used to measure the similarity between the new and old policies. Its calculation formula is as follows:
[0170]
[0171] In the formula, the numerator is the probability of selecting action in the new policy when in state , and the denominator is the probability of selecting action in the old policy when in state .
[0172] Finally, define the loss function, and its definition formula is as follows:
[0173]
[0174] Among them, clip is the clipping function, represents the clipping threshold, which is used to limit the amplitude of policy update; usually takes 0.1 or 0.2; the usage here is that when the value of is within , , take the original value, when is less than , the output result is , when is greater than , the output result is , and the final result takes the smaller value of and . This definition of the loss function limits the amplitude of parameter update by introducing the clipping function. Specifically, in the case where the clipping function does not play a role, the policy gradient is directly calculated according to and A ( , ) in the loss function. When the clipping function plays a role, it means that the difference between the new and old policies is too large, and the loss function becomes a constant without , resulting in a gradient of 0, which limits the gradient update amplitude and avoids drastic policy updates.
[0175] Specific update of the parameters in the policy function:
[0176] (1) Take the partial derivative of all the parameters of the policy function in the loss function to find the rate of change of the loss function with respect to each parameter ;
[0177] (2) Update the parameters using the calculated gradient, and the update formula is as follows:
[0178]
[0179] When the loss function of the algorithm finally converges and drops below the convergence criterion, the training ends. By combining the performance of the model and the convergence status and convergence value of the reward function, it is possible to determine whether the training effect of the model meets the requirements. If the effect is not good, the parameters can be adjusted to finally achieve the ideal effect.
[0180] By improving the reinforcement learning algorithm, the present invention embeds the long short-term memory network (LSTM) into the policy network architecture of the PPO algorithm, achieving the purpose of making more efficient decisions at unsignalized intersections using the historical state data of shuttle vehicles; by constructing a Markov decision model for the unsignalized intersection scenario that cooperates with the algorithm, the understanding ability of the policy model for tasks is improved; by adopting the method of online reinforcement learning (collecting data in real time and learning and updating during the interaction with the environment), the obtained model has the ability of self-learning, can continuously improve the traffic efficiency during the interaction, enhance the exploration and risk tolerance of the decision-making model for the unsignalized intersection scenario; at the same time, it can improve the model convergence speed, save storage and computing resources, and reduce the model computing overhead. Finally, it achieves the purpose of improving the traffic strategy of vehicles at intersections and reducing the conservatism of the decision-making process.
[0181] Although the embodiments of the present invention have been disclosed above, it is not limited to the applications listed in the specification and embodiments. It can be fully applied to various fields suitable for the present invention. For those familiar with the field, additional modifications can be easily made. Therefore, without departing from the general concept defined by the claims and their equivalents, the present invention is not limited to the specific details and the examples shown and described herein.
Claims
1. A method for intelligent vehicle decision making at an unsignalized intersection, characterized in that: include: Constructing an intelligent vehicle decision model for an unsignalized intersection, the decision model comprising: a strategy network and a value network; Wherein, the strategy network and the value network are both composed of LSTM neural networks and fully connected layers; The information of the ego vehicle from time tn to time t, the information of the participating vehicles in the traffic, the relative position information of the traffic intersection and the ego vehicle, and the relative position information of the target task point and the ego vehicle form the state parameters as the input variables of the strategy network, and the throttle control signal, steering wheel angle control signal and brake control signal corresponding to the state parameters at time t form the action parameters as the output variables of the strategy network; The state parameter and the action parameter are input into the value network, and the value network outputs the value corresponding to the state parameter and the action parameter; The PPO algorithm is used to construct the loss functions of the policy network and the value network respectively, and the decision model is trained. The parameters of the policy network and the value network are optimized respectively with the minimum loss function as the optimization goal to obtain the optimal decision model; When the vehicle drives to an unsignaled intersection, state parameters are obtained in real time and input into the optimal decision model, and the optimal decision model outputs an optimal throttle control signal, a steering wheel angle control signal, and a brake control signal; The reward function used in the value network is: R t =R1+R2+R3+R4+R5+R6; Among them, R1 is the task completion reward value, R2 is the collision penalty value, R3 is the progressive distance reward value, R4 is the collision risk progressive penalty value, R5 is the reasonable speed reward value, and R6 is the stability reward value; The calculation formula of the progressive distance reward is: R3=c*(1-d nom )-timestep*d; Among them, timestep is the value of the current time step, d nom is the normalized distance, d nom =d t / d0, c and d are weight coefficients; d0 is the distance from the vehicle to the target position at the initial moment, d t is the distance from the vehicle to the target position at time t; The method for calculating the progressive penalty value of the collision risk comprises the following steps: Step 1: Based on the position of the vehicle (x0, y0) and the position of the participating vehicles (x i ,y i ) calculates the relative position vector According to the vehicle speed (V 0x , V 0y ) and the speed of the participating vehicles (V ix , V iy ) Calculate the relative velocity vector Step 2: Calculate the relative speed v of the vehicle along the relative position direction d : Step 3: Calculate the estimated collision time t p : Step 4: Determine the progressive penalty value of collision risk based on the expected collision time: Among them, e, f, g, and h are adjustment parameters, and they satisfy e>f>g>h.
2. The intelligent vehicle decision-making method for an unsignalized intersection according to claim 1, characterized in that: The loss function of the policy network is: L policy =E[min(r t (θ)A(s t ,a t ),clip(r t (θ), 1-ε, 1+ε)A(s t ,a t ))]; Among them, clip is the clipping function, r t (θ) is the importance sampling ratio, θ is the policy network parameter, A(s t , a t ) is the strategy advantage function, s t and a t are the state and the corresponding action at time t, ε represents the clipping threshold, and E represents the expectation.
3. The intelligent vehicle decision-making method for an unsignalized intersection according to claim 2 is characterized in that: The loss function of the value network is: Among them, V θ (s t ) is the state value function, for the input state S t , the output is the value of the current state; is the actual value at time t.
4. The intelligent vehicle decision-making method for an unsignalized intersection according to claim 3 is characterized in that: The actual value is calculated as follows: in, is the value advantage function; δ t , δ t+1 , δ t+2 , δ t+T are the TD errors at time t, time t+1, time t+2, and time t+T respectively; The calculation formula of TD error at time t is: δ t =R t +γV θ (s t+1 )-V θ (s t ); In the formula, R t is the actual reward at time t, V θ (s t+1 ) is the corresponding state S of the value network output t+1 The value of , γ is the attenuation factor.
5. The intelligent vehicle decision-making method at an unsignalized intersection according to claim 4 is characterized in that: The LSTM neural network used by the strategy network and the value network both has 1 layer and 128 hidden units.
6. The intelligent vehicle decision-making method at an unsignalized intersection according to claim 5, characterized in that: The number of neurons in the fully connected layer used by the strategy network and the value network is 64.
7. The intelligent vehicle decision-making method at an unsignalized intersection according to claim 6 is characterized in that: The activation function used by the fully connected layer of the strategy network is Tanh, and the activation function used by the fully connected layer of the value network is ReLU.
Citation Information
Patent Citations
Collaborative decision-making control method for passage of autonomous vehicles at non-signalized intersection
CN115909778A
Networked vehicle cooperative control method based on deep reinforcement learning
CN118690786A
Multi-agent unmanned driving decision-making method and system for intersection scene without signal lamp
CN118918720A
Network information age optimization method and system based on meta-deep reinforcement learning
CN119135551A