Vehicle decision optimization method based on reinforcement learning and related equipment
By obtaining lane-changing vehicle data and using reinforcement learning models to determine and optimize lane-changing actions, the problem of low accuracy in decision-making of autonomous vehicles is solved, and more efficient vehicle behavior analysis and decision-making optimization are achieved.
Patent Information
- Application Number
- CN202510418134.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-08
AI Technical Summary
The existing decision-making methods for autonomous driving vehicles have low accuracy in dealing with traffic simulation environments, mainly due to the ineffective combination of deep learning and reinforcement learning library, the lack of flexibility and versatility of hard-coded methods, the in-depth processing of vehicle trajectory and metadata, insufficient analysis of vehicle characteristics and behavior patterns, strong inadaptation of keyframe range determination, and difficult to accurately capture vehicle behavior.
By obtaining lane change data of multiple lane change vehicles, the reinforcement learning model is used to determine the target vehicle for dangerous lane change behavior, calculate the reward value of predicted lane change actions, and optimize the reinforcement learning model to obtain the final lane change action, and control the vehicle to perform lane change.
The accuracy of decision-making of autonomous vehicles is improved, and the reliability of information management is improved by analyzing vehicle behavior patterns, calculating reward values for optimization, and a reasonable final lane change action is obtained, which improves the accuracy and safety of vehicle decision-making.
Smart Images

Figure CN120270268A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent transportation technology, and particularly relates to a vehicle decision-making optimization method based on reinforcement learning and related devices. Background Art
[0002] In the field of autonomous driving, existing research has many deficiencies in dealing with vehicle reinforcement learning decisions in traffic simulation environments. Traditional methods often rely only on a small number of toolkits, and do not effectively combine deep learning and reinforcement learning libraries with traffic simulation software, which limits the processing ability for complex scenarios. For environmental variables, hard coding is mostly used, lacking flexibility and generality, and compatibility problems are likely to occur, affecting the stability of subsequent operations. The processing of vehicle trajectories and metadata is only for simple storage, without in-depth conversion and preprocessing, and the data format is difficult to meet the requirements of complex operations. The grouped management of vehicle data is relatively rough, without considering vehicle characteristics and behavior patterns, and it is impossible to deeply analyze the performance of different vehicles. The determination of the key frame range uses fixed rules or simple time intervals, lacking adaptability to vehicles and scenarios, and it is difficult to accurately capture key information, which limits the accurate simulation and analysis of vehicle behavior. It can be seen that the current vehicle decision-making method has the problem of low accuracy in autonomous vehicle decision-making. Summary of the Invention
[0003] This application provides a vehicle decision-making optimization method based on reinforcement learning and related devices, which can solve the problem of low accuracy in autonomous vehicle decision-making.
[0004] In the first aspect, an embodiment of this application provides a vehicle decision-making optimization method based on reinforcement learning. The vehicle decision-making optimization method includes:
[0005] Obtain lane-changing data of multiple lane-changing vehicles; the lane-changing data is used to describe the original lane and the lane after lane-changing of the lane-changing vehicle, and the positions and speeds of the lane-changing vehicle itself and surrounding vehicles before and after lane-changing;
[0006] Determine a target lane-changing vehicle that performs dangerous lane-changing behavior from all lane-changing vehicles based on all lane-changing data;
[0007] Using a reinforcement learning model, obtain the predicted lane-changing action of each target lane-changing vehicle based on the lane-changing data of all target lane-changing vehicles, and calculate the reward value corresponding to the predicted lane-changing action; the predicted lane-changing action is used to describe whether the lane-changing vehicle changes lanes;
[0008] Optimize the reinforcement learning model based on all reward values to obtain an optimized reinforcement learning model;
[0009] Use the optimized reinforcement learning model to obtain the final lane-changing action of the target vehicle, and control the target vehicle to change lanes according to the final lane-changing action.
[0010] Optionally, determining target lane-changing vehicles with dangerous lane-changing behaviors from all lane-changing vehicles based on all lane-changing data includes:
[0011] For each lane-changing vehicle respectively, perform the following steps:
[0012] Calculate the minimum time to collision (TTC) between the lane-changing vehicle and each first surrounding vehicle before the lane change and the TTC between the lane-changing vehicle and each second surrounding vehicle after the lane change according to the lane-changing data of the lane-changing vehicle;
[0013] Calculate the original lane danger assessment index based on all the minimum TTCs before the lane change and calculate the lane danger assessment index after the lane change based on all the minimum TTCs after the lane change;
[0014] Determine whether the original lane danger assessment index and the lane danger assessment index after the lane change meet the lane-changing danger conditions;
[0015] If so, determine the lane-changing vehicle as the target lane-changing vehicle performing dangerous lane-changing behaviors.
[0016] Optionally, calculating the minimum time to collision (TTC) between the lane-changing vehicle and each surrounding vehicle before the lane change and the TTC between the lane-changing vehicle and each first surrounding vehicle after the lane change according to the lane-changing data of the lane-changing vehicle includes:
[0017] Through the formula:
[0018]
[0019] Calculate the minimum time to collision TTC between the lane-changing vehicle and the first surrounding vehicle before the lane change;
[0020] Wherein, L represents the distance between the lane-changing vehicle and the first surrounding vehicle before the lane change. When the lane-changing vehicle is behind the first surrounding vehicle before the lane change, v f represents the speed of the lane-changing vehicle before the lane change, and v p represents the speed of the first surrounding vehicle before the lane change. When the lane-changing vehicle is in front of the first surrounding vehicle before the lane change, v f represents the speed of the first surrounding vehicle before the lane change, and v p represents the speed of the lane-changing vehicle before the lane change;
[0021] Through the formula:
[0022]
[0023] Calculate the minimum time to collision TTC' between the lane-changing vehicle and the second surrounding vehicle after the lane change;
[0024] Wherein, L' represents the distance between the lane-changing vehicle and the second surrounding vehicle after the lane change. When the lane-changing vehicle is behind the second surrounding vehicle after the lane change, vf Represents the speed of the lane-changing vehicle after lane change, v p Represents the speed of the surrounding vehicles in the second week after lane change. When the lane-changing vehicle is in front of the surrounding vehicles in the second week after lane change, v f Represents the speed of the surrounding vehicles in the second week after lane change, v p Represents the speed of the lane-changing vehicle after lane change.
[0025] Optionally, calculate the original lane risk assessment index based on the minimum time to collision before all lane changes, including:
[0026] Take the minimum time to collision before lane change with the smallest value as the first minimum time to collision;
[0027] Judge whether the first minimum time to collision is less than or equal to the time threshold;
[0028] If so, assign the original lane risk assessment index as 1;
[0029] Otherwise, assign the original lane risk assessment index as 0.
[0030] Optionally, calculate the lane risk assessment index after lane change based on the minimum time to collision after all lane changes, including:
[0031] Take the minimum time to collision after lane change with the smallest value as the second minimum time to collision;
[0032] Judge whether the second minimum time to collision is less than or equal to the time threshold;
[0033] If so, assign the lane risk assessment index after lane change as 1;
[0034] Otherwise, assign the lane risk assessment index after lane change as 0.
[0035] Optionally, the lane change risk condition is:
[0036] The value of the original lane risk assessment index is 1 or the value of the lane risk assessment index after lane change is 1.
[0037] Optionally, calculate the reward value corresponding to the predicted lane change action, including:
[0038] Through the formula:
[0039] R t = w1·R position + w2·R collision + w3·R comfort + w4·R over-speed
[0040] Calculate the reward value R t ;
[0041] Among them, w1, w2, w3, and w4 all represent weight coefficients, and R position represents the position evaluation value, and R collision represents the collision risk value, and R comfort represents the driving comfort evaluation value, and R over-speed represents the speeding evaluation value:
[0042]
[0043] R over-speed = -k4·(v - v max ) 2 ·I(v > v max )
[0044] Among them, x represents the position of the target lane-changing vehicle before performing the predicted lane-changing action, x2 represents the position of the adjacent vehicle in front of the target lane-changing vehicle before the target lane-changing vehicle performs the predicted lane-changing action, x2 represents the position of the adjacent vehicle behind the target lane-changing vehicle before the target lane-changing vehicle performs the predicted lane-changing action, k1 and k2 represent the coefficients of the collision risk value, Δv represents the speed difference between the target lane-changing vehicle and the adjacent vehicle in front of the target lane-changing vehicle before the target lane-changing vehicle performs the predicted lane-changing action, and d min represents the minimum distance between the target lane-changing vehicle and the adjacent vehicle in front of the target lane-changing vehicle before the target lane-changing vehicle performs the predicted lane-changing action, k3 represents the coefficient of the driving comfort evaluation value, and a current represents the acceleration of the target lane-changing vehicle when performing the predicted lane-changing action, aprevious represents the historical acceleration of the target lane-changing vehicle, and a max represents the maximum acceleration, k4 represents the coefficient of the speeding evaluation value, v represents the speed of the target lane-changing vehicle, and v max represents the maximum speed limit, and I(v > v max ) represents the indicator function. When v > v max , I(v > v max ) = 1. When v > v max is not satisfied, I(v > v max ) = 0.
[0045] In a second aspect, an embodiment of the present application provides a vehicle decision-making optimization device based on reinforcement learning, including:
[0046] An acquisition module, configured to acquire lane-changing data of multiple lane-changing vehicles; the lane-changing data is used to describe the original lane and the lane after lane-changing of the lane-changing vehicle, and the positions and speeds of the lane-changing vehicle itself and surrounding vehicles before and after lane-changing;
[0047] A determination module, configured to determine a target lane-changing vehicle that performs a dangerous lane-changing behavior from all lane-changing vehicles based on all lane-changing data;
[0048] A calculation module, configured to use a reinforcement learning model to obtain a predicted lane-changing action of each target lane-changing vehicle based on the lane-changing data of all target lane-changing vehicles, and calculate a reward value corresponding to the predicted lane-changing action; the predicted lane-changing action is used to describe whether the lane-changing vehicle changes lanes.
[0049] An optimization module, configured to optimize the reinforcement learning model based on all the reward values to obtain an optimized reinforcement learning model.
[0050] A control module, configured to use the optimized reinforcement learning model to obtain the final lane-changing action of the target vehicle, and control the target vehicle to change lanes according to the final lane-changing action.
[0051] In a third aspect, an embodiment of the present application provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the above-mentioned vehicle decision optimization method based on reinforcement learning is implemented.
[0052] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned vehicle decision optimization method based on reinforcement learning is implemented.
[0053] The above solution of the present application has the following beneficial effects:
[0054] In the embodiment of the present application, by obtaining the lane-changing data of multiple lane-changing vehicles, then determining target lane-changing vehicles that perform dangerous lane-changing behaviors from all lane-changing vehicles based on all the lane-changing data, and then using a reinforcement learning model to obtain a predicted lane-changing action of each target lane-changing vehicle based on the lane-changing data of all target lane-changing vehicles, and calculating a reward value corresponding to the predicted lane-changing action, and then optimizing the reinforcement learning model based on all the reward values to obtain an optimized reinforcement learning model, and finally using the optimized reinforcement learning model to obtain the final lane-changing action of the target vehicle, and controlling the target vehicle to change lanes according to the final lane-changing action. Among them, determining the target vehicles that perform dangerous lane-changing behaviors can analyze the behavior patterns of the vehicles and improve the reliability of vehicle information management. Calculating the reward value corresponding to the predicted lane-changing action can describe the advantages and disadvantages of the predicted lane-changing action. Optimizing based on the reward value improves the performance of the optimized reinforcement learning model. The rationality of the final lane-changing action obtained by using the reinforcement learning model with high performance is improved. Changing lanes according to the reasonable final lane-changing action effectively improves the accuracy of the decision-making of autonomous driving vehicles.
[0055] Other beneficial effects of the present application will be described in detail in the subsequent specific implementation part. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0057] Figure 1 It is a flowchart of a vehicle decision-making optimization method based on reinforcement learning provided by an embodiment of the present application;
[0058] Figure 2 It is a schematic structural diagram of a vehicle decision-making optimization device provided by an embodiment of the present application;
[0059] Figure 3 It is a schematic structural diagram of a terminal device provided by an embodiment of the present application. Detailed implementation manners
[0060] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system structures and technologies are presented to thoroughly understand the embodiments of the present application. However, those skilled in the art should understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0061] It should be understood that when used in the specification and claims of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0062] It should also be understood that the term "and / or" used in the specification and claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0063] As used in the specification and claims of the present application, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" depending on the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" depending on the context.
[0064] In addition, in the description of the specification and the appended claims of this application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be construed as indicating or implying relative importance.
[0065] The reference to "one embodiment" or "some embodiments" etc. described in the specification of this application means that in one or more embodiments of this application, specific features, structures or characteristics described in connection with that embodiment are included. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments" etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprise", "include", "have" and their variants all mean "include but not limited to", unless otherwise specifically emphasized in other ways.
[0066] Aiming at the problem of low accuracy of existing autonomous vehicle decision-making, the embodiments of this application provide a vehicle decision optimization method based on reinforcement learning. This vehicle decision optimization method obtains the lane-changing data of multiple lane-changing vehicles, then determines the target lane-changing vehicles that perform dangerous lane-changing behaviors from all lane-changing vehicles based on all the lane-changing data, and then uses a reinforcement learning model to obtain the predicted lane-changing actions of each target lane-changing vehicle based on the lane-changing data of all target lane-changing vehicles, and calculates the reward values corresponding to the predicted lane-changing actions. Then, based on all the reward values, the reinforcement learning model is optimized to obtain an optimized reinforcement learning model. Finally, the optimized reinforcement learning model is used to obtain the final lane-changing action of the target vehicle, and the target vehicle is controlled to change lanes according to the final lane-changing action. Among them, determining the target vehicles that perform dangerous lane-changing behaviors can analyze the behavior patterns of the vehicles and improve the reliability of vehicle information management. Calculating the reward values corresponding to the predicted lane-changing actions can express the advantages and disadvantages of the predicted lane-changing actions. Optimizing based on the reward values improves the performance of the optimized reinforcement learning model. The rationality of the final lane-changing action obtained by using the reinforcement learning model with high performance is improved, and changing lanes according to the reasonable final lane-changing action effectively improves the accuracy of autonomous vehicle decision-making.
[0067] Next, an exemplary description of the vehicle decision optimization method based on reinforcement learning provided by this application is given.
[0068] As Figure 1 shown, the vehicle decision optimization method based on reinforcement learning provided by this application includes the following steps:
[0069] Step 11, obtain the lane-changing data of multiple lane-changing vehicles.
[0070] The above lane-changing vehicle is the vehicle observed to change lanes, and the above lane-changing data is used to describe the original lane and the lane after lane-changing of the lane-changing vehicle, as well as the positions and speeds of the lane-changing vehicle itself and the surrounding vehicles before and after lane-changing.
[0071] In some embodiments of the present application, the lane-changing data of the lane-changing vehicle can be obtained by accessing a public dataset (such as a high-definition driving dataset).
[0072] Step 12, determine a target lane-changing vehicle that performs a dangerous lane-changing behavior from all lane-changing vehicles based on all lane-changing data.
[0073] In some embodiments of the present application, the step of determining a target lane-changing vehicle that performs a dangerous lane-changing behavior from all lane-changing vehicles based on all lane-changing data includes:
[0074] For each lane-changing vehicle, perform the following steps:
[0075] The first step is to calculate the minimum time to collision (TTC) between the lane-changing vehicle and each first surrounding vehicle before lane-changing and the minimum time to collision (TTC') between the lane-changing vehicle and each second surrounding vehicle after lane-changing according to the lane-changing data of the lane-changing vehicle.
[0076] The first surrounding vehicle is the vehicle adjacent to the lane-changing vehicle before the lane-changing vehicle changes lanes. The second surrounding vehicle is the vehicle adjacent to the lane-changing vehicle after the lane-changing vehicle completes the lane-changing.
[0077] Specifically, through the formula:
[0078]
[0079] Calculate the minimum time to collision (TTC) between the lane-changing vehicle and the first surrounding vehicle before lane-changing.
[0080] Where L represents the distance between the lane-changing vehicle and the first surrounding vehicle before lane-changing. When the lane-changing vehicle is behind the first surrounding vehicle before lane-changing, vf represents the speed of the lane-changing vehicle before lane-changing, vp represents the speed of the first surrounding vehicle before lane-changing. When the lane-changing vehicle is in front of the first surrounding vehicle before lane-changing, vf represents the speed of the first surrounding vehicle before lane-changing, vp represents the speed of the lane-changing vehicle before lane-changing.
[0081] Through the formula:
[0082]
[0083] Calculate the minimum time to collision (TTC') between the lane-changing vehicle and the second surrounding vehicle after lane-changing.
[0084] Wherein, L' represents the distance between the lane-changing vehicle and the second surrounding vehicle after lane changing. When the lane-changing vehicle is behind the second surrounding vehicle after lane changing, vf represents the speed of the lane-changing vehicle after lane changing, and vp represents the speed of the second surrounding vehicle after lane changing. When the lane-changing vehicle is in front of the second surrounding vehicle after lane changing, vf represents the speed of the second surrounding vehicle after lane changing, and vp represents the speed of the lane-changing vehicle after lane changing.
[0085] Step 2: Calculate the original lane risk assessment index based on all the minimum collision times before lane changing, and calculate the lane risk assessment index after lane changing based on all the minimum collision times after lane changing.
[0086] First, take the minimum collision time before lane changing with the smallest value as the first minimum collision time; determine whether the first minimum collision time is less than or equal to the time threshold; if so, assign the original lane risk assessment index as 1; otherwise, assign the original lane risk assessment index as 0.
[0087] Then, take the minimum collision time after lane changing with the smallest value as the second minimum collision time; determine whether the second minimum collision time is less than or equal to the time threshold; if so, assign the lane risk assessment index after lane changing as 1; otherwise, assign the lane risk assessment index after lane changing as 0.
[0088] Step 3: Determine whether the original lane risk assessment index and the lane risk assessment index after lane changing meet the lane-changing risk condition.
[0089] If so, determine that the lane-changing vehicle is the target lane-changing vehicle performing a dangerous lane-changing behavior.
[0090] Otherwise, determine that the lane-changing vehicle does not have a dangerous lane-changing behavior and do not mark the lane-changing vehicle.
[0091] It should be noted that the lane-changing risk condition is: the value of the original lane risk assessment index is 1 or the value of the lane risk assessment index after lane changing is 1.
[0092] Step 13: Use the reinforcement learning model to obtain the predicted lane-changing actions of each target lane-changing vehicle based on the lane-changing data of all target lane-changing vehicles, and calculate the reward values corresponding to the predicted lane-changing actions.
[0093] The above predicted lane-changing actions are used to describe whether the lane-changing vehicle changes lanes. For example, the value of the predicted lane-changing action is 0 or 1. When the value is 1, it means that the lane-changing vehicle changes lanes at the next moment. When the value is 0, it means that the lane-changing vehicle does not change lanes at the next moment.
[0094] In some embodiments of the present application, the step of using the reinforcement learning model to obtain the predicted lane-changing actions of each target lane-changing vehicle based on the lane-changing data of all target lane-changing vehicles and calculating the reward values corresponding to the predicted lane-changing actions is specifically as follows:
[0095] First step, organize the lane-changing data of each target lane-changing vehicle.
[0096] Extract the relevant data (such as vehicle position, speed, driving angle, etc.) of the vehicle that is in the same lane as the target lane-changing vehicle and in front of the target lane-changing vehicle before lane-changing, as well as the relevant data of the vehicle that is in the same lane as the target lane-changing vehicle and behind the target lane-changing vehicle from the lane-changing data.
[0097] Integrate the relevant data of the target lane-changing vehicle, the vehicle in front before lane-changing, and the vehicle behind after lane-changing into one data.
[0098] Second step, input all the data obtained in the previous step into the reinforcement learning model, and use the reinforcement learning model to obtain the predicted lane-changing actions of each target lane-changing vehicle.
[0099] It should be noted that the reinforcement learning algorithm in the reinforcement learning model is, for example, the Proximal Policy Optimization (PPO) algorithm. Use a simulation software (such as the Simulation of Urban MObility (SUMO) software) to simulate all the data obtained in the previous step; and for each data, take the positions, speeds, and driving angles of the target lane-changing vehicle, the vehicle in front before lane-changing, and the vehicle behind after lane-changing in the data as the state space in the reinforcement learning algorithm. The action space is (0, 1), where 1 means the lane-changing vehicle changes lanes at the next moment, and 0 means the lane-changing vehicle does not change lanes at the next moment. Select a value from the action space as the value of the predicted lane-changing action according to the policy in the reinforcement learning algorithm (such as selecting 0 or 1 according to the distance between the target lane-changing vehicle and the vehicle in front).
[0100] Exemplarily, control the target lane-changing vehicle to execute the predicted lane-changing action in the simulation. If a collision occurs to the target lane-changing vehicle in the simulation, or no vehicle collision occurs after executing the predicted lane-changing action, or the target lane-changing vehicle drives out of the simulation range and still has not completed the lane change, the simulation of the target lane-changing vehicle ends. The distance between vehicles can be calculated according to the Euclidean formula, and it can be judged whether a collision occurs between vehicles according to the distance. When the distance is less than a preset value (such as 4), it is considered that a collision has occurred.
[0101] Third step, calculate the reward value corresponding to the predicted lane-changing action.
[0102] Specifically, through the formula:
[0103] R t = w1·R position + w2·R collision + w3·R comfort + w4·Rover-speed
[0104] Calculate the reward value R t ;
[0105] where w1, w2, w3, and w4 all represent weight coefficients, R position represents the position evaluation value, R collision represents the collision risk value, R comfort represents the driving comfort evaluation value, R over-spped represents the speeding evaluation value:
[0106]
[0107] R over-spped = -k4·(v - v max ) 2 ·I(v > v max )
[0108] where x represents the position of the target lane-changing vehicle before performing the predicted lane-changing action, x2 represents the position of the adjacent vehicle in front of the target lane-changing vehicle before performing the predicted lane-changing action, x2 represents the position of the adjacent vehicle behind the target lane-changing vehicle before performing the predicted lane-changing action, k1 and k2 represent the coefficients of the collision risk value, Δv represents the speed difference between the target lane-changing vehicle and the adjacent vehicle in front of it before performing the predicted lane-changing action, d min represents the minimum distance between the target lane-changing vehicle and the adjacent vehicle in front of it before performing the predicted lane-changing action, k3 represents the coefficient of the driving comfort evaluation value, a current represents the acceleration of the target lane-changing vehicle when performing the predicted lane-changing action, a previous represents the historical acceleration of the target lane-changing vehicle, a max represents the maximum acceleration, k4 represents the coefficient of the speeding evaluation value, v represents the speed of the target lane-changing vehicle, v max represents the maximum speed limit, I(v > v max ) represents the indicator function. When v > v max , I(v > v max ) = 1; when v > v max is not satisfied, I(v > v max ) = 0.
[0109] It is worth mentioning that the formula for calculating the reward value above includes values in four aspects: position evaluation value, collision risk value, driving comfort evaluation value, and speeding evaluation value, enabling the reward value to describe the position, collision risk, driving comfort, and speeding of vehicles, and improving the rationality and accuracy of the reward value.
[0110] Step 14: Optimize the reinforcement learning model based on all reward values to obtain an optimized reinforcement learning model.
[0111] Exemplarily, the model training strategy in the reinforcement learning algorithm (such as the training process in the PPO algorithm, and the above reward values are the rewards in this algorithm) can be used to optimize the reinforcement learning model to obtain an optimized reinforcement learning model. For example, in the PPO algorithm, set its policy network architecture to net_arch = [64, 64], the learning rate to 5e-4, the batch size to 32, the discount factor to 0.99, and use to calculate the state value function, r t as the immediate reward at time t to optimize the policy of the model.
[0112] Step 15: Use the optimized reinforcement learning model to obtain the final lane-changing action of the target vehicle, and control the target vehicle to change lanes according to the final lane-changing action.
[0113] The above target vehicle is an autonomous vehicle that needs to change lanes.
[0114] Specifically, obtain the relevant information (speed, position, driving direction, etc.) of the target vehicle, the vehicle in front of the target vehicle, and the vehicle behind the target vehicle at the current moment, and use the relevant information of these three vehicles as the state space in the optimized reinforcement learning model. The action space is (0, 1). Use the optimized reinforcement learning model to select the final lane-changing action from the action space, and through the control system of the autonomous vehicle, control the target vehicle to change lanes at the next moment of the current moment according to the final lane-changing action.
[0115] It is worth mentioning that identifying the target vehicle that performs dangerous lane-changing behavior can analyze the behavior pattern of the vehicle and improve the reliability of vehicle information management. Calculating the reward value corresponding to the predicted lane-changing action can describe the pros and cons of the predicted lane-changing action. Optimizing based on the reward value improves the performance of the optimized reinforcement learning model. The rationality of the final lane-changing action obtained by using the reinforcement learning model with high performance is improved, and lane-changing is performed according to the reasonable final lane-changing action, effectively improving the accuracy of the decision-making of autonomous vehicles.
[0116] In addition, the specific advantages of the method of this application are as follows:
[0117] 1. Comprehensiveness and accuracy of the reinforcement learning environment
[0118] Traditional autonomous vehicle control methods, such as simply using simple rules or a single control algorithm, often ignore the interactions and influences among multiple vehicles in complex traffic scenarios. However, the reinforcement learning environment based on SUMO simulation constructed in this application incorporates multiple factors such as vehicle position, speed, lane change, and collision detection into a unified analysis and control framework, fully considering the mutual influences among different vehicle behaviors. This environment not only retains the basic state information of the vehicle but also introduces complex reward mechanisms and end conditions, making the environment more similar to real traffic scenarios, capable of more accurately simulating the operating rules of the autonomous driving system, and providing a more comprehensive and accurate data basis for vehicle decision-making.
[0119] 2. Precision of Collision Detection
[0120] In collision detection, this application considers the precise position relationships among multiple vehicles and uses the Euclidean distance formula to calculate the distance between vehicles to accurately determine whether a collision occurs. This precise collision detection takes into account the precise position information of the vehicles and the dynamic changes of different vehicles, making the collision detection more in line with the actual situation. At the same time, incorporating collision situations into the consideration of reward values and end conditions can provide more accurate data and information sources for the safety assessment of the autonomous driving system, not only making the autonomous driving system safer and more reliable but also helping to provide more reasonable decision-making bases for vehicle path planning and speed control.
[0121] 3. Multi-dimensional Consideration in Decision-making
[0122] In traditional autonomous driving control methods, vehicle decision-making is often limited to simple speed control or a single lane change rule. However, this application realizes multi-dimensional consideration in decision-making by constructing a complete reinforcement learning environment. It not only considers vehicle decision-making at different positions and speeds but also takes into account multiple factors such as lane change, time consumption, and interaction with other vehicles. This multi-dimensional consideration in decision-making not only takes into account the importance and influence of the vehicle's state at the current moment but also considers the vehicle's long-term behavior in different scenarios and its contribution to the entire traffic system. Therefore, the decisions made are more comprehensive and accurate, and can provide more valuable references for improving the overall performance of the autonomous driving system.
[0123] 4. Efficiency and Scalability of the Algorithm
[0124] With the expansion of traffic scenarios and the increase in complexity, traditional autonomous driving control algorithms may face problems such as low decision-making efficiency and poor scalability. The reinforcement learning algorithm adopted in this application, such as the PPO algorithm, as well as the carefully designed network architecture and parameter settings, improve the computational efficiency and scalability of vehicle decision-making. These algorithms and parameter settings can not only make vehicle decisions quickly and accurately in complex traffic scenarios, but also be flexibly extended and optimized according to different scenario requirements, such as by adjusting parameters such as roads, rewards, and end conditions to adapt to different traffic environments and experimental requirements.
[0125] The method of this application will be exemplarily described below in combination with a specific example.
[0126] The method of this application is experimentally verified using a computer.
[0127] I. Conditions required for the experiment:
[0128] Dataset: The high-definition driving dataset (HighD) is used, which contains vehicle trajectory data and metadata files, stored in the path.. / trajectory_data_analysis / HighD / processed / , and the specific files are all_tracks.csv (vehicle trajectory data) and all_tracksMeta.csv (vehicle metadata).
[0129] Software environment: SUMO traffic simulation software, version 1.8.0, and its SUMO_HOME environment variable needs to be set to. / SUMO / sumo-1.8.0 to ensure the correct invocation of relevant tools.
[0130] The Python programming language and its related libraries, including but not limited to tensorflow, stable_baselines3, pandas, traci, gym, numpy, etc. Different libraries are used for different functions. For example, tensorflow can assist in deep learning calculations, pandas is used for data processing, traci is used to interact with SUMO, and gym provides a basis for building a reinforcement learning environment, etc.
[0131] Hardware environment: There are no special hardware requirements, but for better performance, it is recommended to use a computer with a certain computing power, such as at least 4GB of memory and a multi-core processor.
[0132] II. Experiment parameter settings:
[0133] A series of parameters are set through the argparse tool, as follows:
[0134] SUMO Configuration:
[0135] show_gui: Defaults to False and is used to control whether to display the SUMO graphical interface.
[0136] sumocfgfile: Defaults to highD.sumocfg and is the path to the SUMO configuration file.
[0137] svID: Defaults to "SV" and is the ID of the ego vehicle.
[0138] clvID: Defaults to "CLV" and is the ID of the leading vehicle.
[0139] tfvID: Defaults to "TFV" and is the ID of the following vehicle (i.e., the target lane-changing vehicle).
[0140] start_time: Defaults to 10 and is the number of simulation steps before learning starts.
[0141] num_action: Defaults to 2 and is the number of action spaces.
[0142] lane_change_time: Defaults to 5 and is the time for lane change.
[0143] Road Configuration:
[0144] min_x_position: Defaults to 0.0 and is the minimum lateral position of the vehicle.
[0145] max_x_position: Defaults to 440.0 and is the maximum lateral position of the vehicle.
[0146] min_y_position: Defaults to 23.0 and is the minimum longitudinal position of the vehicle.
[0147] max_y_position: Defaults to 31.0 and is the maximum longitudinal position of the vehicle.
[0148] min_speed: Defaults to 0.0 and is the minimum longitudinal speed of the vehicle.
[0149] max_speed: Defaults to 100.0 and is the maximum longitudinal speed of the vehicle.
[0150] count: Defaults to 10 and is the length of one training episode.
[0151] collision: Defaults to False and is the ego vehicle collision flag.
[0152] sleep: Defaults to True and is the sleep flag for each simulation.
[0153] Reward Configuration:
[0154] w_speed: The default value is 0.1, which is the weight of the desired speed reward.
[0155] R_time: The default value is -1, which is the reward for time consumption.
[0156] R_collision: The default value is -400, which is the negative reward for ego vehicle collision.
[0157] End Condition Configuration:
[0158] target_lane_id: The default value is 1, which is the ID of the target lane.
[0159] max_count: The default value is 500, which is the maximum length of a training episode.
[0160] III. Experiment Process:
[0161] Data Reading and Processing:
[0162] Use pandas to read the vehicle trajectory data and metadata files in the HighD dataset, and convert the data types of some data columns. For example, convert the edge column to string type, and convert the departLane and arrivalLane columns to integer type.
[0163] Through the VehicleDataSelector class, randomly select a group_id from the data, filter out the corresponding vehicle data, and at the same time determine the initial and final frame ranges of the SV vehicle, as well as the maximum and minimum number of frames of the entire group.
[0164] Vehicle Management and Simulation Environment Setup:
[0165] Use the Vehicles class to manage vehicles, including adding and removing vehicles, updating the position, speed, and lane information of vehicles, performing lane changes for vehicles according to count or actions, and at the same time handling the start time, duration, and control mode switching of lane changes.
[0166] Build a reinforcement learning environment through the SumoGym class, initialize various parameters, including parameters of the road, reward, end, and observation space. In the step method, update the action space, perform lane changes according to actions, update vehicle information, check for collisions, calculate rewards based on the current state, and determine whether to end the simulation. In the reset method, the simulation environment can be reset to start a new simulation.
[0167] Experiment Phase:
[0168] Environment Testing:
[0169] Create a SumoGym environment, generate random actions using env.action_space.sample(), and continuously update vehicle information and calculate rewards in multiple episodes (e.g., 10 episodes) to observe the vehicle behavior and reward changes under different random actions.
[0170] Model training:
[0171] Use the PPO algorithm in stable_baselines3, use MlpPolicy as the policy, set the network architecture to [64, 64], learning rate to 5e-4, batch size to 32, gamma to 0.99, enable detailed output, store the training logs in the "logs / " path, use model.learn(int(2e4)) to train the model, and save the trained model to the.. / models / ppo2 path.
[0172] Model testing:
[0173] Load the trained model, such as loading from a specified path (e.g.,. / models / ppo1), and in multiple episodes (e.g., 10 episodes), update vehicle information using the actions predicted by the model, calculate rewards, and evaluate the model performance.
[0174] IV. Experimental results:
[0175] Environmental test results:
[0176] In the environmental test using random actions, the vehicle behavior under different actions can be observed, including lane changes, speed adjustments, and interactions with other vehicles. At the same time, the fluctuations of rewards can be seen. According to different combinations of actions and vehicle states, the reward values will change accordingly, which helps to evaluate whether the environmental settings are reasonable.
[0177] Model training results:
[0178] After a certain number of training iterations (e.g., 2e4 times), the loss of the model gradually decreases, indicating that the model continuously optimizes the policy during the learning process to achieve higher rewards.
[0179] Model testing results:
[0180] In the model testing stage, after loading the trained model, using the actions predicted by the model can make the vehicle show better behavior in the simulation environment, such as more reasonable lane changes and speed control, and obtain higher rewards, indicating that the model has learned a certain decision-making strategy and can make better decisions according to the environmental state, improving the performance of the autonomous driving system in this simulation environment.
[0181] Through simulation experiments using the HighD dataset, the method of this application has demonstrated feasibility in the environmental testing, model training, and model testing phases, and can simulate and optimize different traffic scenarios and conditions by adjusting different parameters and experimental settings, providing an effective tool and platform for the development and evaluation of autonomous driving systems.
[0182] Next, an exemplary description of the vehicle decision-making optimization device based on reinforcement learning provided by this application will be given.
[0183] As Figure 2 shown, an embodiment of this application provides a vehicle decision-making optimization device based on reinforcement learning. The vehicle decision-making optimization device 200 based on reinforcement learning includes:
[0184] An acquisition module 201, configured to acquire lane-changing data of multiple lane-changing vehicles; the lane-changing data is used to describe the original lane and the lane after lane-changing of the lane-changing vehicle, and the positions and speeds of the lane-changing vehicle itself and surrounding vehicles before and after lane-changing.
[0185] A determination module 202, configured to determine a target lane-changing vehicle performing a dangerous lane-changing behavior from all lane-changing vehicles based on all lane-changing data.
[0186] A calculation module 203, configured to use a reinforcement learning model to obtain a predicted lane-changing action for each target lane-changing vehicle based on the lane-changing data of all target lane-changing vehicles, and calculate a reward value corresponding to the predicted lane-changing action; the predicted lane-changing action is used to describe whether the lane-changing vehicle changes lanes.
[0187] An optimization module 204, configured to optimize the reinforcement learning model based on all reward values to obtain an optimized reinforcement learning model.
[0188] A control module 205, configured to use the optimized reinforcement learning model to obtain the final lane-changing action of the target vehicle, and control the target vehicle to change lanes according to the final lane-changing action.
[0189] It should be noted that for the information interaction, execution process, etc. between the above-mentioned devices / units, since they are based on the same concept as the method embodiment of this application, their specific functions and the technical effects brought can be specifically referred to in the method embodiment part, and will not be elaborated here.
[0190] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.
[0191] As Figure 3 shown, an embodiment of the present application provides a terminal device. The terminal device D10 in this embodiment includes: at least one processor D100 ( Figure 3 only one processor is shown in the figure), a memory D101, and a computer program D102 stored in the memory D101 and executable on the at least one processor D100. When the processor D100 executes the computer program D102, the steps in any of the foregoing method embodiments are implemented.
[0192] Specifically, when the processor D100 executes the computer program D102, by obtaining the lane-changing data of multiple lane-changing vehicles, then determining the target lane-changing vehicles that perform dangerous lane-changing behaviors from all lane-changing vehicles based on all the lane-changing data, and then using a reinforcement learning model, obtaining the predicted lane-changing actions of each target lane-changing vehicle based on the lane-changing data of all target lane-changing vehicles, calculating the reward values corresponding to the predicted lane-changing actions, then optimizing the reinforcement learning model based on all the reward values to obtain an optimized reinforcement learning model, and finally using the optimized reinforcement learning model to obtain the final lane-changing action of the target vehicle, and controlling the target vehicle to change lanes according to the final lane-changing action. Among them, determining the target vehicles that perform dangerous lane-changing behaviors can analyze the behavior patterns of the vehicles and improve the reliability of vehicle information management. Calculating the reward values corresponding to the predicted lane-changing actions can express the advantages and disadvantages of the predicted lane-changing actions. Optimizing based on the reward values improves the performance of the optimized reinforcement learning model. The rationality of the final lane-changing action obtained by using the reinforcement learning model with high performance is improved. Changing lanes according to the reasonable final lane-changing action effectively improves the accuracy of the decision-making of autonomous driving vehicles.
[0193] The so-called processor D100 may be a central processing unit (CPU, Central Processing Unit), and this processor D100 may also be other general-purpose processors, digital signal processors (DSP, Digital Signal Processor), application-specific integrated circuits (ASIC, Application Specific Integrated Circuit), field-programmable gate arrays (FPGA, Field-Programmable Gate Array), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.
[0194] In some embodiments, the memory D101 may be an internal storage unit of the terminal device D10, such as the hard disk or memory of the terminal device D10. In some other embodiments, the memory D101 may also be an external storage device of the terminal device D10, such as a plug-in hard disk equipped on the terminal device D10, a smart media card (SMC, SmartMedia Card), a secure digital (SD, Secure Digital) card, a flash card (Flash Card), etc. Further, the memory D101 may also include both the internal storage unit of the terminal device D10 and the external storage device. The memory D101 is used to store an operating system, application programs, a boot loader (BootLoader), data, and other programs, such as the program code of the computer program, etc. The memory D101 may also be used to temporarily store data that has been output or will be output.
[0195] The embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned various method embodiments can be implemented.
[0196] The embodiments of the present application provide a computer program product. When the computer program product runs on a terminal device, the terminal device can be made to execute the steps in the above-mentioned various method embodiments.
[0197] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of this application, a computer program can be used to instruct the relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the vehicle decision optimization method device / terminal device based on reinforcement learning, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, RandomAccess Memory), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.
[0198] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0199] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0200] The above is the preferred implementation manner of this application. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle described in this application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of this application.
Claims
1. A vehicle decision-making optimization method based on reinforcement learning, characterized in that, Including: Obtain the lane-changing data of multiple lane-changing vehicles; The lane-changing data is used to describe the original lane and the lane after lane-changing of the lane-changing vehicle, and the positions and speeds of the lane-changing vehicle itself and surrounding vehicles before and after lane-changing; Based on all the lane-changing data, determine the target lane-changing vehicles that perform dangerous lane-changing behaviors from all the lane-changing vehicles; Using a reinforcement learning model, obtain the predicted lane-changing actions of each target lane-changing vehicle based on the lane-changing data of all the target lane-changing vehicles, and calculate the reward values corresponding to the predicted lane-changing actions; the predicted lane-changing actions are used to describe whether the lane-changing vehicle changes lanes; Optimize the reinforcement learning model based on all the reward values to obtain an optimized reinforcement learning model; Use the optimized reinforcement learning model to obtain the final lane-changing action of the target vehicle, and control the target vehicle to change lanes according to the final lane-changing action.
2. The vehicle decision optimization method according to claim 1, wherein The determining the target lane-changing vehicles that perform dangerous lane-changing behaviors from all the lane-changing vehicles based on all the lane-changing data includes: For each lane-changing vehicle, perform the following steps: According to the lane-changing data of the lane-changing vehicle, calculate the minimum time to collision (TTC) between the lane-changing vehicle and each first surrounding vehicle before lane-changing and the minimum TTC between the lane-changing vehicle and each second surrounding vehicle after lane-changing; Calculate the original lane danger assessment index based on all the minimum TTCs before lane-changing, and calculate the lane danger assessment index after lane-changing based on all the minimum TTCs after lane-changing; Judge whether the original lane danger assessment index and the lane danger assessment index after lane-changing meet the lane-changing danger condition; If so, determine the lane-changing vehicle as the target lane-changing vehicle that performs dangerous lane-changing behaviors.
3. The vehicle decision optimization method according to claim 2, wherein, The calculating the minimum TTC between the lane-changing vehicle and each surrounding vehicle before lane-changing and the minimum TTC between the lane-changing vehicle and each first surrounding vehicle after lane-changing according to the lane-changing data of the lane-changing vehicle includes: Through the formula: Calculate the minimum TTC ttc between the lane-changing vehicle and the first surrounding vehicle before lane-changing; Among them, L represents the distance between the lane-changing vehicle and the first surrounding vehicle before lane change. When the lane-changing vehicle is behind the first surrounding vehicle before lane change, v f represents the speed of the lane-changing vehicle before lane change, v p represents the speed of the first surrounding vehicle before lane change. When the lane-changing vehicle is in front of the first surrounding vehicle before lane change, v f represents the speed of the first surrounding vehicle before lane change, v p represents the speed of the lane-changing vehicle before lane change; Through the formula: Calculate the minimum TTC ttc' between the lane-changing vehicle and the second surrounding vehicle after lane-changing; Among them, L' represents the distance between the lane-changing vehicle and the second surrounding vehicle after the lane change. When the lane-changing vehicle is behind the second surrounding vehicle after the lane change, v f represents the speed of the lane-changing vehicle after the lane change, v p represents the speed of the second surrounding vehicle after the lane change. When the lane-changing vehicle is in front of the second surrounding vehicle after the lane change, v f represents the speed of the second surrounding vehicle after the lane change, v p represents the speed of the lane-changing vehicle after the lane change.
4. The vehicle decision optimization method according to claim 3, characterized in that The calculating the original lane danger assessment index based on all the minimum TTCs before lane-changing includes: Take the minimum TTC before lane-changing with the smallest value as the first minimum TTC; Judge whether the first minimum TTC is less than or equal to the time threshold; If so, assign the value 1 to the original lane danger assessment index; Otherwise, assign the value 0 to the original lane danger assessment index.
5. The vehicle decision-making optimization method according to claim 3, wherein The calculating the lane danger assessment index after lane-changing based on all the minimum TTCs after lane-changing includes: Take the minimum TTC after lane-changing with the smallest value as the second minimum TTC; Judge whether the second minimum TTC is less than or equal to the time threshold; If so, assign the value 1 to the lane danger assessment index after lane-changing; Otherwise, assign the value 0 to the lane danger assessment index after lane-changing.
6. The vehicle decision optimization method according to claim 5, wherein The lane-changing danger condition is: The value of the original lane danger assessment index is 1 or the value of the lane danger assessment index after lane-changing is 1.
7. The vehicle decision optimization method according to claim 1, characterized in that The calculating the reward value corresponding to the predicted lane-changing action includes: Through the formula: R t = w1·R position + w2·R collision + w3·R comfort + w4·R over-speed Calculate the reward value R t ; Among them, w1, w2, w3, and w4 all represent weight coefficients, and R position represents the position evaluation value, and R collision represents the collision risk value, and R comfort represents the driving comfort evaluation value, and R over-speed represents the speeding evaluation value: P over-speed = -k4·(v - v max ) 2 ·I(v > v max ) Among them, x represents the position of the target lane-changing vehicle before performing the predicted lane-changing action, x2 represents the position of the adjacent vehicle in front of the target lane-changing vehicle before the target lane-changing vehicle performs the predicted lane-changing action, x2 represents the position of the adjacent vehicle behind the target lane-changing vehicle before the target lane-changing vehicle performs the predicted lane-changing action, k1 and k2 represent the coefficients of the collision risk value, Δv represents the speed difference between the target lane-changing vehicle and the adjacent vehicle in front of the target lane-changing vehicle before the target lane-changing vehicle performs the predicted lane-changing action, d min represents the minimum distance between the target lane-changing vehicle and the adjacent vehicle in front of the target lane-changing vehicle before the target lane-changing vehicle performs the predicted lane-changing action, k3 represents the coefficient of the driving comfort evaluation value, a current represents the acceleration of the target lane-changing vehicle when performing the predicted lane-changing action, a previous represents the historical acceleration of the target lane-changing vehicle, a max represents the maximum acceleration, k4 represents the coefficient of the speeding evaluation value, v represents the speed of the target lane-changing vehicle, v max represents the maximum speed limit, I(v>v max ) represents the indicator function. When v>v max is satisfied, I(v>v max ) = 1. When v>v max is not satisfied, I(v>v max ) = 0.
8. A vehicle decision-making optimization device based on reinforcement learning, characterized in that, Including: An acquisition module, configured to acquire the lane-changing data of multiple lane-changing vehicles; The lane-changing data is used to describe the original lane and the lane after lane-changing of the lane-changing vehicle, as well as the positions and speeds of the lane-changing vehicle itself and the surrounding vehicles before and after lane-changing; A determination module, configured to determine a target lane-changing vehicle performing a dangerous lane-changing behavior from all lane-changing vehicles based on all the lane-changing data; A calculation module, configured to use a reinforcement learning model to obtain a predicted lane-changing action of each target lane-changing vehicle based on the lane-changing data of all the target lane-changing vehicles, and calculate a reward value corresponding to the predicted lane-changing action; the predicted lane-changing action is used to describe whether the lane-changing vehicle changes lanes; An optimization module, configured to optimize the reinforcement learning model based on all the reward values to obtain an optimized reinforcement learning model; A control module, configured to use the optimized reinforcement learning model to obtain a final lane-changing action of the target vehicle, and control the target vehicle to change lanes according to the final lane-changing action.
9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the vehicle decision optimization method based on reinforcement learning according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the vehicle decision optimization method based on reinforcement learning according to any one of claims 1 to 7.