Automatic driving intelligent decision generation method
By acquiring traffic scene information and historical data and using deep reinforcement learning models to generate the driving intentions and decisions of autonomous vehicles, the problem of existing technologies that are difficult to balance safety, comfort and social compatibility is solved, and flexible and efficient autonomous driving decisions are achieved.
Patent Information
- Application Number
- CN202510880952.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-09-12
AI Technical Summary
Existing autonomous driving decision-making methods find it difficult to simultaneously balance safety, comfort, and social compatibility in complex traffic scenarios, resulting in inflexible vehicle behavior and even possible traffic accidents. Deep reinforcement learning methods are also prone to falling into local optimality.
By acquiring image information and historical driving data in traffic scenes, determining the global state vector and road right area, and using deep reinforcement learning models combined with multi-objective adaptation functions, flexible driving intentions and decisions are generated, comprehensively considering safety, comfort and social compatibility.
It achieves safe, efficient and social interaction norm-compliant driving decisions for autonomous vehicles in complex traffic scenarios, avoids local optimal problems, and improves decision-making flexibility and overall performance.
Smart Images

Figure CN120636154A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of autonomous driving, and in particular to a method for generating intelligent decisions for autonomous driving. Background Art
[0002] With the rapid development of autonomous driving technology, improving vehicle decision-making capabilities has become a key research direction. Its core goal is to achieve safe and efficient autonomous decision-making in complex traffic environments. However, the dynamic changes in traffic scenarios and the uncertainty of participant behavior pose significant challenges to autonomous driving systems' decision-making. Therefore, how to provide autonomous driving systems with flexible decision-making capabilities while ensuring safety has become a core research issue.
[0003] Currently, autonomous driving decision-making methods primarily fall into three categories: rule-based, deep learning-based, and human-like driving. In complex traffic scenarios, these methods struggle to balance safety, comfort, and social compatibility, resulting in inflexible vehicle behavior and even the potential for accidents. While deep reinforcement learning methods can optimize strategies through environmental interaction, they are prone to falling into local optimality. Summary of the Invention
[0004] This application provides a method for generating intelligent decisions for autonomous driving, which solves the technical problem that current autonomous driving vehicles have difficulty in generating flexible decisions that meet safety, comfort, and social compatibility, and are prone to falling into local optimality.
[0005] To achieve the above objectives, this application adopts the following technical solutions:
[0006] In a first aspect, a method for generating intelligent decisions for autonomous driving is provided, comprising: obtaining image information and historical driving data of each vehicle in a traffic scene in which the autonomous driving vehicle is located; determining a global state vector of the traffic scene based on the image information and historical driving data of each vehicle, wherein the global state vector is used to characterize dynamic information in the traffic scene; determining a right-of-way area of each vehicle based on the historical driving data of each vehicle in the traffic scene, wherein the right-of-way area is used to characterize areas around the vehicle that require priority avoidance; determining a right-of-way violation score of each vehicle based on the degree of overlap of the right-of-way areas of each vehicle; inputting the global state vector and the right-of-way violation score into a deep reinforcement learning model to obtain an output result, wherein the output result is used to characterize at least one driving intention of the autonomous driving vehicle and the confidence level of each driving intention; and determining the driving intention corresponding to the highest confidence score as the current driving decision.
[0007] In conjunction with the first aspect above, in one possible implementation, determining a global state vector of a traffic scene includes:
[0008] Collect traffic scene images and determine the visual feature matrix V of the traffic scene imagest ;
[0009] Collect historical trajectory data of each traffic participant in the traffic scene, wherein the historical trajectory data includes at least one of the following: position p t-n:t , speed v t-n:t , acceleration a t-n:t ;
[0010] Based on the formula: Identify time series behavior characteristics in, is the dynamic behavior feature vector of the i-th participant; FC represents the fully connected neural network module; p i,t-n:t is the position of the i-th participant; v i,t-n:t is the speed of the i-th participant; a i,t-n:t is the acceleration of the i-th participant;
[0011] Based on the formula: Obtaining a vehicle area vector in the traffic scene, where the vehicle area vector is used to characterize the dynamic behavior trend of each participant at a current moment;
[0012] The vehicle region vector is weighted by spatiotemporal attention to obtain the global state vector.
[0013] In combination with the first aspect above, in one possible implementation, spatiotemporal attention weighting is performed on the vehicle region vector to obtain a global state vector, including:
[0014] Vehicle area vector Perform spatial attention weighting based on the formula: Determine the context vector z t ;in, is the attention weight of the i-th region at time step t; L represents the total number of regions involved in the calculation;
[0015] For the context vector z t Perform time attention weighting based on the formula: Determine the global state vector C T ; where i is each time step; T is the total number of time steps; w t is the scalar weight, w t Satisfies the following formula:
[0016] w t =Softmax(v t ·h t )t=1,2,...,T
[0017] Where t is the time series in the entire time length of T; v tis the feature vector obtained during the training process; h t is the hidden state vector corresponding to time step t; Softmax() is the normalization function.
[0018] In conjunction with the first aspect above, in one possible implementation, determining the right-of-way violation score of each vehicle includes:
[0019] Determining the right-of-way area of each vehicle based on historical vehicle travel data, where the historical vehicle travel data includes at least one of the following: position, speed, acceleration, and traffic density;
[0020] Based on the overlap of the right-of-way areas of each vehicle, a right-of-way violation score is determined for each vehicle. The score is used to characterize the compatibility of autonomous driving behavior with social traffic norms.
[0021] The right-of-way area A-ROW of each vehicle satisfies the following formula:
[0022] A-ROW={(x,y)|x min ≤x≤x max ,y min ≤y≤y max}
[0023] in:
[0024] x min =x,x max =x+L·cosθ
[0025]
[0026] Among them, (x, y) represents the current position of the vehicle; x min The starting horizontal coordinate of the right-of-way area is aligned horizontally with the current position of the vehicle; ω is the width of the vehicle; y min The longitudinal extent of the right-of-way area is determined by the width of the vehicle; θ is the vehicle's driving direction angle; and the stopping distance L satisfies the following formula:
[0027]
[0028] Where v is the speed of the vehicle; a max is the maximum deceleration of the vehicle; ρ is the density of vehicles on the road; k ρ is the density adjustment factor;
[0029] In combination with the first aspect above, in a possible implementation, the road right violation score SC T Satisfies the following formula:
[0030] SC T =αT i+β(A-ROW) i
[0031] Among them, T i The total duration of vehicles in the traffic scene that violate the right of way of other vehicles (A-ROW); i It is the cumulative overlapping area of the right-of-way area violated by other vehicles in the traffic scene; α and β are weighting coefficients.
[0032] In conjunction with the first aspect above, in one possible implementation, determining at least one driving intention of the autonomous driving vehicle and the confidence level of each driving intention includes:
[0033] Based on the global state vector C T and road right-of-way violation score SC T Get the joint input vector s joint =[C T ,SC T ];
[0034] Based on the formula: π(s joint ;θ)=[p keep ,p left ,p right ], determine the driving intention probability vector π(s joint ; θ); the driving intention probability vector is the confidence level corresponding to each driving intention, and the driving intention includes at least one of the following: keeping lane, changing lanes left, changing lanes right; where θ is a network parameter; p keep To maintain lane confidence; p left is the confidence level of the left lane change; p right is the confidence level of right lane change.
[0035] In combination with the first aspect above, in a possible implementation, the training process of the deep reinforcement learning model includes:
[0036] Combined input vector s joint =[C T ,SC T ] is input into the deep reinforcement learning model to obtain the driving intention and the confidence level corresponding to each driving intention;
[0037] Based on the formula: A=argmaxπ(s joint ;θ), determine the driving decision A corresponding to the current maximum confidence; where argmax is π(s joint ; θ) corresponds to the driving intention when it reaches the maximum value;
[0038] Generate corresponding continuous control parameters for driving decision A, which include at least one of the following: steering angle control and acceleration control; wherein the steering angle control range is [-0.5, 0.5] rad, which is used to control the lateral displacement of the vehicle; the acceleration control range is [-5, 5] m / s 2 Continuous control parameters can be directly applied to autonomous vehicles.
[0039] enabling the autonomous vehicle to execute continuous control parameters;
[0040] The network parameters θ are updated using a multi-objective adaptation function; the deep reinforcement learning model is based on the driving decision A, the confidence level corresponding to the driving decision A, the current global state vector, and the road right violation score SC T After training, the model returns the decision corresponding to the highest confidence and the reward corresponding to the decision. The reward serves as feedback to adjust the confidence.
[0041] In combination with the first aspect above, in one possible implementation, the multi-objective adaptation function Fitness satisfies the following formula:
[0042] Fitness=ω1·f safe +ω2·f eff +ω3·f conf +ω4·f soc
[0043] Among them, f safe is the safety index; f eff is the traffic efficiency index; f conf is the comfort index; f soc is the social compatibility index; ω1, ω2, ω3, ω4 are the weight coefficients of each index;
[0044] f safe Satisfies the following formula:
[0045]
[0046] Among them, d t is the distance between the autonomous vehicle and the nearest vehicle in front, d min is the safety threshold, t is the time step, and T is the total time it takes to complete the autonomous driving decision execution;
[0047] and\or, f eff Satisfies the following formula:
[0048]
[0049] Among them, v avg is the average speed of the autonomous vehicle, v des The expected speed of the autonomous vehicle;
[0050] and\or, f conf Satisfies the following formula:
[0051]
[0052] Among them, a t is the longitudinal acceleration, ω t is the angular velocity, α and β are weighted coefficients for adjusting the riding experience; t is the time step, and T is the total time it takes for the autonomous driving decision to be executed;
[0053] and\or, f soc Satisfies the following formula:
[0054]
[0055] in, represents the space occupied by the vehicle at time t; It represents the absolute right-of-way area of the jth surrounding vehicle in the traffic scene at that moment; Area() is the area calculation function; t is the time step, and T is the total time to complete the autonomous driving decision execution.
[0056] In combination with the first aspect above, in a possible implementation, the reward R t Satisfies the following formula:
[0057]
[0058] in, It is a safety reward item; It is an efficiency reward item; It is a comfort bonus item; is the social compatibility reward item; ω1, ω2, ω3, ω4 are the weight coefficients of each sub-item;
[0059] Satisfies the following formula:
[0060]
[0061] Among them, d min Indicates the minimum headway distance between the vehicle and the nearest vehicle in the current time step; ∈ takes 10 -6 ;
[0062] Satisfies the following formula:
[0063]
[0064] Among them, v t is the current speed; v des is the expected speed;
[0065] Satisfies the following formula:
[0066]
[0067] Among them, a t is the linear acceleration, ω t is the angular velocity, α and β are the empirical coefficients for adjusting the weight;
[0068] Satisfies the following formula:
[0069]
[0070] Among them, A overlap is the spatial overlapping area between the absolute right-of-way area of the vehicle and the absolute right-of-way area of the infringed vehicle; γ is the penalty coefficient.
[0071] In a second aspect, a device for generating intelligent decisions for autonomous driving includes: a communication unit and a processing unit: the communication unit is configured to obtain image information and historical driving data of each vehicle in a traffic scene in which the autonomous driving vehicle is located;
[0072] The processing unit is used to determine the global state vector of the traffic scene based on the image information and historical driving data of each vehicle, and the global state vector is used to represent the dynamic information in the traffic scene; determine the right-of-way area of each vehicle based on the historical driving data of each vehicle in the traffic scene, and the right-of-way area is used to represent the area around the vehicle that requires priority avoidance; determine the right-of-way violation behavior score of each vehicle based on the overlap of the right-of-way areas of each vehicle; input the global state vector and the right-of-way violation behavior score into the deep reinforcement learning model to obtain an output result, and the output result is used to represent at least one driving intention of the autonomous driving vehicle, as well as the confidence level of each driving intention; determine the driving intention corresponding to the highest confidence score as the current driving decision.
[0073] In a third aspect, the present application provides a method for generating intelligent decisions for autonomous driving, comprising: a processor and a storage medium; the storage medium comprises instructions, and the processor is used to execute the instructions to implement the method described in the first aspect and any possible implementation of the first aspect.
[0074] In a fourth aspect, the present application provides a computer-readable storage medium, which stores instructions. When the instructions are executed on an electronic device, the electronic device executes the method described in the first aspect and any possible implementation of the first aspect.
[0075] In a fifth aspect, the present application provides a computer program product comprising instructions, which, when executed on an electronic device, enables the electronic device to execute the method as described in the first aspect and any possible implementation manner of the first aspect.
[0076] The present application provides a method for generating intelligent decisions for autonomous driving, which obtains image information and historical driving data of each vehicle in a traffic scene in which the autonomous driving vehicle is located; determines a global state vector of the traffic scene based on the image information and historical driving data of each vehicle, and the global state vector is used to represent dynamic information in the traffic scene; determines the right-of-way area of each vehicle based on the historical driving data of each vehicle in the traffic scene, and the right-of-way area is used to represent the area around the vehicle that requires priority avoidance; determines the right-of-way violation behavior score of each vehicle based on the overlap of the right-of-way areas of each vehicle; inputs the global state vector and the right-of-way violation behavior score into a deep reinforcement learning model to obtain an output result, and the output result is used to represent at least one driving intention of the autonomous driving vehicle, as well as the confidence level of each driving intention; and determines that the driving intention corresponding to the highest confidence score is the current driving decision.
[0077] Based on this, the present application provides a method for generating intelligent decisions for autonomous driving that integrates the flexibility, globality, and comprehensive performance of autonomous driving vehicle decisions, enabling safe, efficient, and socially interactive driving decisions to be made in complex traffic scenarios without falling into local optimality.
[0078] It should be understood that the description of technical features, technical solutions, beneficial effects or similar language in this application does not imply that all features and advantages can be realized in any single embodiment. On the contrary, it is understood that the description of a feature or beneficial effect means that a specific technical feature, technical solution or beneficial effect is included in at least one embodiment. Therefore, the description of a technical feature, technical solution or beneficial effect in this specification does not necessarily refer to the same embodiment. Furthermore, the technical features, technical solutions and beneficial effects described in the present embodiment can also be combined in any appropriate manner. Those skilled in the art will understand that the embodiment can be implemented without one or more specific technical features, technical solutions or beneficial effects of a specific embodiment. In other embodiments, additional technical features and beneficial effects can also be identified in specific embodiments that do not embody all embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0079] Figure 1 A system architecture diagram of a system for generating intelligent decisions for autonomous driving provided in an embodiment of the present application;
[0080] Figure 2 A flowchart of a method for generating intelligent decisions for autonomous driving provided in an embodiment of the present application;
[0081] Figure 3 A flowchart of another method for generating intelligent autonomous driving decisions provided in an embodiment of the present application;
[0082] Figure 4 A schematic diagram of the structure of a method and apparatus for generating intelligent decisions for autonomous driving provided in an embodiment of the present application;
[0083] Figure 5 A schematic diagram of the hardware structure of a device for generating an intelligent decision for autonomous driving provided in an embodiment of the present application. DETAILED DESCRIPTION
[0084] In the description of this application, unless otherwise specified, " / " means "or", for example, A / B can mean A or B. "And / or" in this article is merely a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, "at least one" means one or more, and "a plurality" means two or more. Words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not limit them to be necessarily different.
[0085] It should be noted that, in this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described in this application as "exemplary" or "for example" should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0086] The method for generating an autonomous driving intelligent decision-making provided in the embodiment of the present application can be applied to Figure 1 In the generation system of an autonomous driving intelligent decision, as shown in Figure 1 As shown, the communication system includes: a communication device 101 and an electronic device 102.
[0087] Among them, the communication device 101 is used to obtain image information and historical driving data of each vehicle in the traffic scene where the autonomous driving vehicle is located.
[0088] The electronic device 102 is used to determine the global state vector of the traffic scene based on the image information and historical driving data of each vehicle, the global state vector is used to represent the dynamic information in the traffic scene, and the right-of-way area of each vehicle is determined based on the historical driving data of each vehicle in the traffic scene. The right-of-way area is used to represent the area around the vehicle that requires priority avoidance, and the right-of-way violation behavior score of each vehicle is determined based on the overlap of the right-of-way areas of each vehicle. The global state vector and the right-of-way violation behavior score are input into the deep reinforcement learning model to obtain an output result. The output result is used to represent at least one driving intention of the autonomous driving vehicle and the confidence of each driving intention, and the driving intention corresponding to the highest confidence score is determined as the current driving decision.
[0089] With the rapid development of autonomous driving technology, improving vehicle decision-making capabilities has become a key research direction. Its core goal is to achieve safe and efficient autonomous decision-making in complex traffic environments. However, the dynamic changes in traffic scenarios and the uncertainty of participant behavior pose significant challenges to autonomous driving systems' decision-making. Therefore, how to provide autonomous driving systems with flexible decision-making capabilities while ensuring safety has become a core research issue.
[0090] Autonomous driving decision-making methods primarily fall into three categories: rule-based, deep learning-based, and human-like driving. In complex traffic scenarios, these methods struggle to balance safety, comfort, and social compatibility, resulting in inflexible vehicle behavior and even the possibility of accidents. While deep reinforcement learning methods can optimize strategies through environmental interaction, they are prone to falling into local optimality.
[0091] To solve the technical problem that current autonomous driving vehicles are difficult to generate flexible decisions that meet the requirements of safety, comfort and social compatibility and are prone to falling into local optimality. An embodiment of the present application provides a method for generating intelligent decisions for autonomous driving, which includes: obtaining image information and historical driving data of each vehicle in the traffic scene where the autonomous driving vehicle is located; determining the global state vector of the traffic scene based on the image information and historical driving data of each vehicle, the global state vector is used to represent the dynamic information in the traffic scene; determining the right-of-way area of each vehicle based on the historical driving data of each vehicle in the traffic scene, the right-of-way area is used to represent the area around the vehicle that needs to be avoided first; determining the right-of-way violation behavior score of each vehicle based on the overlap of the right-of-way areas of each vehicle; inputting the global state vector and the right-of-way violation behavior score into a deep reinforcement learning model to obtain an output result, the output result is used to represent at least one driving intention of the autonomous driving vehicle, and the confidence of each driving intention; determining the driving intention corresponding to the highest confidence score as the current driving decision.
[0092] like Figure 2As shown, the method for generating an intelligent decision for autonomous driving provided in an embodiment of the present application includes:
[0093] S201: Determine a global state vector of a traffic scene.
[0094] Among them, the traffic scene is the scene around the autonomous driving vehicle captured by its own sensors; the global state vector is used to represent the dynamic information in the traffic scene.
[0095] In one possible implementation, the global state vector C T Satisfies the following formula:
[0096]
[0097] Where i is each time step; T is the total number of time steps; z t is the context vector; w t is the scalar weight, w t Satisfies the following formula:
[0098] w t =Softmax(v t ·h t )t=1,2,...,T
[0099] Where t is the time series in the entire time length of T; v t is the feature vector obtained during the training process; h t is the hidden state vector corresponding to time step t; Softmax() is the normalization function.
[0100] S202: Determine the road right violation score of each vehicle.
[0101] Among them, right of way is used to characterize the areas in the traffic scene that each vehicle needs to give priority to avoid; the right of way violation behavior score is used to characterize the compatibility between autonomous driving behavior and social traffic norms. The higher the right of way violation behavior score, the lower the compatibility of the vehicle's current behavior with social rules; conversely, the lower the right of way violation behavior score, the higher the compatibility of the vehicle's current behavior with social rules.
[0102] In one possible implementation, the road right violation score SC T Satisfies the following formula:
[0103] SC T =αT i +β(A-ROW) i
[0104] Among them, T i The total duration of vehicles in the traffic scene that violate the right of way of other vehicles (A-ROW);i It is the cumulative overlapping area of the right-of-way area violated by other vehicles in the traffic scene; α and β are weighting coefficients.
[0105] S203: Determine at least one driving intention of the autonomous driving vehicle and a confidence level of each driving intention.
[0106] The driving intention includes at least one of the following: keeping lane, changing lanes left, and changing lanes right.
[0107] Through the above global state vector C T and right-of-way violation score SC T Determine joint =[C T ,SC T ], based on the formula: π(s joint ;θ)=[p keep ,p left ,p right ], determine the driving intention probability vector π(s joint ;θ).
[0108] Among them, θ is the policy network parameter; p keep To maintain lane confidence; p left is the confidence level of the left lane change; p right is the confidence level of right lane change.
[0109] S204: Determine the driving intention corresponding to the highest confidence score as the current driving decision.
[0110] Among them, the decision intention corresponding to the highest confidence in the driving intention probability vector obtained each time is taken as the current driving decision.
[0111] Based on the above technical solution, in an embodiment of the present application, by obtaining image information and historical driving data of each vehicle in the traffic scene in which the autonomous driving vehicle is located; based on the image information and historical driving data of each vehicle, the global state vector of the traffic scene is determined, and the global state vector is used to characterize the dynamic information in the traffic scene; based on the historical driving data of each vehicle in the traffic scene, the right-of-way area of each vehicle is determined, and the right-of-way area is used to characterize the area around the vehicle that needs priority avoidance; based on the overlap of the right-of-way areas of each vehicle, the right-of-way violation behavior score of each vehicle is determined; the global state vector and the right-of-way violation behavior score are input into the deep reinforcement learning model to obtain an output result, and the output result is used to characterize at least one driving intention of the autonomous driving vehicle, as well as the confidence of each driving intention; the driving intention corresponding to the highest confidence score is determined as the current driving decision.
[0112] In some embodiments, combined Figure 2 ,like Figure 3As shown, the process of determining the global state vector of the traffic scene in step 201 can be specifically implemented by the following steps 301 to 304:
[0113] S301. Obtain image information and historical driving data of each vehicle in the traffic scene where the autonomous driving vehicle is located.
[0114] Among them, image information is converted into a sequence of images with continuous time steps by devices on autonomous driving vehicles, such as videos captured by cameras; historical driving data is the driving data of each traffic participant in the traffic scene obtained based on the collected video, and the driving data includes at least one of the following: position, speed, acceleration, and traffic density.
[0115] S302: Based on the image information and historical driving data of each vehicle, determine the time series behavior characteristics and obtain the visual feature matrix of the traffic scene image.
[0116] Among them, the time series behavior feature is to encode the traffic scene images in each time step and obtain the visual feature map of each image. After the convolution network processing, the visual feature matrix V of each image is obtained. t , V t Satisfies the following formula: Among them, H and W represent the number of spatial regions into which the image is divided, and C is the channel dimension.
[0117] S303: Determine the regional feature vector based on the visual feature matrix.
[0118] In one possible implementation, based on the formula: Identify time series behavior characteristics in, is the dynamic behavior feature vector of the i-th participant; FC represents the fully connected neural network module; p i,t-n:t is the position of the i-th participant; v i,t-n:t is the speed of the i-th participant; a i,t-n:t is the acceleration of the i-th participant.
[0119] Based on the formula: Determine the regional feature vector in the traffic scene. The regional feature vector is used to characterize the dynamic behavior trend of each participant at the current moment. t is the above visual feature matrix; is the behavioral feature of the above time series; φ() is the feature fusion function.
[0120] It should be pointed out that fusion refers to splicing or mapping the input features into the same space.
[0121] S304: The regional features are weighted in the spatiotemporal attention to determine the global state vector.
[0122] Among them, spatiotemporal attention is the temporal attention network and the spatial attention network; the global state vector is used to represent the dynamic information in the traffic scene.
[0123] In one possible implementation, based on the formula: Vehicle area vector Perform spatial attention weighting to determine the context vector z t ;in, is the attention weight of the i-th region at time step t; L represents the total number of regions involved in the calculation.
[0124] Based on the formula: For the context vector z t Perform time attention weighting to determine the global state vector C T ; Where i is each time step; T is the total time; w t is the scalar weight, w t Satisfies the following formula:
[0125] w t =Softmax(v t ·h t )t=1,2,...,T
[0126] Where t is the time series in the entire time length of T; v t is the feature vector obtained during the training process; h t is the hidden state vector corresponding to time step t; Softmax() is the normalization function.
[0127] In some embodiments, combined Figure 2 ,like Figure 3 As shown, the process of determining the road right violation score of each vehicle in step 201 can be specifically implemented through the following steps 305-306:
[0128] S305 : Determine the right-of-way area of each vehicle based on the image information and historical driving data of each vehicle.
[0129] The image information and historical driving data of each vehicle are obtained in step S301 ; the right-of-way area is a corresponding right-of-way area generated for all vehicles in the traffic scene.
[0130] In one possible implementation, the vehicle's right-of-way area A-ROW satisfies the following formula:
[0131] A-ROW={(x,y)|x min≤x≤x max ,y min ≤y≤y max}
[0132] in:
[0133] x min =x,x max =x+L·cosθ
[0134]
[0135] Among them, (x, y) represents the current position of the vehicle; x min The starting horizontal coordinate of the right-of-way area is aligned horizontally with the current position of the vehicle; ω is the width of the vehicle; y min The longitudinal extent of the right-of-way area is determined by the width of the vehicle; θ is the vehicle's driving direction angle; and the stopping distance L satisfies the following formula:
[0136]
[0137] Where v is the speed of the vehicle; a max is the maximum deceleration of the vehicle; ρ is the density of vehicles on the road; k ρ is the density adjustment factor;
[0138] The right-of-way area A-ROW of each vehicle is (x, y) and meets the requirements of vehicle x min ≤x≤x max and ≤y≤y max If the right-of-way area A-ROW of the vehicle overlaps, it is determined that there is a right-of-way violation.
[0139] For example, vehicle speed v, maximum deceleration a max , traffic density ρ and density adjustment coefficient k ρ They are all set according to experimental or simulated roads; among them, the density adjustment coefficient k ρ The value is generally 0.5-2.0.
[0140] S306 : Determine a right-of-way violation score for each vehicle based on the overlap of the right-of-way areas of each vehicle.
[0141] The degree of overlap is determined by the overlapping areas of the right-of-way areas of each vehicle and the duration of the right-of-way infringement.
[0142] In some embodiments, combined Figure 2 ,like Figure 3 As shown, the process of determining the driving intention corresponding to the highest score confidence as the current driving decision in step 204 can be specifically implemented by the following steps 307-308:
[0143] Step 307: The global state vector C T and road right violation score SC T Input into the deep reinforcement learning model to get the output result.
[0144] Among them, the deep reinforcement learning model adopts a combination of strategy network and genetic algorithm architecture, in which the input of the strategy network is based on the global state vector C T and road right-of-way violation score SC T The obtained joint input vector s joint =[C T ,SC T The output is a driving intention probability vector. The genetic algorithm provides feedback for policy optimization, using a multi-objective adaptation function to optimize the parameters of the deep reinforcement learning model. Based on the input driving decision, it returns a reward and a new driving decision.
[0145] In one possible implementation, based on the formula: π(s joint ;θ)=[p keep ,p left ,p right ], determine the driving intention probability vector π(s joint ; θ); the driving intention probability vector is the confidence level corresponding to each driving intention. Driving intention includes at least one of the following: keeping lane, changing lanes left, changing lanes right; where θ is a network parameter; p keep To maintain lane confidence; p left is the confidence level of the left lane change; p right is the confidence level of right lane change.
[0146] Step 308: Training of deep reinforcement learning model.
[0147] The training process uses the genetic algorithm mentioned above; the input driving decision is the driving intention with the highest confidence among the driving intentions generated by the deep reinforcement learning model each time. Based on the formula: A = argmaxπ(s joint ;θ), determine the driving decision A corresponding to the current maximum confidence; where argmax is π(s joint The driving intention corresponding to the maximum value of θ) is generated from the driving decision A to generate corresponding continuous control parameters, which include at least one of the following: steering angle control and acceleration control; wherein the steering angle control range is [-0.5, 0.5] rad, which is used to control the lateral displacement of the vehicle; the acceleration control range is [-5, 5] m / s 2 .
[0148] It's important to note that continuous control parameters can be directly applied to the autonomous vehicle, causing it to execute driving decision A corresponding to the continuous control parameters. A multi-objective adaptation function is used to update the network parameter θ. A deep reinforcement learning model is trained based on the current driving decision A and the confidence level. The model returns a new driving decision and the corresponding reward, which serves as feedback to adjust the confidence level. If the executed driving decision A results in a higher reward, the model will increase the confidence level corresponding to driving decision A; otherwise, it will decrease the confidence level corresponding to driving decision A. The timing of driving decision execution is pre-set, for example, executing the next driving decision updated by the genetic algorithm every 1 second.
[0149] In one possible implementation, the multi-objective fitness function Fitness satisfies the following formula:
[0150] Fitness=ω1·f safe +ω2·f eff +ω3·f conf +ω4·f soc
[0151] Among them, f safe is the safety index; f eff is the traffic efficiency index; f conf is the comfort index; f soc is the social compatibility index; ω1, ω2, ω3, ω4 are the weight coefficients of each index;
[0152] f safe Satisfies the following formula:
[0153]
[0154] Among them, d t is the distance between the autonomous vehicle and the nearest vehicle in front, d min is the safety threshold, t is the time step, and T is the total time it takes to complete the autonomous driving decision execution;
[0155] and\or, f eff Satisfies the following formula:
[0156]
[0157] Among them, v avg is the average speed of the autonomous vehicle, v des The expected speed of the autonomous vehicle;
[0158] and\or, f conf Satisfies the following formula:
[0159]
[0160] Among them, a t is the longitudinal acceleration, ω t is the angular velocity, α and β are weighted coefficients for adjusting the riding experience; t is the time step, and T is the total time it takes for the autonomous driving decision to be executed;
[0161] and\or, f soc Satisfies the following formula:
[0162]
[0163] in, represents the space occupied by the vehicle at time t; It represents the absolute right-of-way area of the jth surrounding vehicle in the traffic scene at that moment; Area() is the area calculation function; t is the time step, and T is the total time to complete the autonomous driving decision execution.
[0164] It should be pointed out that the above multi-objective adaptation function can be linearly weighted. In each iteration of the algorithm, the weight coefficient is set by the driving decision made and the importance of the four indicators in the experimental or simulation environment.
[0165] In one possible implementation, the reward R t Satisfies the following formula:
[0166]
[0167] in, It is a safety bonus item; It is an efficiency reward item; It is a comfort bonus item; is the social compatibility reward item; ω1, ω2, ω3, ω4 are the weight coefficients of each sub-item;
[0168] Satisfies the following formula:
[0169]
[0170] Among them, d min Indicates the minimum headway distance between the vehicle and the nearest vehicle in the current time step; ∈ takes 10 -6 ;
[0171] Satisfies the following formula:
[0172]
[0173] Among them, v t is the current speed; v des is the expected speed;
[0174] Satisfies the following formula:
[0175]
[0176] Among them, a t is the linear acceleration, ω t is the angular velocity, α and β are the empirical coefficients for adjusting the weight;
[0177] Satisfies the following formula:
[0178]
[0179] Among them, A overlap is the spatial overlapping area between the absolute right-of-way area of the vehicle and the absolute right-of-way area of the infringed vehicle; γ is the penalty coefficient.
[0180] In the embodiment of the present application, the image information and historical driving data of each vehicle in the traffic scene where the autonomous driving vehicle is located are obtained; based on the image information and historical driving data of each vehicle, the global state vector of the traffic scene is determined, and the global state vector is used to represent the dynamic information in the traffic scene, which solves the technical problem that current autonomous driving vehicles are difficult to generate flexible decisions that meet safety, comfort and social compatibility, and are prone to falling into local optimality. Based on the historical driving data of each vehicle in the traffic scene, the right-of-way area of each vehicle is determined, and the right-of-way area is used to represent the area around the vehicle that requires priority avoidance; based on the overlap of the right-of-way areas of each vehicle, the right-of-way violation behavior score of each vehicle is determined; the global state vector and the right-of-way violation behavior score are input into the deep reinforcement learning model to obtain the output result, and the output result is used to represent at least one driving intention of the autonomous driving vehicle, as well as the confidence level of each driving intention, so that the driving decision of the autonomous driving vehicle meets safety, comfort and social compatibility.
[0181] The above mainly introduces the scheme of the embodiment of the present application from the perspective of device implementation. It is understandable that each device, for example, the generation device of autonomous driving intelligent decision, includes at least one of the hardware structure and software modules corresponding to the execution of each function in order to realize the above functions. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0182] The embodiment of the present application can divide the functional units of the generation device of the autonomous driving intelligent decision according to the above method example. For example, each functional unit can be divided according to each function, or two or more functions can be integrated into one processing unit. The above integrated unit can be implemented in the form of hardware or in the form of software functional units. It should be noted that the division of units in the embodiment of the present application is schematic and is only a logical functional division. There may be other division methods in actual implementation.
[0183] In the case of an integrated unit, Figure 4 A possible structural diagram of a device for generating an autonomous driving intelligent decision (denoted as a device for generating an autonomous driving intelligent decision 40) involved in the above-mentioned embodiment is shown. The device for generating an autonomous driving intelligent decision 40 includes a processing unit 401 and a communication unit 402, and may also include a storage unit 403. Figure 4 The structural diagram shown can be used to illustrate the structure of the automatic driving intelligent decision-making generation device involved in the above embodiments.
[0184] when Figure 4 The structural schematic diagram shown is used to illustrate the structure of the autonomous driving intelligent decision-making generating device involved in the above-mentioned embodiment. The processing unit 401 is used to control and manage the actions of the autonomous driving intelligent decision-making generating device, the communication unit 402 is used for the autonomous driving intelligent decision-making generating device to communicate with other devices, and the storage unit 403 is used to store the program code and data of the autonomous driving intelligent decision-making generating device.
[0185] For example, the communication unit 401 is used to obtain image information and historical driving data of each vehicle in the traffic scene where the autonomous driving vehicle is located;
[0186] Processing unit 401 is used to determine the global state vector of the traffic scene based on the image information and historical driving data of each vehicle, and the global state vector is used to represent the dynamic information in the traffic scene; determine the right-of-way area of each vehicle based on the historical driving data of each vehicle in the traffic scene, and the right-of-way area is used to represent the area around the vehicle that requires priority avoidance; determine the right-of-way violation behavior score of each vehicle based on the overlap of the right-of-way areas of each vehicle; input the global state vector and the right-of-way violation behavior score into the deep reinforcement learning model to obtain an output result, and the output result is used to represent at least one driving intention of the autonomous driving vehicle, as well as the confidence of each driving intention; determine the driving intention corresponding to the highest confidence score as the current driving decision.
[0187] In one possible implementation, the road right violation score SC T Satisfies the following formula:
[0188] SC T =αT i +β(A-ROW) i
[0189] Among them, T i The total duration of vehicles in the traffic scene that violate the right of way of other vehicles (A-ROW); i It is the cumulative overlapping area of the right-of-way area violated by other vehicles in the traffic scene; α and β are weighting coefficients.
[0190] In a possible implementation, the processing unit 401 is further configured to determine a visual feature matrix of the traffic scene image; and determine a time series behavior feature.
[0191] In one possible implementation, the time series behavior characteristics Satisfies the following formula:
[0192]
[0193] in, is the dynamic behavior feature vector of the i-th participant; FC represents the fully connected neural network module; p i,t-n:t is the position of the i-th participant; v i,t-n:t is the speed of the i-th participant; a i,t-n:t is the acceleration of the i-th participant.
[0194] In a possible implementation, the processing unit 401 is further configured to obtain a regional feature vector in the traffic scene.
[0195] In one possible implementation, the region feature vector Satisfies the following formula:
[0196]
[0197] In a possible implementation, the processing unit 401 is further configured to perform spatiotemporal attention weighting on the regional feature vector to obtain a global state vector.
[0198] In one possible implementation, the global state vector C T Satisfies the following formula:
[0199]
[0200] Where i is each time step; T is the total number of time steps; z t is the context vector; w t is the scalar weight, w t Satisfies the following formula:
[0201] wt =Softmax(ν t ·h t )t=1,2,...,T
[0202] Where t is the time series in the entire time length of T; ν t is the feature vector obtained during the training process; h t is the hidden state vector corresponding to time step t; Softmax() is the normalization function.
[0203] In one possible implementation, the context vector z t Satisfies the following formula:
[0204]
[0205] in, is the attention weight of the i-th region at time step t; L represents the total number of regions involved in the calculation
[0206] In a possible implementation, the processing unit 401 is further configured to determine the right-of-way area of each vehicle.
[0207] In one possible implementation, the right-of-way area A-ROW of each vehicle satisfies the following formula:
[0208] A-ROW={(x,y)|x min ≤x≤x max ,y min ≤y≤y max}
[0209] in:
[0210] x min =x,x max =x+L·cosθ
[0211]
[0212] Among them, (x, y) represents the current position of the vehicle; x min The starting horizontal coordinate of the right-of-way area is aligned horizontally with the current position of the vehicle; ω is the width of the vehicle; y min The longitudinal extent of the right-of-way area is determined by the width of the vehicle; θ is the vehicle's driving direction angle; and the stopping distance L satisfies the following formula:
[0213]
[0214] Where v is the speed of the vehicle; a max is the maximum deceleration of the vehicle; ρ is the density of vehicles on the road; k ρ is the density adjustment factor.
[0215] In one possible implementation, the processing unit 401 is further configured to determine a driving intention probability vector, where the driving intention probability vector is a confidence level corresponding to each driving intention, and the driving intention includes at least one of the following: keeping lane, changing lanes left, and changing lanes right.
[0216] In one possible implementation, the driving intention probability vector π(s joint ; θ) satisfies the following formula:
[0217] π(s joint ;θ)=[p keep ,p left ,p right ]
[0218] Among them, θ is the network parameter; p keep To maintain lane confidence; p left is the confidence level of the left lane change; p right is the confidence level of right lane change.
[0219] In one possible implementation, the processing unit 401 is further configured to train a deep reinforcement learning model. The driving decision A is converted into corresponding continuous control parameters, and the autonomous vehicle is caused to execute the continuous control parameters. The network parameter θ is updated using a multi-objective adaptation function. The deep reinforcement learning model is configured to train the driving decision A, the confidence level corresponding to the driving decision A, the current global state vector, and the road right violation score SC. T After training, the model returns the decision corresponding to the highest confidence and the reward corresponding to the decision. The reward serves as feedback to adjust the confidence.
[0220] In one possible implementation, the multi-objective fitness function Fitness satisfies the following formula:
[0221] Fitness=ω1·f safe +ω2·f eff +ω3·f conf +ω4·f soc
[0222] Among them, f safe is the safety index; f eff is the traffic efficiency index; f conf is the comfort index; f soc is the social compatibility index; ω1, ω2, ω3, ω4 are the weight coefficients of each index;
[0223] f safe Satisfies the following formula:
[0224]
[0225] Among them, d t is the distance between the autonomous vehicle and the nearest vehicle in front, d min is the safety threshold, t is the time step, and T is the total time it takes to complete the autonomous driving decision execution;
[0226] and\or, f eff Satisfies the following formula:
[0227]
[0228] Among them, v avg is the average speed of the autonomous vehicle, v des The expected speed of the autonomous vehicle;
[0229] and\or, f conf Satisfies the following formula:
[0230]
[0231] Among them, a t is the longitudinal acceleration, ω t is the angular velocity, α and β are weighted coefficients for adjusting the riding experience; t is the time step, and T is the total time it takes for the autonomous driving decision to be executed;
[0232] and\or, f soc Satisfies the following formula:
[0233]
[0234] in, represents the space occupied by the vehicle at time t; It represents the absolute right-of-way area of the jth surrounding vehicle in the traffic scene at that moment; Area() is the area calculation function; t is the time step, and T is the total time to complete the autonomous driving decision execution.
[0235] In one possible implementation, the reward R t Satisfies the following formula:
[0236]
[0237] in, It is a safety reward item; It is an efficiency reward item; It is a comfort bonus item; is the social compatibility reward item; ω1, ω2, ω3, ω4 are the weight coefficients of each sub-item;
[0238] Satisfies the following formula:
[0239]
[0240] Among them, d min Indicates the minimum headway distance between the vehicle and the nearest vehicle in the current time step; ∈ takes 10 -6 ;
[0241] Satisfies the following formula:
[0242]
[0243] Among them, v t is the current speed; v des is the expected speed;
[0244] Satisfies the following formula:
[0245]
[0246] Among them, a t is the linear acceleration, ω t is the angular velocity, α and β are the empirical coefficients for adjusting the weight;
[0247] Satisfies the following formula:
[0248]
[0249] Among them, A overlap is the spatial overlapping area between the absolute right-of-way area of the vehicle and the absolute right-of-way area of the infringed vehicle; γ is the penalty coefficient.
[0250] Among them, the processing unit 401 can be a processor or a controller, and the communication unit 402 can be a communication interface, a transceiver, a transceiver, a transceiver circuit, a transceiver device, etc. Among them, the communication interface is a general term and can include one or more interfaces. The storage unit 403 can be a memory. When the generating device 40 for autonomous driving intelligent decision is a chip, the processing unit 401 can be a processor or a controller, and the communication unit 402 can be an input interface and / or output interface, a pin or a circuit, etc. The storage unit 403 can be a storage unit within the chip (for example, a register, a cache, etc.), or it can be a storage unit located outside the chip (for example, a read-only memory (ROM), a random access memory (RAM), etc.).
[0251] Among them, the communication unit can also be called a transceiver unit. The antenna and control circuit with transceiver functions in the autonomous driving intelligent decision-making generation device 40 can be regarded as the communication unit 402 of the communication device, and the processor with processing function can be regarded as the processing unit 401 of the autonomous driving intelligent decision-making generation device 40. Optionally, the device used to implement the receiving function in the communication unit 402 can be regarded as a communication unit, and the communication unit is used to perform the receiving steps in the embodiment of the present application. The communication unit can be a receiver, a receiver, a receiving circuit, etc. The device used to implement the sending function in the communication unit 402 can be regarded as a sending unit, and the sending unit is used to perform the sending steps in the embodiment of the present application. The sending unit can be a transmitter, a transmitter, a sending circuit, etc.
[0252] Figure 4 If the integrated units are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the various embodiments of the present application. The storage medium for storing computer software products includes various media that can store program codes, such as USB flash drives, mobile hard drives, read-only memories, random access memories, magnetic disks or optical disks.
[0253] Figure 4 A unit in a can also be called a module, for example, a processing unit can be called a processing module.
[0254] The embodiment of the present application also provides a hardware structure diagram of a device for generating an autonomous driving intelligent decision (denoted as a device for generating an autonomous driving intelligent decision 50), see Figure 5 The device 50 for generating intelligent decision for autonomous driving includes a processor 501 and, optionally, a memory 502 connected to the processor 501.
[0255] In the first possible implementation, see Figure 5, the generating device 50 for intelligent decision making for autonomous driving further includes a transceiver 503. The processor 501, the memory 502, and the transceiver 503 are connected via a bus. The transceiver 503 is used to communicate with other devices or a communication network. Optionally, the transceiver 503 may include a transmitter and a receiver. The device used to implement the receiving function in the transceiver 503 may be regarded as a receiver, and the receiver is used to perform the receiving step in the embodiment of the present application. The device used to implement the sending function in the transceiver 503 may be regarded as a transmitter, and the transmitter is used to perform the sending step in the embodiment of the present application.
[0256] Based on the first possible implementation, Figure 5 The structural diagram shown can be used to illustrate the structure of the automatic driving intelligent decision-making generation device involved in the above embodiments.
[0257] in, Figure 5 Alternatively, the system chip in the device for generating intelligent autonomous driving decisions may be illustrated. In this case, the actions performed by the device for generating intelligent autonomous driving decisions may be implemented by the system chip. The specific actions performed are described above and will not be repeated here.
[0258] During implementation, each step of the method provided in this embodiment can be completed by hardware integrated logic circuits in a processor or by software instructions. The steps of the method disclosed in the embodiments of this application can be directly implemented as execution by a hardware processor, or as a combination of hardware and software modules in a processor.
[0259] The processor in this application may include, but is not limited to, at least one of the following: a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a microcontroller unit (MCU), or an artificial intelligence processor, and other types of computing devices that run software. Each computing device may include one or more cores for executing software instructions to perform operations or processing. The processor may be a separate semiconductor chip, or it may be integrated into a semiconductor chip together with other circuits. For example, it may form an SoC (system on a chip) with other circuits (such as a codec circuit, a hardware acceleration circuit, or various bus and interface circuits), or it may be integrated into the ASIC as a built-in processor of the ASIC. The ASIC with the integrated processor may be packaged separately or together with other circuits. In addition to the core for executing software instructions to perform operations or processing, the processor may further include necessary hardware accelerators, such as a field programmable gate array (FPGA), a PLD (programmable logic device), or a logic circuit that implements dedicated logic operations.
[0260] The memory in the embodiments of the present application may include at least one of the following types: read-only memory (ROM) or other types of static storage devices that can store static information and instructions, random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or electrically erasable programmable read-only memory (EEPROM). In some scenarios, the memory may also be a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited to this.
[0261] An embodiment of the present application also provides a computer-readable storage medium, comprising instructions, which, when executed on a computer, enables the computer to execute any of the above methods.
[0262] An embodiment of the present application also provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute any of the above methods.
[0263] An embodiment of the present application also provides a chip, which includes a processor and an interface circuit, the interface circuit is coupled to the processor, the processor is used to run a computer program or instruction to implement the above method, and the interface circuit is used to communicate with other modules outside the chip.
[0264] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using a software program, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more media that can be integrated. The available media may be magnetic media (eg, floppy disks, hard disks, magnetic tapes), optical media (eg, DVDs), or semiconductor media (eg, solid state disks (SSDs)).
[0265] Although the present application is described herein in conjunction with various embodiments, in the process of implementing the claimed application, those skilled in the art may understand and implement other variations of the disclosed embodiments by reviewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude multiple situations. A single processor or other unit may implement several functions listed in the claims. Certain measures are recorded in different dependent claims, but this does not mean that these measures cannot be combined to produce good results.
[0266] Although the present application has been described with reference to specific features and embodiments thereof, it is apparent that various modifications and combinations may be made thereto without departing from the spirit and scope of the present application. Accordingly, this specification and the drawings are merely illustrative of the present application as defined by the appended claims and are deemed to cover any and all modifications, variations, combinations or equivalents within the scope of the present application. Obviously, those skilled in the art may make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, the present application is intended to include such modifications and variations as fall within the scope of the claims of the present application and their equivalents.
Claims
1. A method for generating intelligent decision-making for autonomous driving, characterized in that: include: Obtain image information and historical driving data of each vehicle in the traffic scene where the autonomous vehicle is located; Determining a global state vector of the traffic scene based on the image information and historical driving data of each vehicle, wherein the global state vector is used to represent dynamic information in the traffic scene; Determining a right-of-way area for each vehicle based on historical driving data of each vehicle in the traffic scene, wherein the right-of-way area is used to represent an area around the vehicle that requires priority avoidance; Determining a right-of-way violation score for each vehicle based on the overlap of the right-of-way areas of each vehicle; Inputting the global state vector and the right-of-way violation score into a deep reinforcement learning model to obtain an output result, wherein the output result is used to represent at least one driving intention of the autonomous driving vehicle and a confidence level of each driving intention; The driving intention corresponding to the highest confidence score is determined as the current driving decision.
2. The method according to claim 1, characterized in that Determining a global state vector of the traffic scene includes: Collect the traffic scene image and determine the visual feature matrix V of the traffic scene image t ; Collect historical trajectory data of each traffic participant in the traffic scene, wherein the historical trajectory data includes at least one of the following: position p t-n:t , speed v t-n:t , acceleration a t-n:t ; Based on the formula: Characterizing time series behavior in, is the dynamic behavior feature vector of the i-th participant; FC represents the fully connected neural network module; p i,t-n:t is the position of the i-th participant; v i,t-n:t is the speed of the i-th participant; a i,t-n:t is the acceleration of the i-th participant; Based on the formula: Obtaining a regional feature vector in the traffic scene, wherein the regional feature vector is used to characterize the dynamic behavior trend of each participant at a current moment; Perform spatiotemporal attention weighting on the vehicle region feature vector to obtain a global state vector.
3. The method according to claim 2, characterized in that Perform spatiotemporal attention weighting on the regional feature vector to obtain a global state vector, including: The regional feature vector Perform spatial attention weighting based on the formula: Determine the context vector z t ;in, is the attention weight of the i-th region at time step t; L represents the total number of regions involved in the calculation; For the context vector z t Perform time attention weighting based on the formula: Determine the global state vector C T ; Where i is each time step; T is the total time; w t is the scalar weight, w t Satisfies the following formula: w t =Softmax(v t ·h t )t=1,2,...,T Where t is the time series in the entire time length of T; v t is the feature vector obtained during the training process; h t is the hidden state vector corresponding to time step t; Softmax() is the normalization function.
4. The method according to claim 1, wherein Determine a right-of-way violation score for each vehicle, including: Determining the right-of-way area of each vehicle based on the historical driving data of the vehicle; Determining a right-of-way violation score for each vehicle based on the overlap of the right-of-way areas of each vehicle, the score being used to characterize the compatibility between the autonomous driving behavior and social traffic norms; The right-of-way area A_ROW of each vehicle satisfies the following formula: A_ROW={(x,y)∣x min ≤x≤x max ,y min ≤y≤y max } in: x min =x,x max =x+L·cosθ Among them, (x, y) represents the current position of the vehicle; x min The starting horizontal coordinate of the right-of-way area is aligned horizontally with the current position of the vehicle; ω is the width of the vehicle; y min The longitudinal extent of the right-of-way area is determined by the width of the vehicle; θ is the vehicle's driving direction angle; and the stopping distance L satisfies the following formula: Where v is the speed of the vehicle; a max is the maximum deceleration of the vehicle; ρ is the density of vehicles on the road; k ρ is the density adjustment factor.
5. The method according to claim 4, characterized in that The road right violation score SC T Satisfies the following formula: SC T =αT i +β(A_ROW) i Among them, T i The total duration of vehicles in the traffic scene that violate the right of way for other vehicles (A_ROW) i It is the cumulative overlapping area of the right-of-way area violated by other vehicles in the traffic scene; α and β are weighting coefficients.
6. The method according to claim 1, characterized in that Determining at least one driving intention of the autonomous vehicle and a confidence level of each driving intention, including: Based on the global state vector C T The right-of-way violation score SC T Get the joint input vector s joint =[C T ,SC T ]; Based on the formula: π(s joint ;θ)=[p keep ,p left ,p right ], determine the driving intention probability vector π(s joint ; θ); the driving intention probability vector is the confidence level corresponding to each driving intention, wherein the driving intention includes at least one of the following: keeping lane, changing lanes left, changing lanes right; wherein θ is a network parameter; p keep To maintain lane confidence; p left is the confidence level of the left lane change; p right is the confidence level of right lane change.
7. The method according to any one of claims 1 to 6, characterized in that The training process of the deep reinforcement learning model includes: The joint input vector s joint =[C T ,SC T ] is input into the deep reinforcement learning model to obtain driving intention and the confidence level corresponding to each driving intention; Based on the formula: A=argmaxπ(s joint ;θ), determine the driving decision A corresponding to the current maximum confidence; where argmax is π(s joint ; θ) corresponds to the driving intention when it reaches the maximum value; Generate corresponding continuous control parameters for driving decision A, which include at least one of the following: steering angle control and acceleration control; wherein the range of steering angle control is [-0.5, 0.5] rad, which is used to control the lateral displacement of the vehicle; the range of acceleration control is [-5, 5] m / s 2 ; The continuous control parameters can directly act on the autonomous driving vehicle; causing the autonomous vehicle to execute the continuous control parameters; The network parameters θ are updated using a multi-objective adaptation function; the deep reinforcement learning model is based on the driving decision A, the confidence level corresponding to the driving decision A, the current global state vector, and the road right violation score SC T After training, the model returns the decision corresponding to the highest confidence and the reward corresponding to the decision. The reward serves as feedback to adjust the confidence.
8. The method according to claim 7, characterized in that The multi-objective fitness function Fitness satisfies the following formula: Fitness=ω1·f safe +ω2·f eff +ω3·f conf +ω4·f soc Among them, f safe is the safety index; f eff is the traffic efficiency index; f conf is the comfort index; f soc is the social compatibility index; ω1, ω2, ω3, ω4 are the weight coefficients of each index; f safe Satisfies the following formula: Among them, d t is the distance between the autonomous vehicle and the nearest vehicle in front, d min is the safety threshold, t is the time step, and T is the total time it takes to complete the autonomous driving decision execution; and\or, f eff Satisfies the following formula: Among them, v avg is the average speed of the autonomous vehicle, v des The expected speed of the autonomous vehicle; and\or, f conf Satisfies the following formula: Among them, a t is the longitudinal acceleration, ω t is the angular velocity, α and β are weighted coefficients for adjusting the riding experience; t is the time step, and T is the total time it takes for the autonomous driving decision to be executed; and\or, f soc Satisfies the following formula: in, represents the space occupied by the vehicle at time t; It represents the absolute right-of-way area of the jth surrounding vehicle in the traffic scene at that moment; Area() is the area calculation function; t is the time step, and T is the total time to complete the autonomous driving decision execution.
9. The method according to claim 7, characterized in that Reward R t Satisfies the following formula: in, It is a safety reward item; It is an efficiency reward item; It is a comfort bonus item; is the social compatibility reward item; ω1, ω2, ω3, ω4 are the weight coefficients of each sub-item; Satisfies the following formula: Among them, d min Indicates the minimum headway distance between the vehicle and the nearest vehicle in the current time step; ∈ takes 10 -6 ; Satisfies the following formula: Among them, v t is the current speed; v des is the expected speed; Satisfies the following formula: Among them, a t is the linear acceleration, ω t is the angular velocity, α and β are the empirical coefficients for adjusting the weight; Satisfies the following formula: Among them, A overlap is the spatial overlapping area between the absolute right-of-way area of the vehicle and the absolute right-of-way area of the infringed vehicle; γ is the penalty coefficient.
10. A device for generating intelligent decisions for autonomous driving, characterized in that: include: Communication unit and processing unit, The communication unit is used to obtain image information and historical driving data of each vehicle in the traffic scene where the autonomous driving vehicle is located; The processing unit is used to determine the global state vector of the traffic scene based on the image information and historical driving data of each vehicle, and the global state vector is used to represent the dynamic information in the traffic scene; determine the right-of-way area of each vehicle based on the historical driving data of each vehicle in the traffic scene, and the right-of-way area is used to represent the area around the vehicle that needs to be avoided first; determine the right-of-way violation behavior score of each vehicle based on the overlap of the right-of-way areas of each vehicle; input the global state vector and the right-of-way violation behavior score into the deep reinforcement learning model to obtain an output result, and the output result is used to represent at least one driving intention of the autonomous driving vehicle, as well as the confidence of each driving intention; determine the driving intention corresponding to the highest confidence score as the current driving decision.