Vehicle decision-making control model training method and apparatus, vehicle decision-making control method and apparatus, and device

Through the method of stratified reinforcement learning, the feature extraction network and behavioral decision-making control model are used to solve the behavioral decision-making problems of autonomous driving systems in complex traffic scenarios, achieving more efficient and flexible vehicle control, and improving safety and energy efficiency.

WO2025156939A1PCT designated stage Publication Date: 2025-07-31JIUZHI (SUZHOU) INTELLIGENT TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/144152
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-24
Filing Date
2024-12-31
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

The existing autonomous driving behavior decision-making methods have limitations in dealing with complex traffic scenarios and uncertainties, making it difficult to achieve flexible and efficient vehicle behavior control.

Method used

The hierarchical reinforcement learning method is adopted to characterize the vehicle state space through the feature extraction network, and combine the upper-level behavior decision-making sub-model and the lower-level behavior control sub-model to perform hierarchical decision-making and control, reducing dependence on manual rules and improving the adaptability and flexibility of the system.

Benefits of technology

It improves the adaptability and flexibility of autonomous driving systems in complex traffic environments, and enhances the safety, stability and energy efficiency of vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024144152_31072025_PF_FP_ABST
    Figure CN2024144152_31072025_PF_FP_ABST
Patent Text Reader

Abstract

A vehicle decision-making control model training method and apparatus, a vehicle decision-making control method and apparatus, and a device. The vehicle decision-making control model training method comprises: using a feature extraction network to perform feature encoding on a sample vehicle state space of a sample vehicle, so as to obtain a sample state encoding feature (S110); using an upper-level behavior decision-making sub-model to perform behavior decision-making on the sample state encoding feature, so as to obtain an upper-level behavior decision-making prediction result (S120); using a lower-level behavior control sub-model to perform behavior control on the sample state encoding feature, so as to obtain sample vehicle state information of the sample vehicle (S130); according to the sample vehicle state information, a sample obstacle position, road information, a sample vehicle speed limit and the maximum vehicle acceleration, determining a lower-level behavior control single-step loss (S140); according to the upper-level behavior decision-making prediction result, the lower-level behavior control single-step loss and an upper-level behavior decision-making single-step cycle, determining an upper-level behavior decision-making loss (S150); and, according to the upper-level behavior decision-making loss, training a vehicle decision-making control model (S160).
Need to check novelty before this filing date? Find Prior Art

Description

Vehicle decision control model training and vehicle decision control method, device and equipment

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on January 24, 2024, with application number 202410101337.5, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the fields of artificial intelligence technology and autonomous driving technology, for example, to a vehicle decision control model training and vehicle decision control method, device and equipment. Background Art

[0003] Autonomous driving technology has become a hot topic in future transportation research, with behavioral decision-making being one of the core technologies of autonomous driving systems. Behavioral decision-making involves a vehicle's action plan in various traffic scenarios, including whether to change lanes, when to slow down, and when to overtake. Effective behavioral decision-making is crucial to ensuring vehicle safety and passenger comfort.

[0004] There are various technical solutions for behavioral decision-making in the field of autonomous driving. These solutions are often based on traditional rule-based decision-making, such as traffic regulations and vehicle perception information. However, this approach has limitations when dealing with complex traffic scenarios and uncertainties. Therefore, a more flexible and adaptable approach to behavioral decision-making is needed to improve the performance of autonomous driving systems. Summary of the Invention

[0005] The present application provides a vehicle decision control model training and vehicle decision control method, device and equipment to improve the adaptability and flexibility of autonomous driving vehicles.

[0006] An embodiment of the present application provides a training method for a vehicle decision-making and control model, the method comprising: using a feature extraction network to perform feature encoding on a sample vehicle state space of a sample vehicle to obtain a sample state encoding feature; the sample vehicle state space includes sample vehicle information; using an upper-layer behavior decision sub-model to perform behavior decision on the sample state encoding feature to obtain an upper-layer behavior decision prediction result; using a lower-layer behavior control sub-model to perform behavior control on the sample state encoding feature to obtain sample vehicle state information of the sample vehicle; determining a lower-layer behavior control single-step loss based on the sample vehicle state information, sample obstacle position, road information, sample vehicle speed limit, and maximum vehicle acceleration; determining an upper-layer behavior decision loss based on the upper-layer behavior decision prediction result, the lower-layer behavior control single-step loss, and the upper-layer behavior decision single-step period; and training the vehicle decision-making and control model based on the upper-layer behavior decision loss.

[0007] An embodiment of the present application provides a vehicle decision-making control method, which includes: obtaining a target vehicle state space of a target autonomous driving vehicle; the target vehicle state space is represented by a grid graph; the target vehicle state space is input into a vehicle decision-making control model to obtain a target behavior control result of the target autonomous driving vehicle; wherein the vehicle decision-making control model is trained by the vehicle decision-making control model training method described in any embodiment of the present application; and the target autonomous driving vehicle is controlled using the target behavior control result.

[0008] An embodiment of the present application provides a training device for a vehicle decision-making and control model, which includes: a sample state feature determination module, configured to use a feature extraction network to perform feature encoding on a sample vehicle state space of a sample vehicle to obtain a sample state encoding feature; the sample vehicle state space includes sample vehicle information; an upper-level decision prediction module, configured to use an upper-level behavior decision sub-model to perform behavior decision on the sample state encoding feature to obtain an upper-level behavior decision prediction result; a downlink behavior control module, configured to use a lower-level behavior control sub-model to perform behavior control on the sample state encoding feature to obtain the sample vehicle state information of the sample vehicle; a lower-level control loss determination module, configured to determine the lower-level behavior control single-step loss based on the sample vehicle state information, sample obstacle location, road information, sample vehicle speed limit and maximum vehicle acceleration; an upper-level decision loss determination module, configured to determine the upper-level behavior decision loss based on the upper-level behavior decision prediction result, the lower-level behavior control single-step loss and the upper-level behavior decision single-step period; and a model training module, configured to train the vehicle decision-making and control model based on the upper-level behavior decision loss.

[0009] An embodiment of the present application provides a vehicle decision-making control device, which includes: a target state space determination module, configured to obtain a target vehicle state space of a target autonomous driving vehicle; the target vehicle state space is represented by a grid graph; a target control result determination module, configured to input the target vehicle state space into a vehicle decision-making control model to obtain a target behavior control result of the target autonomous driving vehicle; wherein the vehicle decision-making control model is trained by the vehicle decision-making control model training method described in any embodiment of the present application; and a vehicle control module, configured to use the target behavior control result to control the target autonomous driving vehicle.

[0010] An embodiment of the present application provides an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the vehicle decision control model training method or the vehicle decision control method described in any embodiment of the present application.

[0011] An embodiment of the present application provides a computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable a processor to implement the vehicle decision control model training method or vehicle decision control method described in any embodiment of the present application when executed. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] FIG1A is a flowchart of a method for training a vehicle decision control model according to a first embodiment of the present application;

[0013] FIG1B is a schematic diagram of a grid diagram representing a sample vehicle state space according to the first embodiment of the present application;

[0014] FIG1C is a schematic diagram showing a matrix representation of a drivable area according to the first embodiment of the present application;

[0015] FIG2 is a flow chart of a method for training a vehicle decision control model according to a second embodiment of the present application;

[0016] FIG3 is a flow chart of a vehicle decision-making control method provided according to a third embodiment of the present application;

[0017] FIG4 is a schematic structural diagram of a vehicle decision control model training device provided according to a third embodiment of the present application;

[0018] FIG5 is a schematic structural diagram of a vehicle decision control device provided according to a third embodiment of the present application;

[0019] Figure 6 is a structural diagram of an electronic device that implements the vehicle decision control model training method or vehicle decision control method of an embodiment of the present application. DETAILED DESCRIPTION

[0020] The following will describe the technical solutions in the embodiments of this application in conjunction with the drawings in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0021] It should be noted that the terms "first", "second", "sample", "target", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0022] In addition, it should be noted that in the technical solution of this application, the collection, storage, use, processing, transmission, provision and disclosure of relevant data such as the sample vehicle state space and the target vehicle state space are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0023] Example 1

[0024] FIG1A is a flow chart of a method for training a vehicle decision-making control model according to the first embodiment of the present application. This embodiment is applicable to situations where autonomous vehicles make vehicle behavior decisions in complex traffic environments. The method can be performed by a training device for a vehicle decision-making control model. The device can be implemented in the form of hardware and / or software and can be integrated into an electronic device that carries the training function of the vehicle decision-making control model, such as a server. As shown in FIG1A , the method includes the following steps.

[0025] S110 , using a feature extraction network to perform feature encoding on a sample vehicle state space of the sample vehicle to obtain a sample state encoding feature.

[0026] In this embodiment, the sample vehicle refers to the source vehicle of the data used for training the vehicle decision-making control model. The so-called sample vehicle state space refers to the spatial data for the sample vehicle to make behavioral decisions; optionally, the sample vehicle state space includes sample vehicle information, sample obstacle information, and a sample drivable area. The so-called sample vehicle information refers to the relevant information of the sample vehicle; optionally, the sample vehicle information includes the sample vehicle position information, sample vehicle speed, sample vehicle acceleration, and sample vehicle heading angle. The so-called sample obstacle information refers to the obstacle information around the sample vehicle; optionally, the sample obstacle information includes the sample obstacle position, sample obstacle speed, sample obstacle acceleration, and sample obstacle heading angle. The so-called sample drivable area refers to the drivable area of ​​the sample vehicle on the road.

[0027] For example, the sample vehicle state space can be represented by a grid diagram, and then the grid diagram can be represented by a matrix. The sample vehicle state space can be represented by a grid diagram, and the area 50 meters forward, 20 meters backward, and 10 meters left and right of the sample vehicle is defined as the behavior decision range. The space in this range is divided into 0.1*0.1 grids, as shown in Figure 1B. The sample drivable area is represented by a grid-equal-size matrix Sa, which is 700*200 in size. ij 0 / 1 filling is used to indicate whether the position of the i-th row and j-th column in the grid map is a feasible area, where 0 indicates infeasible and 1 indicates feasible, as shown in FIG1C .

[0028] The sample vehicle and obstacle information can also be represented using four matrices of equal size, each representing position, velocity, acceleration, and heading angle. The total matrix size is 4*700*200, denoted as Ss. This can be done by using an occupancy grid approach: the corresponding matrix positions for occupied grids are filled with 1s or 0s; the corresponding matrix positions for velocity, acceleration, and heading angle are filled with corresponding values. In summary, the sample vehicle state space can be represented using a 5*700*200 matrix.

[0029] The so-called feature extraction network refers to the network used to extract features from the vehicle state space; optionally, the feature extraction network can be composed of a deformable attention mechanism and a convolutional neural network. It is understood that the use of a deformable attention mechanism allows the model to dynamically adjust its attention based on the driving situation, better adapting to different traffic scenarios, such as urban roads, highways, or complex intersections. At the same time, the use of a convolutional neural network structure can capture rich spatial and temporal features, helping the model better understand the environment around the vehicle, such as road conditions, vehicle position and speed, and obstacle distribution.

[0030] The so-called sample state encoding feature refers to the feature obtained after encoding the sample vehicle state space, which can be expressed in the form of a matrix or vector.

[0031] The matrix representation of the sample vehicle state space of the sample vehicle is input into the feature extraction network, and the network performs feature extraction to obtain the sample state encoding feature.

[0032] S120. Use the upper-level behavior decision sub-model to make a behavior decision on the sample state encoding features to obtain an upper-level behavior decision prediction result.

[0033] In this embodiment, the upper-level behavior decision submodel is used to make vehicle behavior decisions, namely, whether the vehicle needs to perform high-level actions such as lane changes. The optional upper-level behavior decision submodel is composed of a deep Q-network. The upper-level behavior decision prediction results include high-level actions such as maintaining the current lane and lane changes, where lane changes include left or right lane changes. For example, a one-hot type output can be used to represent the behavior decision results.

[0034] The sample state encoding features can be input into the upper-level behavior decision sub-model, and the upper-level behavior decision prediction results can be obtained through reinforcement learning of the model.

[0035] S130 , using the lower-layer behavior control sub-model to perform behavior control on the sample state coding features to obtain sample vehicle state information of the sample vehicle.

[0036] In this embodiment, the lower-level behavior control submodel is used to output the vehicle's motion state, including acceleration and curvature. Optionally, the lower-level behavior control submodel is composed of a deep Q-network. It should be noted that the upper-level behavior decision submodel and the lower-level behavior control submodel have the same structure but different parameters.

[0037] The so-called sample vehicle state information refers to relevant information such as the position, speed, acceleration, heading angle and curvature of the sample vehicle in each action state; optionally, the sample vehicle state information includes the sample vehicle state speed, the sample vehicle state acceleration, the sample vehicle state curvature, the sample vehicle state heading angle and the sample vehicle state position; wherein, the sample vehicle state acceleration refers to the acceleration of the sample vehicle in each future action state (for example, the state every second); the sample vehicle state speed refers to the speed of the sample vehicle in each future action state (for example, the state every second); the sample vehicle state curvature refers to the curvature of the sample vehicle in each future action state (for example, the state every second); the sample vehicle state heading angle refers to the heading angle of the sample vehicle in each future action state (for example, the state every second); the sample vehicle state position refers to the position of the sample vehicle in each future action state (for example, the state every second).

[0038] The sample state encoding features can be input into the lower-level behavior control sub-model, and after model learning, the sample vehicle state information of the sample vehicle can be obtained.

[0039] An optional method is to use a lower-level behavior control sub-model to perform behavior control on the sample state encoding features to obtain the sample vehicle state acceleration and sample vehicle state curvature in the sample vehicle state information of the sample vehicle; based on the sample vehicle state acceleration and the sample vehicle state curvature, as well as the sample vehicle information, determine the sample vehicle state speed, sample vehicle state heading angle and sample vehicle state position in the sample vehicle state information.

[0040] The sample state encoding features can be input into the lower-level behavior control sub-model. After model learning processing, the sample vehicle state acceleration and sample vehicle state curvature in the sample vehicle state information of the sample vehicle can be obtained. In this way, the sample vehicle state acceleration and sample vehicle state curvature of the sample vehicle in different states can be obtained. Then, based on the trajectory inversion formula, the sample vehicle state acceleration and sample vehicle state curvature of different states (such as adjacent states) and the initial sample vehicle information (sample vehicle position, sample vehicle speed, sample vehicle acceleration, sample vehicle heading angle and sample vehicle curvature) can be used to determine the sample vehicle state speed, sample vehicle state heading angle and sample vehicle state position. The trajectory inversion formula is as follows: k =v k-1 +0.5*(a k +a k-1 )*Δt ω k =0.5*(kappa k +kappa k-1 )*0.5(v k +v k-1 ) θ k =θ k-1 +ω k (x k ,y k )=(x k-1 ,y k-1 )+0.5(v k +v k-1 )(sin(0.5(θ k +θ k-1 )),cos(0.5(θ k +θ k-1 )))

[0041] Among them, k represents the kth state (for example, the kth moment), and k is a natural number greater than 1; v represents the sample vehicle state speed; a represents the sample vehicle state acceleration; Δt represents the time difference between adjacent states; ω represents the heading angle change; kappa represents the sample vehicle state curvature; θ represents the sample vehicle state heading angle; (x, y) represents the sample vehicle state position.

[0042] In order to ensure the feasibility of the vehicle trajectory and meet the kinematic constraints, after obtaining the sample vehicle state acceleration and sample vehicle state curvature output by the model, a tanh operation is performed on the sample vehicle state acceleration and sample vehicle state curvature. That is, a layer of tanh function (with coefficients and offset) is added after the output layer of the lower-level behavior control sub-model to strictly limit the sample vehicle state acceleration and sample vehicle state curvature to between the maximum acceleration and deceleration and the maximum positive and negative curvature.

[0043] In order to ensure that the vehicle speed is within a reasonable range, the sample vehicle state speed that is greater than the vehicle's maximum speed limit is truncated, that is, the vehicle's maximum speed limit is used as the sample vehicle state speed.

[0044] S140 , determining a lower-level behavior control single-step loss based on the sample vehicle state information, the sample obstacle location, the road information, the sample vehicle speed limit, and the maximum vehicle acceleration.

[0045] In this embodiment, the sample vehicle speed limit refers to the maximum permissible speed of the sample vehicle. The maximum vehicle acceleration refers to the maximum permissible acceleration of the sample vehicle. The lower-level behavior control single-step loss refers to the training loss of the lower-level behavior control submodel. Optionally, the lower-level behavior control single-step loss includes lane keeping loss and lane changing loss. The lane keeping loss refers to the loss corresponding to the lane-keeping purpose, and the lane changing loss refers to the loss corresponding to the lane changing purpose.

[0046] Alternatively, a preset loss function can be used to determine the lower-level behavior control single-step loss based on sample vehicle state information, sample obstacle locations, road information, sample vehicle speed limits, and maximum vehicle acceleration. It should be noted that this embodiment does not limit the preset loss function.

[0047] S150 , determining the upper-level behavior decision loss according to the upper-level behavior decision prediction result, the lower-level behavior control single-step loss, and the upper-level behavior decision single-step cycle.

[0048] In this embodiment, the upper-level behavior decision single-step cycle refers to the duration of the current behavior decision, in seconds. To prevent the upper-level behavior decision sub-model from diverging and the invalidity of overly planned behaviors, the upper-level behavior decision single-step cycle has a maximum value of 6 seconds.

[0049] The so-called upper-level behavior decision loss refers to the loss used to train the upper-level behavior decision sub-model, which is used to evaluate the vehicle's actions. The purpose is to enable the vehicle to take safe and efficient actions in different situations. For example, safe behavior will receive positive rewards, while dangerous behavior will receive negative rewards.

[0050] An optional method is to determine the upper-layer associated lower-layer single-step loss from the lower-layer behavior control single-step loss based on the upper-layer behavior decision prediction result; and determine the upper-layer behavior decision loss based on the upper-layer associated lower-layer single-step loss and the upper-layer behavior decision single-step period.

[0051] The upper-layer associated lower-layer single-step loss refers to the lower-layer behavior control loss corresponding to the upper-layer behavior decision.

[0052] If the upper-layer behavior decision prediction result is to keep the current lane, the upper-layer associated lower-layer single-step loss is the lane keeping loss; if the upper-layer behavior decision prediction result is to change lanes, the upper-layer associated lower-layer single-step loss is the lane changing loss. The upper-layer behavior decision loss can then be determined based on the upper-layer associated lower-layer single-step loss and the upper-layer behavior decision single-step period using the following formula:

[0053] Among them, Ru represents the upper-level behavior decision loss; Δd represents the upper-level behavior decision single-step cycle; Rl represents the upper-level associated lower-level single-step loss.

[0054] It is understandable that compared to ordinary reinforcement learning, the cleverness of hierarchical reinforcement learning lies in the fact that the reward (loss) of the upper-level decision is evaluated by the entire lower-level control process.

[0055] S160: Train the vehicle decision control model based on the upper-layer behavior decision loss.

[0056] In this embodiment, the vehicle decision control model is used to control the driving of the autonomous driving vehicle; optionally, the vehicle decision control model includes a feature extraction network, an upper-layer behavior decision sub-model and a lower-layer behavior control sub-model.

[0057] According to the upper-level behavior decision loss, the vehicle decision control model is trained until the upper-level behavior decision loss tends to be stable and the training is stopped.

[0058] The technical solution of the embodiment of the present application adopts a feature extraction network to perform feature encoding on the sample vehicle state space of the sample vehicle to obtain the sample state encoding feature; the sample vehicle state space includes sample vehicle information, and then adopts the upper-level behavior decision sub-model to perform behavior decision on the sample state encoding feature to obtain the upper-level behavior decision prediction result, and adopts the lower-level behavior control sub-model to perform behavior control on the sample state encoding feature to obtain the sample vehicle state information of the sample vehicle, and then determines the lower-level behavior control single-step loss based on the sample vehicle state information, the sample obstacle position, the road information, the sample vehicle speed limit and the maximum vehicle acceleration, and determines the upper-level behavior decision loss based on the upper-level behavior decision prediction result, the lower-level behavior control single-step loss and the upper-level behavior decision single-step period, and finally trains the vehicle decision control model based on the upper-level behavior decision loss. The above technical solution realizes vehicle behavior control through hierarchical reinforcement learning, i.e., adopting the upper-level behavior decision sub-model and the lower-level behavior control sub-model, provides a framework that can make autonomous decisions and reduces the dependence on artificial rules; at the same time, it improves the adaptability and flexibility of the autonomous driving system in complex traffic environments, thereby improving the safety, stability and energy efficiency of the vehicle.

[0059] Example 2

[0060] Figure 2 is a flow chart of a vehicle decision-making control model training method according to Example 2 of this application. Building on the previous example, this example illustrates "determining the single-step loss of underlying behavioral control based on sample vehicle state information, sample obstacle locations, road information, sample vehicle speed limits, and maximum vehicle acceleration," and provides an alternative implementation. As shown in Figure 2, the method includes the following steps.

[0061] S210: Using a feature extraction network, feature encoding is performed on the sample vehicle state space of the sample vehicle to obtain a sample state encoding feature.

[0062] The sample vehicle state space includes sample vehicle information.

[0063] S220: Using the upper-level behavior decision sub-model, perform behavior decision on the sample state encoding features to obtain an upper-level behavior decision prediction result.

[0064] S230: Use the lower-layer behavior control sub-model to perform behavior control on the sample state coding features to obtain sample vehicle state information of the sample vehicle.

[0065] S240: Determine a lower-level behavior control single-step loss based on the sample vehicle state information, the sample obstacle location, the road information, the sample vehicle speed limit, and the maximum vehicle acceleration.

[0066] An optional approach can determine lane keeping loss based on sample vehicle state information, sample obstacle locations, road boundaries and speed limits in road information, sample vehicle speed limits, and maximum vehicle acceleration; determine lane changing loss based on sample vehicle state positions and lane centerline positions in road information; and determine the lower-level behavioral control single-step loss based on the lane keeping loss and lane changing loss. Lane centerline position refers to the lateral position of the lane centerline.

[0067] Lane keeping loss can be determined based on a preset loss function, sample vehicle state information, sample obstacle locations, road boundaries and speed limits in the road information, sample vehicle speed limits, and maximum vehicle acceleration. It should be noted that this embodiment does not limit the preset loss function. Lane changing loss can then be determined based on the sample vehicle state location and lane centerline location in the road information using the following formula: Where le represents the current lateral position of the vehicle; c represents the centerline position of the lane in the road information; R k Denotes the lane changing loss. The lane keeping loss and lane changing loss are used as the underlying behavior to control the single-step loss.

[0068] Exemplarily, lane keeping loss is determined based on sample vehicle state information, sample obstacle positions, road boundaries and road speed limits in road information, sample vehicle speed limits, and maximum vehicle acceleration, including: determining obstacle loss based on the sample vehicle state position and the sample obstacle position; wherein the obstacle loss includes static obstacle loss and / or dynamic obstacle loss; determining boundary loss based on the sample vehicle state position and the road boundary; determining efficiency loss based on the sample vehicle state speed, the road speed limit, and the sample vehicle speed limit; determining smoothness loss based on the sample vehicle state acceleration and the maximum vehicle acceleration; determining emission loss based on the sample vehicle state speed and the sample vehicle state acceleration based on an energy consumption evaluation model; and determining lane keeping loss based on the obstacle loss, boundary loss, efficiency loss, smoothness loss, and emission loss.

[0069] First, obstacle loss is determined based on the sample vehicle state positions and sample obstacle positions. Obstacle loss includes static obstacle loss and / or dynamic obstacle loss. For static obstacle loss, the distance between the sample vehicle state position and the sample obstacle position is determined. For example, the minimum distance between the sample obstacle polygon and the sample vehicle's bounding box can be used as the distance between the sample vehicle state position and the sample obstacle position. Obstacle loss can then be determined based on the following formula:

[0070] Among them, R s represents the static obstacle loss; k s represents the static loss coefficient, which can be fine-tuned manually; dist represents the distance between the sample vehicle state position and the sample obstacle position.

[0071] Second, for dynamic obstacle loss, the determination process is similar to that of static obstacle loss. First, the sample vehicle state position at each future moment, that is, each state, is calculated, and a bounding box is generated by expanding the sample vehicle state position. The sample obstacle polygon at each moment is determined accordingly, and the distance between the sample vehicle state position and the sample obstacle at each moment is calculated. The dynamic obstacle loss is then determined based on the distance. For example, the distance between the sample vehicle state position and the sample obstacle corresponding to 2 meters in the future can be used to determine the dynamic obstacle loss. Due to the uncertainty of future moment predictions, a discount factor needs to be set, such as 0.9. That is, Among them, R d Indicates dynamic obstacle loss.

[0072] Third, the boundary loss is determined based on the sample vehicle's position and the road boundary. The boundary loss refers to the cost of approaching the road boundary. Because crossing lanes during detours increases risk, an additional penalty is imposed for approaching lane boundaries. When the vehicle does not cross the road boundary, the boundary loss is zero. When the vehicle crosses the road boundary, the distance between the sample vehicle's position and the road boundary is determined. The boundary loss can then be determined based on the sample vehicle's position and the road boundary using the following formula:

[0073] Among them, R b represents the boundary loss; k b It represents the boundary coefficient, which can be fine-tuned manually; boundary_dist represents the distance between the sample vehicle state position and the road boundary.

[0074] Fourth, determine the efficiency loss based on the sample vehicle state speed, the road speed limit, and the sample vehicle speed limit. Efficiency loss refers to the efficiency used to incentivize vehicle travel, hoping that the vehicle can travel a longer distance in the same amount of time. The maximum speed that the vehicle can travel is determined from the road speed limit and the sample vehicle speed limit, that is, the maximum speed of the vehicle itself. The ratio between the sample vehicle state speed and the maximum speed of the vehicle itself is then used as the efficiency loss. For example, it can be determined using the following formula:

[0075] Among them, R e represents efficiency loss; k e Indicates the efficiency loss coefficient, which can be fine-tuned manually; v ego Indicates the state speed of the sample vehicle; v l_mam Indicates road speed limit; v e_max Indicates the speed limit of the sample vehicle.

[0076] Fifth, determine the smoothness loss based on the sample vehicle state acceleration and the maximum vehicle acceleration. The smoothness loss is used to ensure vehicle driving smoothness. Determine the absolute value of the vehicle acceleration change value of the sample vehicle state acceleration in adjacent states, and determine the maximum vehicle acceleration of the sample vehicle on the road. The ratio between the absolute value of the vehicle acceleration change value and the maximum vehicle acceleration is then used as the smoothness loss. For example, this can be determined using the following formula:

[0077] Among them, R c represents steady loss; k c Indicates the steady loss coefficient, which can be fine-tuned manually; represents the sample vehicle state acceleration at time t (state); represents the sample vehicle state acceleration at time t-1 (state); a max_accIndicates the maximum acceleration allowed for the vehicle; a max_dcc Indicates the maximum acceleration allowed by the road.

[0078] Sixth, based on the energy consumption assessment model, determine the emission loss based on the sample vehicle state speed and sample vehicle state acceleration. The energy consumption assessment model can be VT-Micro 2. Emission loss refers to the loss used to assess vehicle energy consumption. Based on the sample vehicle state speed and sample vehicle state acceleration, the energy consumption table in the energy consumption assessment model can be used to obtain the vehicle energy consumption. Furthermore, based on the vehicle energy consumption, the emission loss can be determined. For example, it can be determined using the following formula:

[0079] Among them, R t represents the emission loss, j2 represents the energy consumption table in the energy consumption evaluation model; v represents the sample vehicle state speed; a represents the sample vehicle state acceleration.

[0080] Finally, the lane keeping loss is determined based on the obstacle loss, boundary loss, efficiency loss, smoothness loss, and emission loss. For example, the lane keeping loss R can be determined by the following formula: lk : R lk =-R s -R d -R b +R e -R c -R t

[0081] S250: Determine the upper-level behavior decision loss according to the upper-level behavior decision prediction result, the lower-level behavior control single-step loss, and the upper-level behavior decision single-step cycle.

[0082] S260: Train the vehicle decision control model based on the upper-layer behavior decision loss.

[0083] The technical solution of the embodiment of the present application adopts a feature extraction network to perform feature encoding on the sample vehicle state space of the sample vehicle to obtain the sample state encoding feature; the sample vehicle state space includes sample vehicle information, and then adopts the upper-level behavior decision sub-model to perform behavior decision on the sample state encoding feature to obtain the upper-level behavior decision prediction result, and adopts the lower-level behavior control sub-model to perform behavior control on the sample state encoding feature to obtain the sample vehicle state information of the sample vehicle, and then determines the lower-level behavior control single-step loss based on the sample vehicle state information, the sample obstacle position, the road information, the sample vehicle speed limit and the maximum vehicle acceleration, and determines the upper-level behavior decision loss based on the upper-level behavior decision prediction result, the lower-level behavior control single-step loss and the upper-level behavior decision single-step period, and finally trains the vehicle decision control model based on the upper-level behavior decision loss. The above technical solution realizes vehicle behavior control through hierarchical reinforcement learning, i.e., adopting the upper-level behavior decision sub-model and the lower-level behavior control sub-model, provides a framework that can make autonomous decisions and reduces the dependence on artificial rules; at the same time, it improves the adaptability and flexibility of the autonomous driving system in complex traffic environments, thereby improving the safety, stability and energy efficiency of the vehicle.

[0084] Example 3

[0085] Figure 3 is a flow chart of a vehicle decision-making control method provided in accordance with Example 3 of the present application. This embodiment is applicable to situations where autonomous vehicles make vehicle behavior decisions in complex traffic environments. The method can be executed by a vehicle decision-making control device, which can be implemented in hardware and / or software and integrated into an electronic device that carries the vehicle decision-making control function, such as an autonomous vehicle. As shown in Figure 3, the method includes the following steps.

[0086] S310: Obtain a target vehicle state space of a target autonomous driving vehicle.

[0087] In this embodiment, the target autonomous vehicle refers to an autonomous vehicle that requires real-time vehicle control. The so-called target vehicle state space refers to the spatial data for the target autonomous vehicle to perform behavior control; optionally, the target vehicle state space includes target vehicle information, target obstacle information, and target drivable area. The so-called target vehicle information refers to relevant information of the target autonomous vehicle; optionally, the target vehicle information includes target vehicle position information, target vehicle speed, target vehicle acceleration, and target vehicle heading angle. The so-called target obstacle information refers to obstacle information around the target autonomous vehicle; optionally, the target obstacle information includes target obstacle position, target obstacle speed, target obstacle acceleration, and target obstacle heading angle. The target vehicle state space is represented by a grid graph, and the grid graph can be represented by a matrix.

[0088] The target vehicle state space of the target autonomous driving vehicle can be obtained in real time.

[0089] S320: Input the target vehicle state space into the vehicle decision control model to obtain the target behavior control result of the target autonomous driving vehicle.

[0090] The vehicle decision control model is trained using the vehicle decision control model training method provided in any embodiment of the present application. The so-called target behavior control result refers to the vehicle control instructions of the autonomous driving vehicle, including maintaining the current lane or changing lanes.

[0091] The target vehicle state space can be input into the vehicle decision control model, and after model processing, the target behavior control result of the target autonomous driving vehicle can be obtained.

[0092] S330: Control the target autonomous driving vehicle using the target behavior control result.

[0093] The target behavior control results can be used to control the target autonomous driving vehicle.

[0094] The technical solution provided by the embodiments of this application obtains the target vehicle state space of the target autonomous vehicle; the target vehicle state space is represented by a grid graph, and then the target vehicle state space is input into the vehicle decision-making control model to obtain the target behavior control result of the target autonomous vehicle. The target behavior control result is then used to control the target autonomous vehicle. This technical solution can improve the adaptability and flexibility of autonomous vehicle control.

[0095] Example 4

[0096] Figure 4 is a schematic diagram of the structure of a vehicle decision-making and control model training device provided in accordance with Example 3 of the present application. This embodiment is applicable to situations where autonomous vehicles make vehicle behavior decisions in complex traffic environments. The device can be implemented in hardware and / or software and can be integrated into an electronic device that carries the vehicle decision-making and control model training function, such as a server. As shown in Figure 4, the device includes: a sample state feature determination module 410, which is configured to use a feature extraction network to feature encode the sample vehicle state space of the sample vehicle to obtain a sample state encoding feature; the sample vehicle state space includes sample vehicle information; an upper-level decision prediction module 420, which is configured to use an upper-level behavior decision sub-model to perform behavior decision on the sample state encoding feature to obtain an upper-level behavior decision prediction result; a downlink behavior control module 430, which is configured to use a lower-level behavior control sub-model to perform behavior control on the sample state encoding feature to obtain the sample vehicle state information of the sample vehicle; a lower-level control loss determination module 440, which is configured to determine the lower-level behavior control single-step loss based on the sample vehicle state information, the sample obstacle position, the road information, the sample vehicle speed limit and the maximum vehicle acceleration; an upper-level decision loss determination module 450, which is configured to determine the upper-level behavior decision loss based on the upper-level behavior decision prediction result, the lower-level behavior control single-step loss and the upper-level behavior decision single-step period; and a model training module 460, which is configured to train the vehicle decision control model based on the upper-level behavior decision loss.

[0097] The technical solution of the embodiment of the present application adopts a feature extraction network to perform feature encoding on the sample vehicle state space of the sample vehicle to obtain the sample state encoding feature; the sample vehicle state space includes sample vehicle information, and then adopts the upper-level behavior decision sub-model to perform behavior decision on the sample state encoding feature to obtain the upper-level behavior decision prediction result, and adopts the lower-level behavior control sub-model to perform behavior control on the sample state encoding feature to obtain the sample vehicle state information of the sample vehicle, and then determines the lower-level behavior control single-step loss based on the sample vehicle state information, the sample obstacle position, the road information, the sample vehicle speed limit and the maximum vehicle acceleration, and determines the upper-level behavior decision loss based on the upper-level behavior decision prediction result, the lower-level behavior control single-step loss and the upper-level behavior decision single-step period, and finally trains the vehicle decision control model based on the upper-level behavior decision loss. The above technical solution realizes vehicle behavior control through hierarchical reinforcement learning, i.e., adopting the upper-level behavior decision sub-model and the lower-level behavior control sub-model, provides a framework that can make autonomous decisions and reduces the dependence on artificial rules; at the same time, it improves the adaptability and flexibility of the autonomous driving system in complex traffic environments, thereby improving the safety, stability and energy efficiency of the vehicle.

[0098] Optionally, the downlink behavior control module 430 is configured to: use the lower-level behavior control sub-model to perform behavior control on the sample state coding features to obtain the sample vehicle state acceleration and the sample vehicle state curvature in the sample vehicle state information of the sample vehicle; determine the sample vehicle state speed, the sample vehicle state heading angle and the sample vehicle state position in the sample vehicle state information based on the sample vehicle state acceleration and the sample vehicle state curvature, as well as the sample vehicle information.

[0099] Optionally, the lower-layer control loss determination module 440 includes: a lane keeping loss determination unit, configured to determine the lane keeping loss based on sample vehicle state information, sample obstacle positions, road boundaries and road speed limits in road information, sample vehicle speed limits, and maximum vehicle acceleration; a lane changing loss determination unit, configured to determine the lane changing loss based on the sample vehicle state position and the lane centerline position in the road information; and a lower-layer control loss determination unit, configured to determine the lower-layer behavior control single-step loss based on the lane keeping loss and the lane changing loss.

[0100] Optionally, the lane keeping loss determination unit is configured to: determine the obstacle loss based on the sample vehicle state position and the sample obstacle position; wherein the obstacle loss includes static obstacle loss and / or dynamic obstacle loss; determine the boundary loss based on the sample vehicle state position and the road boundary; determine the efficiency loss based on the sample vehicle state speed, the road speed limit and the sample vehicle speed limit; determine the smoothness loss based on the sample vehicle state acceleration and the maximum vehicle acceleration; determine the emission loss based on the sample vehicle state speed and the sample vehicle state acceleration based on the energy consumption evaluation model; determine the lane keeping loss based on the obstacle loss, boundary loss, efficiency loss, smoothness loss and emission loss.

[0101] Optionally, the upper-layer decision loss determination module 450 is configured to: determine the upper-layer associated lower-layer single-step loss from the lower-layer behavior control single-step loss based on the upper-layer behavior decision prediction result; and determine the upper-layer behavior decision loss based on the upper-layer associated lower-layer single-step loss and the upper-layer behavior decision single-step cycle.

[0102] Optionally, the sample vehicle state space also includes sample obstacle information and a sample drivable area; wherein the sample vehicle information includes sample vehicle position information, sample vehicle speed, sample vehicle acceleration and sample vehicle heading angle; the sample obstacle information includes sample obstacle position, sample obstacle speed, sample obstacle acceleration and sample obstacle heading angle; the sample vehicle state space is represented by a grid graph.

[0103] Optionally, the upper-layer behavior decision sub-model and the lower-layer behavior control sub-model have the same structure but different parameters; the upper-layer behavior decision sub-model is composed of a deep Q network.

[0104] The vehicle decision-making and control model training device provided in the embodiments of the present application can execute the vehicle decision-making and control model training method provided in any embodiment of the present application, and has functional modules corresponding to the execution method.

[0105] Example 5

[0106] FIG5 is a schematic diagram of the structure of a vehicle decision control device provided according to the third embodiment of the present application. This embodiment can be applied to situations where autonomous vehicles make vehicle behavior decisions in complex traffic environments. The device can be implemented in the form of hardware and / or software and can be integrated into electronic devices that carry vehicle decision control functions, such as autonomous vehicles. As shown in FIG5 , the device includes: a target state space determination module 510, which is configured to obtain a target vehicle state space of a target autonomous vehicle; the target vehicle state space is represented by a grid graph; a target control result determination module 520, which is configured to input the target vehicle state space into a vehicle decision control model to obtain a target behavior control result of the target autonomous vehicle; wherein the vehicle decision control model is trained by the vehicle decision control model training method provided in any embodiment of the present application; and a vehicle control module 530, which is configured to control the target autonomous vehicle using the target behavior control result.

[0107] The technical solution provided by the embodiments of this application obtains the target vehicle state space of the target autonomous vehicle; the target vehicle state space is represented by a grid graph, and then the target vehicle state space is input into the vehicle decision-making control model to obtain the target behavior control result of the target autonomous vehicle. The target behavior control result is then used to control the target autonomous vehicle. This technical solution can improve the adaptability and flexibility of autonomous vehicle control.

[0108] Optionally, the target vehicle state space includes target vehicle information, target obstacle information and target drivable area; wherein, the target vehicle information includes target vehicle position information, target vehicle speed, target vehicle acceleration and target vehicle heading angle; the target obstacle information includes target obstacle position, target obstacle speed, target obstacle acceleration and target obstacle heading angle.

[0109] The vehicle decision control device provided in the embodiments of the present application can execute the vehicle decision control method provided in any embodiment of the present application, and has functional modules corresponding to the execution method.

[0110] Example 6

[0111] Figure 6 is a structural diagram of an electronic device that implements a training method for a vehicle decision control model or a vehicle decision control method according to an embodiment of the present application; Figure 6 shows a structural diagram of an electronic device 10 that can be used to implement an embodiment of the present application. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or required herein.

[0112] As shown in FIG6 , the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 and a random access memory (RAM) 13, that is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor, and the processor 11 can perform various appropriate actions and processes according to the computer program stored in the ROM 12 or the computer program loaded from the storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0113] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0114] The processor 11 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors that run machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the vehicle decision control model training method or the vehicle decision control method.

[0115] In some embodiments, the training method of the vehicle decision control model or the vehicle decision control method may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the training method of the vehicle decision control model or the vehicle decision control method described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to execute the training method of the vehicle decision control model or the vehicle decision control method in any other appropriate manner (for example, by means of firmware).

[0116] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard parts (ASSPs), system-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0117] Computer programs for implementing the methods of the present application may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0118] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. A computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM) or a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0119] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a cathode ray tube (CRT) or a liquid crystal display (LCD) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0120] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0121] A computing system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. The client-server relationship arises through computer programs running on the respective computers and establishing a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, a host product within a cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosts and virtual private server (VPS) services.

[0122] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the multiple steps described in this application can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of this application can be achieved. This is not limited herein.

Claims

1. A training method for a vehicle decision control model, comprising: Using a feature extraction network to perform feature encoding on the sample vehicle state space of a sample vehicle to obtain sample state encoding features; The sample vehicle state space includes sample vehicle information; Using an upper-layer behavior decision sub-model to perform behavior decision on the sample state encoding features to obtain an upper-layer behavior decision prediction result; Using a lower-layer behavior control sub-model to perform behavior control on the sample state encoding features to obtain the sample vehicle state information of the sample vehicle; According to the sample vehicle state information, sample obstacle position, road information, sample vehicle speed limit, and maximum vehicle acceleration, determining the single-step loss of the lower-layer behavior control; According to the upper-layer behavior decision prediction result, the single-step loss of the lower-layer behavior control, and the single-step period of the upper-layer behavior decision, determining the loss of the upper-layer behavior decision; Training the vehicle decision control model according to the loss of the upper-layer behavior decision.

2. The method according to claim 1, wherein, The step of using a lower-layer behavior control sub-model to perform behavior control on the sample state encoding features to obtain the sample vehicle state information of the sample vehicle includes: Using a lower-layer behavior control sub-model to perform behavior control on the sample state encoding features to obtain the sample vehicle state acceleration and sample vehicle state curvature in the sample vehicle state information of the sample vehicle; According to the sample vehicle state acceleration, the sample vehicle state curvature, and the sample vehicle information, determining the sample vehicle state speed, sample vehicle state heading angle, and sample vehicle state position in the sample vehicle state information.

3. The method according to claim 2, wherein The step of determining the single-step loss of the lower-layer behavior control according to the sample vehicle state information, sample obstacle position, road information, sample vehicle speed limit, and maximum vehicle acceleration includes: Determining a lane keeping loss according to the sample vehicle state information, sample obstacle position, road boundary and road speed limit in the road information, sample vehicle speed limit, and maximum vehicle acceleration; Determining a lane change loss according to the sample vehicle state position and the lane centerline position in the road information; Determining the single-step loss of the lower-layer behavior control according to the lane keeping loss and the lane change loss.

4. The method according to claim 3, wherein The step of determining the lane keeping loss according to the sample vehicle state information, sample obstacle position, road boundary and road speed limit in the road information, sample vehicle speed limit, and maximum vehicle acceleration includes: Determining an obstacle loss according to the sample vehicle state position and the sample obstacle position; wherein, the obstacle loss includes at least one of a static obstacle loss and a dynamic obstacle loss; Determining a boundary loss according to the sample vehicle state position and the road boundary; Determining an efficiency loss according to the sample vehicle state speed, road speed limit, and sample vehicle speed limit; Determining a smoothness loss according to the sample vehicle state acceleration and the maximum vehicle acceleration; Based on an energy consumption evaluation model, determining an emission loss according to the sample vehicle state speed and the sample vehicle state acceleration; Determining the lane keeping loss according to the obstacle loss, the boundary loss, the efficiency loss, the smoothness loss, and the emission loss.

5. The method according to claim 1, wherein, Determining the upper-layer behavior decision loss according to the upper-layer behavior decision prediction result, the single-step loss of the lower-layer behavior control, and the single-step period of the upper-layer behavior decision includes: Determining the upper-layer associated lower-layer single-step loss from the single-step loss of the lower-layer behavior control according to the upper-layer behavior decision prediction result; Determining the upper-layer behavior decision loss according to the upper-layer associated lower-layer single-step loss and the single-step period of the upper-layer behavior decision.

6. The method according to any one of claims 1-5, wherein The sample vehicle state space further includes sample obstacle information and a sample drivable area; wherein, the sample vehicle information includes sample vehicle position information, sample vehicle speed, sample vehicle acceleration, and sample vehicle heading angle; the sample obstacle information includes sample obstacle position, sample obstacle speed, sample obstacle acceleration, and sample obstacle heading angle; the sample vehicle state space is represented by a grid map.

7. The method according to any one of claims 1-5, wherein, The upper-layer behavior decision sub-model and the lower-layer behavior control sub-model have the same structure but different parameters; the upper-layer behavior decision sub-model is composed of a deep Q network.

8. A vehicle decision control method includes: Obtaining the target vehicle state space of the target autonomous vehicle; The target vehicle state space is represented by a grid map; Inputting the target vehicle state space into the vehicle decision control model to obtain the target behavior control result of the target autonomous vehicle; wherein, the vehicle decision control model is trained by the training method of the vehicle decision control model according to any one of claims 1-7; Controlling the target autonomous vehicle by using the target behavior control result.

9. The method according to claim 8, wherein The target vehicle state space includes target vehicle information, target obstacle information, and a target drivable area; wherein, the target vehicle information includes target vehicle position information, target vehicle speed, target vehicle acceleration, and target vehicle heading angle; the target obstacle information includes target obstacle position, target obstacle speed, target obstacle acceleration, and target obstacle heading angle.

10. A training device for a vehicle decision control model includes: A sample state feature determination module configured to perform feature encoding on the sample vehicle state space of the sample vehicle by using a feature extraction network to obtain sample state encoding features; The sample vehicle state space includes sample vehicle information; An upper-layer decision prediction module configured to perform behavior decision on the sample state encoding features by using an upper-layer behavior decision sub-model to obtain an upper-layer behavior decision prediction result; A downlink behavior control module configured to perform behavior control on the sample state encoding features by using a lower-layer behavior control sub-model to obtain the sample vehicle state information of the sample vehicle; A lower-layer control loss determination module configured to determine the single-step loss of the lower-layer behavior control according to the sample vehicle state information, the sample obstacle position, the road information, the sample vehicle speed limit, and the maximum vehicle acceleration; An upper-layer decision loss determination module configured to determine the upper-layer behavior decision loss according to the upper-layer behavior decision prediction result, the single-step loss of the lower-layer behavior control, and the single-step period of the upper-layer behavior decision; A model training module, configured to train a vehicle decision control model according to the upper-level behavior decision loss.

11. A vehicle decision control device, comprising: A target state space determination module, configured to obtain the target vehicle state space of a target autonomous vehicle; The target vehicle state space is represented by a grid map; A target control result determination module, configured to input the target vehicle state space into a vehicle decision control model to obtain a target behavior control result of the target autonomous vehicle; wherein, the vehicle decision control model is trained by the training method of the vehicle decision control model according to any one of claims 1-7. A vehicle control module, configured to control the target autonomous vehicle by using the target behavior control result.

12. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the training method of the vehicle decision control model according to any one of claims 1-7, or the vehicle decision control method according to any one of claims 8-9.

13. A computer-readable storage medium, storing computer instructions, the computer instructions being used to implement the training method of the vehicle decision control model according to any one of claims 1-7, or the vehicle decision control method according to any one of claims 8-9 when executed by a processor.

Citation Information

Patent Citations

  • Automatic driving decision-making method and device, vehicle and storage medium

    CN115140091A

  • Decision-making method for safe driving of large commercial vehicle in urban low-speed environment

    CN115257819A

  • Model training method for lane changing decision, and target lane determination method and device

    CN115743168A

  • Intelligent automobile lane changing decision-making method and device, electronic equipment and storage medium

    CN115782880A

  • Vehicle energy-saving motion planning model and method based on deep reinforcement learning

    CN115935780A