Intelligent automobile decision-making and multi-scale updating method and system under multiple scenes

By setting corresponding decision makers for different scenarios and adopting the dual-end multi-scale update mechanism of the vehicle cloud, the problem of insufficient adaptability of driverless cars in complex and unfamiliar scenarios is solved, and decision-making ability and self-learning ability are improved.

CN119928909APending Publication Date: 2025-05-06SOUTH CHINA UNIV OF TECH

Patent Information

Application Number
CN202510017771.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing autonomous vehicle decision-making system has poor adaptability in complex and unfamiliar scenarios, making it difficult to achieve full driving scenario coverage.

Method used

By setting corresponding decision makers for different scenarios, the matching between scenarios and decisions is improved, and the dual-end multi-scale update mechanism of Cheyun is adopted to improve the self-learning ability of smart cars.

Benefits of technology

It improves the decision-making ability of smart cars in complex and special scenarios, enhances self-learning ability, and reduces the burden of decision-making and computing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119928909A_ABST
    Figure CN119928909A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-scene intelligent automobile decision-making and multi-scale updating method and system, the system comprises a scene classification module, a classification decision-making module and a multi-scale updating module, the scene classification module provides scene classification basis for a general scene, a complex scene and a special scene for the classification decision-making module; the classification decision-making module provides vehicle state data for the multi-scale updating module, and the multi-scale updating module provides an updated scene classification model and a decision-making device for the scene classification module and the classification decision-making module. In the classification decision-making module, general scenes correspond to classification decision-making devices in a one-to-one mode, and complex scenes and special scenes correspond to large model decision-making devices. According to the method, through scene classification and classification decision making, the matching performance of intelligent automobile decision making and scenes is effectively enhanced, and meanwhile, through multi-scale updating at the two ends of the automobile cloud, the learning ability of the intelligent automobile is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent transportation, and in particular to a method and system for intelligent vehicle decision-making and multi-scale updating in multiple scenarios. Background Art

[0002] At present, it is still difficult for driverless cars to be implemented in real scenarios. At present, most decision-making systems in driverless cars either adopt a unified decision-making strategy in all driving scenarios, or formulate decision-making strategies for a specific scenario, and simulate and test them in common high-frequency scenarios. Existing decision-making systems and methods often find it difficult to effectively cover all driving scenarios, especially unfamiliar and complex driverless scenarios, resulting in poor adaptability of driverless car decision-making systems.

[0003] The following is a reinforcement learning multi-lane driving decision method for dynamic traffic environment (CN119117006A), which does not perform scene classification, uses a single decision maker to deal with all driving scenes, and does not have a multi-scale update mechanism for both the vehicle and the cloud. A reinforcement learning-based method for intelligent merging into a highway ramp for autonomous driving vehicles (CN202311451680) is targeted at a single scene, only uses a reinforcement learning method, and does not involve a multi-scale update mechanism for both the vehicle and the cloud. The driving environment of autonomous driving vehicles varies, and it is difficult to cope with complex and changing dynamic environments using a unified decision-making framework. Summary of the invention

[0004] The present invention provides a method and system for intelligent vehicle decision-making and multi-scale updating in multiple scenarios, which improves the matching between scenarios and decisions by setting corresponding decision makers for different scenarios, and empowers intelligent vehicles through multi-scale updating on both ends of the vehicle and the cloud, thereby solving the problems of poor generalization of existing intelligent vehicle decisions and lack of self-learning ability.

[0005] The present invention is achieved by at least one of the following technical solutions.

[0006] A multi-scenario intelligent vehicle decision-making and multi-scale updating method comprises the following steps:

[0007] (1) Extracting the driving environment characteristics of the smart car, dividing the driving scenes into general scenes, complex scenes and special scenes according to the environmental characteristics, selecting special scenes from the driving scenes and storing them in a special scene storage library;

[0008] (2) Different decision-makers are selected to make decisions for different driving scenarios, and then the trajectory generated by the decision is optimized based on the road surface and vehicle dynamics parameter estimates, and the reference trajectory is converted into a control variable and transmitted to the vehicle;

[0009] (3) Based on the decision-making effect evaluation and special scene repository, a dataset is established, and the decision maker and scene classifier are updated at multiple scales on the vehicle side and the cloud side respectively.

[0010] Furthermore, it includes a scene classification module, a classification decision module and a multi-scale update module; the scene classification module provides the classification decision module with a classification basis for different scenes, the classification decision module provides the multi-scale update module with vehicle status data, and the multi-scale update module is used to update the scene classification model of the scene classification module and the decision maker of the classification decision module.

[0011] Further, the scene classification module includes an environment perception model, a feature extraction submodule, a feature combination submodule, a scene classification model, and a special scene filter;

[0012] The environmental perception model inputs sensor data and uses the spatial and temporal features of the sensor data to output environmental features to the feature extraction submodule. The feature extraction submodule outputs environmental features to the feature combination submodule. The feature combination submodule outputs the combined environmental features, namely the scene features, to the scene classification model. The scene classification model outputs scene features and scene labels for general scenes, complex scenes and special scenes. The special scene filter stores the scene features of special scenes in the special scene repository.

[0013] Furthermore, the environmental characteristics include four layers of environmental elements: road structure, traffic signals, traffic participants and weather conditions.

[0014] Furthermore, the scene classification model consists of two parts: an unsupervised clustering algorithm and a supervised classification algorithm. The unsupervised clustering algorithm first clusters the scene features to form multiple scene clusters, i.e., multiple categories of scenes;

[0015] The special scene filter will filter out special scenes: if the distance between certain scene features and the surrounding cluster centers in the feature space is greater than the threshold, these scene features will be labeled as special scenes, and their environmental feature information will be stored in the special scene repository; for scene features whose distance from the cluster center is less than the threshold, they will be classified into general scenes and complex scenes by manually setting the scene occurrence frequency threshold and environmental element complexity threshold.

[0016] An unsupervised clustering algorithm is used to assign corresponding general scene, complex scene and special scene labels to scene features, and a data set with one-to-one correspondence between scene features and scene labels is constructed to train a supervised classification algorithm and form a complete scene classification model.

[0017] Further, the classification decision module includes a classification decision maker, a large model decision maker, a road surface parameter estimation submodule, a vehicle dynamics parameter estimation submodule and a trajectory optimization submodule;

[0018] The scene features of the general scene are input into the corresponding classification decision maker, and the scene features of the complex scene or the special scene are input into the large model decision maker; the classification decision maker or the large model decision maker outputs the path points to the trajectory optimization submodule; the road surface parameter estimation submodule and the vehicle dynamics parameter estimation submodule estimate the road surface parameters and the vehicle dynamics parameters respectively according to the on-board sensor information, and input the estimation results into the trajectory optimization submodule; the trajectory optimization submodule outputs the control amount of acceleration, deceleration and steering wheel angle to the vehicle, so that the vehicle state changes.

[0019] Furthermore, the classification decision maker is constructed by a reinforcement learning algorithm, which includes a near-end policy optimization algorithm, a deep deterministic policy gradient, and a soft actor-critic algorithm. The classification decision maker consists of a fully connected training mode of a general scenario and multiple classification decisions. A general scenario requires training multiple classification decision makers, and finally the classification decision maker that should be adopted when facing the general scenario in the actual operation process is determined according to its operating efficiency and decision score;

[0020] The decision score is the decision score of the classification decision maker after the training is stabilized. The operation efficiency and decision score jointly constitute the evaluation score of the classification decision maker:

[0021] s x =λ η η x +λ score score x ;

[0022] In the formula, s x represents the classification decision evaluation score of the classification decision maker x, η x Indicates the operating efficiency of the classification decision maker x, score x The decision score of the classification decision maker x, λ η and λ score Respectively represent the weight factors of operation efficiency and decision score, which are used to balance the numerical difference between operation efficiency and decision score; the classification decision maker with the highest classification decision maker evaluation score is used as the classification decision maker corresponding to the general scenario i in the actual operation process;

[0023] The large model decision maker is a large model decision maker based on the Transformer model;

[0024] Road surface parameter estimation submodule: The road scene image in front of the vehicle is captured by the on-board camera, and the road surface features are extracted through the lightweight semantic segmentation network ENet. MobileNetV3 is used as the classification network for the road surface image. Finally, the adhesion coefficient corresponding to the classified road surface category is obtained according to the reference value table of the vehicle longitudinal slip adhesion coefficient.

[0025] Vehicle dynamics parameter estimation submodule: simulates vehicle dynamics based on radial basis neural network, uses vehicle dynamics model based on radial basis neural network to estimate and model the internal dynamic parameters of the vehicle in real time, collects the lateral speed, longitudinal speed, lateral acceleration, longitudinal acceleration, yaw rate state data at this moment, and the throttle, brake and steering control data at this moment through sensors, inputs these data into the radial basis neural network, and predicts the changes in the lateral speed, longitudinal speed, lateral acceleration, longitudinal acceleration, yaw rate state data of the vehicle at the next moment;

[0026] Trajectory optimization submodule: The trajectory optimization submodule includes an input data module, a constraint module, and an optimization objective function module; the data of the input data module include the trajectory points output by the above decision maker, the road adhesion coefficient, and the lateral speed, longitudinal speed, lateral acceleration, longitudinal acceleration, and yaw rate state data at the next moment; the constraint module is designed based on the road adhesion coefficient and the lateral speed, longitudinal speed, lateral acceleration, longitudinal acceleration, and yaw rate state data at the next moment, and the lateral acceleration constraint is:

[0027] a y ≤μ·g;

[0028] Among them, a y is the lateral acceleration, μ is the road adhesion coefficient, and g is the gravitational acceleration;

[0029] The objective function of the optimization objective function module is defined as a multi-objective optimization problem, including smoothness optimization and target tracking optimization:

[0030] Trajectory smoothness objective function:

[0031]

[0032] Among them, J smoot is the trajectory smoothness objective function, N is the number of trajectory points, is the square of the change in acceleration.

[0033] Trajectory tracking objective function:

[0034]

[0035] Among them, J tracking is the trajectory tracking objective function, N is the number of trajectory points, and p i is the actual trajectory point, is the reference trajectory point.

[0036] Furthermore, the multi-scale update module includes a decision effect evaluation submodule, a vehicle-side short-cycle self-learning submodule, and a cloud-side long-cycle retraining submodule;

[0037] The input of the decision effect evaluation submodule is the state of the car after following the path point generated by the decision maker in a special scenario, and the output is the decision score;

[0038] The vehicle-side short-cycle self-learning submodule is used to update the large model decision maker of the classification decision module: collect the scene features and corresponding decision scores of special scenes to form a user data set, which only includes special scene data collected by car users during driving; use the user data set to train the large model decision maker until the large model decision maker can obtain stable decision scores in all special scenes of the user data set;

[0039] The cloud-based long-cycle retraining submodule is used to update the decision makers of the scene classification model and the classification decision module: the driving scenes of the enterprise data set are reclassified according to the scene classification model. After reclassification, it is determined whether the position and range of the new and old scene clusters in the feature space have changed, and whether the new and old scene clusters overlap in the feature space. If a new scene cluster does not overlap with all old scene clusters, the scene cluster is regarded as a general scene of a new type, and the fully connected training mode is used for this general scene to obtain the best classification decision maker; if overlap occurs, the intersection-over-union formula is used to calculate the overlap rate. When the overlap rate is greater than or equal to 90%, the scene cluster is regarded as a general scene of the same type as the overlapping old scene cluster, and the corresponding classification decision maker is used as the best classification decision maker for the new scene cluster, and this classification decision maker is trained based on the original parameters; when the overlap rate is less than 90%, the scene cluster is also regarded as a general scene of a new type, and the above-mentioned fully connected training mode is also used to obtain the best classification decision maker.

[0040] Furthermore, in addition to the special scenario data collected by automobile users, the enterprise dataset also has general scenario data, complex scenario data and special scenario data collected by the enterprise itself. The user dataset is a subset of the enterprise dataset.

[0041] Furthermore, the decision effect evaluation function of the decision effect evaluation submodule is as follows:

[0042] score=λ safe score safe +λ stable score stable +λ comfort score comfort +λ efficient score efficient +λ eco score eco ;

[0043] In the formula, score represents the decision score, score safe、score stable 、score comfort 、score efficient 、score eco They represent safety score, stability score, comfort score, efficiency score, and economic score respectively, and λ safe , stable , comfort , efficient , eco They represent the weight factors corresponding to the safety score, stability score, comfort score, efficiency score, and economic score respectively.

[0044] Compared with the prior art, the present invention has the following beneficial effects:

[0045] (1) Establish a scene classification model, divide the driving environment into several general scenes, complex scenes and special scenes, and adopt a classification decision-making mechanism so that each of the above scenes has a corresponding decision maker. The decision-making system will adopt targeted decision makers and decision-making behaviors for different scenes. A lightweight classification decision maker is used for general scenes, and a large model decision maker is used for complex scenes and special scenes. This is conducive to reducing the decision-making computing burden of smart cars and improving the decision-making ability of smart cars in complex scenes and special scenes.

[0046] (2) A vehicle-cloud dual-end multi-scale update framework is adopted. The spare computing power of the vehicle computing unit is used to perform short-cycle training and update of the large model decision maker. The large computing power of the cloud server is used to perform long-cycle training and update of the scene classification model and all decision makers including the classification decision maker and the large model decision maker. This approach will help smart cars quickly form decision makers for special scenarios when they encounter them, thereby improving their self-learning capabilities. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 It is a schematic diagram of the overall structural framework of the present invention;

[0048] Figure 2 This is a schematic diagram of the structural framework of the scene classification module of the present invention;

[0049] Figure 3 It is a schematic diagram of the structural framework of the environment perception model of the present invention;

[0050] Figure 4 It is a schematic diagram of the structural framework of the fusion feature encoder of the present invention;

[0051] Figure 5 Schematic diagram of general scenes, complex scenes and special scenes in feature space of the present invention;

[0052] Figure 6 This is a schematic diagram of the structural framework of the classification decision module of the present invention;

[0053] Figure 7 A schematic diagram of a fully connected training mode for a general scenario of the present invention;

[0054] Figure 8 It is a schematic diagram of the structural framework of the multi-scale update module of the present invention;

[0055] Fig. 9 The present invention is a flowchart of a smart car decision-making and multi-scale updating method in multiple scenarios. DETAILED DESCRIPTION

[0056] To make the technical solution and advantages of the present application clearer, the implementation methods of the present application will be further described in detail below in conjunction with the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0057] like Figure 1-Figure 8 As shown, this embodiment provides a smart car decision-making and multi-scale update system in multiple scenarios, including a scene classification module, a classification decision module and a multi-scale update module. The scene classification module provides the classification decision module with scene classification basis for general scenes, complex scenes and special scenes, the classification decision module provides the multi-scale update module with vehicle status data, and the multi-scale update module is used to update the scene classification model of the scene classification module and the decision maker of the classification decision module.

[0058] like Figure 2 As shown, the scene classification module includes an environment perception model, a feature extraction submodule, a feature combination submodule, a scene classification model, and a special scene filter.

[0059] like Figure 3-Figure 4 As shown, the environment perception model includes a camera image encoder, a lidar point cloud encoder, a multi-view camera transformation module, a point cloud feature flattening module, a fusion feature encoder and multiple task heads;

[0060] The camera image encoder is composed of a Swin Transformer model or other models suitable for image encoding;

[0061] The lidar point cloud encoder consists of a VoxelNet model or other models that encode point clouds in voxel form;

[0062] The multi-view camera transformation module adopts the LSS algorithm and uses the paradigm of depth estimation-view cone point cloud generation-coordinate transformation to map the multi-view camera image features after the camera image encoder to unified image features using the estimated depth information and the known camera internal and external parameter matrix;

[0063] The point cloud feature flattening module is a pooling layer that pools the point cloud features passed through the lidar point cloud encoder along the z-axis.

[0064] The features output by the multi-view camera transformation module and the point cloud feature flattening module are stacked along the z-axis to obtain stacked features.

[0065] The fusion feature encoder is a Transformer architecture, which consists of a temporal self-attention mechanism, a residual connection and regularization component, a spatial cross-attention mechanism, and a feedforward neural network, and is used to extract spatial and temporal features from sensor data. Specifically, the fusion feature query Q and the historical fusion feature B t-1 After the temporal self-attention mechanism, the fused feature query Q is subjected to residual connection and regularization 1; the features obtained after residual connection and regularization 1 and the stacked features are subjected to spatial cross-attention mechanism, and then the features obtained after residual connection and regularization 1 are subjected to residual connection and regularization 2; the features obtained after residual connection and regularization 2 are subjected to residual connection and regularization 3 with the features obtained after residual connection and regularization 2 after passing through the feedforward neural network; the above steps are repeated 6 times to finally obtain the fused feature B at the current moment t ; Historical fusion feature B t-1 That is, the fusion feature B of the previous moment t , the fused feature query Q is a set of randomly initialized self-learnable parameters.

[0066] The multi-task head includes a 3D object detection head, a semantic segmentation head, etc., which can output different environmental features. The environmental features include four layers of environmental elements: road structure, traffic signals, traffic participants, and weather conditions.

[0067] As one of the embodiments, the environmental features of the road structure layer include lane width, lane line type, lane line curvature, road slope, road defects, road static obstacles, road surface material and other information; the environmental features of the traffic signal layer include road surface indication arrows, traffic cones, pedestrian crosswalks, traffic lights, various signs and other information; the environmental features of the traffic participant layer include the three-dimensional size, type label, speed direction and other information of traffic participants such as other cars or pedestrians; the environmental features of the weather condition layer include light intensity, rain, snow and fog, air humidity and other information. It is worth noting that the SwinTransformer model, VoxelNet model, LSS algorithm, pooling layer, Transformer architecture, temporal self-attention mechanism, residual connection and regularization component, spatial cross attention mechanism and feedforward neural network are all existing models, algorithms or components that can be checked, and the three-dimensional target detection head and semantic segmentation head in the multi-task head can also be implemented through the existing task head that can be checked.

[0068] The environmental perception model outputs the captured environmental features to the feature extraction submodule. The feature extraction submodule uses the following rule-based method to screen important environmental features. Specifically, for the environmental features of the road structure layer, the relevant information of the vehicle lane and the left and right lanes are screened out as important environmental features; for the environmental features of the traffic signal layer, the different types of traffic signals closest to the vehicle are screened out as important environmental features, with the vehicle as the center and observed along the longitudinal front and rear directions of the lane; for the environmental features of the traffic participant layer, the relevant information of the traffic participant closest to the vehicle in the eight directions is screened out as important environmental features, where the eight directions include left front, front, right front, left, right, left rear, rear and right rear; for the environmental features of the weather condition layer, the information that affects the perception and decision of the vehicle, such as light intensity, rain, snow and fog, and air humidity, is screened out as important environmental features.

[0069] The feature extraction submodule outputs the selected important environmental features to the feature combination submodule. The feature combination submodule is a series of regularized data processing steps. Specifically, the important environmental features of each layer above are each formed into a feature vector, and then a feature matrix is ​​formed in the order of road structure layer, traffic signal layer, traffic participant layer and weather condition layer. Finally, all values ​​in the feature matrix are normalized to obtain scene features.

[0070] The feature combination submodule outputs the scene features to the scene classification model. The scene classification model consists of two parts: an unsupervised clustering algorithm and a supervised classification algorithm. The unsupervised clustering algorithm first clusters the scene features to form multiple scene clusters, that is, multiple categories of scenes.

[0071] The special scene filter will filter out special scenes: if the distance between some scene features and the center of the surrounding clusters in the feature space is greater than the threshold, these scene features will be labeled as special scenes, and their environmental feature information will be stored in the special scene repository.

[0072] For scene features whose distance from the cluster center is less than the threshold, they are classified into general scenes and complex scenes by manually setting the frequency threshold of scene occurrence and the complexity threshold of environmental elements, such as Figure 5 shown.

[0073] An unsupervised clustering algorithm is used to assign corresponding general scene, complex scene and special scene labels to scene features, and a data set with one-to-one correspondence between scene features and scene labels is constructed. This is used to train a supervised classification algorithm to form a complete scene classification model, enabling it to acquire the ability to identify scene categories online.

[0074] The scene classification model outputs scene features and scene labels for general scenes, complex scenes and special scenes, and the special scene filter stores the scene features of special scenes in the special scene repository.

[0075] As one embodiment, the unsupervised clustering algorithm may adopt the K-means algorithm, and the number of scene categories may be determined by the within-cluster sum of squared errors (SSE) and the elbow rule. The SSE calculation formula is as follows:

[0076]

[0077] In the formula, n represents the number of scenes, m represents the number of scene features in this type of scene, and μ (j) represents the cluster center of the i-th type of scene, x (j) Represents the jth scene sample of the i-th type of scene. By taking different scene numbers n, calculating the corresponding SSE, drawing a curve with the scene number n as the horizontal axis and the SSE as the vertical axis, the number of scenes n corresponding to the curve elbow data point is determined as the optimal number of scenes according to the elbow rule; the supervised classification algorithm can be constructed using multi-layer perceptron and one-hot encoding, and one-hot encoding is set for each type of scene. The input dimension size of the multi-layer perceptron is the dimension size of the scene sample, and the output dimension size of the multi-layer perceptron and the dimension size of the one-hot encoding are exactly the number of scene types n. It is worth noting that the number of scene types refers to the number of scene clusters obtained after the unsupervised clustering algorithm, which is also the number of general scenes (general sub-scenes), and does not include complex scenes and special scenes.

[0078] In this embodiment, a general scenario refers to a scenario in which a car appears frequently during driving, the environmental factors are not complex, and the decision-making ability requirements are low. A complex scenario refers to a scenario in which a car appears in a medium to low frequency during driving, the environmental factors are complex, and the decision-making ability requirements are high. A special scenario refers to a scenario in which a car appears extremely rarely during driving, and the above three scenarios can be distinguished by the thresholds set by the scene classification model. The scene refers to the instantaneous comprehensive reflection of the four layers of environmental factors, namely, road structure, traffic signals, traffic participants, and weather conditions, within the detectable range around the car. The general scenario includes n general sub-scenarios, and the specific number is determined by the scene classification model.

[0079] like Figure 6 As shown, the classification decision module includes a classification decision maker, a large model decision maker, a road surface parameter estimation submodule, a vehicle dynamics parameter estimation submodule and a trajectory optimization submodule.

[0080] The number of classification decision makers is the same as the number of general scenes, which is n classification decision makers. General scenes and classification decision makers correspond one to one. For example, when the car detects the first general scene 1, only the first classification decision maker 1 can be used, and other sub-classification decision makers or large model decision makers such as the second classification decision maker 2 and the third classification decision maker 3 cannot be used. In addition, when the car detects a complex scene or a special scene, the large model decision maker is used.

[0081] In addition, when a car can obtain a stable decision score through the decision-making effect evaluation sub-module in a special scenario, the special scenario can be deleted from the special scenario library and classified as a general scenario or a complex scenario according to the complexity of its environmental factors and the requirements for decision-making capabilities.

[0082] The classification decision maker or the large model decision maker outputs the trajectory points to the trajectory optimization submodule. The road surface parameter estimation submodule and the vehicle dynamics parameter estimation submodule estimate the road surface parameters and vehicle dynamics parameters respectively according to the information from the on-board sensors, and input the estimation results to the trajectory optimization submodule. The trajectory optimization submodule outputs the control amount of acceleration, deceleration and steering wheel angle to the vehicle, causing the vehicle state to change.

[0083] The classification decision maker is constructed by a reinforcement learning algorithm, which includes the proximal policy optimization algorithm (PPO), deep deterministic policy gradient (DDPG), soft actor-critic algorithm (SAC), etc. For different general scenarios, the specific reinforcement learning algorithm type and the parameters of the learnable parameters of the network model in the reinforcement learning algorithm are jointly determined by the operating efficiency and the decision score of the subsequent decision effect evaluation submodule. The determination process of the classification decision maker consists of a fully connected training mode of a class of general scenarios and multiple classification decisions. Various classification decision makers need to be trained for a class of general scenarios. Finally, the classification decision maker that should be adopted when facing the general scenario in the actual operation process is determined based on its operating efficiency and decision score. Figure 7As shown in FIG. 1 , the process of the fully connected training mode is as follows: according to the different learnable parameters of the network model in the reinforcement learning algorithm, the first SAC classification decision maker 1, the second SAC classification decision maker 2 and the third SAC classification decision maker 3, as well as the first DDPG classification decision maker 1 and the second DDPG classification decision maker 2 are initialized; in general scenario i, the first SAC classification decision maker 1, the second SAC classification decision maker 2 and the third SAC classification decision maker 3, as well as the first DDPG classification decision maker 1 and the second DDPG classification decision maker 2 are trained simultaneously; it should be noted that in order to ensure the comparability of the different classification decision makers constructed using the reinforcement learning algorithm, the different classification decision makers are required to have the same state space, action space and reward function, the state space is the feature space of the scene feature, the action space is the path point position of the next 5 time steps, and the reward function is the decision score described later; after the reward function of all classification decision makers in general scenario i tends to be stable or reaches a sufficient number of training steps, the operation efficiency and decision score of different classification decision makers are statistically analyzed, wherein the operation efficiency is measured by the existing evaluation index floating point operations (Floating Point Operations Per Second, FLOPs) calculation, that is, the number of floating-point operations that the computing model needs to perform to complete a forward propagation or inference, and the decision score is the reward function (decision score) of the classification decision maker after training is stabilized. The operating efficiency and decision score jointly constitute the evaluation score of the classification decision maker:

[0084] s x =λ η ηx x +λs score score x ;

[0085] In the formula, s x represents the classification decision evaluation score of the classification decision maker x, η x Indicates the operating efficiency of the classification decision maker x, score x The decision score of the classification decision maker x, λ η and λ score The weight factors of the operating efficiency and decision score are used to balance the numerical difference between the operating efficiency and the decision score. Finally, the classification decision maker with the highest classification decision maker evaluation score is used as the classification decision maker corresponding to the general scenario i in the actual operation process. The advantage of this fully connected training mode is that the car can select a decision algorithm with high operating efficiency as the sub-classification decision maker under the premise of ensuring a certain decision score, thereby enhancing the matching of the decision and the scenario while reducing the computational burden of the decision system as much as possible.

[0086] The large model decision maker is a large model decision maker based on the Transformer model. It is constructed based on the Transformer model. The number of parameters of the large model decision maker is much larger than the reinforcement learning algorithm in the above-mentioned classification decision maker, and it has a high level of logical reasoning ability. Specifically, the input dimension of the Transformer model is adjusted to match the dimension of the scene feature, and the output dimension of the Transformer model is adjusted to match the dimension formed by the path point position in the next 5 time steps, that is, the construction of the large model decision maker based on the Transformer model is completed; the reinforcement learning training paradigm is also used to train the large model decision maker, and the large model decision maker based on the Transformer model is defined as the policy in reinforcement learning, the output path point is defined as the action space in reinforcement learning, and the input scene feature is defined as the observation space in reinforcement learning. The decision score is used as the reward function. After defining the basic elements of the above-mentioned reinforcement learning (strategy, action space, observation space, reward function), training can be carried out.

[0087] Road surface parameter estimation submodule: Combined with road surface recognition technology based on computer vision, the road conditions are monitored in real time through on-board cameras and lidar. Specifically, the road surface recognition technology based on computer vision includes four steps: image acquisition, semantic segmentation, road surface image classification, and adhesion coefficient matching. First, the road scene image in front of the vehicle is captured by the on-board camera, and the road surface features are extracted through the efficient and lightweight semantic segmentation network ENet (Efficient Neural Network). Then, MobileNetV3 is used as the classification network of the road surface image, and the road surface classification network is trained based on the million-sample road surface image dataset RSCD (Road surface classification dataset) launched by Tsinghua University in 2022. Finally, the adhesion coefficient corresponding to the classified road surface category is obtained according to the reference value table of the vehicle longitudinal slip adhesion coefficient.

[0088] Vehicle dynamics parameter estimation submodule: simulate vehicle dynamics based on radial basis function network (RBF network), use vehicle dynamics model based on radial basis function network (RBF network) to estimate and model the internal dynamic parameters of the vehicle in real time, collect lateral speed, longitudinal speed, lateral acceleration, longitudinal acceleration, yaw rate state data at this moment, and throttle, brake and steering control data at this moment through sensors, input these data into the neural network model, and predict the changes in the lateral speed, longitudinal speed, lateral acceleration, longitudinal acceleration and yaw rate state data of the vehicle at the next moment.

[0089] Trajectory optimization submodule: The trajectory optimization submodule includes an input data module, a constraint module, and an optimization objective function module. The data in the input data module include the trajectory points output by the above decision maker, the road adhesion coefficient, and the lateral speed, longitudinal speed, lateral acceleration, longitudinal acceleration, and yaw rate state data at the next moment. The constraint module is designed based on the road adhesion coefficient and the lateral speed, longitudinal speed, lateral acceleration, longitudinal acceleration, and yaw rate state data at the next moment. The lateral acceleration constraint:

[0090] A y ≤μ·g;

[0091] Among them, a y is the lateral acceleration, μ is the road adhesion coefficient, and g is the gravitational acceleration.

[0092] The objective function of the optimization objective function module is defined as a multi-objective optimization problem, including smoothness optimization and target tracking optimization:

[0093] Trajectory smoothness objective function:

[0094]

[0095] Among them, J smoot is the trajectory smoothness objective function, N is the number of trajectory points, is the square of the change in acceleration.

[0096] Trajectory tracking objective function:

[0097]

[0098] Among them, J tracking is the trajectory tracking objective function, N is the number of trajectory points, and p i is the actual trajectory point, is the reference trajectory point.

[0099] like Figure 8 As shown in the figure, the multi-scale update module includes a decision effect evaluation submodule, a vehicle-side short-cycle self-learning submodule, and a cloud-side long-cycle retraining submodule.

[0100] The input of the decision effect evaluation submodule is the vehicle state after the car follows the path point generated by the decision maker in a special scenario, and the output is the decision score. Specifically, the decision effect evaluation submodule constructs a decision effect evaluation function from five evaluation indicators: safety, stability, comfort, efficiency, and economy, so as to obtain a decision score closely related to these five evaluation indicators. The decision effect evaluation function is shown in the following formula:

[0101] score=λ safe score safe +λ stable score stable +λ comfort score comfort +λ efficient score efficient +λ eco score eco ;

[0102] In the formula, score represents the decision score, score safe 、score stable 、score comfort 、score efficient 、score eco They represent safety score, stability score, comfort score, efficiency score, and economic score respectively, and λ safe , stable , comfort , efficient , eco They represent the weight factors corresponding to the safety score, stability score, comfort score, efficiency score, and economic score respectively.

[0103] Further:

[0104]

[0105] score stable =e -|β| / π ;

[0106] score comfort =0.4(1-|a x | / 2g)+0.6(1-|a y | / g);

[0107]

[0108] In the formula, pego represents the position coordinates of the vehicle, p obstacle represents the position coordinates of the obstacle closest to the vehicle, including other vehicles, pedestrians, road boundaries, etc.; β represents the side slip angle of the vehicle's center of mass, in radians; a x and a y represents the longitudinal acceleration and lateral acceleration of the vehicle, g represents the acceleration due to gravity; p i represents the i-th path point generated by the decision maker, i∈[1,5]; p goal Represents the position coordinates of the target point, which is provided by the global route; the position coordinates of the vehicle, the sideslip angle of the center of mass, and the longitudinal / lateral acceleration are collectively referred to as the vehicle state.

[0109] The scene features and corresponding decision scores of special scenes are collected to form a user data set. The user data set only includes special scene data collected by car users during driving. The user data set is a subset of the enterprise data set. In addition to the special scene data collected by car users, the enterprise data set also has general scene data, complex scene data and special scene data collected by the enterprise itself. The user data set is used for short-term self-learning on the vehicle side, and the enterprise data set is used for long-term retraining on the cloud side. The purpose of distinguishing the two data sets is to make the data collected by users focus only on special scenes, avoiding the additional costs brought by the collection of general scenes and complex scenes. In addition, the enterprise data set including general scenes, complex scenes and special scenes is needed to provide data for updating the scene classification model in the long-term retraining submodule on the cloud side.

[0110] After the above-mentioned user data set obtains data, the short-cycle self-learning sub-module on the vehicle side only updates the large model decision maker; after the above-mentioned enterprise data set obtains data, the long-cycle retraining sub-module on the cloud side updates the scene classification model and all decision makers including the classification decision makers and the large model decision makers.

[0111] Specifically, in the short-cycle self-learning sub-module on the vehicle side, the large model decision maker is trained using the user data set until the large model decision maker can obtain stable decision scores in all special scenarios of the user data set. The above process is completed on the on-board computing unit and storage device.

[0112] In the long-term retraining submodule on the cloud, the cloud server first reclassifies the driving scenes for the enterprise data set according to the above-mentioned scene classification model. After reclassification, it is determined whether the position and range of the new and old scene clusters in the feature space have changed. For scene clusters with slight changes, the original corresponding classification decision maker is used. For scene clusters with large changes, it is necessary to obtain the corresponding optimal classification decision maker according to the above-mentioned fully connected training mode. For complex and special scenes, the large model decision maker needs to be retrained. Specifically, first determine whether the new and old scene clusters overlap in the feature space. If a new scene cluster does not overlap with all the old scene clusters, the scene cluster is regarded as a general scene of a new type, and the above-mentioned fully connected training mode is used for the general scene to obtain the best classification decision maker; if overlap occurs, the intersection-over-union formula is used to calculate the overlap rate. When the overlap rate is greater than or equal to 90%, the scene cluster is regarded as a general scene of the same type as the overlapping old scene cluster, and the corresponding classification decision maker is used as the best classification decision maker for the new scene cluster, and this classification decision maker is trained based on its original parameters; when the overlap rate is less than 90%, the scene cluster is also regarded as a general scene of a new type, and the above-mentioned fully connected training mode is also used to obtain the best classification decision maker.

[0113] The above short cycle generally refers to an update cycle within a week, which can be flexibly adjusted according to actual needs. That is, when the car encounters a special scene during actual driving and obtains the corresponding environmental characteristics and decision scores, the car can carry out short-term self-learning tasks on the vehicle side when the on-board computing unit has sufficient computing power. The long cycle generally refers to an update cycle of a week or more, which is relatively fixed. The cloud server needs to collect special scene data from multiple vehicles and then carry out long-term retraining tasks on the cloud within a relatively fixed long period. After the training is completed, the scene classification model and the weight parameters of various decision makers are packaged for smart cars to update the scene classification model and decision makers of the car through over-the-air download technology.

[0114] like Fig. 9 As shown, the above-mentioned intelligent vehicle decision-making and multi-scale updating methods in multiple scenarios are implemented, and the specific steps are as follows:

[0115] S1. Extracting driving environment characteristics of smart cars, dividing driving scenes into general scenes, complex scenes and special scenes according to the environmental characteristics, screening special scenes from driving scenes and storing them in a special scene repository;

[0116] S2. Select different decision makers to make decisions for different driving scenarios, optimize the trajectory generated by the decision according to the road surface and vehicle dynamics parameter estimation, and convert the reference trajectory into control quantity and transmit it to the car;

[0117] S3. Establish a data set based on the decision effect evaluation and special scene repository, and perform multi-scale updates on the decision maker and scene classifier on the vehicle side and the cloud side respectively.

[0118] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be equivalent replacement methods and are included in the protection scope of the present invention.

Claims

1. A multi-scenario intelligent vehicle decision-making and multi-scale updating method, characterized in that: The following steps are involved: (1) Extracting the driving environment characteristics of the smart car, dividing the driving scenes into general scenes, complex scenes and special scenes according to the environmental characteristics, selecting special scenes from the driving scenes and storing them in a special scene storage library; (2) Different decision-makers are selected to make decisions for different driving scenarios, and then the trajectory generated by the decision is optimized based on the road surface and vehicle dynamics parameter estimates, and the reference trajectory is converted into a control variable and transmitted to the vehicle; (3) Based on the decision-making effect evaluation and special scene repository, a dataset is established, and the decision maker and scene classifier are updated at multiple scales on the vehicle side and the cloud side respectively.

2. A system for implementing the multi-scenario intelligent vehicle decision-making and multi-scale updating method according to claim 1, characterized in that: It includes a scene classification module, a classification decision module and a multi-scale update module; the scene classification module provides the classification decision module with the classification basis of different scenes, the classification decision module provides the multi-scale update module with vehicle status data, and the multi-scale update module is used to update the scene classification model of the scene classification module and the decision maker of the classification decision module.

3. The multi-scenario intelligent vehicle decision-making and multi-scale update system according to claim 1 is characterized in that: The scene classification module includes an environment perception model, a feature extraction submodule, a feature combination submodule, a scene classification model, and a special scene filter; The environmental perception model inputs sensor data and uses the spatial and temporal features of the sensor data to output environmental features to the feature extraction submodule. The feature extraction submodule outputs environmental features to the feature combination submodule. The feature combination submodule outputs the combined environmental features, namely the scene features, to the scene classification model. The scene classification model outputs scene features and scene labels for general scenes, complex scenes and special scenes. The special scene filter stores the scene features of special scenes in the special scene repository.

4. The multi-scenario intelligent vehicle decision-making and multi-scale updating system according to claim 3 is characterized in that: The environmental characteristics include four layers of environmental elements: road structure, traffic signals, traffic participants and weather conditions.

5. The multi-scenario intelligent vehicle decision-making and multi-scale updating system according to claim 3 is characterized in that: The scene classification model consists of two parts: an unsupervised clustering algorithm and a supervised classification algorithm. The unsupervised clustering algorithm first clusters the scene features to form multiple scene clusters, that is, scenes of multiple categories; The special scene filter will filter out special scenes: if the distance between certain scene features and the surrounding cluster centers in the feature space is greater than the threshold, these scene features will be labeled as special scenes, and their environmental feature information will be stored in the special scene repository; for scene features whose distance to the cluster center is less than the threshold, they will be classified into general scenes and complex scenes by manually setting the scene occurrence frequency threshold and environmental factor complexity threshold; An unsupervised clustering algorithm is used to assign corresponding general scene, complex scene and special scene labels to scene features, and a data set with one-to-one correspondence between scene features and scene labels is constructed to train a supervised classification algorithm and form a complete scene classification model.

6. The multi-scenario intelligent vehicle decision-making and multi-scale update system according to claim 1, characterized in that: The classification decision module includes a classification decision maker, a large model decision maker, a road surface parameter estimation submodule, a vehicle dynamics parameter estimation submodule and a trajectory optimization submodule; The scene features of the general scene are input into the corresponding classification decision maker, and the scene features of the complex scene or special scene are input into the large model decision maker; The classification decision maker or the large model decision maker outputs the path points to the trajectory optimization submodule; The road surface parameter estimation submodule and the vehicle dynamics parameter estimation submodule estimate the road surface parameters and vehicle dynamics parameters respectively according to the on-board sensor information, and input the estimation results into the trajectory optimization submodule; The trajectory optimization submodule outputs the control values ​​of acceleration, deceleration and steering wheel angle to the vehicle, causing the vehicle state to change.

7. The multi-scenario intelligent vehicle decision-making and multi-scale update system according to claim 6, characterized in that: The classification decision maker is constructed by a reinforcement learning algorithm, which includes a near-end policy optimization algorithm, a deep deterministic policy gradient, and a soft actor-critic algorithm. The classification decision maker consists of a fully connected training mode of a general scenario and multiple classification decisions. A general scenario requires training multiple classification decision makers. Finally, the classification decision maker that should be adopted when facing the general scenario in the actual operation process is determined based on its operating efficiency and decision score. The decision score is the decision score of the classification decision maker after the training is stabilized. The operation efficiency and decision score jointly constitute the evaluation score of the classification decision maker: s x =λ η or x +λ score score x ; In the formula, s x represents the classification decision evaluation score of the classification decision maker x, η x Indicates the operating efficiency of the classification decision maker x, score x The decision score of the classification decision maker x, λ η and λ score Respectively represent the weight factors of operation efficiency and decision score, which are used to balance the numerical difference between operation efficiency and decision score; The classification decision maker that obtains the highest classification decision maker evaluation score is used as the classification decision maker corresponding to the general scenario i in the actual operation process; The large model decision maker is a large model decision maker based on the Transformer model; Road Parameter Estimation Submodule: The vehicle-mounted camera captures the road scene image in front of the vehicle, and the road features are extracted through the lightweight semantic segmentation network ENet. MobileNetV3 is used as the classification network of the road image. Finally, the adhesion coefficient corresponding to the classified road type is obtained according to the reference value table of the longitudinal slip adhesion coefficient of the vehicle. Vehicle dynamics parameter estimation submodule: simulates vehicle dynamics based on radial basis neural network, uses vehicle dynamics model based on radial basis neural network to estimate and model the internal dynamic parameters of the vehicle in real time, collects the lateral speed, longitudinal speed, lateral acceleration, longitudinal acceleration, yaw rate state data at this moment, and the throttle, brake and steering control data at this moment through sensors, inputs these data into the radial basis neural network, and predicts the changes in the lateral speed, longitudinal speed, lateral acceleration, longitudinal acceleration, yaw rate state data of the vehicle at the next moment; Trajectory optimization submodule: The trajectory optimization submodule includes an input data module, a constraint module, and an optimization objective function module; the data of the input data module include the trajectory points output by the above decision maker, the road adhesion coefficient, and the lateral speed, longitudinal speed, lateral acceleration, longitudinal acceleration, and yaw rate state data at the next moment; the constraint module is designed based on the road adhesion coefficient and the lateral speed, longitudinal speed, lateral acceleration, longitudinal acceleration, and yaw rate state data at the next moment, and the lateral acceleration constraint is: a y ≤μ·g; Among them, a y is the lateral acceleration, μ is the road adhesion coefficient, and g is the gravitational acceleration; The objective function of the optimization objective function module is defined as a multi-objective optimization problem, including smoothness optimization and target tracking optimization: Trajectory smoothness objective function: Among them, J smoot is the trajectory smoothness objective function, N is the number of trajectory points, is the square of the change in acceleration; Trajectory tracking objective function: Among them, J tracking is the trajectory tracking objective function, N is the number of trajectory points, and p i is the actual trajectory point, is the reference trajectory point.

8. The multi-scenario intelligent vehicle decision-making and multi-scale update system according to claim 1, characterized in that: The multi-scale update module includes a decision effect evaluation submodule, a vehicle-side short-cycle self-learning submodule, and a cloud-side long-cycle retraining submodule; The input of the decision effect evaluation submodule is the state of the car after following the path point generated by the decision maker in a special scenario, and the output is the decision score; The vehicle-side short-cycle self-learning submodule is used to update the large model decision maker of the classification decision module: collect the scene features and corresponding decision scores of special scenes to form a user data set, which only includes special scene data collected by car users during driving; use the user data set to train the large model decision maker until the large model decision maker can obtain stable decision scores in all special scenes of the user data set; The cloud-based long-cycle retraining submodule is used to update the decision makers of the scene classification model and the classification decision module: the driving scenes of the enterprise data set are reclassified according to the scene classification model. After reclassification, it is determined whether the position and range of the new and old scene clusters in the feature space have changed, and whether the new and old scene clusters overlap in the feature space. If a new scene cluster does not overlap with all old scene clusters, the scene cluster is regarded as a general scene of a new type, and the fully connected training mode is used for this general scene to obtain the best classification decision maker; if overlap occurs, the intersection-over-union formula is used to calculate the overlap rate. When the overlap rate is greater than or equal to 90%, the scene cluster is regarded as a general scene of the same type as the overlapping old scene cluster, and the corresponding classification decision maker is used as the best classification decision maker for the new scene cluster, and this classification decision maker is trained based on the original parameters; when the overlap rate is less than 90%, the scene cluster is also regarded as a general scene of a new type, and the above-mentioned fully connected training mode is also used to obtain the best classification decision maker.

9. The multi-scenario intelligent vehicle decision-making and multi-scale update system according to claim 8, characterized in that: In addition to the special scenario data collected by automobile users, the enterprise dataset also includes general scenario data, complex scenario data and special scenario data collected by the enterprise itself. The user dataset is a subset of the enterprise dataset.

10. The multi-scenario intelligent vehicle decision-making and multi-scale update system according to claim 8, characterized in that: The decision effect evaluation function of the decision effect evaluation submodule is shown as follows: score=λ safe score safe +λ stable score stable +λ comfort score comfort +λ efficient score efficient +λ eco score eco ; In the formula, score represents the decision score, score safe 、score stable 、score comfort 、score efficient 、score eco They represent safety score, stability score, comfort score, efficiency score, and economic score respectively, and λ safe , stable , comfort , efficient , eco They represent the weight factors corresponding to the safety score, stability score, comfort score, efficiency score, and economic score respectively.

Citation Information

Patent Citations

  • Automatic driving vehicle high-speed ramp intelligent afflux method based on reinforcement learning

    CN117227761A

  • Reinforcement learning multi-lane driving decision-making method for dynamic traffic environment

    CN119117006A

Cited By

  • Model updating method, device and system and storage medium

    CN121212244A

  • Expressway multi-expert autonomous behavior decision-making method based on scene perception constraint

    CN121617266A

  • A Multi-Expert Autonomous Behavior Decision-Making Method for Highways Based on Scene-Aware Constraints

    CN121617266B