An Autonomous Driving Adversarial Training Method and Device
By generating feasible trajectories through the bicycle point cloud data and performing trajectory resampling and classification training, the problem of insufficient robustness and generalization capabilities of autonomous driving technology in complex scenarios is solved, and more efficient and safe autonomous driving is achieved.
Patent Information
- Application Number
- CN202510551073.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-04-29
Smart Images

Figure CN120071305B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of autonomous driving, and in particular to an autonomous driving adversarial training method and device. Background Art
[0002] With the rapid development of artificial intelligence technology, autonomous driving technology has become a research hotspot in the fields of the automotive industry and intelligent transportation systems. The core of autonomous driving technology lies in achieving autonomous navigation and safe driving of vehicles in complex road environments through key links such as perception, decision-making, and control. However, in the face of changing traffic scenarios, uncertain dynamic obstacles, and complex road structures, how to ensure the robustness and safety of autonomous driving systems is a key problem that needs to be solved urgently. Especially in dealing with massive point cloud data, performing efficient path planning, and coping with adversarial attacks, existing autonomous driving technologies still face many challenges.
[0003] Currently, autonomous driving technology mainly relies on technical means such as deep learning, computer vision, and sensor fusion. At the perception level, environmental information is obtained through sensors such as lidar and cameras, and deep learning algorithms are used to perform tasks such as object detection and semantic segmentation to understand the surrounding environment. At the decision-making level, rule-based methods or reinforcement learning algorithms are used to generate feasible driving strategies. However, these methods are often limited by data quality, model generalization ability, and vulnerability to adversarial attacks. Especially in complex scenarios such as urban intersections and highway merges, the performance of autonomous driving systems still needs to be improved.
[0004] Trajectory planning, as a key link in autonomous driving, its accuracy directly affects the driving safety and efficiency of vehicles. Traditional trajectory planning methods often rely on rules or optimization algorithms and are difficult to handle dynamically changing environments and complex driving scenarios. At the same time, although existing deep learning models have made certain progress in trajectory prediction and planning, their robustness under adversarial attacks is still insufficient, and they are easily affected by noise interference, resulting in performance degradation.
[0005] In summary, the existing autonomous driving technology in the prior art has problems of poor robustness and generalization ability in complex scenarios. Summary of the Invention
[0006] To solve the above problems, embodiments of the present application provide an autonomous driving adversarial training method and device to solve the problems of poor robustness and generalization ability of autonomous driving technology in complex scenarios existing in the prior art.
[0007] The technical solution of the present application is implemented as follows:
[0008] On the one hand, embodiments of the present application provide an autonomous driving adversarial training method, and the method includes:
[0009] Obtain multi-frame ego vehicle point cloud data;
[0010] Generate a feasible trajectory based on the ego vehicle point cloud data, and obtain occupancy information along the trajectory through trajectory resampling to generate a trajectory with occupancy information;
[0011] Classify the trajectory with occupancy information based on a truncated diffusion model;
[0012] Fuse noise with the classified trajectory for training to complete the training of the end-to-end autonomous driving model.
[0013] Optionally, the steps of generating a feasible trajectory based on the ego vehicle point cloud data and obtaining occupancy information along the trajectory through trajectory resampling to generate a trajectory with occupancy information include:
[0014] Convert the ego vehicle point cloud data of each frame from the vehicle coordinate system to the world coordinate system, and remove dynamic objects, only retaining the point cloud data of the static environment;
[0015] Extract the drivable area excluding obstacles based on the point cloud data of the static environment;
[0016] Under the dynamic constraints of the vehicle, generate multiple feasible trajectories based on the drivable area;
[0017] Resample the original ego vehicle point cloud data based on the generated multiple feasible trajectories, and extract the occupancy information along the trajectory to generate a trajectory with occupancy information.
[0018] Optionally, the steps of converting the ego vehicle point cloud data of each frame from the vehicle coordinate system to the world coordinate system and removing dynamic objects, only retaining the point cloud data of the static environment include:
[0019] Extract the position information in the ego vehicle point cloud data of each frame;
[0020] According to the formula
[0021] Pworld = Tvehicle→world * Pvehicled
[0022] Convert the position information to the world coordinate system; where, Pworld represents the converted world coordinate system, Tvehicle→world represents the transformation matrix, and Pvehicled represents the homogeneous coordinate determined according to the position information;
[0023] Classify the ego vehicle point cloud data after coordinate transformation according to semantic labels or classification algorithms, and identify the static environment and dynamic objects.
[0024] Optionally, the steps of extracting the drivable area excluding obstacles based on the point cloud data of the static environment include:
[0025] Divide the point cloud data of the static environment into multiple voxels, and each of the voxels marks a small cube;
[0026] Classify each of the voxels based on a classification function;
[0027] Extract the boundary of the drivable area based on the classification result.
[0028] Optionally, under the dynamic constraints of the vehicle, the steps of generating multiple feasible trajectories based on the drivable area include:
[0029] Randomly sample the starting and ending positions of the trajectory;
[0030] Combine a path planning algorithm to generate a trajectory connecting the starting and ending positions within the drivable area;
[0031] According to the dynamic constraints, screen out available paths from the generated trajectories, and represent the available paths using an attitude transformation matrix.
[0032] Optionally, the steps of resampling the original ego-vehicle point cloud data according to the generated multiple feasible trajectories, extracting the occupancy information along the trajectories, and generating trajectories with occupancy information include:
[0033] Calculate the nearest neighbor points of each trajectory point in the ego-vehicle point cloud data of the original point cloud;
[0034] For each of the trajectory points, judge whether the trajectory point is feasible based on the state of the nearest neighbor points to generate a trajectory with occupancy information.
[0035] Optionally, the steps of classifying the trajectories with occupancy information according to a truncated diffusion model include:
[0036] Extract a set of anchor points from the trajectories with occupancy information through K-Means clustering, and add Gaussian noise to each anchor point, where the anchor points represent the key features of the driving trajectory;
[0037] Gradually denoise the noisy anchor points until they are restored to be close to the real driving trajectory;
[0038] Use a diffusion decoder to take the denoised anchor points as input and output the predicted classification scores and denoised trajectories;
[0039] Evaluate the output of the diffusion decoder according to a total loss function, and optimize the parameters of the diffusion decoder according to the total loss function until the total loss value is less than a set value.
[0040] Optionally, the steps of training the classified trajectory fusion noise to complete the training of the end-to-end autonomous driving model include:
[0041] Inject noise into different modules of the autonomous driving system and define the results of the classified trajectory after injecting noise;
[0042] Determine the loss ratio of each module at the current time step according to the results after noise injection;
[0043] Update the normalized weight of each module according to the loss ratio;
[0044] Determine the total loss at the next time step according to the updated normalized weight, and optimize the working parameters of each module according to the gradient of the total loss and the normalized weight of each module until the training of the end-to-end autonomous driving model is completed.
[0045] Optionally, updating the normalized weight of each module satisfies the formula:
[0046] .
[0047] Where is the normalized weight of module j at time step t; is the average value of the loss ratios of all modules at time step t; Norm represents the normalization function; exp represents the exponential function with base e; K is the total number of modules, represents the loss ratio of module j at time step t; represents the loss ratio of module K at time step t.
[0048] On the other hand, the embodiment of the present application also provides an autonomous driving adversarial training device, and the device includes:
[0049] A data acquisition module for acquiring multiple frames of ego vehicle point cloud data;
[0050] A data processing module for generating a feasible trajectory according to the ego vehicle point cloud data, and obtaining occupancy information along the trajectory through trajectory resampling to generate a trajectory with occupancy information;
[0051] The data processing module is further configured to classify the trajectory with occupancy information according to the truncated diffusion model;
[0052] A training module for training the classified trajectory fusion noise to complete the training of the end-to-end autonomous driving model.
[0053] Compared with the prior art, the present application has the following technical effects:
[0054] An autonomous driving adversarial training method and device provided by this application first obtains multiple frames of ego vehicle point cloud data, then generates a feasible trajectory based on the ego vehicle point cloud data, and obtains occupancy information along the trajectory through trajectory resampling to generate a trajectory with occupancy information, and then classifies the trajectory with occupancy information according to a truncated diffusion model; finally, the classified trajectory is fused with noise for training to complete the training of the end-to-end autonomous driving model. Since in the autonomous driving adversarial training method provided by this application, a trajectory with occupancy information is generated by means of trajectory resampling, high-quality data input can be provided for subsequent trajectory classification and reconstruction. By classifying the trajectory, further processing of the trajectory is realized, ensuring the accuracy and diversity of the trajectory. Finally, by fusing the classified trajectory with noise for training, the generalization ability and robustness of the model are improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the embodiments of this specification. For those of ordinary skill in the art, other drawings can also be obtained according to these drawings.
[0056] Figure 1 It is a schematic diagram of the modules of the electronic device provided by the embodiment of this application.
[0057] Figure 2 It is an exemplary flowchart of the autonomous driving adversarial training method provided by the embodiment of this application.
[0058] Figure 3 It is an exemplary flowchart of the sub-steps of S104 provided by the embodiment of this application.
[0059] Figure 4 It is an exemplary flowchart of the sub-steps of S106 provided by the embodiment of this application.
[0060] Figure 5 It is an exemplary flowchart of the sub-steps of S108 provided by the embodiment of this application.
[0061] Figure 6 It is a schematic diagram of the modules of the autonomous driving adversarial training device provided by the embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0062] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Apparently, the described embodiments are some but not all of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the scope of protection of this application.
[0063] To facilitate the understanding of the embodiments of this application, the following will further explain and illustrate with specific embodiments in conjunction with the accompanying drawings. The embodiments do not limit the embodiments of this application.
[0064] As described in the background art, currently, there are problems with poor robustness and generalization ability in the complex scenarios of autonomous driving technology. In view of this, to solve this problem, this application provides an autonomous driving adversarial training method, which can improve the generalization ability and robustness of the model by enhancing the quality of input data, processing and optimizing the data, and performing adversarial noise fusion training.
[0065] It should be noted that the autonomous driving adversarial training method provided by this application can be applied to an electronic device. Optionally, Figure 1 FIG. 11 shows a schematic structural block diagram of an electronic device 100 provided by an embodiment of this application. The electronic device 100 includes a memory 102, a processor 101, and a communication interface 103. The memory 102, the processor 101, and the communication interface 103 are directly or indirectly electrically connected to each other to achieve data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses or signal lines.
[0066] The memory 102 can be used to store software programs and modules, such as program instructions or modules corresponding to the autonomous driving adversarial training device provided by an embodiment of this application. The processor 101 executes various functional applications and data processing by executing the software programs and modules stored in the memory 102, and then executes the steps of the positioning method provided by an embodiment of this application. The communication interface 103 can be used to communicate signaling or data with other node devices.
[0067] Among them, the memory 102 can be, but is not limited to, a Random Access Memory (RAM), a Read Only Memory (ROM), a Programmable Read-Only Memory 102 (PROM), an Erasable Programmable Read-Only Memory (EPROM), an Electric Erasable Programmable Read-Only Memory (EEPROM), etc.
[0068] The processor 101 can be an integrated circuit chip with signal processing capabilities. The processor 101 can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0069] It can be understood that Figure 1 the structure shown is only schematic, and the electronic device 100 can also include more or fewer components than Figure 1 shown therein, or have a configuration different from Figure 1 that shown. Figure 1 Each component shown therein can be implemented by hardware, software, or a combination thereof.
[0070] An exemplary description of the autonomous driving adversarial training method provided in this application will be given below. As an implementation, please refer to Figure 2 and the autonomous driving adversarial training method includes:
[0071] S102, obtaining multiple frames of ego-vehicle point cloud data.
[0072] S104, generating a feasible trajectory based on the ego-vehicle point cloud data, and obtaining occupancy information along the trajectory through trajectory resampling to generate a trajectory with occupancy information.
[0073] S106, classifying the trajectory with occupancy information according to the truncated diffusion model.
[0074] S108. Train the classified trajectory fusion noise to complete the training of the end-to-end autonomous driving model.
[0075] It should be noted that in end-to-end autonomous driving technology, the ego vehicle point cloud data refers to the three-dimensional environmental information centered on the ego vehicle collected by sensors such as lidar (LiDAR). These data are presented in the form of point clouds, and each point contains three-dimensional coordinates , and sometimes also contains information such as reflection intensity, etc., to support core functions such as perception, positioning, and planning through the ego vehicle point cloud data.
[0076] In S104, the trajectory resampling method is used to fill in the blanks in the original data, thereby providing high-quality input data for subsequent trajectory reconstruction and classification. It mainly transforms the ego vehicle point cloud data from the vehicle coordinate system to the world coordinate system through coordinate transformation, removes dynamic objects, only retains the static environment, then uses voxelization and semantic segmentation to extract the drivable area, and conducts path planning, considering the dynamic constraints of the vehicle, generates multiple feasible trajectories, and finally generates an occupancy grid by resampling the occupancy information along the trajectories.
[0077] The following is a detailed description of S104:
[0078] Among them, please refer to Figure 3 , S104 includes:
[0079] S1041. Convert the ego vehicle point cloud data of each frame from the vehicle coordinate system to the world coordinate system, and remove dynamic objects, only retaining the point cloud data of the static environment.
[0080] S1042. Extract the drivable area excluding obstacles based on the point cloud data of the static environment.
[0081] S1043. Generate multiple feasible trajectories according to the drivable area under the dynamic constraints of the vehicle.
[0082] S1044. Resample the original ego vehicle point cloud data according to the multiple generated feasible trajectories, extract the occupancy information along the trajectories, so as to generate a trajectory with occupancy information.
[0083] In S1041, first take the ego vehicle point cloud data as the original input, and process and apply it based on the ego vehicle point cloud data of each frame, and then extract the occupancy sequence. Among them, the ego vehicle data of each frame usually includes the position P vehicle , speed V vehicle and acceleration A vehicle , the speed V vehicle and acceleration A vehicle are generally related to the state of dynamic objects, while the position P vehicle is mainly related to the static environment.
[0084] After that, coordinate transformation of the point cloud is performed, that is, the point cloud data of the ego vehicle needs to be transformed from the vehicle coordinate system to the world coordinate system. The point cloud data of each frame is a set , where represents the spatial coordinates of the i-th point cloud.
[0085] In the specific coordinate transformation process, it is necessary to rely on the formula
[0086] P world =T vehicle→world *P vehicled
[0087] to transform the position information to the world coordinate system; where P world represents the transformed world coordinate system, T vehicle→world represents the transformation matrix, and P vehicled represents the homogeneous coordinates determined according to the position information. Generally, the transformation matrix T vehicle→world is a 4×4 transformation matrix. P vehicled is the matrix representation of the above set , and only need to expand each point cloud into homogeneous coordinates [xi, yi, zi, 1] T , and arrange all points in a row to form an n×4 matrix. Therefore, P vehicled can be represented as an n×4 matrix, and each row represents the homogeneous coordinates [xi, yi, zi, 1] of a point T .
[0088] After that, classify the point cloud data of the ego vehicle after coordinate transformation according to semantic labels or classification algorithms, and identify static environments and dynamic objects, and then the effect of removing dynamic objects and only retaining static environments can be achieved. For example, a classification function can be used to distinguish static environments and dynamic objects, and retain points (i.e., the point cloud of the static environment).
[0089] After obtaining the static environment, it is necessary to remove obstacles and obtain the drivable area. For example, the drivable area includes areas such as roads and parking spaces. Obstacles include dynamic obstacles such as pedestrians or static obstacles such as roadblocks.
[0090] Specifically, first, divide the point cloud data of the static environment into multiple voxels, and each of the voxels marks a small cube.
[0091] For a given point cloud , it can be divided into N voxels through the voxelization process , the size of the voxel is predetermined, and each voxel represents a small area, and the size can be defined by the voxel size parameter v.
[0092] 。
[0093] Among them, the voxel size is the set voxel dimension, and P world is the point cloud after coordinate transformation.
[0094] After that, each voxel is classified to determine whether it belongs to categories such as roads, obstacles, or sidewalks. The classification function is used to assign a label to each voxel V i :
[0095] 。
[0096] Through semantic segmentation, a semantic label is assigned to each voxel Vi, and drivable areas such as roads and parking spaces can be identified, and obstacles are removed. For example, only voxels classified as roads will be marked as drivable areas, while voxels assigned categories such as sidewalks and obstacles are marked as non-drivable areas.
[0097] It can be understood that based on the classification results, non-drivable areas can be removed, and only drivable areas are retained, and then the boundaries of the drivable areas can be extracted. For example, the drivable area consists of a set of voxels and then the bounding box of these voxels can be extracted to form the topological structure of the drivable area.
[0098] After that, under dynamic constraints, diverse and feasible trajectories are generated based on the drivable area, that is, multiple paths are generated within the drivable area.
[0099] The way to generate the trajectory is as follows: first select a starting point and a target point in the drivable area, and this target point is the end position.
[0100] After that, a smooth path planning algorithm ( algorithm) is used to generate a trajectory connecting the starting point and the target point within the drivable area. These algorithms are based on graph search and heuristic search and can find the shortest path between each node (voxel) in the graph. The goal of the algorithm is to minimize the cost function:
[0101] 。
[0102] Among them, is the actual cost (path length) from the starting point to the current node; is the heuristic estimated cost (Euclidean distance) from the current node to the target node.
[0103] The finally generated trajectory is a series of points , each representing a position in the trajectory.
[0104] It should be noted that in the generated trajectory, the available paths need to be screened in combination with the dynamic constraints of the vehicle, such as the maximum speed , the maximum acceleration and the maximum steering angle . By restricting the movement of the vehicle on the trajectory, it is ensured that the trajectory is feasible.
[0105] .
[0106] where is the speed of the i-th point in the trajectory; is the acceleration; is the steering angle of the vehicle.
[0107] After determining the available paths, a series of pose transformation matrices can be used to represent the available paths, where each T i is a 4×4 transformation matrix used to describe the motion state of the vehicle on the trajectory.
[0108] .
[0109] where, R i is the rotation matrix, t i is the translation vector, describing the pose of the vehicle at the trajectory point Pi.
[0110] Finally, the original ego-vehicle point cloud data is resampled to extract the occupancy information along the trajectory, achieving the effect of improving the quality of the input data. Specifically, when resampling the trajectory points, the nearest neighbor points of each trajectory point Pi in the original point cloud need to be calculated. Data structures such as KD-tree or Ball-tree can be used to accelerate the nearest neighbor search. Set a radius r, and find the point closest to the trajectory point Pi through the nearest neighbor search and calculate its occupancy information:
[0111] .
[0112] where, represents the point P world in the original point cloud that is closest to the point Pi and has a distance less than r, and P j .
[0113] And, for each trajectory point Pi, based on the nearest neighbor point P jDetermine whether the state judgment trajectory point Pi is feasible, and generate an occupancy grid (or occupancy voxel). The occupancy grid can be represented as a two-dimensional or three-dimensional array, where each grid cell represents whether it is occupied.
[0114] Among them, if the trajectory point P i position is occupied by an obstacle, then O(P i ) = 1; if the trajectory point P i position is free, then O(P i ) = 0.
[0115] Finally, obtain the occupancy rate information along the trajectory to fill in the blanks in the original data, and solve the problems of unbalanced actual collected data (certain types of movements are more common than others) and limited diversity (lack of trajectory data under different conditions in the same scene).
[0116] As an implementation method, please refer to Figure 4 , S106 includes:
[0117] S1061, extract a set of anchor points from the trajectory with occupancy information through K-Means clustering, and add Gaussian noise to each anchor point, where the anchor points represent the key features of the driving trajectory.
[0118] S1062, gradually denoise the noisy anchor points until they are restored to be close to the real driving trajectory.
[0119] S1063, use the diffusion decoder to take the denoised anchor points as input, and output the predicted classification scores and the denoised trajectory.
[0120] S1064, evaluate the output of the diffusion decoder according to the total loss function, and optimize the parameters of the diffusion decoder according to the total loss function until the total loss value is less than the set value.
[0121] Among them, after obtaining the trajectory with occupancy information, it is necessary to classify the trajectory with occupancy information to ensure the accuracy and diversity of the trajectory.
[0122] During the processing, it is first necessary to create anchor points, where the anchor points are extracted from the trajectory with occupancy information output in step S104, and the anchor points are used to generate noisy trajectory data and serve as the input of the truncated diffusion model.
[0123] As an implementation method, this application extracts a set of "anchor points" from the trajectory with occupancy information through K-Means clustering , and these anchor points represent some key features of the driving trajectory. To better simulate the changes that may be encountered in actual driving, for each anchor point Add some Gaussian noise, i.e., random perturbations, to simulate unpredictable changes in the real environment.
[0124] After that, to limit the noise within a suitable range, a "truncated diffusion process" is used to gradually denoise. Specifically, starting from these noisy anchor points, the denoising operation is carried out step by step until it is restored to be close to the real driving trajectory. The mathematical form of this process is as follows:
[0125] 。
[0126] Where, represents the k-th trajectory after the i-th step of denoising, is the truncation coefficient at the i-th step, is the anchor point, is the noise sampled from the standard normal distribution; , is the number of truncated diffusion steps to avoid excessive removal of noise, and 。
[0127] It should be noted that through the processing of truncated diffusion, although the trajectory corresponding to the anchor point is close to the real driving trajectory, there is still noise, but the noise is within a controllable range.
[0128] After that, the denoised anchor points are used as inputs by the diffusion decoder, and the predicted classification scores and denoised trajectories are output. Specifically, during the training process, the diffusion decoder will use the noisy trajectory (i.e., the anchor point with added noise) as the input, and output the predicted classification score and the denoised trajectory :
[0129] 。
[0130] Where z represents conditional information. The purpose of the model is to make the decoded trajectory as close as possible to the real trajectory and make the classification prediction accurate.
[0131] At the same time, to improve the accuracy of classification prediction, this application also evaluates the output of the diffusion decoder according to the total loss function and optimizes the parameters of the diffusion decoder according to the total loss function.
[0132] Specifically, the loss in the training process includes two parts. One is the trajectory reconstruction loss, and the goal is to make the denoised trajectory as similar as possible to the real trajectory; the other is the classification loss. To make the model accurately distinguish positive samples and negative samples, the binary cross-entropy loss BCE is introduced. We use the noisy trajectory corresponding to the ground truth trajectory closest to the anchor point as the positive sample ( ), and the others are used as negative samples ( ).
[0133] .
[0134] Among them, represents the trajectory reconstruction loss, BCE is the binary cross-entropy loss, represents the coefficient used to balance these two losses, L represents the total loss value, represents the ground truth trajectory, that is, the trajectory that the vehicle actually travels in the real driving scenario.
[0135] By determining the total loss value, the performance of the diffusion decoder in the current training state can be evaluated. Moreover, the smaller the total loss value, the closer the prediction result of the model is to the true value. Therefore, the goal of training is to make the value of L as low as possible. The parameters of the diffusion decoder can be optimized according to the total loss value, and finally the total loss value is made less than the set value.
[0136] Finally, the classified trajectory is fused with noise for training, and then through adversarial noise fusion training, the generalization ability and robustness of the model are further improved.
[0137] Among them, please refer to Figure 5 , S108 includes:
[0138] S1081, Inject noise into different modules of the autonomous driving system and define the results of the classified trajectory after injecting noise.
[0139] S1082, Determine the loss ratio of each module at the current time step according to the results after noise injection.
[0140] S1083, Update the normalized weight of each module according to the loss ratio.
[0141] S1084, Determine the total loss at the next time step according to the updated normalized weight, and optimize the working parameters of each module according to the gradient of the total loss and the normalized weight of each module until the training of the end-to-end autonomous driving model is completed.
[0142] In the autonomous driving system, there are multiple functional modules, including a trajectory formation module, a map formation module, a motion prediction module, and a planning module. When injecting noise, for the training model, it is guided by the overall goal rather than the loss of each independent module. Therefore, noise needs to be injected into the inputs of different modules. This method ensures that noise is generated from the overall view of the model, that is, using the overall loss for backpropagation, rather than focusing on the individual module losses that may be contradictory and have a negative impact on the robustness of the overall decision-making.
[0143] It should be noted that the goal of injecting noise in this step is different from that in S106. In S106, the introduction of noise is to simulate the uncertainty in the real environment, and denoising and trajectory reconstruction are carried out through the truncated diffusion model. That is, by adding Gaussian noise to the anchor points, the unpredictable changes that may be encountered in the real driving environment are simulated, and the truncated diffusion model is used to gradually denoise to recover the trajectory data close to the real driving trajectory.
[0144] In this step, however, the introduction of noise is to improve the robustness of the model under adversarial attacks and optimize the overall performance of the model through adversarial noise fusion training. That is, adversarial noise is injected into different end-to-end modules (trajectory formation module, map formation module, motion prediction module, planning module). By optimizing the amount of noise injection, the model can still maintain stable performance under the influence of adversarial noise.
[0145] In comparison, the noise in S106 is to simulate the random interference in the real environment and is a kind of non-targeted noise. The noise in this step, on the other hand, is to simulate the interference in adversarial attacks or complex environments and is a kind of targeted noise.
[0146] Moreover, after injecting the noise, it is necessary to find the optimal amount of noise injection. Therefore, in the implementation process, the model output is first defined, which represents the result after injecting the noise set on the input data
[0147] and is specifically implemented through the following function combination:
[0148] .
[0149] Among them, , , respectively represent the specific adversarial perturbations injected into the m-th perception module, the k-th prediction module, and the planning module. Xn represents the output data, such as sensor data (point cloud, image), vehicle state (speed, position), etc. Planner represents the planning module, Predictor represents the prediction module, and Perceiver represents the perception module.
[0150] The amount of noise injection satisfies the formula, that is:
[0151] .
[0152] Among them, represents the optimal combination of adversarial noise parameters obtained by maximizing the total loss function during the noise injection process; C is the constraint set of the noise; is the total loss function; is the true label.
[0153] To manage the different contributions of modules during training, dynamic weight accumulation adaptation is introduced, which adaptively adjusts the loss weight of each module to the overall objective according to its contribution during the noise injection process. This method introduces a normalization weight function to eliminate the dimensional difference, accelerate the model convergence speed, reduce the impact of outliers on model training, and improve its stability and model generalization ability.
[0154] To extend the concept of multi-tasking to multiple modules, the loss of each module at the current time step t The ratio relative to the previous value is calculated as:
[0155] .
[0156] where is the loss of module j at time step t - 2; is the loss of module j at time step t - 1; is the loss ratio of module j at time step t.
[0157] Then, based on these loss ratios, the normalization weight formula is used to update the weights:
[0158] .
[0159] where is the normalized weight of module j at time step t; is the average of the loss ratios of all modules at time step t; Norm represents the normalization function; exp represents the exponential function with base e; N is the total number of modules; represents the loss ratio of module j at time step t; represents module 's loss ratio at time step t.
[0160] It should be noted that in the above formula, the numerator part represents the loss ratio of the j-th module at time step t, and only calculates the normalized exponential value for module j. The denominator needs to iterate over all modules, so the summation index variable k is introduced. j is the external index representing the target module for which the weight is to be calculated currently. k is the internal index, only used to iterate over all modules and sum.
[0161] Finally, the total loss at time step t + 1 is calculated based on the updated weights:
[0162] .
[0163] This method can ensure that the weights dynamically adapt to the performance of each module over time, thus improving stability and overall performance.
[0164] It should be noted that in the steps of S1081 to S1084, the optimal amount of noise injection is achieved through dynamic adaptive noise optimization. That is, the amount of noise injection is not fixed in advance, but is gradually adjusted through multiple iterations (time steps). After each noise injection, the weight is dynamically calculated according to the loss ratio (degree of performance degradation) of the module, and then the weight affects the noise injection intensity of the next time step. By forming a closed loop between noise injection and weight update, the model can automatically learn the sensitivity of different modules to noise during training and allocate the optimal anti-noise resources. Specifically, after determining the normalized weight of each module, the module with a higher weight is more likely to be noticed, which is equivalent to indirectly adjusting the priority of noise injection - the module with a higher weight may be assigned more stringent adversarial noise constraints in subsequent training. Moreover, the parameters (neural network weights) of each module are optimized through the backpropagation of the gradient of the total loss, rather than being manually set. During backpropagation, the gradient is distributed to different modules according to the weight, thereby adjusting its internal parameters, ultimately achieving the optimal noise injection, and completing the training of the model based on this input.
[0165] Generally speaking, for the autonomous driving adversarial training method provided in this application, on the one hand, by adopting the trajectory resampling strategy, the gaps in the original data are effectively filled, and the problems of data imbalance and limited diversity are solved. At the same time, using the truncated diffusion model for trajectory reconstruction and classification optimization, and adding adversarial noise fusion training, enable the autonomous driving system to more accurately identify road information, predict the movement trajectories of other vehicles and pedestrians, and generate safe and efficient driving paths when facing complex and changeable traffic environments, thereby significantly improving the robustness and accuracy of the autonomous driving system and reducing the risk of traffic accidents.
[0166] On the other hand, noise is injected into different modules in an end-to-end manner, and backpropagation is carried out under the guidance of the overall objective to optimize the amount of noise injection. At the same time, the dynamic weight accumulation adaptive method is adopted to adaptively adjust the loss weight of each module to the overall objective according to its contribution during the noise injection process. This comprehensive training strategy not only accelerates the convergence speed of the model but also enables the model to find the optimal solution faster during training. In addition, by eliminating the dimension difference through the normalized weight function, the stability and efficiency of training are further improved.
[0167] On the third hand, by generating diverse and feasible trajectories, and using the truncated diffusion model for denoising and classification optimization, the model can learn richer driving scenarios and trajectory features. At the same time, the introduction of adversarial noise fusion training enables the model to maintain stable performance when facing unknown or extreme situations. Significantly enhances the generalization ability and adaptability of the model, enabling the autonomous driving system to better adapt to different traffic environments and driving scenarios.
[0168] Based on the above implementation, an embodiment of the present application further provides an autonomous driving adversarial training device 200. Please refer to Figure 6 , the device includes:
[0169] A data acquisition module 210, configured to acquire multiple frames of ego vehicle point cloud data.
[0170] It can be understood that through the data acquisition module 210, the above S102 can be executed.
[0171] A data processing module 220, configured to generate a feasible trajectory based on the ego vehicle point cloud data, and obtain occupancy information along the trajectory through trajectory resampling to generate a trajectory with occupancy information.
[0172] It can be understood that through the data processing module 220, the above S104 can be executed.
[0173] The data processing module 220 is further configured to classify the trajectory with occupancy information based on a truncated diffusion model.
[0174] It can be understood that through the data processing module 220, the above S106 can be executed.
[0175] A training module 230, configured to train the classified trajectory by fusing noise to complete the training of the end-to-end autonomous driving model.
[0176] It can be understood that through the training module 230, the above S108 can be executed.
[0177] In summary, an autonomous driving adversarial training method and device provided by the present application first acquire multiple frames of ego vehicle point cloud data, then generate a feasible trajectory based on the ego vehicle point cloud data, and obtain occupancy information along the trajectory through trajectory resampling to generate a trajectory with occupancy information. Then, classify the trajectory with occupancy information based on a truncated diffusion model; finally, train the classified trajectory by fusing noise to complete the training of the end-to-end autonomous driving model. Since in the autonomous driving adversarial training method provided by the present application, a trajectory with occupancy information is generated by using the trajectory resampling method, high-quality data input can be provided for subsequent trajectory classification and reconstruction. By classifying the trajectory, further processing of the trajectory is realized, ensuring the accuracy and diversity of the trajectory. Finally, by training the classified trajectory by fusing noise, the effect of improving the generalization ability and robustness of the model is achieved.
[0178] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of the devices, methods, and computer program products according to the embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function.
[0179] It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved.
[0180] It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0181] In addition, the various functional modules in the embodiments of the present application can be integrated together to form an independent part, or each module can exist separately, or two or more modules can be integrated to form an independent part.
[0182] If the above functions are implemented in the form of software functional modules and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories, random access memories, magnetic disks, or optical discs that can store program codes.
[0183] Although the present invention is disclosed as above, the present invention is not limited thereto. Any person skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention should be determined by the scope defined by the claims.
Claims
1. An autonomous driving adversarial training method, characterized in that, The method includes: Obtain multiple frames of ego-vehicle point cloud data; Generate a feasible trajectory based on the ego-vehicle point cloud data, and obtain the occupancy information along the trajectory through trajectory resampling to generate a trajectory with occupancy information; Classify the trajectory with occupancy information based on a truncated diffusion model; Fuse noise into the classified trajectory for training to complete the training of the end-to-end autonomous driving model; The steps of generating a feasible trajectory based on the ego-vehicle point cloud data and obtaining the occupancy information along the trajectory through trajectory resampling to generate a trajectory with occupancy information include: Convert the ego-vehicle point cloud data of each frame from the vehicle coordinate system to the world coordinate system, and remove dynamic objects, only retaining the point cloud data of the static environment; Extract the drivable area excluding obstacles based on the point cloud data of the static environment; Generate multiple feasible trajectories based on the drivable area under the dynamic constraints of the vehicle; Calculate the nearest neighbor points of each trajectory point in the ego-vehicle point cloud data of the original point cloud; For each of the trajectory points, determine whether the trajectory point is feasible based on the state of the nearest neighbor points to generate a trajectory with occupancy information; The steps of fusing noise into the classified trajectory for training to complete the training of the end-to-end autonomous driving model include: Inject noise into different modules of the autonomous driving system and define the results of the classified trajectory after noise injection; Determine the loss ratio of each module at the current time step based on the results after noise injection; Update the normalized weights of each module according to the loss ratio; Determine the total loss at the next time step based on the updated normalized weights, and optimize the working parameters of each module according to the gradient of the total loss and the normalized weights of each module until the training of the end-to-end autonomous driving model is completed.
2. The autonomous driving adversarial training method according to claim 1, wherein, The steps of converting the ego-vehicle point cloud data of each frame from the vehicle coordinate system to the world coordinate system and removing dynamic objects, only retaining the point cloud data of the static environment include: Extract the position information in the ego-vehicle point cloud data of each frame; According to the formula P world =T vehicle→world *P vehicled Convert the position information to the world coordinate system; where P world represents the converted world coordinate system, T vehicle→world represents the transformation matrix, P vehicled represents the homogeneous coordinates determined according to the position information; Classify the ego-vehicle point cloud data after coordinate transformation based on semantic labels or classification algorithms, and identify the static environment and dynamic objects.
3. The automatic driving adversarial training method according to claim 1, wherein The steps of extracting the drivable area excluding obstacles based on the point cloud data of the static environment include: Divide the point cloud data of the static environment into multiple voxels, and mark each voxel with a small cube; Classify each of the voxels based on a classification function; Extract the boundary of the drivable area based on the classification results.
4. The autonomous driving adversarial training method according to claim 1, wherein The steps of generating multiple feasible trajectories based on the drivable area under the dynamic constraints of the vehicle include: Randomly sample the starting and ending positions of the trajectory; Combine a path planning algorithm to generate a trajectory connecting the starting and ending positions within the drivable area; According to the dynamic constraints, screen out the available paths from the generated trajectories and represent the available paths using an attitude transformation matrix.
5. The autonomous driving adversarial training method according to claim 1, wherein The steps of classifying the trajectory with occupancy information based on a truncated diffusion model include: Extract a set of anchor points from the trajectory with occupancy information through K-Means clustering, and add Gaussian noise to each anchor point, where the anchor points represent the key features of the driving trajectory; Gradually denoise the noisy anchor points until they are restored to be close to the true driving trajectory; Use a diffusion decoder to take the denoised anchor points as input and output the predicted classification scores and the denoised trajectory; Evaluate the output of the diffusion decoder according to the total loss function, and optimize the parameters of the diffusion decoder according to the total loss function until the total loss value is less than the set value.
6. The method for autonomous driving adversarial training according to claim 1, wherein The normalized weights of each module satisfy the formula: wherein, is the normalized weight of module j at time step t; is the average of the loss ratios of all modules at time step t; Norm represents the normalization function; exp represents the exponential function with base e; N is the total number of modules, represents the loss ratio of module j at time step t; represents the loss ratio of module K at time step t.
7. An autonomous driving adversarial training device, characterized in that, For implementing the method according to any one of claims 1 to 6, the device includes: A data acquisition module for acquiring multi-frame ego vehicle point cloud data; A data processing module for generating a feasible trajectory based on the ego vehicle point cloud data, and obtaining the occupancy information along the trajectory through trajectory resampling to generate a trajectory with occupancy information; The data processing module is further configured to classify the trajectory with occupancy information according to the truncated diffusion model; A training module for training the classified trajectory by fusing noise to complete the training of the end-to-end autonomous driving model.
Citation Information
Patent Citations
Emergency planning and security assurance
CN114175023A
Vehicle trajectory planning method, device, equipment and medium
CN116520839A