Automatic driving confrontation training method and device
By combining trajectory resampling and truncating diffusion models, we generate and classify trajectories with occupied information, and improve the robustness and generalization capabilities of the autonomous driving model through noise training, solving the problem of insufficient robustness and generalization capabilities in complex scenarios in the prior art.
Patent Information
- Application Number
- CN202510551073.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-29
AI Technical Summary
The existing autonomous driving technology has poor robustness and generalization capabilities in complex scenarios, making it difficult to deal with variable traffic scenarios and confrontational attacks.
Trajectory resampling is used to generate trajectories with occupancy information, and classify them by truncating diffusion models. Finally, the classified trajectory is fused with noise for training to improve the generalization ability and robustness of the end-to-end autonomous driving model.
By improving the quality of input data and further processing of trajectories, the accuracy and diversity of trajectories are ensured, the robustness and generalization capabilities of the autonomous driving system are significantly improved, and the vulnerability of adversarial attacks is reduced.
Smart Images

Figure CN120071305A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of autonomous driving, and in particular, to a method and device for autonomous driving adversarial training. Background Art
[0002] With the rapid development of artificial intelligence technology, autonomous driving technology has become a research hotspot in the fields of the automotive industry and intelligent transportation systems. The core of autonomous driving technology lies in achieving autonomous navigation and safe driving of vehicles in complex road environments through key links such as perception, decision-making, and control. However, in the face of changing traffic scenarios, uncertain dynamic obstacles, and complex road structures, how to ensure the robustness and safety of autonomous driving systems is a key problem that needs to be solved urgently. Especially in dealing with massive point cloud data, performing efficient path planning, and coping with adversarial attacks, existing autonomous driving technologies still face many challenges.
[0003] Currently, autonomous driving technology mainly relies on technical means such as deep learning, computer vision, and sensor fusion. At the perception level, environmental information is obtained through sensors such as lidar and cameras, and deep learning algorithms are used to perform tasks such as object detection and semantic segmentation to understand the surrounding environment. At the decision-making level, rule-based methods or reinforcement learning algorithms are used to generate feasible driving strategies. However, these methods are often limited by data quality, model generalization ability, and vulnerability to adversarial attacks. Especially in complex scenarios such as urban intersections and highway merges, the performance of autonomous driving systems still needs to be improved.
[0004] Trajectory planning, as a key link in autonomous driving, its accuracy directly affects the driving safety and efficiency of vehicles. Traditional trajectory planning methods often rely on rules or optimization algorithms and are difficult to handle dynamic environments and complex driving scenarios. At the same time, although existing deep learning models have made certain progress in trajectory prediction and planning, their robustness under adversarial attacks is still insufficient, and they are vulnerable to noise interference, resulting in performance degradation.
[0005] In summary, the existing autonomous driving technologies in the prior art have problems of poor robustness and generalization ability in complex scenarios. Summary of the Invention
[0006] To solve the above problems, embodiments of the present application provide a method and device for autonomous driving adversarial training to solve the problems of poor robustness and generalization ability of existing autonomous driving technologies in complex scenarios.
[0007] The technical solution of the present application is implemented as follows: On the one hand, embodiments of the present application provide an autonomous driving adversarial training method, the method comprising: Obtaining multiple frames of ego-vehicle point cloud data; Generate a feasible trajectory based on the ego vehicle point cloud data, and obtain the occupancy information along the trajectory through trajectory resampling to generate a trajectory with occupancy information. Classify the trajectory with occupancy information according to the truncated diffusion model. Fuse noise into the classified trajectories for training to complete the training of the end-to-end autonomous driving model.
[0008] Optionally, the steps of generating a feasible trajectory based on the ego vehicle point cloud data and obtaining the occupancy information along the trajectory through trajectory resampling to generate a trajectory with occupancy information include: Convert the ego vehicle point cloud data of each frame from the vehicle coordinate system to the world coordinate system, and remove dynamic objects, only retaining the point cloud data of the static environment. Extract the drivable area excluding obstacles based on the point cloud data of the static environment. Under the dynamic constraints of the vehicle, generate multiple feasible trajectories based on the drivable area. Resample the original ego vehicle point cloud data according to the generated multiple feasible trajectories, and extract the occupancy information along the trajectory to generate a trajectory with occupancy information.
[0009] Optionally, the steps of converting the ego vehicle point cloud data of each frame from the vehicle coordinate system to the world coordinate system and removing dynamic objects, only retaining the point cloud data of the static environment include: Extract the position information in the ego vehicle point cloud data of each frame. According to the formula Pworld=Tvehicle→world*Pvehicled Convert the position information to the world coordinate system; where Pworld represents the converted world coordinate system, Tvehicle→world represents the transformation matrix, and Pvehicled represents the homogeneous coordinate determined according to the position information. Classify the ego vehicle point cloud data after coordinate transformation according to semantic labels or classification algorithms, and identify the static environment and dynamic objects.
[0010] Optionally, the steps of extracting the drivable area excluding obstacles based on the point cloud data of the static environment include: Divide the point cloud data of the static environment into multiple voxels, and each voxel marks a small cube. Classify each voxel based on the classification function. Extract the boundary of the drivable area based on the classification results.
[0011] Optionally, the steps of generating multiple feasible trajectories based on the drivable area under the dynamic constraints of the vehicle include: The starting and ending positions of the randomly sampled trajectory; Combined with a path planning algorithm, generate a trajectory connecting the starting and ending positions within the drivable area; According to the dynamic constraints, filter out the available paths from the generated trajectories, and represent the available paths using an attitude transformation matrix.
[0012] Optionally, the steps of resampling the original ego-vehicle point cloud data according to the generated multiple feasible trajectories, extracting the occupancy information along the trajectories, and generating trajectories with occupancy information include: Calculate the nearest neighbor points of each trajectory point in the ego-vehicle point cloud data of the original point cloud; For each of the trajectory points, judge whether the trajectory point is feasible based on the state of the nearest neighbor points, so as to generate a trajectory with occupancy information.
[0013] Optionally, the steps of classifying the trajectories with occupancy information according to a truncated diffusion model include: Extract a set of anchor points from the trajectories with occupancy information through K-Means clustering, and add Gaussian noise to each anchor point, where the anchor points represent the key features of the driving trajectory; Gradually denoise the noisy anchor points until they are restored to be close to the real driving trajectory; Use a diffusion decoder to take the denoised anchor points as input, and output the predicted classification scores and denoised trajectories; Evaluate the output of the diffusion decoder according to the total loss function, and optimize the parameters of the diffusion decoder according to the total loss function until the total loss value is less than the set value.
[0014] Optionally, the steps of training the end-to-end autonomous driving model by fusing the classified trajectories with noise include: Inject noise into different modules of the autonomous driving system, and define the results after injecting noise into the classified trajectories; Determine the loss ratio of each module at the current time step according to the results after noise injection; Update the normalized weights of each module according to the loss ratio; Determine the total loss at the next time step according to the updated normalized weights, and optimize the working parameters of each module according to the gradient of the total loss and the normalized weights of each module until the training of the end-to-end autonomous driving model is completed.
[0015] Optionally, updating the normalized weights of each module satisfies the formula: .
[0016] Where, is the normalized weight of module j at time step t; is the average of the loss ratios of all modules at time step t; Norm represents the normalization function; exp represents the exponential function with base e; K is the total number of modules, represents the loss ratio of module j at time step t; represents the loss ratio of module K at time step t.
[0017] On the other hand, the embodiment of the present application also provides an autonomous driving adversarial training device, and the device includes: A data acquisition module, configured to acquire multiple frames of ego vehicle point cloud data; A data processing module, configured to generate a feasible trajectory based on the ego vehicle point cloud data, and obtain occupancy information along the trajectory through trajectory resampling to generate a trajectory with occupancy information; The data processing module is further configured to classify the trajectory with occupancy information based on a truncated diffusion model; A training module, configured to train the classified trajectory by fusing noise to complete the training of the end-to-end autonomous driving model.
[0018] Compared with the prior art, the present application has the following technical effects: An autonomous driving adversarial training method and device provided by the present application first acquire multiple frames of ego vehicle point cloud data, then generate a feasible trajectory based on the ego vehicle point cloud data, and obtain occupancy information along the trajectory through trajectory resampling to generate a trajectory with occupancy information, and then classify the trajectory with occupancy information based on a truncated diffusion model; finally, train the classified trajectory by fusing noise to complete the training of the end-to-end autonomous driving model. Since in the autonomous driving adversarial training method provided by the present application, a trajectory with occupancy information is generated by using the trajectory resampling method, high-quality data input can be provided for subsequent trajectory classification and reconstruction. By classifying the trajectory, further processing of the trajectory is realized, ensuring the accuracy and diversity of the trajectory. Finally, by training the classified trajectory by fusing noise, the generalization ability and robustness of the model are improved. Description of the Drawings
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present specification or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the embodiments of the present specification. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.
[0020] Figure 1 It is a schematic diagram of the modules of the electronic device provided by the embodiment of the present application.
[0021] Figure 2 Exemplary flowchart of the autonomous driving adversarial training method provided by an embodiment of this application.
[0022] Figure 3 Exemplary flowchart of the sub-steps of S104 provided by an embodiment of this application.
[0023] Figure 4 Exemplary flowchart of the sub-steps of S106 provided by an embodiment of this application.
[0024] Figure 5 Exemplary flowchart of the sub-steps of S108 provided by an embodiment of this application.
[0025] Figure 6 Schematic diagram of the modules of the autonomous driving adversarial training device provided by an embodiment of this application. Detailed implementation manners
[0026] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are some but not all of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of this application without creative efforts shall fall within the protection scope of this application.
[0027] For ease of understanding of the embodiments of this application, the following will further explain with specific embodiments with reference to the accompanying drawings. The embodiments do not constitute a limitation on the embodiments of this application.
[0028] As described in the background art, currently, the autonomous driving technology has problems of poor robustness and generalization ability in complex scenarios. In view of this, to solve this problem, this application provides an autonomous driving adversarial training method, which improves the generalization ability and robustness of the model by improving the quality of input data, processing and optimizing the data, and performing adversarial noise fusion training.
[0029] It should be noted that the autonomous driving adversarial training method provided by this application can be applied to an electronic device. Optionally, Figure 1 FIG. shows a schematic structural block diagram of an electronic device 100 provided by an embodiment of this application. The electronic device 100 includes a memory 102, a processor 101, and a communication interface 103. The memory 102, the processor 101, and the communication interface 103 are directly or indirectly electrically connected to each other to realize data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses or signal lines.
[0030] The memory 102 can be used to store software programs and modules, such as the program instructions or modules corresponding to the automatic driving adversarial training device provided in the embodiments of the present application. The processor 101 executes various functional applications and data processing by executing the software programs and modules stored in the memory 102, and then executes the steps of the positioning method provided in the embodiments of the present application. The communication interface 103 can be used for signaling or data communication with other node devices.
[0031] Among them, the memory 102 can be, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory 102 (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc.
[0032] The processor 101 can be an integrated circuit chip with signal processing capabilities. The processor 101 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0033] It can be understood that Figure 1 the structure shown is only schematic, and the electronic device 100 may also include more or fewer components than those shown Figure 1 in the figure, or have a different configuration from that shown Figure 1 in the figure. Figure 1 Each component shown in the figure can be implemented by hardware, software, or a combination thereof.
[0034] The following provides an exemplary description of the automatic driving adversarial training method provided in the present application. As an implementation manner, please refer to Figure 2 , the automatic driving adversarial training method includes: S102. Obtain multiple frames of ego-vehicle point cloud data.
[0035] S104. Generate a feasible trajectory based on the ego-vehicle point cloud data, and obtain the occupancy information along the trajectory through trajectory resampling to generate a trajectory with occupancy information.
[0036] S106. Classify the trajectory with occupancy information according to the truncated diffusion model.
[0037] S108. Train the classified trajectory by fusing noise to complete the training of the end-to-end autonomous driving model.
[0038] It should be noted that in end-to-end autonomous driving technology, ego-vehicle point cloud data refers to the three-dimensional environmental information centered on the ego-vehicle collected by sensors such as lidar (LiDAR). These data are presented in the form of point clouds, and each point contains three-dimensional coordinates , and sometimes also contains information such as reflection intensity, etc., which supports core functions such as perception, positioning, and planning through ego-vehicle point cloud data.
[0039] In S104, the trajectory resampling method is used to fill in the blanks in the original data, thereby providing high-quality input data for subsequent trajectory reconstruction and classification. It mainly transforms the ego-vehicle point cloud data from the vehicle coordinate system to the world coordinate system through coordinate transformation, removes dynamic objects, and only retains the static environment. Then, it uses voxelization and semantic segmentation to extract the drivable area, and conducts path planning, considering the dynamic constraints of the vehicle, to generate multiple feasible trajectories. Finally, it resamples the occupancy information along the trajectory to generate an occupancy grid.
[0040] The following provides a detailed description of S104: Among them, please refer to Figure 3 , S104 includes: S1041. Transform the ego-vehicle point cloud data of each frame from the vehicle coordinate system to the world coordinate system, and remove dynamic objects, only retaining the point cloud data of the static environment.
[0041] S1042. Extract the drivable area excluding obstacles based on the point cloud data of the static environment.
[0042] S1043. Generate multiple feasible trajectories based on the drivable area under the dynamic constraints of the vehicle.
[0043] S1044. Resample the original ego-vehicle point cloud data according to the generated multiple feasible trajectories, extract the occupancy information along the trajectory, and generate a trajectory with occupancy information.
[0044] In S1041, first, the ego vehicle point cloud data is used as the original input, and processed and utilized based on each frame of the ego vehicle point cloud data to extract the occupancy sequence. Among them, the ego vehicle data of each frame usually includes the position P vehicle , speed V vehicle and acceleration A vehicle , speed V vehicle and acceleration A vehicle are generally related to the state of dynamic objects, while the position P vehicle is mainly associated with the static environment.
[0045] After that, coordinate transformation of the point cloud is performed, that is, the ego vehicle point cloud data needs to be transformed from the vehicle coordinate system to the world coordinate system. The point cloud data of each frame is a set , where represents the spatial coordinates of the i-th point cloud.
[0046] In the specific coordinate transformation process, it is necessary to transform the position information to the world coordinate system according to the formula P world =T vehicle→world *P vehicled ; where, P world represents the transformed world coordinate system, T vehicle→world represents the transformation matrix, and P vehicled represents the homogeneous coordinate determined according to the position information. Generally, the transformation matrix T vehicle→world is a 4×4 transformation matrix. P vehicled is the matrix representation of the above set , and only need to expand each point cloud into the homogeneous coordinate [xi, yi, zi, 1] T , and arrange all the points in a row to form an n×4 matrix. Therefore, P vehicled can be represented as an n×4 matrix, and each row represents the homogeneous coordinate [xi, yi, zi, 1] T .
[0047] After that, classify the ego vehicle point cloud data after coordinate transformation according to semantic labels or classification algorithms, and identify the static environment and dynamic objects, and then the effect of removing dynamic objects and only retaining the static environment can be achieved. For example, a classification function can be used to distinguish the static environment and dynamic objects, and retain points (i.e., the point cloud of the static environment).
[0048] After obtaining the static environment, it is necessary to remove obstacles and obtain the drivable area. For example, the drivable area includes areas such as roads and parking spaces. Obstacles include dynamic obstacles such as pedestrians or static obstacles such as roadblocks.
[0049] Specifically, first, the point cloud data of the static environment is divided into multiple voxels, and each of the voxels marks a small cube.
[0050] For a given point cloud , it can be divided into N voxels through the voxelization process , the size of the voxel is predetermined, and each voxel represents a small area, and the size can be defined by the voxel size parameter v.
[0051] .
[0052] Among them, the voxel size is the set voxel size, and P world is the point cloud after coordinate transformation.
[0053] After that, each voxel is classified to determine whether it belongs to categories such as roads, obstacles, or sidewalks. The classification function is used to assign a label to each voxel V i : .
[0054] Through semantic segmentation, a semantic label is assigned to each voxel Vi, and drivable areas such as roads and parking spaces can be identified, and obstacles are removed. For example, only the voxels classified as roads will be marked as drivable areas, while the voxels assigned categories such as sidewalks and obstacles are marked as non-drivable areas.
[0055] It can be understood that based on the classification results, non-drivable areas can be removed, and only drivable areas are retained, and then the boundaries of the drivable areas can be extracted. For example, the drivable area consists of a set of voxels and then the bounding box of these voxels can be extracted to form the topological structure of the drivable area.
[0056] After that, under dynamic constraints, diverse and feasible trajectories are generated based on the drivable area, that is, multiple paths are generated within the drivable area.
[0057] The way to generate the trajectory is as follows: first, a starting point and a target point are selected within the drivable area, and this target point is the end position.
[0058] After that, a smooth path planning algorithm ( algorithm) is used to generate a trajectory connecting the starting point and the target point within the drivable area. These algorithms are based on graph search and heuristic search and can find the shortest path between each node (voxel) of the graph. The goal of the algorithm is to minimize the cost function: 。
[0059] Among them, is the actual cost (path length) from the starting point to the current node; is the heuristic estimated cost (Euclidean distance) from the current node to the target node.
[0060] The finally generated trajectory is a series of points , each representing a position in the trajectory.
[0061] It should be noted that in the generated trajectory, the available paths need to be screened in combination with the vehicle's dynamic constraints, such as the maximum speed , the maximum acceleration and the maximum steering angle , by restricting the vehicle's movement on the trajectory to ensure that the trajectory is feasible.
[0062] 。
[0063] Among them is the speed of the i-th point in the trajectory; is the acceleration; is the steering angle of the vehicle.
[0064] After determining the available paths, a series of pose transformation matrices can be used to represent the available paths, where each T i is a 4×4 transformation matrix used to describe the vehicle's motion state on the trajectory.
[0065] 。
[0066] Among them, R i is the rotation matrix, and t i is the translation vector, describing the vehicle's pose at the trajectory point Pi.
[0067] Finally, the original ego-vehicle point cloud data is resampled to extract the occupancy information along the trajectory, achieving the effect of improving the quality of the input data. Specifically, when resampling the trajectory points, for each trajectory point Pi, its nearest neighbor point in the original point cloud needs to be calculated. Data structures such as KD-tree or Ball-tree can be used to accelerate the nearest neighbor search. Set a radius r, and find the point closest to the trajectory point Pi through the nearest neighbor search and calculate its occupancy information: 。
[0068] Among them, represents in the original point cloud P worldAmong them, the point P closest to the point Pi and with a distance less than r j .
[0069] Moreover, for each trajectory point Pi, based on the nearest neighbor point P j to determine whether the trajectory point Pi is feasible, and generate an occupancy grid (or occupancy voxel). The occupancy grid can be represented as a two-dimensional or three-dimensional array, where each grid cell represents whether it is occupied.
[0070] Among them, if the position of the trajectory point P i is occupied by an obstacle, then O(P i ) = 1; if the position of the trajectory point P i is free, then O(P i ) = 0.
[0071] Finally, obtain the occupancy rate information along the trajectory to fill in the blanks in the original data, and solve the problems of unbalanced actual collected data (some types of movements are more common than others) and limited diversity (lack of trajectory data under different conditions in the same scene).
[0072] As an implementation, please refer to Figure 4 , S106 includes:[[]] S1061, extracting a set of anchor points from the trajectories with occupancy information through K-Means clustering, and adding Gaussian noise to each anchor point, where the anchor points represent the key features of the driving trajectories.
[0073] S1062, gradually denoising the noisy anchor points until they are restored to be close to the real driving trajectories.
[0074] S1063, using a diffusion decoder to take the denoised anchor points as input, and output the predicted classification scores and the denoised trajectories.
[0075] S1064, evaluating the output of the diffusion decoder according to the total loss function, and optimizing the parameters of the diffusion decoder according to the total loss function until the total loss value is less than the set value.
[0076] Among them, after obtaining the trajectories with occupancy information, it is necessary to classify the trajectories with occupancy information to ensure the accuracy and diversity of the trajectories.
[0077] During the processing, it is first necessary to create anchor points, where the anchor points are extracted from the trajectories with occupancy information output in step S104, and the anchor points are used to generate noisy trajectory data and serve as the input of the truncated diffusion model.
[0078] As an implementation, this application extracts a set of "anchor points" from the trajectories with occupancy information through K-Means clustering , these anchor points represent certain key features of the driving trajectory. To better simulate the variations that may occur in actual driving, some Gaussian noise, i.e., random perturbations, are added to each anchor point to simulate the unpredictable changes in the real environment.
[0079] After that, to limit the noise within a suitable range, a "truncated diffusion process" is used to gradually denoise. Specifically, starting from these noisy anchor points, the denoising operation is carried out step by step until it is restored to be close to the real driving trajectory. The mathematical form of this process is as follows: .
[0080] Among them, represents the k-th trajectory after the i-th step of denoising, is the truncation coefficient at the i-th step, is the anchor point, is the noise sampled from the standard normal distribution; , is the number of truncated diffusion steps to avoid excessive removal of noise, and .
[0081] It should be noted that through the processing of truncated diffusion, although the trajectory corresponding to the anchor point is close to the real driving trajectory, there is still noise, but the noise is within a controllable range.
[0082] After that, the denoised anchor points are used as input by the diffusion decoder, and the predicted classification scores and denoised trajectories are output. Specifically, during the training process, the diffusion decoder will use the noisy trajectory (i.e., the anchor points with added noise) as input and output the predicted classification scores and the denoised trajectory : .
[0083] Among them, z represents the conditional information. The purpose of the model is to make the decoded trajectory as close as possible to the real trajectory and make the classification prediction accurate.
[0084] At the same time, to improve the accuracy of classification prediction, this application also evaluates the output of the diffusion decoder according to the total loss function and optimizes the parameters of the diffusion decoder according to the total loss function.
[0085] Specifically, the loss in the training process includes two parts. One is the trajectory reconstruction loss, and the goal is to make the denoised trajectory as similar as possible to the real trajectory; the other is the classification loss. To make the model accurately distinguish positive and negative samples, the binary cross-entropy loss BCE is introduced. We use the ground truth trajectory closest to the anchor point The corresponding noise trajectory is used as a positive sample ( ), and the others are used as negative samples ( ).
[0086] .
[0087] Among them, represents the trajectory reconstruction loss, BCE is the binary cross - entropy loss, represents the coefficient used to balance these two losses, L represents the total loss value, represents the ground - truth trajectory, that is, the trajectory that the vehicle actually travels in the real driving scenario.
[0088] By determining the total loss value, the performance of the diffusion decoder in the current training state can be evaluated. Moreover, the smaller the total loss value, the closer the prediction result of the model is to the true value. Therefore, the training goal is to make the value of L as low as possible, and the parameters of the diffusion decoder can be optimized according to the total loss value, and finally make the total loss value less than the set value.
[0089] Finally, the classified trajectories are fused with noise for training, and then through adversarial noise fusion training, the generalization ability and robustness of the model are further improved.
[0090] Among them, please refer to Figure 5 , S108 includes: S1081, Inject noise into different modules of the autonomous driving system and define the results of the classified trajectories after injecting noise.
[0091] S1082, Determine the loss ratio of each module at the current time step according to the results after noise injection.
[0092] S1083, Update the normalized weight of each module according to the loss ratio.
[0093] S1084, Determine the total loss at the next time step according to the updated normalized weight, and optimize the working parameters of each module according to the gradient of the total loss and the normalized weight of each module until the training of the end - to - end autonomous driving model is completed.
[0094] In the autonomous driving system, there are multiple functional modules, including a trajectory formation module, a map formation module, a motion prediction module, and a planning module. When injecting noise, for the training model, it is guided by the overall goal rather than the loss of each independent module. Therefore, noise needs to be injected into the inputs of different modules. This method ensures that noise is generated from the overall view of the model, that is, using the overall loss for backpropagation, rather than focusing on the individual module losses that may be contradictory and have a negative impact on the robustness of the overall decision - making.
[0095] It should be noted that the goal of injecting noise in this step is different from that in S106. In S106, the introduction of noise is to simulate the uncertainty in the real environment and perform denoising and trajectory reconstruction through the truncated diffusion model. That is, by adding Gaussian noise to the anchor points, the unpredictable changes that may occur in the real driving environment are simulated, and the truncated diffusion model is used to gradually denoise to recover the trajectory data close to the real driving trajectory.
[0096] In this step, however, the introduction of noise is to improve the robustness of the model under adversarial attacks and optimize the overall performance of the model through adversarial noise fusion training. That is, adversarial noise is injected into different end-to-end modules (trajectory formation module, map formation module, motion prediction module, planning module). By optimizing the amount of noise injection, the model can still maintain stable performance under the influence of adversarial noise.
[0097] In comparison, the noise in S106 simulates random interference in the real environment and is a non-targeted noise. The noise in this step, on the other hand, simulates interference in adversarial attacks or complex environments and is a targeted noise.
[0098] Moreover, after injecting the noise, the optimal amount of noise injection needs to be found. Therefore, in the implementation process, the model output is first defined, which represents the result after injecting the noise set on the input data and is specifically implemented through the following function combination: .
[0099] Among them, , , respectively represent the specific adversarial perturbations injected into the m-th perception module, the k-th prediction module, and the planning module. Xn represents the output data, such as sensor data (point cloud, image), vehicle state (speed, position), etc. Planner represents the planning module, Predictor represents the prediction module, and Perceiver represents the perception module.
[0100] The amount of noise injection satisfies the formula, that is: .
[0101] Among them, represents the optimal combination of adversarial noise parameters obtained by maximizing the total loss function during the noise injection process; C is the constraint set of the noise; is the total loss function; is the true label.
[0102] To manage the different contributions of modules during training, dynamic weight accumulation adaptation is introduced, which adaptively adjusts the loss weight of each module to the overall objective according to its contribution during the noise injection process. This method introduces a normalization weight function to eliminate the dimension difference, accelerate the model convergence speed, reduce the impact of outliers on model training, and improve its stability and model generalization ability.
[0103] To extend the concept of multi-tasking to multiple modules, the loss of each module at the current time step t The ratio relative to the previous value is calculated as: .
[0104] Where is the loss of module j at time step t-2; is the loss of module j at time step t-1; is the loss ratio of module j at time step t.
[0105] Then, based on these loss ratios, the normalization weight formula is used to update the weights: .
[0106] Where is the normalized weight of module j at time step t; is the average of the loss ratios of all modules at time step t; Norm represents the normalization function; exp represents the exponential function with base e; N is the total number of modules; represents the loss ratio of module j at time step t; represents module 's loss ratio at time step t.
[0107] It should be noted that in the above formula, the numerator part represents the loss ratio of the j-th module at time step t, and only calculates the normalized exponential value for module j. The denominator needs to iterate over all modules, so the summation index variable k is introduced. j is an external index representing the target module for which the weight is currently to be calculated. k is an internal index only used to iterate over all modules and sum.
[0108] Finally, the total loss at time step t+1 is calculated based on the updated weights: .
[0109] This method can ensure that the weights dynamically adapt to the performance of each module over time, thereby improving stability and overall performance.
[0110] It should be noted that in the steps of S1081 to S1084, the optimal amount of noise injection is achieved through the method of dynamic adaptive noise optimization. That is, the amount of noise injection is not fixed in advance, but is gradually adjusted through multiple iterations (time steps). After each injection of noise, the weight is dynamically calculated according to the loss ratio (degree of performance degradation) of the module, and then the weight affects the noise injection intensity of the next time step. By forming a closed loop between noise injection and weight update, the model can automatically learn the sensitivity of different modules to noise during training and allocate the optimal anti-noise resources. Specifically, after determining the normalized weight of each module, the module with a higher weight is more likely to be noticed, which is equivalent to indirectly adjusting the priority of noise injection - the module with a higher weight may be assigned more stringent adversarial noise constraints in subsequent training. And, the parameters (neural network weights) of each module are optimized by the gradient backpropagation of the total loss, rather than being set manually. During backpropagation, the gradient is distributed to different modules according to the weight, thereby adjusting its internal parameters, finally achieving the optimal noise injection, and completing the training of the model based on this input.
[0111] Generally speaking, for the autonomous driving adversarial training method provided in this application, on the one hand, by adopting the trajectory resampling strategy, the gaps in the original data are effectively filled, and the problems of data imbalance and limited diversity are solved. At the same time, using the truncated diffusion model for trajectory reconstruction and classification optimization, and adding adversarial noise fusion training, enable the autonomous driving system to more accurately identify road information, predict the movement trajectories of other vehicles and pedestrians, and generate safe and efficient driving paths when facing complex and changeable traffic environments, thereby significantly improving the robustness and accuracy of the autonomous driving system and reducing the risk of traffic accidents.
[0112] On the other hand, noise is injected into different modules in an end-to-end manner, and backpropagation is carried out under the guidance of the overall objective to optimize the amount of noise injection. At the same time, the dynamic weight accumulation adaptive method is adopted to adaptively adjust the loss weight of each module to the overall objective according to its contribution during the noise injection process. This comprehensive training strategy not only accelerates the convergence speed of the model, but also enables the model to find the optimal solution faster during the training process. In addition, by using the normalized weight function to eliminate the dimension difference, the stability and efficiency of the training are further improved.
[0113] On the third hand, by generating diverse and feasible trajectories, and using the truncated diffusion model for denoising and classification optimization, the model can learn richer driving scenarios and trajectory features. At the same time, the introduction of adversarial noise fusion training enables the model to maintain stable performance when facing unknown or extreme situations. Significantly enhances the generalization ability and adaptability of the model, enabling the autonomous driving system to better adapt to different traffic environments and driving scenarios.
[0114] Based on the above implementation, an embodiment of the present application further provides an autonomous driving adversarial training device 200. Please refer to Figure 6 , the device includes: A data acquisition module 210, configured to acquire multiple frames of ego vehicle point cloud data.
[0115] It can be understood that through the data acquisition module 210, the above S102 can be executed.
[0116] A data processing module 220, configured to generate a feasible trajectory based on the ego vehicle point cloud data, and obtain occupancy information along the trajectory through trajectory resampling to generate a trajectory with occupancy information.
[0117] It can be understood that through the data processing module 220, the above S104 can be executed.
[0118] The data processing module 220 is further configured to classify the trajectory with occupancy information based on the truncated diffusion model.
[0119] It can be understood that through the data processing module 220, the above S106 can be executed.
[0120] A training module 230, configured to train the classified trajectory by fusing noise to complete the training of the end-to-end autonomous driving model.
[0121] It can be understood that through the training module 230, the above S108 can be executed.
[0122] In summary, an autonomous driving adversarial training method and device provided by the present application first acquire multiple frames of ego vehicle point cloud data, then generate a feasible trajectory based on the ego vehicle point cloud data, and obtain occupancy information along the trajectory through trajectory resampling to generate a trajectory with occupancy information. Then, classify the trajectory with occupancy information based on the truncated diffusion model; finally, train the classified trajectory by fusing noise to complete the training of the end-to-end autonomous driving model. Since in the autonomous driving adversarial training method provided by the present application, a trajectory with occupancy information is generated by using the trajectory resampling method, high-quality data input can be provided for subsequent trajectory classification and reconstruction. By classifying the trajectory, further processing of the trajectory is realized, ensuring the accuracy and diversity of the trajectory. Finally, by training the classified trajectory by fusing noise, the effect of improving the generalization ability and robustness of the model is achieved.
[0123] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of the devices, methods, and computer program products according to the embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of the code, and the module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function.
[0124] It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved.
[0125] It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0126] In addition, the various functional modules in the embodiments of the present application can be integrated together to form an independent part, or each module can exist separately, or two or more modules can be integrated to form an independent part.
[0127] If the functions are implemented in the form of software functional modules and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories, random access memories, magnetic disks, or optical discs that can store program codes.
[0128] Although the present invention is disclosed as above, the present invention is not limited thereto. Any person skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention should be determined by the scope defined by the claims.
Claims
1. An autonomous driving adversarial training method, characterized in that: The method comprises: Obtain multi-frame ego-vehicle point cloud data; Generate a feasible trajectory according to the ego-vehicle point cloud data, and obtain occupancy information along the trajectory by trajectory resampling to generate a trajectory with occupancy information; Classify trajectories with occupancy information based on the truncated diffusion model; The classified trajectories are fused with noise for training to complete the training of the end-to-end autonomous driving model.
2. The autonomous driving adversarial training method according to claim 1, characterized in that: The steps of generating a feasible trajectory according to the vehicle point cloud data and obtaining occupancy information along the trajectory by trajectory resampling to generate a trajectory with occupancy information include: Convert each frame of the ego-vehicle point cloud data from the vehicle coordinate system to the world coordinate system, remove dynamic objects, and only retain the point cloud data of the static environment; Extracting a drivable area after removing obstacles based on the point cloud data of the static environment; Under the dynamic constraints of the vehicle, generating a plurality of feasible trajectories according to the drivable area; The original ego-vehicle point cloud data is resampled according to the generated multiple feasible trajectories, and the occupancy information along the trajectory is extracted to generate a trajectory with occupancy information.
3. The autonomous driving adversarial training method according to claim 2, characterized in that: The steps of converting each frame of the ego-vehicle point cloud data from the vehicle coordinate system to the world coordinate system and removing dynamic objects to retain only the point cloud data of the static environment include: Extract the location information of each frame of the vehicle point cloud data; According to the formula P world =T vehicle→world *P vehicled The position information is converted to the world coordinate system; wherein, P world Represents the transformed world coordinate system, T vehicle→world represents the transformation matrix, P vehicled represents homogeneous coordinates determined according to the position information; The vehicle point cloud data after the coordinate system conversion is classified according to semantic labels or classification algorithms, and static environments and dynamic objects are identified.
4. The autonomous driving adversarial training method according to claim 2, characterized in that: The step of extracting a drivable area after removing obstacles based on the point cloud data of the static environment comprises: Divide the point cloud data of the static environment into a plurality of voxels, each of which is marked with a small cube; classifying each of the voxels based on a classification function; Based on the classification results, the boundaries of the drivable area are extracted.
5. The autonomous driving adversarial training method according to claim 2, characterized in that: Under the dynamic constraints of the vehicle, the step of generating a plurality of feasible trajectories according to the drivable area comprises: The starting and ending positions of the randomly sampled trajectories; In combination with a path planning algorithm, a trajectory connecting the starting point and the end point is generated within the drivable area; According to the dynamic constraints, available paths are screened out from the generated trajectories, and the available paths are represented by a posture transformation matrix.
6. The autonomous driving adversarial training method according to claim 2, characterized in that: The steps of resampling the original vehicle point cloud data according to the generated multiple feasible trajectories and extracting the occupancy information along the trajectory to generate a trajectory with occupancy information include: Based on each trajectory point, calculate its nearest neighbor point in the original point cloud data of the ego vehicle; For each of the trajectory points, whether the trajectory point is feasible is determined based on the state of the nearest neighbor point, so as to generate a trajectory with occupancy information.
7. The autonomous driving adversarial training method according to claim 1, characterized in that: The steps of classifying trajectories with occupancy information according to the truncated diffusion model include: A set of anchor points are extracted from the trajectory with occupancy information by K-Means clustering, and Gaussian noise is added to each anchor point, wherein the anchor point represents a key feature of the driving trajectory; Gradually denoise the noisy anchor points until they are restored to a trajectory close to the actual driving trajectory; Use the diffusion decoder to take the denoised anchor points as input and output the predicted classification score and denoised trajectory; The output of the diffusion decoder is evaluated according to the total loss function, and the parameters of the diffusion decoder are optimized according to the total loss function until the total loss value is less than a set value.
8. The autonomous driving adversarial training method according to claim 1, characterized in that: The steps of training the classified trajectory with noise to complete the training of the end-to-end autonomous driving model include: Inject noise into different modules of the autonomous driving system and define the results of the classified trajectory after the noise injection; Determine the loss ratio of each module at the current time step based on the results after noise injection; Updating the normalized weights of each module according to the loss ratio; The total loss at the next time step is determined based on the updated normalized weights, and the working parameters of each module are optimized based on the gradient of the total loss and the normalized weights of each module until the training of the end-to-end autonomous driving model is completed.
9. The autonomous driving adversarial training method according to claim 1, characterized in that: The normalized weight of each module satisfies the formula: in, is the normalized weight of module j at time step t; is the average loss ratio of all modules at time step t; Norm represents the normalization function; exp represents the exponential function with e as the base; K is the total number of modules, represents the loss ratio of module j at time step t; represents the loss ratio of module K at time step t.
10. An autonomous driving adversarial training device, characterized in that: The device comprises: Data acquisition module, used to acquire multi-frame ego-vehicle point cloud data; A data processing module, used to generate a feasible trajectory based on the ego-vehicle point cloud data, and obtain occupancy information along the trajectory by trajectory resampling to generate a trajectory with occupancy information; The data processing module is also used to classify trajectories with occupancy information based on the truncated diffusion model; The training module is used to fuse the classified trajectories with noise for training to complete the training of the end-to-end autonomous driving model.
Citation Information
Patent Citations
Emergency planning and security assurance
CN114175023A
Vehicle trajectory planning method, device, equipment and medium
CN116520839A
Vehicle trajectory planning method and device, unmanned vehicle and storage medium
CN119126813A