A method for generating autonomous driving simulation data

By marking and rendering the road images with a feasible lane, a high-reality simulation image is generated, which solves the problem of high cost and low efficiency of labeling of 3D target detection data sets of autonomous driving, and provides a high-precision 3D detection box, which improves the training effect of the autonomous driving model.

CN115719439BActive Publication Date: 2025-07-22ZHEJIANG UNIV

Patent Information

Application Number
CN202211484738.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-24
Publication Date
2025-07-22
Estimated Expiration
2042-11-24

AI Technical Summary

Technical Problem

In the prior art, the labeling cost of autonomous driving 3D object detection data set is high, has low efficiency, and has poor accuracy, and the difficulty of aligning the radar and camera hardware time, resulting in poor training results in the detection model.

Method used

By marking the road images without vehicles, using camera parameters back projection to obtain the driving lane in the world coordinate system, filtering point cloud data to fit the road equation, combining the vehicle 3D model and lighting information for rendering, using the style harmony model to generate high-reality simulation images, and performing 3D detection box annotation.

Benefits of technology

It realizes the generation of high-reality autonomous driving data sets at low cost and efficiently, provides a high-precision 3D object detection box, and improves the training effect of the autonomous driving detection model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115719439B_ABST
    Figure CN115719439B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for generating autonomous driving simulation data. First, after marking the drivable lanes on the road image without vehicles, the drivable lanes in the world coordinate system are obtained through back-projection using camera parameters. The point cloud data is filtered by the drivable lanes, and the real road surface equation is fitted using the filtered point cloud set. A vehicle model is selected from the vehicle 3D model library and placed on the real road surface, and the vehicle angle is adjusted according to the road surface equation. After selecting different lighting information from the lighting library, high-fidelity foreground images and mask information can be obtained through rendering using Blender. Through the style harmonization model, using the background image as a reference, the foreground image is subjected to style harmonization operations, and finally, the foreground image and the background image are pasted using the mask information to obtain the final simulation image. The present invention can quickly obtain synthetic image data highly similar to the real world and easily obtain high-precision detection frames.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for generating driving data, and more particularly to a method for generating autonomous driving simulation data. Background Art

[0002] Autonomous driving relies on 3D detection algorithms based on computer vision, including monocular 3D object detection, multi-view bird's-eye view object detection, etc. Although the implementation methods of various 3D object detection algorithms are different, they all rely on real annotation data. The annotation cost of 3D object detection data sets is much higher than that of 2D object detection data sets, and the annotation efficiency is also much lower than that of 2D object detection data sets. At the same time, training an excellent 3D object detection model requires a data set with accurate annotations. The 3D object detection box is annotated in the point cloud, while the training data is an image. The point cloud data is obtained by a lidar, and the image data is obtained by a camera. Since the exposure time of the camera cannot be determined, the refresh frequency of the lidar cannot be accurately aligned with the camera shooting frequency, that is, in most cases, the manually annotated 3D object box is deviated from the target vehicle after being projected onto the 2D image, which has a great negative impact on the training of the detection model in autonomous driving. Summary of the Invention

[0003] In order to overcome the problems of high cost, low efficiency, and poor accuracy in manually collecting and annotating autonomous driving target detection data sets, the present invention provides a method for generating autonomous driving simulation data. The present invention aims to perform drivable lane annotation on a road image without vehicles, and then obtain the drivable lane in the world coordinate system through back-projection using camera parameters. The point cloud data is screened by the drivable lane, and the real road surface equation is fitted using the screened point cloud set. A vehicle model is selected from the vehicle 3D model library and placed on the real road surface, and the vehicle angle is adjusted according to the road surface equation. After selecting different lighting information from the lighting library, Blender can be used for rendering to obtain a foreground image with high realism and mask information. In order to eliminate the style difference between the foreground and the background, through a style harmonization model, using the background image as a reference, the foreground image is subjected to style harmonization operation, and finally the foreground image and the background image are pasted using the mask information to obtain the final simulation image. The present invention can quickly obtain a high-fidelity autonomous driving data set and high-precision 3D object detection box annotation at extremely low cost.

[0004] The technical solution adopted by the present invention is as follows:

[0005] 1) The road images without vehicles are screened through a vehicle detection model and recorded as target background images. Then, the drivable lanes of the target background images are labeled by a drivable lane detection module to obtain lane labeling information. Finally, the lane labeling information is back-projected through camera parameters to obtain the drivable lanes in the world coordinate system;

[0006] 2) The point cloud coordinate set in the world coordinate system corresponding to the target background image is screened using the drivable lanes in the world coordinate system, and the lane plane equation is fitted based on the screened point cloud coordinate set in the world coordinate system;

[0007] 3) The X-axis coordinate X w and Y-axis coordinate Y w in the world coordinate system are randomly generated in the drivable lanes. The Z-axis coordinate Z w corresponding to the X-axis coordinate X w and Y-axis coordinate Y w is calculated through the lane plane equation. The reference coordinate (X w , Y w , Z w ) is composed of the X-axis coordinate X w , Y-axis coordinate Y w , and Z-axis coordinate Z w ;

[0008] 4) The yaw angle is calculated based on the center axis of the drivable lanes, and then the pitch angle and roll angle are calculated based on the fitted lane plane equation. The reference rotation angle (θ Pitch , θ Roll , θ Yaw ) is composed of the yaw angle, pitch angle, and roll angle;

[0009] 5) After randomly extracting a vehicle model from the vehicle 3D model library and importing it into the image rendering engine, the reference coordinate in 3) is used as the bottom geometric center coordinate of the vehicle model, and the reference rotation angle (θ Pitch , θ Roll , θ Yaw ) in 4) is used as the rotation angle of the vehicle model. Then, a file is randomly selected from the high dynamic range imaging library as the lighting information, and the set vehicle model is rendered through the image rendering engine to obtain the foreground rendered image and the mask information;

[0010] 6) Based on the target background image, the foreground rendered image, and the mask information, image simulation is performed using a style harmonization model to obtain a simulated image;

[0011] 7) The 3D detection box annotation information projected onto the simulated image is calculated based on the size information of the current vehicle model and the camera parameters in the image rendering engine; The simulated image and the corresponding 3D detection box annotation information form the simulation data.

[0012] In the above (1), the lane annotation information is the four vertex coordinates of a quadrilateral. According to the camera parameters, the four vertex coordinates are back-projected using the following formula to obtain the drivable lane in the world coordinate system:

[0013]

[0014]

[0015]

[0016] M2 = R inv T

[0017]

[0018]

[0019]

[0020] Among them, each drivable lane is represented by 4 vertices, represents the back-projection coefficient of the i-th vertex of the drivable lane, u i , v i respectively represent the horizontal and vertical coordinate values of the i-th vertex of the drivable lane in the pixel coordinate system, R represents the rotation matrix in the external camera parameters, T represents the translation matrix in the external camera parameters, K represents the internal camera matrix, R inv represents the inverse matrix of the rotation matrix, K inv represents the inverse matrix of the internal camera matrix, represents the three coordinate values of the i-th vertex in the world coordinate system, M1 represents the rotation-related coefficient matrix in the back-projection, M2 represents the translation-related coefficient matrix in the back-projection, M1[0, 0] represents the 0th row and 0th column of the rotation-related coefficient matrix M1 in the back-projection, and M2[2, 0] represents the 2nd row and 0th column of the translation-related coefficient matrix M2 in the back-projection.

[0021] The above (2) is specifically as follows:

[0022] First, determine the coordinate ranges of the drivable lane on the X-axis and Y-axis in the world coordinate system. Then, screen the point cloud coordinate set corresponding to the target background image in the world coordinate system according to the coordinate ranges of the drivable lane on the X-axis and Y-axis. Next, use the following formula to fit the lane plane equation based on the screened point cloud coordinate set in the world coordinate system. The formula is as follows:

[0023] min||An||

[0024] ||n|| = 1

[0025] Spoints = {p1, p2, …, p m}

[0026] p i = (x i , y i , z i )

[0027] p0 = (x0, y0, z0)

[0028]

[0029]

[0030] n = [a b c] T

[0031] where min|| || represents the minimum value operation, S points represents the set of point cloud coordinates in the world coordinate system obtained by screening, with a total of m points, A represents the vector matrix of all points in the point cloud set pointing to the mean point, n represents the plane normal vector of the lane, p i represents the i-th point in the set of point clouds corresponding to the target background image in the world coordinate system, x i , y i , z i represent the three coordinate values of the i-th point in the point cloud set, p0 represents the coordinate mean point of all points in the point cloud set, x0, y0, z0 respectively represent the means of the three coordinate values corresponding to all points in the point cloud set, and a, b, c represent the cubic coefficient, quadratic coefficient, and linear coefficient in the plane equation.

[0032] In the above 4), the angle between the plane normal vector of the lane plane equation and the x-axis of the world coordinate system is denoted as the pitch angle, and the angle between the plane normal vector of the lane plane equation and the y-axis of the world coordinate system is denoted as the roll angle.

[0033] The style harmonization model includes a first feature extraction module, a second feature extraction module, and a generation network. The target background image, the foreground rendering image, and the mask information are input into the style harmonization model. The target background image is used as the input of the first feature extraction module, the foreground rendering image is used as the input of the second feature extraction module, the foreground rendering image is also used as the first input of the generation network, the outputs of the first feature extraction module and the second feature extraction module are respectively used as the second input and the third input of the generation network, the target background image is also used as the fourth input of the generation network, the mask information is used as the fifth input of the generation network, and the output of the generation network is used as the output of the style harmonization model.

[0034] The first feature extraction module and the second feature extraction module have the same structure. The first feature extraction module includes 5 feature encoding modules, 4 feature decoding modules, 4 layers of max pooling layers, and 4 layers of transposed convolution layers. The input of the first feature extraction module serves as the input of the first feature encoding module. The first feature encoding module is connected to the fourth transposed convolution layer after passing through the first max pooling layer, the second feature encoding module, the second max pooling layer, the third feature encoding module, the third max pooling layer, the fourth feature encoding module, the fourth max pooling layer, and the fifth feature encoding module in sequence. The output of the fourth feature encoding module is concatenated with the output of the fourth transposed convolution layer and then input into the fourth feature decoding module. The fourth feature decoding module is connected to the third transposed convolution layer. The output of the third feature encoding module is concatenated with the output of the third transposed convolution layer and then input into the third feature decoding module. The third feature decoding module is connected to the second transposed convolution layer. The output of the second feature encoding module is concatenated with the output of the second transposed convolution layer and then input into the second feature decoding module. The second feature decoding module is connected to the first transposed convolution layer. The output of the first feature encoding module is concatenated with the output of the first transposed convolution layer and then input into the first feature decoding module. The outputs of the second feature decoding module to the fourth feature decoding module are respectively subjected to bilinear interpolation and then added together with the output of the first feature decoding module to obtain the final feature map as the output of the first feature extraction module.

[0035] The generation network includes 2 adaptive instance normalization layers and 6 fully connected layers. The second input and the third input of the generation network respectively serve as the first input and the second input of the first adaptive instance normalization layer. The first adaptive instance normalization layer is connected to the second adaptive instance normalization layer after passing through 6 fully connected layers in sequence. The first input of the generation network serves as the first input of the second adaptive instance normalization layer. The output of the second adaptive instance normalization layer is denoted as the harmonized foreground image. After performing a texture mapping operation on the harmonized foreground image with the fourth input and the fifth input of the generation network, a simulated image is obtained as the output of the generation network.

[0036] The specific content of item (7) is as follows:

[0037] Taking the bottom geometric center of the current vehicle model as the origin, 8 vertices of the minimum circumscribed cuboid of the current vehicle model are calculated according to the size information of the current vehicle model. The 8 vertices of the minimum circumscribed cuboid are projected to obtain the 3D detection frame annotation information. The calculation formula is as follows:

[0038]

[0039] where, K r is the camera internal parameter in the image rendering engine, R r is the rotation matrix of the camera external parameter in the image rendering engine, T r is the translation matrix of the camera external parameter in the image rendering engine. The three coordinate values of the k-th vertex of the minimum bounding cuboid in the world coordinate system The two pixel coordinate values of the projection of the k-th vertex of the minimum bounding cuboid

[0040] The beneficial effects of the present invention are as follows:

[0041] The invention can quickly obtain synthetic image data highly similar to the real world, with high richness and authenticity in aspects such as vehicle type, position, and lighting information. Compared with traditional acquisition and manual annotation methods, the synthetic data method has a lower cost and a much higher data generation efficiency than the traditional method. At the same time, due to the difficult time alignment problem between radar and camera hardware during the annotation of the autonomous driving dataset, it is difficult to obtain standard detection frames. Relatively speaking, the invention can easily obtain standard and high-precision detection frames, which provides great help for the training of various autonomous driving detection models. Description of the Drawings

[0042] Figure 1 is the flowchart of the method in the present invention.

[0043] Figure 2 is the structural diagram of the feature extraction module in the style harmonization model of the present invention.

[0044] Figure 3 is the structural diagram of the generation network in the style harmonization model of the present invention.

[0045] Figure 4 is the unharmonized synthetic image generated by the present invention.

[0046] Figure 5 is the harmonized synthetic image generated by the present invention. Detailed Embodiments

[0047] To more clearly illustrate the purpose and technical solutions of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings.

[0048] As Figure 1 shown, the present invention includes the following steps:

[0049] 1) Screen and obtain a road image without vehicles through a vehicle detection model and record it as the target background image. Then, perform drivable lane annotation on the target background image through a drivable lane detection module to obtain lane annotation information. Finally, perform back-projection on the lane annotation information through camera parameters to obtain the drivable lane in the world coordinate system;

[0050] 1), the lane annotation information is the four vertex coordinates of a quadrilateral. According to the camera parameters, the four vertex coordinates are back-projected using the following formula to obtain the drivable lanes in the world coordinate system:

[0051]

[0052]

[0053]

[0054] M2 = R inv T

[0055]

[0056]

[0057]

[0058] Among them, each drivable lane is represented by 4 vertices. represents the back-projection coefficient of the i-th vertex of the drivable lane, u i , v i respectively represent the horizontal and vertical coordinate values of the i-th vertex of the drivable lane in the pixel coordinate system, R represents the rotation matrix in the camera extrinsic parameters, T represents the translation matrix in the camera extrinsic parameters, K represents the camera intrinsic matrix, R inv represents the inverse matrix of the rotation matrix, K inv represents the inverse matrix of the camera intrinsic matrix. represents the three coordinate values of the i-th vertex in the world coordinate system, M1 represents the rotation-related coefficient matrix in the back-projection, M2 represents the translation-related coefficient matrix in the back-projection, M1[0, 0] represents the 0th row and 0th column of the rotation-related coefficient matrix M1 in the back-projection, and M2[2, 0] represents the 2nd row and 0th column of the translation-related coefficient matrix M2 in the back-projection.

[0059] 2) Use the drivable lanes in the world coordinate system to filter the point cloud coordinate set corresponding to the target background image in the world coordinate system, and fit the lane plane equation according to the filtered point cloud coordinate set in the world coordinate system;

[0060] 2) Specifically:

[0061] First, determine the coordinate ranges [X′ min , X′ max and [Y′ min , Y′ max of the drivable lanes on the X-axis and Y-axis according to the drivable lanes in the world coordinate system. Specifically, the vertex set S available of each drivable lane is expressed as:

[0062] S available = {P1, P2, P3, P4}

[0063] P i = (X iw ,Y iw ,0)

[0064] X 1w ,X 4w <X 2w ,X 3w

[0065] Y 1w ,Y 2w <Y 3w ,Y 4w

[0066] Among them, P i represents the coordinates of vertex i, X iw ,Y iw respectively represent the X-axis and Y-axis coordinate values of vertex i, X 1w ,X 2w ,X 3w ,X 4w respectively represent the X-axis coordinate values of vertices 1-4, Y 1w ,Y 2w ,Y 3w ,Y 4w respectively represent the Y-axis coordinate values of vertices 1-4. Select the larger value from Y 1w ,Y 2w as Y′ min ,Y 3w ,Y 4w Select the smaller value from as Y′ max ,X 1w ,X 2w Select the larger value from as X′ min ,X 3w ,X 4w Select the larger value from as X′ min .

[0067] Next, according to the coordinate ranges of the drivable lanes on the X-axis and Y-axis, filter the point cloud coordinate set of the target background image corresponding to the world coordinate system, that is, select the points that satisfy X w ∈ [X′ min ,X′ max , Y w ∈ [Y′ min ,Y′ max . Then, according to the filtered point cloud coordinate set of the world coordinate system, fit the lane plane equation using the following formula. The formula is as follows:

[0068] min||An||

[0069] ||n|| = 1

[0070] Sp oints = {p1, p2,..., p m}

[0071] p i = (x i , y i , z i )

[0072] p0 = (x0, y0, z0)

[0073]

[0074]

[0075] n = [a b c] T

[0076] Among them, min|| || represents the minimum value operation, S points represents the set of point cloud coordinates in the world coordinate system obtained by screening, with a total of m points. A represents the vector matrix of all points in the point cloud set pointing to the mean point, n represents the plane normal vector of the lane, p i represents the i-th point in the set of point clouds corresponding to the target background image in the world coordinate system, x i , y i , z i represent the three coordinate values of the i-th point in the point cloud set, p0 represents the coordinate mean point of all points in the point cloud set, x0, y0, z0 respectively represent the means of the three coordinate values corresponding to all points in the point cloud set, and a, b, c represent the cubic coefficient, quadratic coefficient, and linear coefficient in the plane equation. Assume that all points in S points are on the plane and should satisfy An = 0. However, in actual situations, most of the point coordinates fluctuate above and below the target plane. Therefore, the objective function is min||An||, and the constraint condition is ||n|| = 1. Perform singular value decomposition on matrix A, A = UDV T , where D is a diagonal matrix, and U and V are unitary matrices. Then, ||An|| = ||UDV T n|| = ||DV T n||, where V T n is a column matrix, and ||V T n|| = ||n|| = 1. When and only when , ||An|| reaches the minimum value. From this, we can obtain The optimal solution under the constraint of min||An|| is n = (a, b, c) = (v1, v2, v3), and d = -(ax0 + by0 + cz0).

[0077] 3) Randomly generate the X-axis coordinate X in the world coordinate system within the drivable lane w and the Y-axis coordinate Y w , and calculate the Z-axis coordinate Z w corresponding to the X-axis coordinate X w and the Y-axis coordinate Y w (i.e., height). The X-axis coordinate X w , the Y-axis coordinate Y w and the Z-axis coordinate Z w form the reference coordinate (X w , Y w , Z w );

[0078] Randomly generate coordinates X w ∈[X′ min , X′ max and Y w ∈[Y′ min , Y′ max . According to the fitted plane equation aX w + bY w + cZ w + d = 0, the corresponding height coordinate Z w can be calculated.

[0079] 4) Calculate the yaw angle according to the central axis of the drivable lane, and then calculate the pitch angle and roll angle according to the fitted lane plane equation. The yaw angle, pitch angle and roll angle form the reference rotation angle (θ Pitch , θ Roll , θ Yaw );

[0080] In specific implementation, for the vertex set S available = {P1, P2, P3, P4} of each drivable lane, calculate the midpoint P near of P1 and P2 and the midpoint P far of P3 and P4. The yaw angle θ Yaw can be calculated using the following formula:

[0081]

[0082] where X far , Y far represent the X and Y axis coordinate values of the midpoint P far , and X near , Y near represent the midpoint Pnear The X and Y axis coordinate values.

[0083] Denote the included angle between the plane normal vector n=(a, b, c) of the lane plane equation and the unit vector n x =(1, 0, 0) on the x-axis of the world coordinate system as the pitch angle, and denote the included angle between the unit vector n=(a, b, c) where the plane normal vector of the lane plane equation lies and the y-axis n y =(0, 1, 0) of the world coordinate system as the roll angle.

[0084] 5) Randomly select at least one vehicle model from a pre-prepared vehicle 3D model library and import it into the image rendering engine blender. The vehicle models include sedans, vans, special vehicles, and SUVs. Use the reference coordinates (X w , Y w , Z w ) in 3) as the bottom geometric center coordinates of the vehicle model. The bottom geometric center is the origin of the vehicle. If multiple vehicle models are selected, multiple reference coordinates need to be generated in 3). Use the reference rotation angles (θ Pitch , θ Roll , θ Yaw ) in 4) as the rotation angles of the vehicle model; then randomly select a file from the High Dynamic Range Imaging (HDRI) library as the lighting information. The lighting information library consists of high dynamic range imaging (HDRI) in multiple different scenarios, and the scenarios include early morning, noon, afternoon, evening, and night. Render the set vehicle model through the cycles renderer of the image rendering engine blender to obtain the foreground rendered image and the mask information;

[0085] 6) According to the target background image, the foreground rendered image, and the mask information, use the style harmonization model for image simulation to obtain a highly realistic simulation image; in the style harmonization model, first perform background style harmonization on the foreground rendered image to obtain the harmonized foreground image. As Figure 5 shown, compared with Figure 4 , the foreground rendered image in Figure 5 is more harmonious with the background style. Then use the mask information to perform texture synthesis on the harmonized foreground image and the foreground rendered image to obtain a highly realistic simulation image.

[0086] The style harmonization model has a U-shaped structure, including a first feature extraction module, a second feature extraction module, and a generation network. The target background image, the foreground rendering image, and the mask information are input into the style harmonization model. The target background image is used as the input of the first feature extraction module, the foreground rendering image is used as the input of the second feature extraction module, and the foreground rendering image is also used as the first input of the generation network. The outputs of the first feature extraction module and the second feature extraction module are used as the second input and the third input of the generation network respectively. The target background image is also used as the fourth input of the generation network, and the mask information is used as the fifth input of the generation network. The output of the generation network is used as the output of the style harmonization model.

[0087] As Figure 2 shown, the structures of the first feature extraction module and the second feature extraction module are the same. The first feature extraction module includes 5 feature encoding modules, 4 feature decoding modules, 4 layers of max pooling layers, and 4 layers of transposed convolution layers. The input of the first feature extraction module is used as the input of the first feature encoding module. After passing through the first max pooling layer, the second feature encoding module, the second max pooling layer, the third feature encoding module, the third max pooling layer, the fourth feature encoding module, the fourth max pooling layer, and the fifth feature encoding module in sequence, it is connected to the fourth transposed convolution layer. The output of the fourth feature encoding module is concatenated with the output of the fourth transposed convolution layer and then input into the fourth feature decoding module. The fourth feature decoding module is connected to the third transposed convolution layer. The output of the third feature encoding module is concatenated with the output of the third transposed convolution layer and then input into the third feature decoding module. The third feature decoding module is connected to the second transposed convolution layer. The output of the second feature encoding module is concatenated with the output of the second transposed convolution layer and then input into the second feature decoding module. The second feature decoding module is connected to the first transposed convolution layer. The output of the first feature encoding module is concatenated with the output of the first transposed convolution layer and then input into the first feature decoding module. The outputs of the second feature decoding module - the fourth feature decoding module are bilinearly interpolated respectively and then added together with the output of the first feature decoding module to obtain the final feature map as the output of the first feature extraction module. The structures of the feature encoding module and the feature decoding module are the same. Each feature encoding module contains two CBL modules, where C represents the convolutional layer, B represents the batch normalization layer, and L represents the Leaky ReLU activation layer. The convolutional layer, the batch normalization layer, and the Leaky ReLU activation layer are connected in sequence.

[0088] As Figure 3As shown, the generation network includes 2 Adaptive Instance Normalization (AdaIN) layers and 6 fully connected layers. The second input and the third input of the generation network serve as the first input and the second input of the first Adaptive Instance Normalization layer respectively. The first Adaptive Instance Normalization layer is connected to the second Adaptive Instance Normalization layer after passing through 6 fully connected layers in sequence, that is, the output of the sixth fully connected layer serves as the second input of the second Adaptive Instance Normalization layer. The first input of the generation network serves as the first input of the second Adaptive Instance Normalization layer. The output of the second Adaptive Instance Normalization layer is denoted as the harmonized foreground image. After performing a texturing operation on the harmonized foreground image with the fourth input (i.e., the target background image) and the fifth input (i.e., the mask information) of the generation network, a simulation image is obtained as the output of the generation network.

[0089] Two adjacent feature encoding modules are connected by a max pooling layer, and there are a total of 4 max pooling layers. Each layer of the feature extraction module of encoder G will obtain feature maps of different scales, and the formula is:

[0090] G(I) = {F1, F2, F3, F4, F5}

[0091]

[0092] (W, H) represents the size of the feature map. The decoding neural network E is composed of 4 identical decoding modules. The decoding module has the same structure as the encoding module. Adjacent decoding modules are connected by a transposed convolutional layer. There is also a transposed convolutional layer (DConv) connecting the fifth encoding module and the fourth decoding module, with a total of 4 transposed convolutional layers. Each transposed convolutional layer performs an upsampling operation on the features of the previous layer. The specific formula during the decoding process is:

[0093]

[0094] F′4 = Concat(DConv(F5), F4)

[0095]

[0096] where F′ i represents the feature map output by the i-th decoding module D i , and Concat represents the matrix concatenation operation.

[0097] 7) Calculate the 3D detection box annotation information projected onto the simulation image (i.e., the 2D image) based on the size information of the currently pre-annotated vehicle model and the camera parameters in the image rendering engine; The simulation data is composed of the simulation image and the corresponding 3D detection box annotation information, and is used to train the autonomous driving vision perception model to achieve autonomous driving of the vehicle.

[0098] 7) Specifically:

[0099] Take the geometric center of the bottom surface of the current vehicle model as the origin, calculate the 8 vertices of the minimum circumscribed cuboid of the current vehicle model according to the size information (w, l, h) of the current vehicle model, project the 8 vertices of the minimum circumscribed cuboid to obtain the 3D detection frame annotation information, and the calculation formula is as follows:

[0100]

[0101] Among them, K r is the camera internal parameter in the image rendering engine, R r is the rotation matrix of the camera external parameter in the image rendering engine, T r is the translation matrix of the camera external parameter in the image rendering engine, are the three coordinate values of the k-th vertex of the minimum circumscribed cuboid in the world coordinate system, are the two pixel coordinate values of the projection of the k-th vertex of the minimum circumscribed cuboid.

Claims

1. A method for generating autonomous driving simulation data, characterized in that, Including the following steps: 1) Screen the road images without vehicles through the vehicle detection model and record them as target background images. Then, perform drivable lane annotation on the target background images through the drivable lane detection module to obtain lane annotation information. Finally, perform back-projection on the lane annotation information according to the camera parameters to obtain the drivable lanes in the world coordinate system; 2) Use the drivable lanes in the world coordinate system to screen the point cloud coordinate set corresponding to the target background image in the world coordinate system, and fit the lane plane equation according to the screened point cloud coordinate set in the world coordinate system; 3) Randomly generate the X-axis coordinate X in the world coordinate system within the drivable lane w and the Y-axis coordinate Y w , and calculate the Z-axis coordinate Z corresponding to the X-axis coordinate X w and the Y-axis coordinate Y w through the lane plane equation; w Compose the reference coordinate (X w , Y w , Z w ) from the X-axis coordinate X w , the X-axis coordinate Y w , and the Z-axis coordinate Z w ; 4) Calculate the yaw angle based on the central axis of the drivable lane, and then calculate the pitch angle and roll angle according to the fitted lane plane equation. The reference rotation angles (θ Pitch , θ Roll , θ Yaw ) are composed of the yaw angle, pitch angle and roll angle; 5) Randomly extract a vehicle model from the vehicle 3D model library and import it into the image rendering engine. Use the reference coordinates in 3) as the bottom geometric center coordinates of the vehicle model, and use the reference rotation angles (θ Pitch , θ Roll , θ Yaw ) as the rotation angles of the vehicle model; then randomly select a file from the high-dynamic range imaging library as the lighting information, and render the set vehicle model through the image rendering engine to obtain the foreground rendered image and the mask information; 6) According to the target background image, foreground rendering image, and mask information, perform image simulation using the style harmonization model to obtain a simulated image; 7) Calculate the 3D detection box annotation information projected onto the simulated image according to the size information of the current vehicle model and the camera parameters in the image rendering engine; The simulated data consists of the simulated image and the corresponding 3D detection box annotation information.

2. The method for generating autonomous driving simulation data according to claim 1, wherein In the above 1), the lane annotation information is the four vertex coordinates of a quadrilateral. According to the camera parameters, the four vertex coordinates are back-projected using the following formula to obtain the drivable lanes in the world coordinate system: M2 = R inv T Among them, each drivable lane is represented by 4 vertices, representing the reverse projection coefficient of the i-th vertex of the drivable lane, u i , v i respectively representing the horizontal and vertical coordinate values of the i-th vertex of the drivable lane in the pixel coordinate system, R represents the rotation matrix in the external camera parameters, T represents the translation matrix in the external camera parameters, K represents the internal camera matrix, R inv represents the inverse matrix of the rotation matrix, K inv represents the inverse matrix of the internal camera matrix, representing the three coordinate values of the i-th vertex in the world coordinate system, M1 represents the rotation-related coefficient matrix in the reverse projection, M2 represents the translation-related coefficient matrix in the reverse projection, M1[0, 0] represents the 0th row and 0th column of the rotation-related coefficient matrix M1 in the reverse projection, and M2[2, 0] represents the 2nd row and 0th column of the translation-related coefficient matrix M2 in the reverse projection.

3. A method for generating autonomous driving simulation data according to claim 1, characterized in that The above 2) is specifically: First, determine the coordinate ranges of the drivable lanes on the X-axis and Y-axis according to the drivable lanes in the world coordinate system. Then, screen the point cloud coordinate set corresponding to the target background image in the world coordinate system according to the coordinate ranges of the drivable lanes on the X-axis and Y-axis. Next, fit the lane plane equation according to the screened point cloud coordinate set in the world coordinate system using the following formula, and the formula is as follows: min‖An‖ ‖n‖=1 S points = {p1, p2,..., p m} p i = (x i , y i , z i ) p0 = (x0, y0, z0) n = [a b c] T Among them, min‖ ‖ represents the minimum value operation, and S points represents the set of point cloud coordinates in the world coordinate system obtained by screening, with a total of m points. A represents the vector matrix of all points in the point cloud set pointing to the mean point, n represents the plane normal vector of the lane, and p i represents the i-th point in the point cloud set corresponding to the target background image in the world coordinate system, x i , y i , z i represent the three coordinate values of the i-th point in the point cloud set. p0 represents the coordinate mean point of all points in the point cloud set, and x0, y0, z0 respectively represent the means of the three coordinate values corresponding to all points in the point cloud set. a, b, c represent the cubic coefficient, quadratic coefficient, and linear coefficient in the plane equation.

4. A method for generating autonomous driving simulation data according to claim 1, characterized in that, In the above 4), the angle between the plane normal vector of the lane plane equation and the x-axis of the world coordinate system is denoted as the pitch angle, and the angle between the plane normal vector of the lane plane equation and the y-axis of the world coordinate system is denoted as the roll angle.

5. A method for generating autonomous driving simulation data according to claim 1, wherein, The style harmonization model includes a first feature extraction module, a second feature extraction module, and a generation network. The target background image, foreground rendering image, and mask information are input into the style harmonization model. The target background image is used as the input of the first feature extraction module, the foreground rendering image is used as the input of the second feature extraction module, and the foreground rendering image is also used as the first input of the generation network. The outputs of the first feature extraction module and the second feature extraction module are respectively used as the second input and the third input of the generation network. The target background image is also used as the fourth input of the generation network, and the mask information is used as the fifth input of the generation network. The output of the generation network is used as the output of the style harmonization model.

6. The method for generating autonomous driving simulation data according to claim 5, wherein The first feature extraction module and the second feature extraction module have the same structure. The first feature extraction module includes 5 feature encoding modules, 4 feature decoding modules, 4 max pooling layers and 4 deconvolution layers. The input of the first feature extraction module serves as the input of the first feature encoding module. The first feature encoding module is connected to the fourth deconvolution layer after passing through the first max pooling layer, the second feature encoding module, the second max pooling layer, the third feature encoding module, the third max pooling layer, the fourth feature encoding module, the fourth max pooling layer and the fifth feature encoding module in sequence. The output of the fourth feature encoding module is concatenated with the output of the fourth deconvolution layer and then input into the fourth feature decoding module. The fourth feature decoding module is connected to the third deconvolution layer. The output of the third feature encoding module is concatenated with the output of the third deconvolution layer and then input into the third feature decoding module. The third feature decoding module is connected to the second deconvolution layer. The output of the second feature encoding module is concatenated with the output of the second deconvolution layer and then input into the second feature decoding module. The second feature decoding module is connected to the first deconvolution layer. The output of the first feature encoding module is concatenated with the output of the first deconvolution layer and then input into the first feature decoding module. The outputs of the second feature decoding module - the fourth feature decoding module are bilinearly interpolated respectively and then added together with the output of the first feature decoding module, and the final feature map is output and used as the output of the first feature extraction module.

7. A method for generating autonomous driving simulation data according to claim 5, characterized in that, The generation network includes 2 adaptive instance normalization layers and 6 fully connected layers. The second input and the third input of the generation network serve as the first input and the second input of the first adaptive instance normalization layer respectively. The first adaptive instance normalization layer is connected to the second adaptive instance normalization layer after passing through 6 fully connected layers in sequence. The first input of the generation network serves as the first input of the second adaptive instance normalization layer. The output of the second adaptive instance normalization layer is denoted as the harmonized foreground image. After performing a texturing operation on the harmonized foreground image with the fourth input and the fifth input of the generation network, a simulation image is obtained as the output of the generation network.

8. A method for generating autonomous driving simulation data according to claim 1, characterized in that, The 7) specifically is: Taking the bottom geometric center of the current vehicle model as the origin, 8 vertices of the minimum circumscribed cuboid of the current vehicle model are calculated according to the size information of the current vehicle model. The 8 vertices of the minimum circumscribed cuboid are projected to obtain the 3D detection frame annotation information. The calculation formula is as follows: Among them, K r is the internal camera parameter in the image rendering engine, R r is the rotation matrix of the external camera parameter in the image rendering engine, T r is the translation matrix of the external camera parameter in the image rendering engine, are the three coordinate values of the k-th vertex of the minimum bounding box in the world coordinate system, are the two pixel coordinate values of the projection of the k-th vertex of the minimum bounding box.

Citation Information

Patent Citations

  • Vehicular blind spot visualization method, device, terminal and system and vehicle

    CN107554430A

  • Autonomous driving-oriented lane self-checking algorithm in-loop simulation control method and system

    CN109947110A

Cited By

  • Lane line real-time identification method and system, electronic equipment and storage medium

    CN121482736A