Pre-fusion detection method and system for depth estimation optimization, vehicle and medium
By deeply optimizing and feature fusion of camera image data, combined with lidar point cloud data, the problem of insufficient depth estimation optimization in the existing technology is solved, and the target detection accuracy and performance of the autonomous driving system are improved.
Patent Information
- Application Number
- CN202510085150.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-16
AI Technical Summary
The existing object detection scheme that uses camera and laser multimodal pre-fusion to optimize depth estimation has less optimization, resulting in insufficient depth information processing and affecting the accuracy and performance of object detection.
By acquiring the point cloud data of the lidar and the image data of the camera, preprocessing is performed to obtain point cloud features and image features, then the image features are deeply optimized, and the trained depth estimation network is input to improve the accuracy of the depth prediction results, and the depth prediction results are converted into bird's-eye view features and point cloud features for feature fusion for target detection.
It improves the camera's depth perception ability to the environment, improves the accuracy and consistency of depth estimation, and enhances the model performance and recognition accuracy of target detection.
Smart Images

Figure CN120014625A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of autonomous driving perception technology, and in particular to a front fusion detection method, system, vehicle and medium for depth estimation optimization. Background Art
[0002] The forward fusion of LiDAR (Light Laser Detection and Ranging) and cameras plays a key role in autonomous driving, providing more comprehensive and accurate environmental perception information by comprehensively utilizing the advantages of the two sensors. LiDAR and cameras each have their own advantages in detection principles and data characteristics, and forward fusion achieves stronger environmental perception capabilities by comprehensively processing these data at an early stage.
[0003] In practical applications, cameras can capture high-resolution two-dimensional images, which can provide rich color, texture and detail information, and are suitable for identifying visual features such as traffic signs, lane lines and pedestrians. However, the depth perception capability of cameras is limited, and the performance will degrade in low light or bad weather conditions. LiDAR generates high-precision three-dimensional point cloud data by emitting laser pulses and measuring the reflection time. It can accurately measure the distance, shape and size of objects and is not affected by lighting conditions. However, the resolution and detail capture capability of LiDAR are not as good as cameras. Through the front fusion of LiDAR and camera, that is, combining the precise distance measurement data of LiDAR with the rich visual information of the camera, more accurate target detection and classification can be achieved. For example, the point cloud data of LiDAR is used to supplement the depth perception capability of the camera to improve the positioning and size estimation of objects; the image information of the camera can provide more detailed semantic information for the LiDAR point cloud, thereby improving the classification accuracy. This multi-sensor front fusion method not only improves the robustness and reliability of the autonomous driving system, but also provides more comprehensive environmental perception in complex environments, providing stronger support for autonomous driving decision-making and path planning. By making full use of the complementary advantages of the two sensors, front fusion technology is of great significance in the development of autonomous driving.
[0004] In addition, depth estimation is crucial in 3D target detection, and the accuracy of depth information directly affects the 3D positioning and detection accuracy of the target. However, depth estimation faces the following major challenges. The first is the sparsity of the depth map, that is, the point cloud data of LiDAR is usually sparse, and it is difficult to directly use this data for depth estimation; the second is the depth jump problem, that is, the depth of the edge area of the object changes dramatically, and the traditional method is prone to "depth jump" in this area, resulting in blurred target boundaries; finally, the depth difference between different objects, that is, the depth difference between objects is large, which increases the complexity of depth estimation.
[0005] In summary, the existing target detection schemes through multi-modal front fusion of cameras and lasers have little optimization for depth estimation, and no deep optimization of image data features, which makes it easy for features to be misaligned during feature fusion, and thus fails to achieve the optimal effect; and there is a lack of targeted processing of depth information, making it difficult to achieve the best effect of the model. Summary of the invention
[0006] In view of this, the present invention provides a front fusion detection method, system, vehicle and medium for optimizing depth estimation, so as to solve the problem proposed in the above technical background that the existing front fusion detection lacks optimization of depth estimation, resulting in many defects, which in turn seriously affects the depth estimation accuracy and the overall performance of target detection.
[0007] In a first aspect, the present invention provides a pre-fusion detection method for depth estimation optimization, the method comprising:
[0008] Obtain the point cloud data of the laser radar and the image data of the camera, and pre-process the point cloud data and the image data respectively to obtain the point cloud features and the image features respectively;
[0009] Perform deep optimization on image features to obtain deep image features;
[0010] Input the depth image features into the trained depth estimation network to obtain the depth prediction result, wherein the depth estimation network is obtained by supervised training of the camera image data based on the point cloud data of the lidar;
[0011] The bird's-eye view features are determined according to the depth prediction results, and the bird's-eye view features and point cloud features are fused to obtain fused features for target detection.
[0012] The present invention can enhance the camera's depth perception of the surrounding environment by optimizing the depth of image features; inputting the depth image features obtained by depth optimization into a depth estimation network obtained by supervised training of camera image data with point cloud data can further improve the depth estimation accuracy; converting the output of the depth estimation network into a bird's-eye view feature and combining it with the point cloud feature for feature fusion, helps to ensure the validity and accuracy of the depth information in the fused feature, thereby improving the model performance and recognition accuracy of subsequent target detection.
[0013] In an optional implementation, performing deep optimization on image features to obtain deep image features includes:
[0014] Analyze the depth relationship of the image features and generate at least one depth map accordingly, wherein the depth relationship is used to characterize the depth change of each feature area in the image features;
[0015] All depth maps are weighted separately to obtain multiple weighted feature maps;
[0016] All weight feature maps are processed at multiple scales to obtain multiple scale weight maps;
[0017] Based on a first preset feature fusion method, feature fusion is performed on all scale weight maps to obtain a fused feature map, and feature enhancement is performed on the fused feature map to obtain a deep image feature, wherein the first preset feature fusion method includes a bidirectional fusion strategy and feature cross-layer connection.
[0018] The present invention dynamically allocates feature weights according to the depth changes of each feature area in the image features, and performs bidirectional fusion and cross-layer connection on each weight feature map, which can realize adaptive adjustment of feature importance distribution according to depth changes, and at the same time fully utilizes the complementarity between shallow and deep features, further improves the robustness and accuracy of depth estimation, helps to ensure that the fused features have better depth estimation performance, and provides high-quality input for subsequent target detection.
[0019] In an optional implementation, the supervised training process of the depth estimation network includes:
[0020] Obtain point cloud dataset and image dataset;
[0021] Obtain the corresponding dense depth image and edge image for each point cloud data in the point cloud data set, and construct a supervision data set based on all the dense depth images and edge images;
[0022] Preprocess and deeply optimize each image data in the image dataset respectively, correspond to the deep image features, and construct the input dataset based on all the deep image features;
[0023] The input data set is used as the input of the depth estimation network, and the depth estimation network is supervised and trained using the supervised data set to obtain a trained depth estimation model; wherein, during the supervised training process, a gradient weight map is generated based on input data of different scales in the input data set, and the gradient weight map is used to adjust the supervised data in the supervised data set.
[0024] The present invention utilizes input data of different scales in the input data set to generate a gradient weight map, and based on the gradient weight map, adaptively adjusts the dense depth image and edge image generated by the point cloud data at different depth scales. It can realize separate supervision of each scale of the depth image features, thereby capturing the global and local depth details, solving the depth inconsistency problem caused by sparse projection to a certain extent, and ensuring the comprehensiveness and accuracy of depth estimation.
[0025] In an optional implementation, determining a bird's-eye view feature according to the depth prediction result, and fusing the bird's-eye view feature with the point cloud feature to obtain a fused feature for target detection, including:
[0026] Project the depth prediction result into the bird's-eye view space to obtain the bird's-eye view feature;
[0027] Based on a second preset feature fusion method, the bird's-eye view feature and the point cloud feature are fused to obtain a fused feature, wherein the second preset feature fusion method at least includes feature splicing and feature weighted averaging;
[0028] The fused features are input into the preset detection model for target detection.
[0029] The present invention projects the depth prediction results into the bird's-eye view space to achieve automatic alignment of different features, and combines the point cloud features to perform feature fusion to obtain fused features. The fused features can be better utilized for target detection, which helps to improve the accuracy and reliability of the detection model.
[0030] In an optional implementation, a dense depth image is generated based on point cloud data, and the process includes:
[0031] Get the calibration parameters, which include the laser radar parameters and the camera's internal and external parameters.
[0032] Based on the calibration parameters, the point cloud data is unified into the camera coordinate system through coordinate transformation;
[0033] Project the point cloud data in the camera coordinate system to the camera perspective and generate a corresponding depth image;
[0034] Divide the depth image into blocks to obtain multiple small block images;
[0035] Determine the target depth value of each small image block respectively, and fill the target depth value into the corresponding small image block, and obtain a plurality of filled images accordingly;
[0036] Generate a dense depth image based on all padded images.
[0037] The present invention can effectively avoid the problem of missing depth maps due to the sparsity of point cloud data by performing operations such as coordinate transformation, projection, blocking and depth filling on point cloud data, and ensure that point cloud data generates a dense depth image.
[0038] In an optional implementation, the edge image is obtained by performing edge detection on the dense depth image, and the process includes:
[0039] Compute the gradient of a dense depth image;
[0040] Determine edge information based on gradient;
[0041] Edge information is extracted from the dense depth image to obtain an extracted image, and the extracted image is normalized to obtain an edge image.
[0042] The present invention identifies edge positions in images through gradients and performs image normalization operations, thereby ensuring that edge images of different viewing angles are at the same scale and achieving optimal fusion of subsequent features.
[0043] In an optional implementation, the point cloud data and the image data are preprocessed respectively to obtain point cloud features and image features respectively, including:
[0044] Performing preset point cloud processing on the point cloud data, and extracting features from the processed point cloud data to obtain point cloud features, wherein the preset point cloud processing at least includes filtering, downsampling, and voxelization processing;
[0045] Performing preset image processing on the image data, and performing feature extraction on the processed image data to obtain image features, wherein the preset image processing at least includes denoising processing, color correction and enhancement processing.
[0046] The present invention can ensure the quality of point cloud features and image features through the preprocessing process of point cloud data and image data, which helps to improve depth estimation accuracy and target detection performance.
[0047] In a second aspect, the present invention provides a pre-fusion detection system for depth estimation optimization, the system comprising:
[0048] The data processing module is used to obtain the point cloud data of the laser radar and the image data of the camera, and pre-process the point cloud data and the image data respectively to obtain the point cloud features and the image features respectively;
[0049] A deep optimization module is used to perform deep optimization on image features to obtain deep image features;
[0050] The depth supervision module is used to input the depth image features into the trained depth estimation network to obtain the depth prediction result, wherein the depth estimation network is obtained by supervising the camera image data based on the laser radar point cloud data;
[0051] The feature fusion module is used to determine the bird's-eye view features according to the depth prediction results, and to fuse the bird's-eye view features with the point cloud features to obtain fused features for target detection.
[0052] The depth estimation optimized pre-fusion detection system of the present invention can ensure the validity and accuracy of the depth information in the fusion features, help to improve the depth estimation accuracy, and further improve the model performance and recognition accuracy of subsequent target detection.
[0053] In a third aspect, the present invention provides a vehicle, comprising a controller, the controller comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, computer instructions being stored in the memory, and the processor executing the computer instructions to execute a depth estimation optimized front fusion detection method according to the first aspect or any corresponding embodiment thereof.
[0054] In a fourth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute a depth estimation optimized front fusion detection method according to the first aspect or any corresponding embodiment thereof.
[0055] The depth estimation optimized pre-fusion detection method and system of the present invention can greatly enhance the camera's depth perception capability of the surrounding environment by optimizing the depth of image features; the depth estimation accuracy can be further improved by inputting the depth image features obtained by depth optimization into a depth estimation network obtained by supervised training of camera image data using point cloud data; in addition, converting the output of the depth estimation network into a bird's-eye view feature and combining it with the point cloud feature for feature fusion can help ensure the validity and accuracy of the depth information in the fused feature, thereby improving the model performance and recognition accuracy of subsequent target detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0057] Figure 1 is a flow chart of a pre-fusion detection method for depth estimation optimization according to an embodiment of the present invention;
[0058] Figure 2 is a flow chart of another depth estimation optimized front fusion detection method according to an embodiment of the present invention;
[0059] Figure 3 It is a schematic diagram of the structure of the front fusion detection framework;
[0060] Figure 4 It is a schematic diagram of various depth maps;
[0061] Figure 5 is a structural block diagram of a front fusion detection system for depth estimation optimization according to an embodiment of the present invention;
[0062] Figure 6It is a schematic diagram of the structure of a controller of a vehicle according to an embodiment of the present invention. DETAILED DESCRIPTION
[0063] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0064] An embodiment of the present invention provides an embodiment of a pre-fusion detection method for depth estimation optimization. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0065] In this embodiment, a depth estimation optimized front fusion detection method is provided. Figure 1 is a flow chart of a pre-fusion detection method for depth estimation optimization according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:
[0066] Step S101, obtaining the point cloud data of the laser radar and the image data of the camera, and preprocessing the point cloud data and the image data respectively to obtain corresponding point cloud features and image features.
[0067] In this embodiment, the specific acquisition means and preprocessing methods of the point cloud data and image data are not limited here, and can be adaptively determined according to actual needs. For example, acquiring image data through a camera installed on a vehicle and performing image denoising on the image data are only exemplary.
[0068] It should be noted that the point cloud features in this embodiment are also referred to as point cloud 3D features, and the image features are also referred to as image 2D features.
[0069] Step S102, performing depth optimization on the image features to obtain deep image features.
[0070] It should be noted that the depth optimization in this embodiment is intended to optimize the depth information of the image to improve the depth expression capability of the image.
[0071] Step S103, inputting the depth image features into a trained depth estimation network to obtain a depth prediction result, wherein the depth estimation network is obtained by supervised training of the camera image data based on the point cloud data of the laser radar.
[0072] It should be noted that the depth prediction result in this embodiment represents the depth of each pixel in the image data collected by the camera.
[0073] Step S104, determining the bird's-eye view features according to the depth prediction results, and fusing the bird's-eye view features with the point cloud features to obtain fused features for target detection.
[0074] It should be noted that the bird's-eye view feature in this embodiment is the BEV feature. In practical applications, with the development of autonomous driving technology, it has become an industry consensus to use the fusion of multiple sensors to achieve perception of the surrounding environment; among them, BEV (Bird's Eye View) provides a common feature representation for the fusion of multiple sensors. Whether it is based on different cameras or radars, as long as the features are projected in the BEV space, the alignment of different features can be achieved, and then the aligned features are sent to the subsequent network to perform the corresponding perception tasks, that is, the classification and detection functions.
[0075] The pre-fusion detection method for depth estimation optimization in an embodiment of the present invention can enhance the camera's depth perception capability of the surrounding environment and thereby improve the depth estimation accuracy by optimizing the depth of image features and inputting the depth image features obtained by the depth optimization into a depth estimation network obtained by supervised training of camera image data using point cloud data. In addition, by converting the output of the depth estimation network into a bird's-eye view feature and combining it with point cloud features for feature fusion, it helps to ensure the validity and accuracy of the depth information in the fused feature, thereby improving the model performance and recognition accuracy of subsequent target detection.
[0076] In this embodiment, a depth estimation optimized front fusion detection method is provided. Figure 2 FIG. 1 is a flow chart of another depth estimation optimized front fusion detection method according to an embodiment of the present invention. Figure 2 As shown, the process includes the following steps:
[0077] Step S201, obtaining the point cloud data of the laser radar and the image data of the camera, and preprocessing the point cloud data and the image data respectively to obtain corresponding point cloud features and image features.
[0078] Specifically, in the above step S201, the point cloud data and the image data are preprocessed respectively to obtain point cloud features and image features correspondingly, including:
[0079] Step A1, performing preset point cloud processing on the point cloud data, and performing feature extraction on the processed point cloud data to obtain point cloud features, wherein the preset point cloud processing at least includes filtering, downsampling and voxelization processing.
[0080] It should be noted that the specific method of extracting features from point cloud data in this embodiment is not limited here and can be adaptively adjusted according to actual needs, such as using a deep learning model for feature extraction, which is only an exemplary description.
[0081] In this embodiment, a series of preprocessing operations can be performed on the point cloud data to reduce the amount of calculation and improve the data quality. The specific steps of preprocessing include: using point cloud denoising algorithms, such as statistical filtering, radius filtering, etc., to remove noise points in the point cloud; using voxel grid filtering and other methods to downsample the point cloud to reduce the amount of point cloud data, thereby reducing the computational complexity.
[0082] It should be noted that the voxelization of point cloud data in this embodiment is to rasterize the three-dimensional point cloud according to a certain scale. Specifically, voxelization is to divide the three-dimensional space into small cubes (voxels) and map the point cloud data to these voxels to form a voxel grid. This representation method can better process and analyze three-dimensional data.
[0083] Step A2, performing preset image processing on the image data, and performing feature extraction on the processed image data to obtain image features, wherein the preset image processing at least includes denoising processing, color correction and enhancement processing.
[0084] It should be noted that the specific method of extracting features from image data in this embodiment refers to the point cloud data mentioned above and will not be repeated here.
[0085] In this embodiment, a series of preprocessing operations are performed on the image data to improve the quality of the image and the effect of subsequent processing. The specific steps of preprocessing include: using image denoising algorithms, such as Gaussian filtering, median filtering, etc., to remove noise in the image and improve the clarity of the image; using color correction technology to adjust the color balance of the image to ensure that the colors in the image are more realistic and consistent.
[0086] In the embodiment of the present invention, through the preprocessing process of point cloud data and image data, the quality of point cloud features and image features can be ensured, which helps to improve the depth estimation accuracy and target detection performance.
[0087] Step S202, performing depth optimization on the image features to obtain deep image features.
[0088] Specifically, the above step S202 includes:
[0089] Step S2021, analyzing the depth relationship of the image features, and correspondingly generating at least one depth map, wherein the depth relationship is used to characterize the depth change of each feature area in the image features.
[0090] In this embodiment, a dynamic attention mechanism is used to analyze the deep relationship of image features. It should be noted that the dynamic attention mechanism is a mechanism that can dynamically adjust the attention weight when processing input data, so that the model can focus more on key information, thereby improving the performance and robustness of the model. Specifically, the dynamic attention mechanism calculates the correlation or importance between different parts of the input data, assigns different attention weights to the data parts, and then dynamically adjusts the parameters and behavior of the model so that the model can better focus on capturing key information.
[0091] Step S2022: weights are assigned to all depth maps respectively, and a plurality of corresponding weight feature maps are obtained.
[0092] In this embodiment, an adaptive feature weighting strategy is used to assign weights to all depth maps. It should be noted that the adaptive feature weighting strategy is a strategy widely used in the field of machine learning. Its core idea is to dynamically adjust the weights of each part according to input data, model status or task requirements to improve the performance and flexibility of the model. It can be applied to various levels, such as loss functions, neural network layers, and adaptive weights in multi-task learning.
[0093] Step S2023, multi-scale processing is performed on all weight feature maps respectively, and a plurality of corresponding scale weight maps are obtained.
[0094] It should be noted that multi-scale processing refers to the analysis and processing of data or signals at different spatial or temporal scales. It usually uses different filters or decomposition methods to analyze signal structures of different scales from low to high. Specifically, in the field of deep learning, multi-scale processing usually refers to the simultaneous consideration and utilization of scale information of different inputs or feature layers during model training or inference. That is, for different input scales of deep learning models, the network can be allowed to process inputs at different resolutions through methods such as pyramid networks, multi-scale feature pyramids, and image pyramids, thereby capturing different details and levels in the image, and thus improving the performance of the model on various tasks.
[0095] Step S2024, performing feature fusion on all scale weight maps based on a first preset feature fusion method to obtain a fused feature map, and performing feature enhancement on the fused feature map to obtain a deep image feature, wherein the first preset feature fusion method includes a bidirectional fusion strategy and feature cross-layer connection.
[0096] It should be noted that the bidirectional fusion strategy refers to a top-down and bottom-up feature fusion strategy, that is, upsampling shallow features and adding them to deep features, and downsampling deep features and fusing them with shallow features; feature cross-layer connection is jump connection, a technology that directly connects the output and input of the network layer in a neural network. Its basic idea is to directly connect the input to the output in certain layers of the network to allow information to jump between different layers; this connection is usually implemented by addition operations, adding the input and output. Specifically, jump connections allow features at different levels to be fused, which promotes the flow and fusion of information. In addition, by skipping the processing of the intermediate levels, jump connections can pass the underlying detailed information to the higher levels, avoiding information loss and information bottlenecks.
[0097] In the embodiment of the present invention, feature weights are dynamically allocated according to the depth changes of each feature area in the image features, and each weight feature map is bidirectionally fused and cross-layer connected, so as to realize adaptive adjustment of feature importance distribution according to depth changes, and at the same time make full use of the complementarity between shallow and deep features, further improve the robustness and accuracy of depth estimation, help ensure that the fused features have better depth estimation performance, and provide high-quality input for subsequent target detection.
[0098] Step S203, inputting the depth image features into a trained depth estimation network to obtain a depth prediction result, wherein the depth estimation network is obtained by supervised training of the camera image data based on the point cloud data of the laser radar.
[0099] It should be noted that the supervised training process of the depth estimation network in this embodiment includes:
[0100] Step B1, obtaining a point cloud dataset and an image dataset.
[0101] In this embodiment, the specific method of obtaining the point cloud dataset and the image dataset can refer to the relevant content in the previous text, and will not be repeated here.
[0102] Step B2: Obtain corresponding dense depth images and edge images for each point cloud data in the point cloud data set, and construct a supervision data set based on all dense depth images and edge images.
[0103] It should be noted that a dense depth image refers to an image in which the depth value of each pixel is recorded in detail, reflecting the actual distance from each point in the scene to the image collector. This image contains not only the two-dimensional information of the object, but also the three-dimensional depth feature information of the object. Specifically, the dense depth image in this embodiment is generated based on point cloud data, and the process includes:
[0104] Step C1, obtaining calibration parameters, which include laser radar parameters and camera internal and external parameters.
[0105] In this embodiment, the acquisition method and specific content of the calibration parameters can be adaptively determined according to actual needs. For example, the external parameters of the laser radar include its position (translation vector) and posture (rotation matrix) in the vehicle coordinate system, and these parameters are used to convert the point cloud data of the laser radar into the global coordinate system of the vehicle; the external parameters of the camera also include its position and posture, while the internal parameters include parameters such as focal length, principal point coordinates and distortion coefficients, which are used to convert the image data of the camera into the global coordinate system of the vehicle and correct the geometric distortion in the image.
[0106] Step C2: Based on the calibration parameters, the point cloud data is unified into the camera coordinate system through coordinate transformation.
[0107] In this embodiment, the three-dimensional point cloud data of the laser radar needs to be projected to the camera's viewing angle to generate depth maps of multiple views, that is, coordinate transformation processing is performed on the point cloud data.
[0108] Step C3, projecting the point cloud data in the camera coordinate system to the camera viewing angle, and generating a corresponding depth image.
[0109] Step C4, dividing the depth image into blocks to obtain a plurality of small block images.
[0110] In this embodiment, the number of blocks can be adaptively determined according to actual needs and is not limited here. For example, a step size of 2 is used to divide the depth image in sequence.
[0111] Step C5, respectively determining the target depth value of each small image block, and filling the target depth value into the corresponding small image block, and correspondingly obtaining a plurality of filled images.
[0112] In this embodiment, the specific numerical value of the target depth value is not limited here, and is adaptively adjusted according to actual needs. For example, the target depth value is an average depth value, which is only used as an exemplary description.
[0113] Step C6, generating a dense depth image based on all filled images.
[0114] In the embodiment of the present invention, by performing operations such as coordinate transformation, projection, blocking and depth filling on the point cloud data, the problem of missing depth maps due to the sparsity of the point cloud data can be effectively avoided, ensuring that the point cloud data generates a dense depth image.
[0115] It should be noted that the edge image refers to the image obtained by edge extraction of the original image; the edge is the place where the brightness, color or texture and other features in the image change sharply. These changes usually represent the boundaries of different objects in the image. In practical applications, edge detection technology, such as grayscale processing, can be used to extract edge images. The edge image in this embodiment is obtained by edge detection of the dense depth image, and the process includes:
[0116] Step D1, calculate the gradient of the dense depth image.
[0117] In this embodiment, the gradient of each pixel in the image is calculated. It should be noted that in this embodiment, commonly used gradient calculation methods such as Sobel operator, Prewitt operator, etc. can be used for calculation.
[0118] Step D2, determining edge information based on the gradient.
[0119] In this embodiment, the gradient represents the rate of change of pixel values in the image, and the edge regions of the object can be highlighted according to the gradient (because the depth changes in these regions are usually large).
[0120] Step D3, extracting edge information from the dense depth image to obtain an extracted image, and normalizing the extracted image to obtain an edge image.
[0121] In the embodiment of the present invention, the edge position in the image is identified by gradient, and the image is normalized, which can ensure that edge images of different perspectives are at the same scale, and achieve optimal fusion of subsequent features.
[0122] Step B3, preprocessing and depth optimization are performed on each image data in the image data set, corresponding to the deep image features, and an input data set is constructed based on all the deep image features.
[0123] In this embodiment, the preprocessing and depth optimization of the image data may refer to the relevant content in the previous text, and will not be repeated here.
[0124] Step B4, using the input data set as the input of the depth estimation network, and using the supervised data set to perform supervised training on the depth estimation network to obtain a trained depth estimation model; wherein, during the supervised training process, a gradient weight map is generated based on input data of different scales in the input data set, and the gradient weight map is used to adjust the supervised data in the supervised data set.
[0125] In this embodiment, during the supervised training process, a gradient adaptive pooling algorithm is used to process input data of different scales in an input data set, and a gradient weight map is generated accordingly.
[0126] It should be noted that the gradient adaptive pooling algorithm is a technology used in deep learning models, mainly used for feature fusion, that is, by fusing features from different feature layers, each proposed region (ROI, Region of Interest) can obtain more comprehensive and rich feature information; its specific implementation steps include:
[0127] 1. Feature extraction: Extract feature maps from multiple levels in the feature pyramid.
[0128] 2. Feature adjustment: Adjust the feature map to the same size (such as 7x7) through operations such as ROI Align (Region of Interest Align).
[0129] 3. Feature fusion: perform splicing or element-by-element fusion in the channel dimension to form a fused feature map.
[0130] 4. Subsequent processing: The fused feature map is sent to the subsequent classification, regression or mask prediction subnetwork for further processing to output the final detection result.
[0131] In the embodiment of the present invention, input data of different scales in the input data set are used to generate a gradient weight map, and based on the gradient weight map, the dense depth image and edge image generated by the point cloud data are adaptively adjusted at different depth scales, which can achieve separate supervision of each scale of the depth image features, thereby capturing global and local depth details, solving the depth inconsistency problem caused by sparse projection to a certain extent, and ensuring the comprehensiveness and accuracy of depth estimation.
[0132] Step S204, determining the bird's-eye view features according to the depth prediction results, and fusing the bird's-eye view features with the point cloud features to obtain fused features for target detection.
[0133] Specifically, the above step S204 includes:
[0134] Step S2041, projecting the depth prediction result into the bird's-eye view space to obtain the bird's-eye view feature.
[0135] In this embodiment, the bird's-eye view space is the BEV space, and the specific method of converting the features into the bird's-eye view space is not limited here and can be implemented according to conventional means in the art.
[0136] Step S2042: Fusing the bird's-eye view features and the point cloud features based on a second preset feature fusion method to obtain fused features, wherein the second preset feature fusion method at least includes feature splicing and feature weighted averaging.
[0137] Step S2043, inputting the fused features into a preset detection model for target detection.
[0138] In the embodiment of the present invention, the depth prediction result is projected into the bird's-eye view space to realize automatic alignment of different features, and the feature fusion is performed with the point cloud features to obtain the fused features. The fused features can be better utilized for target detection, which helps to improve the accuracy and reliability of the detection model.
[0139] In practical applications, there are still great challenges in the depth estimation of targets in existing pre-fusion target detection schemes, such as the sparsity of depth maps and too many zero values, which increase the difficulty of fitting; the depth jumps in the edge areas of objects lead to blurred target boundaries; and the depth differences between different objects are large, making it difficult to accurately estimate. In response to the above defects, this embodiment proposes a new pre-fusion detection framework, which combines a dynamic scene analysis module and a multi-scale depth supervision module to solve related problems in depth estimation and improve the performance of 3D BEV target detection. In a specific embodiment, Figure 3 It is a schematic diagram of the structure of the front fusion detection framework. It should be noted that Figure 3 The dashed lines in are only used during the model prediction phase.
[0140] Specifically, the front fusion detection framework includes the following steps:
[0141] Step 1: Obtain a surround image through the 6-view surround camera equipment around the vehicle body, obtain laser point cloud data through the top mechanical scanning laser radar, and pre-process and extract features for them respectively.
[0142] In this embodiment, the detailed process of step 1 also includes:
[0143] Step 1.1: Multi-view image acquisition. Specifically, six perspective cameras are installed around the autonomous vehicle, which can cover the entire field of view around the vehicle, thereby acquiring a complete environmental image, namely, a surround image, which is used to provide rich visual information, such as the color, shape, and texture of objects.
[0144] Step 1.2: Image preprocessing: Specifically, image preprocessing includes denoising, color correction, etc. to improve image quality.
[0145] Step 1.3: Image feature extraction. The preprocessed image is input into a deep learning model for feature extraction, such as the SwinTransformer neural network model, which is a transformer-based neural network architecture with powerful feature extraction capabilities. The model extracts multi-scale features of the image layer by layer through a hierarchical transformer module to capture local and global information in the image. The features extracted by the model are represented on the image plane to form a two-dimensional feature map of the image, where the features contain the shape, texture, and other important information of the object.
[0146] Step 1.4: Acquisition of laser point cloud data. Specifically, a mechanical scanning laser radar (LiDAR) is installed on the top of the autonomous driving vehicle to obtain three-dimensional laser point cloud data around the vehicle. The laser radar generates a high-density three-dimensional point cloud by emitting laser beams and receiving reflected signals, providing accurate distance information of the environment.
[0147] Step 1.5: Preprocessing of point cloud data. Similar to image data, laser point cloud data also needs to be preprocessed to reduce the amount of calculation and improve data quality. Specifically, it includes denoising and downsampling to ensure the quality of point cloud data.
[0148] Step 1.6: In order to process the point cloud data more efficiently, it needs to be voxelized.
[0149] Step 1.7: The voxelized point cloud data is input into a deep learning model for feature extraction, such as the VoxelNet neural network model, which is a neural network architecture specifically designed for point cloud data processing. The model extracts spatial features from point cloud data through voxelized input and 3D convolution operations. Specifically, VoxelNet can capture the three-dimensional structural information of point cloud data, including the shape, position, and size of objects; the extracted features are represented in the BEV space to form a three-dimensional feature map of the point cloud, which contains the spatial geometric information in the environment and is essential for target detection.
[0150] Step 2: Obtain the external parameters of the lidar and the calibration parameters such as the external parameters and intrinsic parameters of the camera; based on the calibration parameters, project the input point cloud onto the multi-view depth map, and divide the generated multi-view depth map into multiple blocks according to a certain step size; then, through the expansion operation, fill the entire block with the maximum depth value of each block to generate a dense depth map.
[0151] In this embodiment, the detailed process of step 2 also includes:
[0152] Step 2.1: Obtain calibration parameters.
[0153] Step 2.2: Projection of point cloud data. Specifically, the point cloud data is converted from the laser radar coordinate system to the vehicle coordinate system using the external parameters of the laser radar; the point cloud data is converted from the vehicle coordinate system to the camera coordinate system and pixel coordinate system using the external and internal parameters of the camera; the corresponding depth map is generated according to the pixel coordinates and depth value (i.e., distance) of the point cloud data under the camera's perspective. Among them, these depth maps reflect the distance information of the laser radar point cloud under different camera perspectives.
[0154] Step 2.3: Depth map block processing. Specifically, the generated multi-view depth map is divided into multiple small blocks according to a certain step size, and each small block represents a part of the depth map. Through this block processing, the depth information can be better managed and processed; at the same time, the sparse areas in the depth map can be filled, reducing the problem of missing depth maps caused by the sparsity of point cloud data.
[0155] Step 2.4: Dilation operation and dense depth map generation.
[0156] In this embodiment, in order to generate a dense depth map, each block needs to be expanded. Specifically, the depth values in each block are first counted to find the maximum depth value in the block; the calculated maximum depth value is used to fill the entire block area. In this way, the sparse depth map can be made denser and smoother to reduce holes and discontinuous areas in the depth map.
[0157] Step 3: Based on the dense depth map generated in step 2, calculate the gradients of the multi-view dense depth map on the x-axis and y-axis to extract edge-aware three-dimensional geometric information; then, through maximum pooling and normalization operations, obtain a multi-view edge map to represent the edges of different objects.
[0158] In this embodiment, the detailed process of step 3 also includes:
[0159] Step 3.1: Calculate the gradient of the multi-view dense depth map in the x-axis and y-axis directions. The purpose of this step is to identify the edge position in the image through the depth change rate.
[0160] Step 3.2: In order to enhance the feature representation of the edge map, it is necessary to perform a maximum pooling operation on the edge map. Among them, maximum pooling can compress image data while retaining significant edge features. After the maximum pooling process, the edge map needs to be normalized to ensure that edge maps from different perspectives are compared and fused at the same scale. Through the above steps, normalized multi-view edge maps can be obtained. These edge maps represent the edge information of different objects at different perspectives and are an important basis for realizing edge-perceived three-dimensional geometric information.
[0161] In a specific embodiment, Figure 4is a schematic diagram of multiple depth maps. It should be noted that the "point cloud projection depth map example" in the figure is obtained through the "point cloud data projection" in the above step 2.2; the "fine-grained depth map example" is a dense depth map generated by the above step 2.4; and the "edge depth map example" is a multi-view edge map generated by the above step 3.2. Specifically, Figure 4 It can be seen that after the depth optimization processing of this embodiment, the depth information of the depth map generated by the point cloud data is effectively improved, and a high-quality depth image can be obtained.
[0162] Step 4: Dynamic scene analysis module, based on the image 2D features obtained in step 1, uses dynamic feature fusion strategy to extract more accurate geometric and semantic information from multi-view images, dynamically allocates feature weights through adaptive attention mechanism, especially in edge areas, uses cross-layer fusion technology to combine shallow edge features with deep semantic features, thereby enhancing the model's sensitivity to depth changes.
[0163] In this embodiment, the detailed process of step 4 also includes:
[0164] Step 4.1: The dynamic attention mechanism is the basic component of the dynamic scene analysis module, which is used to analyze the depth changes of the scene and dynamically adjust the feature processing method according to the complexity of each area. First, the multi-head self-attention mechanism is used to globally model the input multi-view image features. This mechanism captures the depth relationship between different objects in the scene by calculating the correlation between each pixel. The module decomposes the features of the multi-view image into several subspaces and then calculates self-attention for each subspace. Specifically, through the attention weights, the module can highlight the important feature areas related to the depth jump, while weakening the influence of the background area; further divide the scene into depth consistent areas and depth mutation areas. By performing gradient analysis and edge detection on the feature map generated by multi-head self-attention, the edge areas with significant depth changes in the scene can be identified. Among them, the depth consistent area uses conventional feature processing, while the depth mutation area adopts a more complex strategy, such as giving higher weights in subsequent feature fusion to cope with rapid depth changes.
[0165] Step 4.2: Assign weights using an adaptive feature weighting strategy.
[0166] In this embodiment, a lightweight feature weighting network is used to assign feature weights. The network contains several convolutional layers, and the input is the feature map calculated in the previous step and its corresponding gradient information. By analyzing the feature map, the network can assign appropriate weights to different regions to ensure that the features of the deep mutation region receive higher attention in subsequent processing. In this way, it is possible to effectively distinguish between depth consistent regions and mutation regions during feature fusion. In addition, considering that depth changes may behave differently at different scales, a multi-scale feature weighting method is used. Specifically, the module processes the input feature map at multiple scales, generates feature maps of multiple scales through different convolution kernels and downsampling operations, and calculates a separate weight map for each scale; subsequently, these weight maps are applied to the corresponding feature maps, and the multi-scale features are fused to obtain a complete scene representation.
[0167] Step 4.3: Use cross-layer fusion technology to make full use of the complementarity between shallow and deep features to improve the accuracy of depth estimation. Among them, shallow features have better edge and texture information, while deep features contain richer semantic information. Specifically, a feature pyramid structure is adopted to organize feature maps of different layers according to spatial resolution and semantic information; at each feature layer, the module fuses features from shallow and deep layers to achieve the complementarity of edge and semantic information. In order to achieve the above process, a bidirectional fusion strategy is used to upsample shallow features and add them to deep features, and to downsample deep features and fuse them with shallow features. At the same time, jump connections are set to further strengthen the fusion of shallow and deep features. Through this connection method, the edge features of the shallow layer can be directly passed to the deep layer, which helps to maintain accurate depth boundary information. The last step is to enhance the features of the fused feature map, such as compressing or expanding the dimension of the feature map to the required size through a series of convolutional layers and nonlinear activation functions, while improving the expressiveness of the features, which can ensure that the fused features have better depth estimation performance and provide high-quality input for subsequent 3D object detection.
[0168] Step 5: Multi-scale depth supervision module, based on the multi-view edge map generated in step 3, provides fine-grained supervision at different depth scales, thereby improving the overall accuracy of depth estimation. Specifically, a multi-scale depth supervision strategy is introduced, combined with the dense depth map information generated in step 2, and the depth inconsistency problem caused by depth projection is alleviated through the gradient adaptive pooling method; at the same time, a dual loss function is used, including global consistency loss and local detail loss to handle global and local depth differences respectively.
[0169] In this embodiment, the detailed process of step 5 also includes:
[0170] Step 5.1: The core idea of the multi-scale deep supervision strategy is to supervise the multi-scale depth feature map input from step 4.3 separately at each scale to capture global and local depth details. Specifically, by optimizing each scale separately, the model can learn deep features at different resolutions. This multi-scale decomposition method can help the model learn global depth information at a coarse-grained scale, while capturing local depth changes at a fine-grained scale, ensuring the comprehensiveness and accuracy of depth estimation. For each scale, the supervision network will use the dense depth feature map obtained from step 2.4 and the edge map obtained in step 3.2 to downsample proportionally to generate multiple depth maps of different scales as supervision signals to guide the learning of the depth estimation network. It should be noted that the depth map should be adjusted to the predicted depth of the current scale through interpolation or completion operations. Figure 1 The resolution is very high.
[0171] Step 5.2: Since a sparse depth map is formed after the 3D point cloud is projected onto a 2D plane, this sparsity makes it difficult for the depth estimation network to capture complete depth features during training. Therefore, the gradient adaptive pooling method is introduced to solve the depth inconsistency problem caused by sparse projection. The gradient of the predicted depth map is calculated at each scale to generate a gradient map to reflect the magnitude of the depth change. The gradient map is used to generate an adaptive weight map to adjust the strength of the supervisory signal. In addition, through the normalization operation, the values in the gradient map are mapped to the range of [0,1] to ensure that high gradient areas (i.e., areas with drastic depth changes) receive higher weights, while low gradient areas (i.e., areas with stable depth) receive lower weights.
[0172] Step 5.3: In order to optimize the depth estimation network at the global and local levels, a dual loss function is used, including global consistency loss and local detail loss. Among them, the global consistency loss is used to constrain the stability of the overall depth estimation and ensure that the depth estimation network learns the global depth distribution in a large range. The loss function uses a smooth L1 loss to reduce the overall error caused by depth differences. The smooth L1 loss approximates the L2 loss when the error is small, and degenerates to the L1 loss when the error is large, balancing the sensitivity and stability of the error. The local detail loss focuses on areas with drastic depth changes, helping the network capture fine depth details. To this end, the local loss is weighted using a gradient adaptive weight map, so that the error in the high gradient area contributes more to the overall loss.
[0173] Step 6: Based on the image feature depth prediction obtained through supervised learning in steps 4 and 5, the 2D features of the image plane obtained in step 1 are projected into the BEV space through the Lift-splat-shot algorithm (LSS, which is used to convert image features from multiple cameras into a unified BEV space to achieve 360° field of view perception) to obtain the image BEV features. The obtained image BEV features are then fused with the point cloud 3D features obtained in step 1 and input into the subsequent detection head module for target detection.
[0174] In this embodiment, the detailed process of step 6 also includes:
[0175] Step 6.1: Use the LSS algorithm to project the 2D features of the image plane into the BEV space. Specifically, in the Lift stage, the depth prediction results are used to "lift" each pixel in the image into the three-dimensional space, that is, the coordinates of each pixel on the image plane are mapped to points in the three-dimensional space through the depth value. In the Splat stage, the points lifted into the three-dimensional space are reprojected into the BEV space. This step projects the three-dimensional point cloud into a unified bird's-eye view to form the BEV features of the image. These features not only contain the two-dimensional information of the image, but also combine the depth information to provide a more comprehensive environmental perception capability. Through the LSS algorithm, the 2D features of the image plane are successfully projected into the BEV space to obtain the BEV image features. These features appear as a unified feature map in the bird's-eye view, which contains image information and depth information obtained from multiple perspectives; the BEV image features provide a rich information basis for subsequent feature fusion and target detection.
[0176] Step 6.2: Fuse BEV image features and point cloud 3D features. Specifically, the purpose of feature fusion is to combine data from different sensors to form a unified feature representation. The fused features contain rich information from the image and point cloud, and can describe the environment more comprehensively.
[0177] Step 6.3: The fused features are input to the subsequent detection head module for target detection. Based on the extracted features, the detection head module classifies and locates the target and outputs the category and three-dimensional coordinates of each target. With this information, the autonomous driving system can accurately identify and track targets in the environment.
[0178] In summary, the pre-fusion detection method for depth estimation optimization in the embodiment of the present invention can enhance the camera's depth perception capability of the surrounding environment, ensure the validity and accuracy of the depth information in the fusion features, further improve the depth estimation accuracy, and thus improve the model performance and recognition accuracy of subsequent target detection.
[0179] In this embodiment, a depth estimation optimized front fusion detection system is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and will not be repeated here. As used below, the term "module" refers to a combination of software and / or hardware that can implement a predetermined function. Although the system described in the following embodiments is preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.
[0180] The present invention provides a pre-fusion detection system with depth estimation optimization, such as Figure 5 As shown, the system includes:
[0181] The data processing module 501 is used to obtain the point cloud data of the laser radar and the image data of the camera, and pre-process the point cloud data and the image data respectively to obtain the point cloud features and the image features accordingly.
[0182] The depth optimization module 502 is used to perform depth optimization on the image features to obtain deep image features.
[0183] The depth supervision module 503 is used to input the depth image features into the trained depth estimation network to obtain the depth prediction result, wherein the depth estimation network is obtained by supervising the camera image data based on the point cloud data of the laser radar.
[0184] The feature fusion module 504 is used to determine the bird's-eye view features according to the depth prediction results, and fuse the bird's-eye view features with the point cloud features to obtain fused features for target detection.
[0185] In some optional embodiments, the data processing module 501 includes: a first processing submodule and a second processing submodule; wherein the first processing submodule is used to perform preset point cloud processing on the point cloud data, and perform feature extraction on the processed point cloud data to obtain point cloud features, wherein the preset point cloud processing includes at least filtering, downsampling and voxelization processing; the second processing submodule is used to perform preset image processing on the image data, and perform feature extraction on the processed image data to obtain image features, wherein the preset image processing includes at least denoising processing, color correction and enhancement processing.
[0186] In some optional embodiments, the depth optimization module 502 includes: a first optimization submodule, a second optimization submodule, a third optimization submodule and a fourth optimization submodule; wherein the first optimization submodule is used to analyze the depth relationship of image features and generate at least one depth map, wherein the depth relationship is used to characterize the depth change of each feature area in the image feature; the second optimization submodule is used to assign weights to all depth maps respectively, and obtain multiple weight feature maps; the third optimization submodule is used to perform multi-scale processing on all weight feature maps respectively, and obtain multiple scale weight maps; the fourth optimization submodule is used to perform feature fusion on all scale weight maps based on the first preset feature fusion method to obtain a fused feature map, and perform feature enhancement on the fused feature map to obtain a deep image feature, wherein the first preset feature fusion method includes a bidirectional fusion strategy and feature cross-layer connection.
[0187] In some optional embodiments, the depth supervision module 503 includes: a first training submodule, a second training submodule, a third training submodule and a fourth training submodule; wherein the first training submodule is used to obtain a point cloud data set and an image data set; the second training submodule is used to obtain corresponding dense depth images and edge images for each point cloud data in the point cloud data set, and construct a supervision data set based on all dense depth images and edge images; the third training submodule is used to preprocess and deeply optimize each image data in the image data set, correspond to the depth image features, and construct an input data set based on all the depth image features; the fourth training submodule is used to use the input data set as the input of the depth estimation network, and use the supervision data set to supervise the depth estimation network to obtain a trained depth estimation model; wherein, during the supervised training process, a gradient weight map is generated based on input data of different scales in the input data set, and the gradient weight map is used to adjust the supervision data in the supervision data set.
[0188] In some optional embodiments, the second training submodule includes: a data generation unit and an edge detection unit; wherein the data generation unit is used to obtain calibration parameters, the calibration parameters include lidar parameters and intrinsic and extrinsic parameters of the camera; based on the calibration parameters, the point cloud data is unified to the camera coordinate system through coordinate transformation; the point cloud data under the camera coordinate system is projected to the camera perspective, and a depth image is generated accordingly; the depth image is divided into blocks to obtain a plurality of small block images; the target depth value of each small block image is determined respectively, and the target depth value is filled into the corresponding small block image, and a plurality of filled images are obtained accordingly; a dense depth image is generated based on all filled images; an edge detection unit is used to calculate the gradient of the dense depth image; edge information is determined based on the gradient; edge information is extracted from the dense depth image to obtain an extracted image, and the extracted image is normalized to obtain an edge image.
[0189] In some optional embodiments, the feature fusion module 504 includes: a first fusion sub-module, a second fusion sub-module and a third fusion sub-module; wherein the first fusion sub-module is used to project the depth prediction result to the bird's-eye view space to obtain the bird's-eye view feature; the second fusion sub-module is used to fuse the bird's-eye view feature and the point cloud feature based on a second preset feature fusion method to obtain a fused feature, wherein the second preset feature fusion method at least includes feature splicing and feature weighted averaging; the third fusion sub-module is used to input the fused feature into a preset detection model for target detection.
[0190] The further functional description of each of the above modules is the same as that of the above corresponding embodiments and will not be repeated here.
[0191] The depth estimation optimized pre-fusion detection system in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.
[0192] The depth estimation optimized pre-fusion detection system of the embodiment of the present invention can ensure the validity and accuracy of the depth information in the fusion features, help improve the depth estimation accuracy, and further improve the model performance and recognition accuracy of subsequent target detection.
[0193] A vehicle is also provided in an embodiment of the present invention, and the vehicle includes a controller. The controller in this embodiment is a vehicle controller, which is used to perform operations such as power supply / power off, sleep and wake up of the sub-controllers and network nodes under it, and each power supply interface thereof can collect the real-time current output. Other controllers with the above functions are applicable.
[0194] Figure 6 is a schematic diagram of the structure of the controller provided by an optional embodiment of the present invention, such as Figure 6As shown, the controller includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. The various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process instructions executed in the controller, including instructions stored in or on the memory to display graphical information of a GUI on an external input / output system (such as a display device coupled to an interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple controllers can be connected, and each controller provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 6 A processor 10 is taken as an example.
[0195] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.
[0196] The memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiment.
[0197] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required by at least one function; the data storage area may store data created according to the use of the controller, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the controller via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0198] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.
[0199] The controller also includes a communication interface 30 for the main control chip to communicate with other devices or a communication network.
[0200] A computer-readable storage medium is also provided in an embodiment of the present invention. The method according to the embodiment of the present invention can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium and downloaded through a network, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state hard disk, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor main control chip or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the method shown in the above embodiment is implemented.
[0201] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the scope defined by the appended claims.
Claims
1. A pre-fusion detection method for depth estimation optimization, characterized in that: The method comprises: Acquire point cloud data of the laser radar and image data of the camera, and preprocess the point cloud data and the image data respectively to obtain point cloud features and image features accordingly; Performing depth optimization on the image features to obtain deep image features; Inputting the depth image features into a trained depth estimation network to obtain a depth prediction result, wherein the depth estimation network is obtained by supervised training of camera image data based on the point cloud data of the laser radar; A bird's-eye view feature is determined according to the depth prediction result, and the bird's-eye view feature and the point cloud feature are fused to obtain a fused feature for target detection.
2. The depth estimation optimized pre-fusion detection method according to claim 1, characterized in that: The step of performing deep optimization on the image features to obtain deep image features includes: Analyze the depth relationship of the image features and generate at least one depth map accordingly, wherein the depth relationship is used to characterize the depth change of each feature area in the image features; All depth maps are weighted separately to obtain multiple weighted feature maps; All weight feature maps are processed at multiple scales to obtain multiple scale weight maps; Based on a first preset feature fusion method, feature fusion is performed on all scale weight maps to obtain a fused feature map, and feature enhancement is performed on the fused feature map to obtain a deep image feature, wherein the first preset feature fusion method includes a bidirectional fusion strategy and feature cross-layer connection.
3. The depth estimation optimized pre-fusion detection method according to claim 2, characterized in that: The supervised training process of the depth estimation network includes: Obtain point cloud dataset and image dataset; Obtaining corresponding dense depth images and edge images for each point cloud data in the point cloud data set, and constructing a supervision data set based on all the dense depth images and edge images; Preprocessing and depth optimization are performed on each image data in the image data set to correspond to deep image features, and an input data set is constructed based on all the deep image features; The input data set is used as the input of the depth estimation network, and the depth estimation network is supervised trained using the supervised data set to obtain a trained depth estimation model; wherein, during the supervised training process, a gradient weight map is generated based on input data of different scales in the input data set, and the supervised data in the supervised data set is adjusted using the gradient weight map.
4. The depth estimation optimized front fusion detection method according to any one of claims 1 to 3, characterized in that: The determining of the bird's-eye view features according to the depth prediction results, and fusing the bird's-eye view features with the point cloud features to obtain fused features for target detection, includes: Project the depth prediction result into the bird's-eye view space to obtain the bird's-eye view feature; Based on a second preset feature fusion method, the bird's-eye view feature and the point cloud feature are subjected to feature fusion to obtain a fusion feature, wherein the second preset feature fusion method at least includes feature splicing and feature weighted averaging; The fused features are input into a preset detection model for target detection.
5. The depth estimation optimized pre-fusion detection method according to claim 3, characterized in that: The dense depth image is generated based on the point cloud data, and the process includes: Acquire calibration parameters, where the calibration parameters include laser radar parameters and camera intrinsic and extrinsic parameters; Based on the calibration parameters, the point cloud data is unified into the camera coordinate system through coordinate transformation; Project the point cloud data in the camera coordinate system to the camera perspective and generate a corresponding depth image; Divide the depth image into blocks to obtain a plurality of small block images; Determine the target depth value of each small image block respectively, and fill the target depth value into the corresponding small image block, so as to obtain a plurality of filled images; Generate a dense depth image based on all padded images.
6. The depth estimation optimized pre-fusion detection method according to claim 5, characterized in that: The edge image is obtained by performing edge detection on the dense depth image, and the process includes: Compute the gradient of a dense depth image; determining edge information based on the gradient; The edge information is extracted from the dense depth image to obtain an extracted image, and the extracted image is normalized to obtain an edge image.
7. The depth estimation optimized pre-fusion detection method according to claim 1, characterized in that: The preprocessing of the point cloud data and the image data respectively to obtain point cloud features and image features accordingly includes: Performing preset point cloud processing on the point cloud data, and performing feature extraction on the processed point cloud data to obtain point cloud features, wherein the preset point cloud processing at least includes filtering, downsampling and voxelization processing; The image data is subjected to a preset image processing, and feature extraction is performed on the processed image data to obtain image features, wherein the preset image processing at least includes denoising processing, color correction and enhancement processing.
8. A pre-fusion detection system with depth estimation optimization, characterized in that: The system comprises: A data processing module is used to obtain the point cloud data of the laser radar and the image data of the camera, and pre-process the point cloud data and the image data respectively to obtain point cloud features and image features accordingly; A depth optimization module, used to perform depth optimization on the image features to obtain deep image features; A depth supervision module, used for inputting the depth image features into a trained depth estimation network to obtain a depth prediction result, wherein the depth estimation network is obtained by supervised training of the camera image data based on the point cloud data of the laser radar; The feature fusion module is used to determine the bird's-eye view features according to the depth prediction results, and fuse the bird's-eye view features with the point cloud features to obtain fused features for target detection.
9. A vehicle, characterized in that: The vehicle includes a controller, and the controller includes: a memory and a processor, the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the depth estimation optimized front fusion detection method described in any one of claims 1 to 7 by executing the computer instructions.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the depth estimation optimized front fusion detection method according to any one of claims 1 to 7.
Citation Information
Cited By
Laser radar and camera pose self-checking method and device based on aerial view
CN121685663A
Light-weight self-supervision monocular depth estimation method and device based on cross-sequence interaction
CN121810756A
Multi-sensor fusion pallet identification method and system based on BEV characteristics
CN122416432A