Bird's eye view based pure vision three-dimensional target detection system and method
Patent Information
- Application Number
- CN202610757024.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-29
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2046-05-29
AI Technical Summary
1.深度预测精度与模型复杂度难以兼顾,高精度模型计算量过大难以车载部署,轻量级模型深度估计精度不足;
1.深度预测精度提升:通过在深度预测前引入可学习的二维语义信息融合机制,利用语义分割网络区分主体与背景区域,通过通道融合方式自适应增强关键信息、抑制背景冗余信息,使深度预测网络能够聚焦于有效区域,在轻量级网络结构下实现深度估计精度的显著提升。
Smart Images

Figure CN122290076B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of environmental perception, and in particular to a method and system for three-dimensional target detection based on a bird's-eye view (BEV) using pure visual input. Background Technology
[0002] The visual perception system of autonomous vehicles needs to extract semantic information from images collected by multiple cameras and fuse this information into a unified bird's-eye view to support subsequent path planning and decision control.
[0003] In existing technologies, pure vision-based 3D object detection methods based on BEVs mainly fall into two categories: explicit transformation and implicit transformation. A typical example of explicit transformation is the LSS (Lift, Splat, Shoot) model, which predicts the depth distribution of image pixels through a depth estimation network, mapping 2D features to 3D space and projecting them onto the BEV plane. However, due to the relatively simple depth prediction network used in the LSS model, its depth estimation accuracy is limited, leading to significant errors in BEV semantic prediction results. A typical example of implicit transformation is the BEVFormer model, which uses a Transformer architecture to implicitly model the projection relationship from 3D space to the 2D image through a self-attention mechanism. While this improves depth prediction accuracy, its complex network structure and massive computational cost make it difficult to deploy and run in real-time on automotive embedded chips.
[0004] Furthermore, when processing multi-view image fusion, existing technologies typically input the original image features directly into the depth prediction network without fully considering the interference of background redundancy information on the depth estimation accuracy. In terms of utilizing temporal information, existing methods either ignore historical frame information, leading to unstable single-frame detection, or employ complex attention mechanisms, resulting in excessive computational overhead, making it difficult to achieve a balance between accuracy and efficiency.
[0005] Therefore, the existing technology has the following technical defects: 1. It is difficult to balance depth prediction accuracy and model complexity. High-precision models have too much computational cost and are difficult to deploy on vehicles, while lightweight models have insufficient depth estimation accuracy. 2. The fusion of features from multiple perspectives of 2D images did not effectively suppress redundant background information, affecting the accuracy of depth estimation; 3. The temporal information fusion method has high computational complexity, or fails to effectively utilize historical frame information to supplement detection stability in occluded or blurred scenes. Summary of the Invention
[0006] The summary of this invention introduces a series of simplified concepts, all of which are simplifications of existing technologies in the field, and will be further explained in detail in the detailed description section. This summary is not intended to limit the key features and essential technical features of the claimed technical solution, nor is it intended to determine the scope of protection of the claimed technical solution.
[0007] The technical problem to be solved by the present invention is to provide a pure visual 3D target detection method and system based on bird's-eye view with 2D semantic fusion and temporal information enhancement, which improves depth prediction accuracy and 3D target detection stability while keeping the model lightweight.
[0008] To address the aforementioned technical problems, this invention provides a pure visual 3D target detection system based on a bird's-eye view, comprising: The two-dimensional image feature extraction module is configured to extract features from the original images of the vehicle from multiple perspectives at the current moment through a two-dimensional image feature extraction network to obtain two-dimensional image features. The two-dimensional semantic information fusion module is configured as follows: The original image is semantically segmented using a semantic segmentation network to obtain a two-dimensional semantic segmentation mask for the subject and background. The two-dimensional semantic segmentation mask and the two-dimensional image features are fused together to obtain two-dimensional features with fused semantic information. The depth feature extraction module is configured to perform depth prediction on the two-dimensional features fused with semantic information through a depth prediction network to obtain a depth estimate of each pixel, and combine the two-dimensional image features to obtain two-dimensional features with depth information. The BEV feature projection module is configured to project the two-dimensional features with depth information onto the BEV feature space based on the camera intrinsic parameters, extrinsic parameters, and the depth estimation value to obtain the BEV features of the current frame. The time-series information fusion module is configured as follows: Obtain historical frame BEV features, and align the historical frame BEV features to the current frame coordinate system according to the vehicle pose transformation matrix; The aligned historical frame BEV features are concatenated and fused with the current frame BEV features in the channel dimension to obtain BEV features with fused temporal information. The three-dimensional target detection module is configured to input the BEV features fused with temporal information into the BEV perception network and output the three-dimensional target detection results.
[0009] Preferably, in a further improvement of the bird's-eye view-based pure visual 3D target detection system, the 2D image feature extraction network is the EfficientNetB0 network.
[0010] Preferably, the pure visual 3D target detection system based on bird's-eye view is further improved, wherein the semantic segmentation network includes a data shape transformation layer, a convolutional layer, a ResNet18 backbone network, and an output shape transformation layer.
[0011] Preferably, the pure visual 3D target detection system based on a bird's-eye view is further improved by fusion of the channels into a single concatenation along the channel dimension, so that the channel weights of the semantic segmentation mask are adaptively learned during training.
[0012] Preferably, in a further improvement of the pure visual 3D target detection system based on a bird's-eye view, the depth prediction network includes a two-dimensional convolutional layer, which expands the input feature channels to C+D channels, where C channels are used to retain the original image feature information and D channels are used for pixel-level depth prediction.
[0013] Preferably, the pure visual 3D target detection system based on bird's-eye view is further improved, wherein the BEV feature space is a BEV grid with the vehicle position as the center, a specified range and a specified grid resolution. In this process, multiple pixel features that fall into the same grid are added together and fused.
[0014] To address the aforementioned technical problems, this invention provides a pure visual 3D target detection method based on a bird's-eye view, which is implemented through 2D semantic fusion and temporal information enhancement, and includes the following steps: S1. Two-dimensional image feature extraction steps: The two-dimensional image feature extraction network is used to extract features from the original images of the vehicle from multiple perspectives at the current moment to obtain two-dimensional image features; S2, Two-dimensional semantic information fusion steps: The original image is semantically segmented using a semantic segmentation network to obtain a two-dimensional semantic segmentation mask for the subject and background. The two-dimensional semantic segmentation mask and the two-dimensional image features are fused together to obtain two-dimensional features with fused semantic information. S3. Deep feature extraction step: The depth prediction network is used to perform depth prediction on the two-dimensional features that fuse semantic information to obtain the depth estimate of each pixel. The two-dimensional image features are then combined to obtain two-dimensional features with depth information. S4, BEV feature projection step: Based on the camera intrinsic parameters, extrinsic parameters and the depth estimation value, project the two-dimensional features with depth information onto the BEV feature space to obtain the current frame BEV features; S5. Timing information fusion steps: Obtain historical frame BEV features, and align the historical frame BEV features to the current frame coordinate system according to the vehicle pose transformation matrix; The aligned historical frame BEV features are concatenated and fused with the current frame BEV features in the channel dimension to obtain BEV features with fused temporal information. S6. Three-dimensional target detection step: Input the BEV features with fused temporal information into the BEV perception network and output the three-dimensional target detection result.
[0015] Preferably, in a further improved method for pure visual 3D target detection based on a bird's-eye view, the 2D image feature extraction network in step S1 is an EfficientNetB0 network, with pre-trained weights loaded for feature extraction.
[0016] Preferably, in a further improvement to the bird's-eye view-based pure visual 3D object detection method, the semantic segmentation network in step S2 includes: A data shape conversion layer is used to adjust the size of input data; Convolutional layers are used to adjust the number of channels; The ResNet18 backbone network is used to extract semantic features; An output shape transformation layer is used to output a semantic segmentation mask that matches the feature space size of the two-dimensional image.
[0017] Preferably, the improved pure visual 3D target detection method based on bird's-eye view further includes the channel fusion in step S2, which involves concatenating the 2D semantic segmentation mask with the 2D image features in the channel dimension, so that the channel weights of the semantic segmentation mask are adaptively learned during training.
[0018] Preferably, in a further improved method for detecting 3D targets based on a bird's-eye view, the depth prediction network in step S3 includes a two-dimensional convolutional layer, which expands the input feature channels to C+D channels. Channel C is used to preserve the original image feature information, and channel D is used for pixel-level depth prediction.
[0019] Preferably, in a further improved method for detecting three-dimensional targets based on a bird's-eye view, the BEV feature space in step S4 is a BEV grid with the vehicle position as the center, a specified range, and a specified grid resolution. In this process, multiple pixel features that fall into the same grid are added together and fused.
[0020] Preferably, in a further improved version of the bird's-eye view-based pure visual 3D target detection method, the method for obtaining the vehicle pose transformation matrix in step S5 includes: Obtain the vehicle pose matrix of the current frame and historical frame self-car pose matrix The vehicle pose matrix includes a rotation matrix R and a translation vector T; Calculate the relative transformation matrix ; The coordinate transformation of the historical frame BEV feature grid is performed using the relative transformation matrix, and the aligned feature values are obtained by bilinear interpolation.
[0021] Preferably, in a further improved version of the bird's-eye view-based pure visual 3D target detection method, the BEV perception network in step S6 includes: The preprocessing layer includes a two-dimensional convolutional layer and a batch normalization layer; There are three residual layers, each containing two basic modules, where the first basic module contains a downsampling layer and the second basic module does not contain a downsampling layer; The BEV probe is used to output the position and category information of three-dimensional targets.
[0022] Preferably, the improved bird's-eye view-based pure visual 3D object detection method further includes a training step: The three-dimensional bounding box annotations of the dataset are converted into two-dimensional semantic segmentation annotations. The conversion includes projecting the three-dimensional bounding boxes onto each camera viewpoint to obtain two-dimensional projection regions, and labeling the projection regions as the subject category and the remaining regions as the background category. The semantic segmentation network is trained under supervision using the converted annotations.
[0023] Compared with the prior art, the present invention has at least the following beneficial effects: 1. Improved depth prediction accuracy: By introducing a learnable two-dimensional semantic information fusion mechanism before depth prediction, the semantic segmentation network is used to distinguish between the subject and background regions. Through channel fusion, key information is adaptively enhanced and background redundant information is suppressed, enabling the depth prediction network to focus on the effective region and achieve a significant improvement in depth estimation accuracy under a lightweight network structure.
[0024] 2. Enhanced detection stability: By fusing the temporal information of historical frames and the current frame in the BEV feature space, and using channel splicing, missing information in occluded and blurred scenes can be effectively supplemented without excessively increasing computational complexity, thereby improving the model's adaptability to complex scenes.
[0025] 3. Controllable computational complexity: The EfficientNetB0 and ResNet18 used in this invention are both lightweight network structures. The temporal fusion adopts channel splicing instead of self-attention mechanism, resulting in fewer overall model parameters. It can be deployed and run in real time on mainstream automotive chips, solving the technical problem that existing high-precision models are difficult to deploy in vehicles.
[0026] 4. End-to-end trainability: Two-dimensional semantic fusion adopts a channel splicing method instead of direct filtering, so that the semantic information fusion weights can be adaptively learned through backpropagation; temporal fusion also adopts a learnable channel splicing method, and the entire system can be jointly optimized end-to-end without the need for staged parameter tuning. Attached Figure Description
[0027] The accompanying drawings are intended to illustrate the general characteristics of the methods, structures, and / or materials used in specific exemplary embodiments of the invention, supplementing the description in the specification. However, the drawings are schematic diagrams not drawn to scale and may not accurately reflect the precise structural or performance characteristics of any of the given embodiments. The drawings should not be construed as limiting or restricting the range of numerical values or properties covered by exemplary embodiments of the invention. The invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.
[0028] Figure 1 This is a schematic diagram showing the division of the working stages of the present invention.
[0029] Figure 2 This is a schematic diagram of the structure of the two-dimensional image feature extraction network (EfficientNetB0) of the present invention.
[0030] Figure 3 This is a schematic diagram of the deep prediction network architecture of the present invention.
[0031] Figure 4 This is a schematic diagram of the semantic segmentation network architecture of the present invention.
[0032] Figure 5 This is a schematic diagram illustrating the principle of projecting two-dimensional image pixels onto the world coordinate system according to the present invention.
[0033] Figure 6 This is a schematic diagram of the mesh division of the BEV feature space according to the present invention.
[0034] Figure 7 This is a schematic diagram illustrating the principle of BEV feature space historical frame alignment in this invention.
[0035] Figure 8 This is a schematic diagram of a network structure that incorporates self-attention mechanisms. Detailed Implementation
[0036] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can fully understand other advantages and technical effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through different specific embodiments, and various details in this specification can also be applied based on different viewpoints, with various modifications or changes made without departing from the overall design concept of the invention. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. The following exemplary embodiments of the present invention can be implemented in many different forms and should not be construed as being limited to the specific embodiments set forth herein. It should be understood that these embodiments are provided to make the disclosure of the present invention thorough and complete, and to fully convey the technical solutions of these exemplary embodiments to those skilled in the art. It should be understood that when an element is referred to as "connected" or "combined" to another element, the element can be directly connected or combined to the other element, or there may be intermediate elements. The difference is that when an element is referred to as "directly connected" or "directly combined" to another element, there are no intermediate elements. Throughout the drawings, the same reference numerals always denote the same elements. Furthermore, it should be understood that although the terms "first," "second," etc., may be used herein to describe different elements, components, regions, layers, and / or parts, these elements, components, regions, layers, and / or parts should not be limited by these terms. These terms are only used to distinguish one element, component, region, layer, or part from another.
[0037] This invention provides a pure visual 3D object detection system based on a bird's-eye view. Based on pure visual input, it improves 3D object detection accuracy through dual optimization of 2D semantic information fusion and temporal information enhancement, while maintaining a lightweight model. This solves the technical problems of insufficient depth prediction accuracy and high model complexity in existing BEV pure visual detection algorithms, which hinder vehicle deployment. The technical solution of this invention is applicable to the field of autonomous driving visual perception, and can achieve end-to-end 3D object detection based on pure visual images from six cameras (front, rear, left front, right front, left rear, and right rear). The model parameters are approximately 2.4G, and it can be directly deployed on mainstream automotive chips.
[0038] First embodiment; The bird's-eye view-based pure visual 3D object detection system provided by this invention includes a 2D image feature extraction module, a 2D semantic information fusion module, a depth feature extraction module, a BEV feature projection module, a temporal information fusion module, and a 3D object detection module connected in sequence. Each module is built on the PyTorch 1.12.1 framework, deployed on a Docker 4.38.0 containerized platform, and runs on an Ubuntu 18.04.6 operating system and a CUDA 11.3 computing platform. It utilizes eight ant8-pcie80 GPU cards for multi-GPU parallel training. Dividing the system into two main stages, the bird's-eye view-based pure visual 3D object detection system of this invention is: a 2D image feature extraction stage and a BEV feature extraction stage. (Refer to...) Figure 1 As shown.
[0039] The two-dimensional image feature extraction module is configured to extract features from the original RGB image of the vehicle from six perspectives at the current moment using the EfficientNetB0 network to obtain two-dimensional image features; the EfficientNetB0 network is the backbone network of the model, and its structure is referenced. Figure 2 As shown, it includes a Conv layer, a BN0 layer, 15 MBconvBlock modules, a BN1 layer, a pooling layer, a dropout layer, a fully connected layer, and a swish activation layer. Pre-trained weights are loaded during training to accelerate model convergence and improve feature extraction efficiency.
[0040] The EfficientNetB0 network uses a composite scaling strategy to uniformly adjust the network depth, width, and input resolution. The input is a six-view original image (the resolution is adapted to the network input requirements), and the output is high-level two-dimensional image features. The number of feature channels matches the input requirements of the subsequent depth prediction network.
[0041] The two-dimensional semantic information fusion module includes a semantic segmentation network and a channel fusion unit. The overall configuration is to perform semantic segmentation on the original image and fuse it with two-dimensional image features, thereby reducing background redundancy and enhancing key subject information. Semantic segmentation network structure reference Figure 4 As shown, it includes, in sequence, a data shape transformation layer, a Conv0 convolutional layer, a ResNet18 backbone network, an Up-1 upsampling layer, an Up-2 upsampling layer, and an output shape transformation layer; among which, the ResNet18 backbone network is a lightweight residual network, including a preprocessing layer (two-dimensional convolutional layer + batch normalization layer) and three residual layers, Layer-1, Layer-2, and Layer-3, which solve the gradient vanishing / exploding problem of deep neural networks through residual connections.
[0042] The semantic segmentation network takes a six-view original RGB image of the vehicle as input and outputs a two-dimensional semantic segmentation mask for the subject and background (dividing only the subject and background into two categories). The spatial size of the output mask is completely matched with the two-dimensional image features output by the two-dimensional image feature extraction module, providing a foundation for subsequent channel fusion.
[0043] The channel fusion unit is configured to concatenate and fuse the two-dimensional semantic segmentation mask and the two-dimensional image features along the channel dimension to obtain two-dimensional features with fused semantic information. The channel fusion does not directly filter features, but concatenates the channels of the semantic segmentation mask and the channels of the two-dimensional image features in the dimension, so that the convolutional layer participates in gradient calculation, and the fusion threshold of semantic information is adaptively adjusted during model training. The optimal degree of influence of semantic information is automatically learned through training iterations.
[0044] The deep feature extraction module is configured to perform pixel-level depth prediction on the two-dimensional features fused with semantic information through a depth prediction network to obtain the depth estimate of each pixel, and combine the two-dimensional image features to obtain two-dimensional features with depth information. The core of the depth prediction network is a two-dimensional convolutional layer, with a structure referenced from... Figure 3 As shown, this convolutional layer expands the number of channels of the two-dimensional features with fused semantic information into C+D channels, where: C channels are used to retain the feature information of the original image, and D channels are dedicated to pixel-level depth prediction; the expanded features are re-fused so that the features of each pixel carry the corresponding depth estimate, thus preserving the integrity of the original image features while modeling depth information.
[0045] The BEV feature projection module includes a camera parameter unit and a spatial projection unit, configured to project two-dimensional features with depth information from the image pixel coordinate system to the BEV bird's-eye view feature space based on the camera intrinsic parameters, extrinsic parameters and depth estimation values, to obtain the BEV features of the current frame. Camera parameter unit: Pre-stores the intrinsic and extrinsic parameters of the vehicle's six-view camera. ; including external parameters This is a 4×4 homogeneous transformation matrix, consisting of the rotation matrix R and translation vector T from the world coordinate system to the camera coordinate system. It describes the transformation from the world coordinate system to the camera coordinate system, and the transformation formula is as follows: ;
[0046] These are the homogeneous coordinates of a 3D point in the world coordinate system. These are the coordinates of a 3D point in the camera coordinate system. It is a rotation matrix describing the transition from world coordinates to camera coordinates. This describes the translation vector from world coordinates to camera coordinates. Combining the above equations, it can be represented by a 4×4 homogeneous transformation matrix. The camera extrinsic parameters are given in the formula by... The expanded formula is as follows: ;
[0047] The camera's intrinsic parameters describe the projection from the camera coordinate system to the pixel coordinate system, denoted by K in the image. The projection formula is as follows: ;
[0048] in, These are 2D pixel coordinates. Mainly composed of coordinates in the image Focal length in pixel coordinate system , And composed of an optical center, and It is usually calculated by converting the physical focal length to the pixel size. Expanding the matrix in the formula yields: ;
[0049] The pixel coordinates are then back-projected back to world coordinates. This step requires the pixel's depth information, which is obtained from the model's depth feature prediction network. The back projection formula is as follows: ; Combining the above formulas, we get: ; ; Spatial projection unit: reference Figure 5 As shown, firstly based on the depth estimate Camera internal parameters Hehe Foreign Reference Through the back projection formula Image pixel coordinates Convert to 3D coordinates in world coordinate system Then, the features in the world coordinate system are projected onto the BEV feature space.
[0050] The BEV feature space structure reference Figure 6 As shown, with the vehicle's position as the origin, the coordinate range of the x-axis and y-axis is from -50 meters to +50 meters, forming a square area with a side length of 100 meters. The grid resolution is 0.5 meters, ultimately forming a 200×200 BEV grid; the height range of the z-axis is from -10 meters to 10 meters, covering the three-dimensional space around the vehicle.
[0051] During the projection process, multiple pixel features falling into the same BEV grid are added and fused. Projection information or invalid information that exceeds the BEV feature space range is directly filtered out. The final output is the current frame BEV feature with dimensions of 200×200×C (C is the number of feature channels).
[0052] The temporal information fusion module includes a pose transformation unit and a temporal fusion unit, configured to fuse the BEV features of historical frames and the current frame, supplementing the uncertainty of single-frame information and improving the completeness and continuity of features: Pose transformation unit: Configured to acquire historical frame BEV features, and align the historical frame BEV features to the current frame coordinate system based on the vehicle pose transformation matrix. The alignment principle is referenced... Figure 7 As shown; the specific steps are as follows: a) Obtain the vehicle's pose matrix in the current frame. and historical frame self-car pose matrix The vehicle pose matrices are all 4×4 homogeneous transformation matrices, satisfying... , where R is the vehicle rotation matrix and T is the vehicle translation vector; b) Calculate the relative transformation matrix between the current frame and the historical frames. ; c) Generate a standardized BEV mesh The coordinate transformation of the mesh is obtained by using a relative transformation matrix. and will Normalize to the interval [-1, 1]; d) Apply the formula to the normalized grid coordinates. , Mapped to image pixel coordinates (W and H are the width and height of the BEV feature map); e) Bilinear interpolation is used to interpolate the mapped pixel coordinates to obtain the aligned historical frame BEV feature values. The bilinear interpolation formula is as follows:
[0053] in, , The decimal part of the pixel coordinates This is a floor operation.
[0054] The temporal fusion unit is configured to concatenate and fuse the aligned historical frame BEV features with the current frame BEV features along the channel dimension to obtain the fused temporal information BEV features; the concatenation and fusion formula is: ,in For the current frame's BEV features, For the aligned historical frame BEV features, concat is a channel-level concatenation operation.
[0055] The preferred channel splicing and fusion method is a learnable fusion method. The convolutional layer weights obtain the optimal values through model training. Compared with weighted averaging and self-attention mechanism fusion, this method ensures the fusion effect while avoiding a significant increase in model computation.
[0056] The three-dimensional target detection module includes a BEV perception network and a detection head unit, and is configured to input BEV features with fused temporal information into the BEV perception network for feature extraction, and output three-dimensional target detection results through the detection head. The BEV perception network structure forms a semantic segmentation network, with an architecture referenced from... Figure 4 As shown, based on ResNet18, it includes: Preprocessing layer: Consists of a two-dimensional convolutional layer and a batch normalization layer, used to initially extract BEV features and standardize the data distribution; The three residual layers are Layer-1, Layer-2, and Layer-3, respectively. Each residual layer contains two basic blocks. One basic block contains a downsampling layer (used to reduce the resolution of the feature map and enhance the global receptive field), while the other basic block does not contain a downsampling layer (used to maintain the original resolution and improve the local feature representation ability). Upsampling layer: Includes two upsampling layers, Up-1 and Up-2, to restore the feature map resolution to 200×200 and match it with the BEV feature space.
[0057] The detection head unit is a BEV detector head, which is connected to the output end of the BEV perception network. The output is the three-dimensional target detection results around the vehicle, including semantic segmentation information such as the position, size, and category of targets such as vehicles and pedestrians, providing spatial semantic support for autonomous driving path planning and behavior decision-making.
[0058] Second embodiment; The present invention provides a pure visual 3D target detection method based on a bird's-eye view, which is implemented based on 2D semantic fusion and temporal information enhancement. It is an end-to-end two-stage detection method. The first stage is 2D image feature processing (including semantic fusion and depth prediction), and the second stage is BEV feature processing (including spatial projection, temporal fusion, and 3D detection). Specifically, it includes the following steps: S1. Two-dimensional image feature extraction steps: The EfficientNetB0 network is used to extract features from the original RGB images of the vehicle at the current time from six perspectives: front, rear, left front, right front, left rear, and right rear, to obtain two-dimensional image features. The EfficientNetB0 network loads pre-trained weights for training and feature extraction. Its structure includes a Conv layer, a BN0 layer, 15 MBconvBlock modules, a BN1 layer, a pooling layer, a dropout layer, a fully connected layer, and a swish activation layer. It is based on a composite scaling strategy to balance computational efficiency and feature extraction accuracy.
[0059] S2. Two-dimensional semantic information fusion steps: Effective fusion of two-dimensional semantic information is achieved through semantic segmentation and channel fusion, specifically including: Semantic segmentation is performed on the original RGB image of the vehicle from six perspectives using a semantic segmentation network to obtain a two-dimensional semantic segmentation mask for the subject and background. The semantic segmentation network consists of a data shape transformation layer, a Conv0 convolutional layer, a ResNet18 backbone network, an Up-1 upsampling layer, an Up-2 upsampling layer, and an output shape transformation layer. The ResNet18 backbone network includes a preprocessing layer and three residual layers (Layer-1, Layer-2, and Layer-3), which solve the gradient vanishing / exploding problem through residual connections. The semantic segmentation classifies only the subject and background into two categories, and the spatial size of the output mask completely matches the two-dimensional image features obtained in step S1.
[0060] The two-dimensional semantic segmentation mask and the two-dimensional image features are concatenated and fused along the channel dimension to obtain two-dimensional features with fused semantic information. The channel fusion allows the convolutional layer to participate in the model's gradient calculation, enabling the semantic information fusion threshold to be adaptively adjusted during training. Through iterative training, the optimal degree of semantic information influence is automatically learned, effectively reducing background redundancy and enhancing key subject information.
[0061] S3. Deep feature extraction step: Perform pixel-level depth prediction on the two-dimensional features that have fused semantic information through a depth prediction network to obtain the depth estimate of each pixel, and combine it with the two-dimensional image features to obtain two-dimensional features with depth information. The core of the depth prediction network is a two-dimensional convolutional layer, which expands the number of channels of the input features to C+D channels. C channels retain the feature information of the original image, and D channels are dedicated to pixel-level depth prediction. The expanded features are then re-fused so that the features of each pixel carry the corresponding depth estimate, thus preserving the integrity of the original image features while modeling depth information.
[0062] S4, BEV Feature Projection Step: Based on the camera intrinsic and extrinsic parameters and the depth estimate obtained in step S3, the two-dimensional features with depth information are projected from the image pixel coordinate system to the BEV bird's-eye view feature space to obtain the current frame's BEV features, specifically including: Obtain the intrinsic and extrinsic parameters from the vehicle's six-view camera. Among them, external parameters Internal Reference ; Based on depth estimates Camera internal parameters Hehe Foreign Reference Through the back projection formula Image pixel coordinates Convert to 3D coordinates in world coordinate system ; Features in the world coordinate system are projected onto the BEV feature space, which has the vehicle position as the origin, the x / y axis range of -50 meters to +50 meters, and a grid resolution of 0.5 meters to form a 200×200 BEV grid. During the projection process, multiple pixel features falling into the same grid are added and fused, and information that exceeds the BEV feature space is directly filtered out to output the BEV features of the current frame.
[0063] S5. Temporal Information Fusion Steps: By merging the BEV features of historical frames and the current frame through pose alignment and channel stitching, the uncertainty of single-frame information is supplemented, specifically including: Obtain historical frame BEV features, and align these features to the current frame coordinate system based on the vehicle pose transformation matrix. Specifically: Obtain the current frame vehicle pose matrix. and historical frame self-car pose matrix ; Calculate the relative transformation matrix The coordinates of the standardized BEV grid are transformed and normalized to [-1,1]. The normalized coordinates are mapped to pixel coordinates, and the aligned historical frame BEV features are obtained by bilinear interpolation.
[0064] The aligned historical frame BEV features are concatenated and fused with the current frame BEV features along the channel dimension to obtain the fused temporal BEV features; the concatenation and fusion formula is as follows: This fusion method is a learnable mode, where the weights of the convolutional layers are trained to obtain optimal values, thus controlling the computational load of the model while ensuring the fusion effect.
[0065] S6. Three-dimensional target detection steps: Input the BEV features with fused temporal information into the BEV perception network for feature extraction, and output the three-dimensional target detection results through the BEV probe. The BEV perception network is based on ResNet18 and includes a preprocessing layer (2D convolution + batch normalization), three residual layers (each containing two basic modules with / without downsampling) and two upsampling layers. The BEV detector outputs semantic information such as the position, size, and category of targets around the vehicle, realizing 3D target detection from a bird's-eye view.
[0066] Optional additional model training steps: To ensure the dataset is compatible with the training requirements of the model in this invention, the invention also includes dataset preprocessing and supervised model training steps, specifically: Dataset annotation transformation: The 3D bounding box annotations of the original dataset are converted into 2D semantic segmentation annotations. The transformation method is as follows: the 3D bounding box is projected onto the image plane of the vehicle's six-view camera to obtain the 2D projection area. The 2D projection area is labeled as the subject category, and the rest of the image is labeled as the background category, generating a semantic segmentation annotation dataset that matches the model. Supervised training: The semantic segmentation network of this invention is trained under supervision using the transformed semantic segmentation annotation dataset. At the same time, a multi-GPU parallel training method is used to train the entire detection model end-to-end. During the training process, the training effect is monitored in real time through the tensorboardX module. The number of parameters of the model after training is about 2.4G, which can be directly deployed on mainstream automotive chips.
[0067] Furthermore, a performance description of the second embodiment model and examples of potential improvements are provided; I. The model performance is described below; The core evaluation metrics of the model in this invention have been significantly improved. The Intersection over Union (IOU) reaches 34.73%, which is 2.66%, 4.73%, and 4.84% higher than existing models such as LSS, FISHING, and OFT, respectively. The accuracy (ACC) reaches 98.06%, and the recall reaches 53.94%, which can capture the features of actual positive examples more comprehensively and reduce missed and false positives. At the same time, the model remains lightweight with approximately 2.4G of parameters, which can be directly deployed on automotive chips, thus resolving the contradiction in existing technologies where "improved accuracy leads to more complex models".
[0068] II. Examples of improvements are explained below; The technical solution of this invention can be flexibly adjusted according to actual application needs. All modifications that do not depart from the core technical features of this invention are within the scope of protection. Specific feasible modifications include: 1. The fusion method is modified to form a new solution: 1.1 The time-series fusion method can be replaced by weighted average fusion. The weighted average formula is as follows: , where α is the weight of the current frame, which can be adaptively learned through model training or manually adjusted according to the actual scene; 1.2 Temporal fusion can be replaced by concatenation, a common and efficient fusion method that combines the feature maps of historical and current frames along the channel dimension to form a higher-dimensional feature map. This method preserves details of both historical and current information, and its formula is expressed as: ; 1.3. The temporal fusion method can be replaced by a self-attention mechanism. The self-attention mechanism dynamically selects important information by weighting the features of historical frames and the current frame, and can capture complex temporal relationships. Its network structure is referenced from [reference needed]. Figure 8 As shown. The self-attention mechanism can dynamically adjust weights based on the relevance of input features, making it ideal for capturing long-term dependencies and important information. It offers high flexibility in modeling temporal information, is applicable to various complex tasks, and is particularly effective in long-term sequence tasks. Its formula simplifies to: This method is particularly suitable for time series modeling of long sequences, and can explicitly model time dependencies, but it has high computational complexity.
[0069] Of the two methods, concatenation and weighted averaging, concatenation fuses information by stitching the feature maps of historical and current frames together along the channel dimension, while weighted averaging adds the BEV features of historical and current frames in a weighted manner. Both methods employ learnable models; the concatenation method's convolutional layer weights and weighting values are trained to obtain optimal values for the most suitable fusion effect.
[0070] Fusion methods based on self-attention mechanisms can model the global dependencies between historical frames and the current frame. By calculating the relationship between each feature and other features, self-attention captures more complex temporal and spatial information, making it particularly suitable for long-term and complex scenarios. However, self-attention mechanisms are computationally expensive, significantly increasing the model's computational load and required parameters; therefore, training and inference can only be performed at the software level.
[0071] After 2D semantic information fusion and temporal fusion, the final BEV features are obtained. These BEV features are then input into the BEV perception network. The BEV perception network is connected to the BEV sensor, which obtains the position and size of vehicles surrounding the vehicle, representing the semantic segmentation result of the BEV. The BEV perception network is constructed using ResNet18. Similar to the semantic segmentation network, a preprocessing layer consisting of 2D convolutional layers and batch normalization layers is first applied. This preprocessing layer is responsible for initially extracting features and standardizing the data distribution. The main network consists of three layers, each containing two basic blocks. One of these basic blocks contains a downsampling layer to reduce the feature map resolution and enhance the global receptive field; the other block does not contain a downsampling layer to maintain the original resolution and improve the local feature representation capability. Both types of basic blocks extract global information while preserving local details.
[0072] 2. New solutions are formed after the backbone network is deformed: While ensuring the lightweight nature of the model, the EfficientNetB0 network can be replaced with other lightweight networks in the EfficientNet series (such as EfficientNetB1), and the ResNet18 network can be replaced with lightweight residual networks in the same series (such as ResNet34). 3. New solutions are formed after the BEV feature space is deformed: The range and resolution of the BEV feature space can be adjusted according to the actual application scenario. For example, the x / y axis range can be adjusted to -30 meters to +30 meters and the grid resolution can be adjusted to 0.3 meters to form a 200×200 BEV grid to adapt to different autonomous driving scenarios. 4. New solutions are formed after changing the position and / or number of camera perspectives: The six-view camera of the present invention can be adjusted to four-view, eight-view, etc. according to the vehicle configuration. The model only needs to adapt the feature extraction stage of the input image, while the core semantic fusion, temporal fusion, BEV projection and other technical solutions remain unchanged.
[0073] In summary, the embodiments provided by this invention possess good versatility and scalability. This invention employs a BEV semantic segmentation head as an evaluation module to verify its spatial understanding and scene perception capabilities under purely visual input conditions. The core function of this invention lies in extracting spatially consistent 3D BEV feature representations from 2D images captured by multiple cameras. By constructing a cross-view feature projection mechanism, this invention can accurately reconstruct the spatial distribution and semantic relationships of objects in a scene.
[0074] After obtaining BEV characteristics, various downstream detectors can be flexibly connected according to different application needs to achieve the identification and localization of multiple targets such as vehicles, pedestrians, lane lines, traffic signs, and static obstacles. Users can perform targeted training according to specific usage scenarios to achieve optimal detection results in specific environments.
[0075] This invention has broad application prospects, particularly suitable as a core visual perception module in autonomous driving perception systems, providing high-quality spatial semantic information support for downstream modules such as path planning and behavior decision-making in intelligent driving vehicles. Simultaneously, this invention also possesses the ability to fuse information with multi-sensor systems (such as millimeter-wave radar and lidar), which can be used to compensate for the shortcomings of other sensors in terms of resolution, susceptibility to adverse weather conditions, etc., thereby improving the robustness of the overall perception system.
[0076] Furthermore, this invention can also be widely applied to the construction of high-precision semantic maps. By segmenting and classifying static scene elements such as road structures, lane boundaries, sidewalks, and parking areas, accurate data support can be provided for the automatic generation and dynamic updating of maps.
[0077] Unless otherwise defined, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. It will also be understood that, unless expressly defined herein, terms such as those defined in a general dictionary shall be interpreted as having the meaning consistent with their meaning in the relevant field context, and not as having an idealized or overly formal meaning.
[0078] The present invention has been described in detail above through specific embodiments and examples, but these are not intended to limit the invention. Many modifications and improvements can be made by those skilled in the art without departing from the principles of the invention, and these should also be considered within the scope of protection of the present invention.
Claims
1. A pure visual 3D target detection system based on a bird's-eye view, characterized in that, include: The two-dimensional image feature extraction module is configured to extract features from the original images of the vehicle from multiple perspectives at the current moment through a two-dimensional image feature extraction network to obtain two-dimensional image features. The two-dimensional semantic information fusion module is configured as follows: The original image is semantically segmented using a semantic segmentation network to obtain a two-dimensional semantic segmentation mask for the subject and background. The two-dimensional semantic segmentation mask and the two-dimensional image features are fused together to obtain two-dimensional features with fused semantic information. The depth feature extraction module is configured to perform depth prediction on the two-dimensional features fused with semantic information through a depth prediction network to obtain a depth estimate of each pixel, and combine the two-dimensional image features to obtain two-dimensional features with depth information. The BEV feature projection module is configured to project the two-dimensional features with depth information onto the BEV feature space based on the camera intrinsic parameters, extrinsic parameters, and the depth estimation value to obtain the BEV features of the current frame. The time-series information fusion module is configured as follows: Obtain historical frame BEV features, and align the historical frame BEV features to the current frame coordinate system according to the vehicle pose transformation matrix; The aligned historical frame BEV features are concatenated and fused with the current frame BEV features in the channel dimension to obtain BEV features with fused temporal information. The 3D target detection module is configured to input the BEV features fused with temporal information into the BEV perception network and output the 3D target detection result.
2. The pure visual 3D target detection system based on a bird's-eye view according to claim 1, characterized in that, The two-dimensional image feature extraction network is the EfficientNetB0 network.
3. The pure visual 3D target detection system based on a bird's-eye view according to claim 1, characterized in that, The semantic segmentation network includes a data shape transformation layer, a convolutional layer, a ResNet18 backbone network, and an output shape transformation layer.
4. The pure visual 3D target detection system based on a bird's-eye view according to claim 1, characterized in that, The channel fusion is performed by concatenating the channels, so that the channel weights of the semantic segmentation mask are adaptively learned during the training process.
5. The pure visual 3D target detection system based on a bird's-eye view according to claim 1, characterized in that: The depth prediction network includes a two-dimensional convolutional layer, which expands the input feature channels to C+D channels, where C channels are used to retain the original image feature information and D channels are used for pixel-level depth prediction.
6. The pure visual 3D target detection system based on a bird's-eye view according to claim 1, characterized in that, The BEV feature space is a BEV grid centered on the vehicle's position, with a specified range and a specified grid resolution; In this process, multiple pixel features that fall into the same grid are added together and fused.
7. A pure visual 3D target detection method based on a bird's-eye view, which is implemented based on 2D semantic fusion and temporal information enhancement, characterized in that... Includes the following steps: S1. Use a two-dimensional image feature extraction network to extract features from the original images of the vehicle from multiple perspectives at the current moment to obtain two-dimensional image features; S2. The original image is semantically segmented using a semantic segmentation network to obtain a two-dimensional semantic segmentation mask for the subject and background. The two-dimensional semantic segmentation mask and the two-dimensional image features are fused together to obtain two-dimensional features with fused semantic information. S3. Perform depth prediction on the two-dimensional features of the fused semantic information through a depth prediction network to obtain the depth estimate of each pixel, and combine the two-dimensional image features to obtain two-dimensional features with depth information. S4. Based on the camera intrinsic parameters, extrinsic parameters, and the depth estimation value, project the two-dimensional features with depth information onto the BEV feature space to obtain the current frame BEV features. S5. Obtain historical frame BEV features and align the historical frame BEV features to the current frame coordinate system according to the vehicle pose transformation matrix. The aligned historical frame BEV features are concatenated and fused with the current frame BEV features in the channel dimension to obtain BEV features with fused temporal information. S6. Input the BEV features with fused temporal information into the BEV perception network and output the three-dimensional target detection results.
8. The pure visual 3D target detection method based on bird's-eye view according to claim 7, characterized in that: The two-dimensional image feature extraction network mentioned in step S1 is the EfficientNetB0 network, which is loaded with pre-trained weights for feature extraction.
9. The pure visual 3D target detection method based on bird's-eye view according to claim 7, characterized in that: The semantic segmentation network mentioned in step S2 includes: A data shape conversion layer is used to adjust the size of input data; Convolutional layers are used to adjust the number of channels; The ResNet18 backbone network is used to extract semantic features; An output shape transformation layer is used to output a semantic segmentation mask that matches the feature space size of the two-dimensional image.
10. The pure visual 3D target detection method based on a bird's-eye view according to claim 7, characterized in that, The channel fusion in step S2 includes: concatenating the two-dimensional semantic segmentation mask with the two-dimensional image features in the channel dimension, so that the channel weights of the semantic segmentation mask are adaptively learned during the training process.
11. The pure visual 3D target detection method based on a bird's-eye view according to claim 7, characterized in that, The depth prediction network in step S3 includes a two-dimensional convolutional layer, which expands the input feature channels to C+D channels; Channel C is used to preserve the original image feature information, and channel D is used for pixel-level depth prediction.
12. The pure visual 3D target detection method based on a bird's-eye view according to claim 7, characterized in that, The BEV feature space mentioned in step S4 is a BEV grid with the vehicle position as the center, a specified range, and a specified grid resolution. In this process, multiple pixel features that fall into the same grid are added together and fused.
13. The pure visual 3D target detection method based on a bird's-eye view according to claim 7, characterized in that, The method for obtaining the vehicle pose transformation matrix in step S5 includes: Obtain the vehicle pose matrix of the current frame and historical frame self-car pose matrix The vehicle pose matrix includes a rotation matrix R and a translation vector T; Calculate the relative transformation matrix ; The coordinate transformation of the historical frame BEV feature grid is performed using the relative transformation matrix, and the aligned feature values are obtained by bilinear interpolation.
14. The pure visual 3D target detection method based on a bird's-eye view according to claim 7, characterized in that, The BEV sensing network mentioned in step S6 includes: The preprocessing layer includes a two-dimensional convolutional layer and a batch normalization layer; There are three residual layers, each containing two basic modules, where the first basic module contains a downsampling layer and the second basic module does not contain a downsampling layer; The BEV probe is used to output the position and category information of three-dimensional targets.
15. The pure visual 3D target detection method based on a bird's-eye view according to claim 7, characterized in that, It also includes training steps: The three-dimensional bounding box annotations of the dataset are converted into two-dimensional semantic segmentation annotations. The conversion includes projecting the three-dimensional bounding boxes onto each camera viewpoint to obtain two-dimensional projection regions, and labeling the projection regions as the subject category and the remaining regions as the background category. The semantic segmentation network is trained under supervision using the converted annotations.
Citation Information
Patent Citations
Multi-view aerial view target detection method based on two-dimensional prior target detection guidance
CN120747468A
SENet-based improved YOLOv8 small target detection method
CN120783028A