Object detection method, device and equipment based on aerial view, medium and product
By integrating the feature extraction and fusion methods of multi-view surround view cameras and lidar data, the problems of task isolation, feature degradation and temporal inconsistency in bird's-eye view perception are solved, multi-task collaborative perception is achieved, and the reliability of three-dimensional object detection and high-precision map segmentation of smart cars is improved.
Patent Information
- Application Number
- CN202510881465.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-10-14
AI Technical Summary
Existing bird's-eye-view perception methods suffer from task isolation, feature degradation, and timing inconsistency, resulting in low reliability of the multi-task environmental perception system of smart cars under limited computing power, and are unable to effectively achieve three-dimensional object detection, high-precision map segmentation, and motion prediction.
By extracting the multi-scale features of the multi-view surround-view camera and the sparse voxel structure of the lidar point cloud data, BEV features are generated. The multi-scale features and BEV features are fused to extract road topology and semantic information, generate a map, determine the three-dimensional bounding box and motion status of traffic participants, and realize multi-task collaborative perception.
It improves the reliability of target detection based on bird's-eye view, is suitable for three-dimensional object detection, high-precision map segmentation and motion prediction, realizes multi-task feature sharing, conflict resolution and temporal consistency maintenance, and improves the perception system performance of smart cars.
Smart Images

Figure CN120783318A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of image processing technology, and in particular to a method, device, equipment, medium and product for target detection based on a bird's-eye view. Background Art
[0002] In recent years, smart cars have become a strategic development direction for the global automotive industry, with a large number of related technologies and products emerging. However, because autonomous driving environmental perception systems are often designed for a single task, such as using independent convolutional neural networks for object detection, U-Net architectures for road segmentation, and temporal long short-term memory (LSTM) models for trajectory prediction, this "task island" model leads to a significant waste of domain controller computing resources. Therefore, it is urgent to develop reliable multi-task fusion optimization methods to fully utilize the limited computing power of smart cars.
[0003] Currently, a large amount of research and practice has been carried out on intelligent multi-task environmental perception. A perception method based on the bird's eye view (BEV) has been proposed. This method can simplify complex three-dimensional environments into two-dimensional images, thereby enabling efficient calculations in real-time systems. Perception based on the bird's eye view has the following problems:
[0004] (1) Task isolation: Detection and segmentation tasks typically use independent network branches, resulting in a split in the information flow between geometric modeling and semantic understanding during feature learning. Dynamic object detection lacks the structured guidance of road topology, while static scene segmentation cannot utilize the spatiotemporal context of moving objects. The two form isolated representations at the feature expression level, weakening the synergistic effect of multi-task joint reasoning.
[0005] (2) Feature degradation: A single BEV feature cannot take into account both geometric modeling and semantic understanding. Traditional BEV space construction methods are prone to losing fine-grained texture information of perspective images during projection, resulting in a decrease in small-scale target detection performance. At the same time, feature generation mechanisms that overly rely on geometric priors weaken high-level semantic associations such as lane curvature and road boundary topology, resulting in a lack of connectivity in the segmentation results and an increase in the misjudgment rate.
[0006] (3) Temporal inconsistency: The temporal features of dynamic target detection and static map segmentation are difficult to update in a coordinated manner. Dynamic target trajectory prediction relies on the continuity modeling of temporal BEV features, while static scene segmentation requires the stability of cross-frame results. Existing methods lack a unified temporal alignment constraint mechanism, resulting in the coexistence of dynamic target trajectory jitter and local mutation of static maps, reducing the reliability of autonomous driving decision-making and planning systems.
[0007] In summary, the current reliability of BEV perception is relatively low. How to develop a multi-task BEV perception algorithm that can be applied to three-dimensional object detection, high-precision map segmentation and motion prediction, and realize multi-task feature sharing, conflict resolution and temporal consistency under the constraints of limited computing power is a core technical problem that needs to be solved in the field of intelligent vehicles. Summary of the Invention
[0008] The present application provides a method, apparatus, device, medium and product for target detection based on a bird's-eye view to improve the reliability of target detection based on a bird's-eye view.
[0009] In a first aspect, an embodiment of the present application provides a method for detecting an object based on a bird's-eye view, comprising:
[0010] Extract multi-scale features from the input images of the multi-view surround camera;
[0011] Convert the LiDAR point cloud data into a sparse voxel structure and generate BEV features based on the voxel features;
[0012] Fusing the multi-scale features and the BEV features to obtain a fused BEV feature;
[0013] extracting road topology and semantic information according to the fused BEV features, and generating a map according to the road topology and the semantic information;
[0014] A three-dimensional bounding box and a motion state of the traffic participant are determined based on the fused BEV features and a prediction result of a position of the traffic participant in the map.
[0015] In a second aspect, an embodiment of the present application further provides a target detection device based on a bird's-eye view, comprising:
[0016] An image feature extraction module is used to extract multi-scale features from the input image of the multi-view surround camera;
[0017] Point cloud feature extraction module, used to convert the point cloud data of the lidar into a sparse voxel structure and generate BEV features based on the voxel features;
[0018] A fusion module, configured to fuse the multi-scale features and the BEV features to obtain a fused BEV feature;
[0019] a map segmentation module, configured to extract road topology and semantic information based on the fused BEV features, and generate a map based on the road topology and the semantic information;
[0020] The target detection module is configured to determine a three-dimensional bounding box and a motion state of the traffic participant based on the fused BEV features and the predicted result of the position of the traffic participant in the map.
[0021] In a third aspect, an electronic device is provided, including:
[0022] one or more processors;
[0023] a storage device storing one or more programs;
[0024] When the one or more programs are executed by the one or more processors, the one or more processors implement the bird's eye view based target detection method as described in the first aspect.
[0025] In a fourth aspect, a computer readable storage medium is provided, having stored thereon a computer program, which when executed by a processor implements the bird's eye view based target detection method as described in the first aspect.
[0026] In a fifth aspect, a computer program product is provided, including a computer program and / or instructions, which when executed by a processor implements the bird's eye view based target detection method as described in any of the above embodiments.
[0027] The embodiments of the present application provide a bird's eye view based target detection method, device, equipment, medium and product, which includes: extracting multi-scale features for input images of a multi-view surround camera; generating a BEV feature map according to the multi-scale feature map; fusing the multi-scale features and the BEV features to obtain fused BEV features; extracting road topology and semantic information according to the fused BEV features, and generating a map according to the road topology and the semantic information; determining a three-dimensional bounding box and a motion state of a traffic participant according to the fused BEV features and a prediction result of a position of the traffic participant in the map. The above technical solution fuses multi-scale geometric features, BEV features and semantic information, and dynamically cooperates timing features of static map segmentation and dynamic target detection, thereby improving the reliability of bird's eye view based target detection, and being suitable for multi-task BEV perception of three-dimensional object detection, high-precision map segmentation and motion prediction. BRIEF DESCRIPTION OF DRAWINGS
[0028] The above and other features, advantages, and aspects of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, like reference numerals are used to represent like elements. It should be understood that the drawings are schematic and elements are not necessarily to scale.
[0029] Figure 1 A flowchart of a bird's eye view based target detection method provided by the embodiments of the present application;
[0030] Figure 2 A schematic diagram of an image feature extraction module provided in an embodiment of the present application is shown in FIG. 1.
[0031] Figure 3 A schematic diagram of a point cloud feature extraction module provided in an embodiment of the present application is shown in FIG. 2.
[0032] Figure 4 A structural schematic diagram of a bird's-eye view based target detection apparatus provided in an embodiment of the present application is shown in FIG. 3.
[0033] Figure 5 A structural schematic diagram of an electronic device provided in an embodiment of the present application is shown in FIG. 4. DETAILED DESCRIPTION
[0034] Before the example embodiments are described in more detail, it should be mentioned that some of the example embodiments are described as processes or methods depicted as flow diagrams. Although the processes are described in a particular sequential order, many of the steps can be performed in parallel, concurrently or in any order. In addition, the order of the steps can be re-arranged. The processes can be terminated when their operations are completed, but can also have additional steps not included in the figure, which can also be performed after the operations of the processes are completed. The processes can correspond to methods, functions, procedures, subroutines, subprograms, etc.
[0035] It should be noted that the terms "first", "second", etc. mentioned in the embodiments of the present application are only used to distinguish different devices, modules, units or other objects, and do not limit the order or interdependence of the functions performed by these devices, modules, units or other objects.
[0036] It should be noted that the terms "first", "second", etc. mentioned in the embodiments of the present application are only used to distinguish different devices, modules, units or other objects, and do not limit the order or interdependence of the functions performed by these devices, modules, units or other objects.
[0037] In addition, the embodiments in the present application and the features in the embodiments can be combined with each other without conflict.
[0038] In the technical solutions of the present application, the acquisition, storage, use, processing, etc. of data all comply with the relevant provisions of national laws and regulations.
[0039] It should be noted that in the embodiments of the present application, some industry existing solutions, components, models, etc. may be mentioned, which should be considered as exemplary, and the purpose is only to illustrate the feasibility of the implementation of the technical solutions of the present application, but does not mean that the applicant has or will necessarily use the relevant content of the solutions.
[0040] Figure 1The present invention provides a flowchart of a method for detecting a target based on a bird's-eye view, which is applicable to situations where traffic participants are perceived based on a bird's-eye view. Specifically, the method for detecting a target based on a bird's-eye view can be performed by a target detection device based on a bird's-eye view, which can be implemented in software and / or hardware and integrated into an electronic device. The electronic device includes, but is not limited to, an electronic control unit (ECU), a microcontroller (MCU), a central processing unit (CPU), a computer, a smart phone, or a server, and other devices with data processing capabilities.
[0041] like Figure 1 As shown, the method specifically includes the following steps:
[0042] S110: extract multi-scale features from the input image of the multi-view surround camera.
[0043] In this embodiment, the RGB image of the multi-view surround camera is used as input. For example, six surround cameras can be used to capture images in the environment to obtain six input images. Exemplarily, a deep residual network (such as ResNet-101-DCN) can be used as the backbone network to extract multi-scale features layer by layer for the multi-channel input images, and the multi-scale features can be extracted through the FPN structure. In this process, a multi-layer residual network can be used to model the semantic flow features of the camera image; convolution and pooling operations can be used to generate multi-scale features containing local details and global semantics. A cross-camera spatial attention mechanism can also be introduced to adaptively fuse features from different perspectives, eliminate geometric deviations caused by perspective differences, ensure the consistency of multi-camera features in spatial alignment, and provide high-precision implicit feature expressions for subsequent modules. Optionally, multi-scale features can be extracted through an image feature extraction module.
[0044] S120, converting the laser radar point cloud data into a sparse voxel structure and generating BEV features based on the voxel features;
[0045] In this embodiment, point cloud data is input and a BEV feature map is output. 3D sparse convolution can be used to extract the geometric flow features contained in the lidar point cloud. For example, a dynamic voxelization method can be used to convert the original disordered point cloud into a regularly arranged sparse voxel structure. A sparse convolutional network is then used to extract multi-scale geometric features from the voxels. A columnar sparse convolution operation is then used to compress the three-dimensional voxel feature map onto a two-dimensional BEV plane, thereby generating a point cloud BEV feature representation containing rich spatial geometric information. Optionally, a BEV feature map can be generated using a point cloud feature extraction module.
[0046] S130, fusing the multi-scale feature and the BEV feature to obtain a fused BEV feature;
[0047] In this embodiment, the multi-scale features and BEV feature maps of each perspective are used as input, the two features are spliced in the channel dimension to form an initial representation of the fused features, and a unified fused BEV feature is output. Exemplarily, multimodal feature fusion can be achieved through a multi-head attention mechanism, that is, multi-scale image features and BEV features are interacted across scales. In this process, a deformable attention mechanism can also be used to dynamically focus on key areas, extract contextual information that is strongly related to the target position from feature maps of different resolutions, and enhance the perceptual robustness of small targets in the distance and occluded scenes. Optionally, a fused feature map can be obtained through a fusion module.
[0048] S140 , extracting road topology and semantic information based on the fused BEV features, and generating a map based on the road topology and the semantic information.
[0049] For example, an encoder-decoder structure can be used to extract road topology and semantic information based on the fusion of BEV features. The encoding part can use strided convolution to gradually downsample features, primarily to capture macroscopic road topology structures such as lane curvature and drivable area boundaries. The decoding part can fuse shallower detail features through jump connections to achieve pixel-level semantic segmentation. In this process, to improve lane connectivity, a topological continuity constraint loss can be introduced to optimize the spatial consistency of the segmentation results and ensure high-precision map generation in complex road scenarios. Optionally, maps can be generated using a map segmentation module.
[0050] S150 : Determine a three-dimensional bounding box and a motion state of the traffic participant based on the fused BEV features and the prediction result of the position of the traffic participant in the map.
[0051] For example, based on the fused feature map, a sparse detection head can be used to parse the decoded BEV features, predict the target center position through the heat map, and regress parameters such as three-dimensional size and heading angle; in this process, a multi-scale feature pyramid can be used to enhance the adaptability to multi-scale targets, and dynamic gradient modulation technology can be combined to balance the learning weights of detection and segmentation tasks to avoid feature conflicts. The final output can cover the three-dimensional bounding boxes and motion states (such as motion direction and speed) of traffic participants such as vehicles and pedestrians, providing accurate perception input for autonomous driving decision planning. Optionally, the three-dimensional bounding boxes and motion states of the traffic participants can be determined by the target detection module.
[0052] The target detection method based on the bird's eye view provided by the embodiment of the application realizes multi-task collaborative perception, and is particularly suitable for a scene of joint optimization of three-dimensional object detection, high-precision map segmentation and motion prediction. The method realizes dynamic feature sharing and conflict resolution among detection, segmentation and prediction tasks by constructing a multi-level BEV feature interaction mechanism, and cooperates the timing features of dynamic target detection and static map segmentation to provide an efficient and unified perception solution for intelligent vehicles.
[0053] In an embodiment, multi-scale features are extracted for each input image of the multi-view surround-view camera, including:
[0054] S1110, for each input image of the multi-view surround-view camera, a feature map of a corresponding resolution is generated by using a convolution operation through a deep residual network, and multi-scale features are generated from the feature map of the corresponding resolution by using a stacked residual block;
[0055] S1120, aligning the multi-scale features of each input image;
[0056] S1130, fusing the multi-scale features of each input image through a feature pyramid network.
[0057] Exemplarily, the extraction of the multi-scale feature map can be implemented by using an image feature extraction module, and the image feature extraction module mainly includes a backbone network based on a deep residual network (ResNet) and a neck network based on an FPN, as shown in Figure 2 .
[0058] As an example, for the backbone network based on the ResNet, the input is the original image of the six cameras of the current frame, and the resolution of each input image can be 1920x1080 pixels. After normalization and size scaling, the input image is converted into a floating-point tensor I∈R 6×3×H×W , where H and W are the image sizes after size scaling. The backbone network can specifically adopt a ResNet-101-DCN structure, and multi-level residual feature extraction is independently performed on each input image. First, a 7x7 convolution of Stage1 can be used to generate a feature map with a resolution reduced by half , and then the Bottleneck residual block stacking of Stage2 to Stage5 is used to gradually output multi-scale features:
[0059]
[0060] where i=1,...,6 represents the i-th camera, C n represents different feature channel numbers.
[0061] As an example, for the FPN-based neck network, a top-down feature pyramid fusion structure is constructed with the multi-scale feature set {F2, F3, F4, F5} output by the ResNet backbone network as input.
[0062] First, convolution is performed on the highest-level features to reduce the dimensionality and obtain feature maps with 256 channels:
[0063] C k =Conv 3×3 (F k ),k=2,3,4,5
[0064] in, Starting from the highest layer, the high-level features are upsampled by bilinear interpolation to keep their spatial resolution consistent with the features of the next layer, and then fused:
[0065] P5=Conv 3×3 (C5)
[0066] P4=Conv 3×3 (C4+UpSample(P5))
[0067] P3=Conv 3×3 (C3+UpSample(P4))
[0068] P2=Conv 3×3 (C2+UpSample(P3))
[0069] Among them, UpSample() represents the bilinear interpolation upsampling operation, and the convolution kernel size is 3×3, which is intended to smooth the features after fusion. The spatial sizes of the feature maps P2, P3, P4, and P5 of each layer are and The number of channels is 256.
[0070] In order to enhance the adaptability of the model to irregular targets, for each fusion feature map P k Applying deformable convolution, formally expressed as:
[0071]
[0072] Where y(p) = represents the feature response of the output feature map at position p, P k () indicates that in the fusion feature map P k The eigenvalue obtained by sampling the specified spatial position, Δp k is the learnable offset, p k Represents the fixed offset coordinates of the standard convolution kernel, K = 9 corresponds to the sampling points of the 3 × 3 convolution kernel, wk is the convolution kernel weight. This design combines shallow high-resolution features with deep semantic information, taking into account both the detection of small targets at long distances and the positioning accuracy of large targets in the near distance.
[0073] Furthermore, the cross-view spatial attention mechanism targets the feature map P2 based on the FPN output (i) (where i = 1, ..., 6 represents the i-th camera) for processing. To reduce computational complexity, this mechanism only performs attention interaction between adjacent camera views to enhance the feature alignment and fusion effect between adjacent views.
[0074] First, P2 for each camera perspective (i) The feature map generates three sets of feature representations: query, key, and value through linear mapping:
[0075]
[0076] Among them, W Q 、W K 、W V is a learnable linear projection matrix. For the i-th camera, the similarity is calculated only with the camera views in the adjacent view set N(i). The attention weight is defined as:
[0077]
[0078] Among them, α ij represents the attention weight between the i-th camera and its adjacent camera j. The Softmax operation is normalized only in the adjacent view dimension, and d is the feature dimension.
[0079] Finally, by weighted summing the value features of adjacent perspectives, we get the output of cross-perspective fusion:
[0080]
[0081] Subsequently, the fused multi-view image features are projected onto a unified BEV plane using a projection mapping function based on the internal and external parameters of each camera. For each pixel position in the BEV space, the features of its mapping position are sampled from the feature map of the corresponding view, and the BEV features are obtained through weighted fusion:
[0082]
[0083] Among them, Π i (p) represents the projection coordinate of pixel p in the BEV space in the image feature of the i-th camera, w i (p) is the corresponding fusion weight, F BEV(p) is the final generated image BEV feature. Through this process, the system effectively integrates multi-view image information and generates semantically rich BEV features, providing high-quality input for subsequent multi-task perception modules.
[0084] In one embodiment, converting laser radar point cloud data into a sparse voxel structure and generating BEV features for voxel features includes:
[0085] S1210, dividing the laser radar point cloud data into a regular voxel grid within a set range according to the three-dimensional coordinates of the point cloud in space;
[0086] S1220 , processing non-empty voxel features based on sparse convolution to obtain sparse features, and mapping the sparse features to a dense BEV space to obtain two-dimensional BEV features.
[0087] For example, the point cloud feature extraction module can be used to process the point cloud data and generate a BEV feature map. The processing process of the point cloud feature extraction module mainly includes two stages: point cloud voxelization and sparse convolution feature extraction. Figure 3 shown.
[0088] As an example, in the first stage, the point cloud data collected by the LiDAR is input and divided into a regular voxel grid covering a certain range around the vehicle based on the three-dimensional coordinates of the point cloud in space. Each voxel corresponds to a small area in the space around the vehicle, and each point in the original point cloud is assigned to a corresponding voxel according to its position. Subsequently, the point features within each voxel are summed and averaged to generate a comprehensive feature representation of that voxel:
[0089]
[0090] Among them, f v represents the aggregated features of voxel v, N v is the number of points inside the voxel, f i is the feature of the i-th point.
[0091] As an example, in the second stage, the non-empty voxel features obtained after voxelization are used as input, and the sparse convolution operation is used to efficiently process the features. The definition of a single-layer sparse convolution is as follows:
[0092]
[0093] Here, p and q are the non-empty point locations in voxel space, N(p) is the neighborhood of point p, and W(p,q) is the convolution kernel weight. Convolution is performed only in the neighborhood of non-empty voxels to reduce computational overhead in empty areas. Instance normalization is applied to the convolution output to stabilize the training process.
[0094] The network structure consists of multiple stacked sparse convolution blocks with residual connections, each of which uses dilated convolution to expand the receptive field:
[0095]
[0096] in, Denotes the expansion rate d l The dilated sparse convolution with increasing dilation rate increases with the number of layers to capture a wider range of context information. The multi-layer sparse features are weightedly fused through the channel attention mechanism, expressed as:
[0097]
[0098] Among them, F fused is the weighted fusion result of multi-layer sparse features, α l is the learnable channel weight of layer l, and Interp() represents unifying the feature resolution to the original voxel grid size through trilinear interpolation.
[0099] Finally, the fused sparse features are mapped to the dense BEV space to form a complete two-dimensional BEV feature map:
[0100] F BEV =SparseToDense(F fused )
[0101] SparseToDense is used to transform the 3D sparse feature projection into a dense BEV feature map. This process efficiently converts sparse 3D features into dense 2D representations, providing rich and structured spatial information for downstream perception tasks.
[0102] In one embodiment, fusing the multi-scale features and the BEV features to obtain a fused BEV feature includes:
[0103] S1310 , performing motion compensation on the BEV features of the historical frame to transform the BEV features of the historical frame into the coordinate system of the current frame;
[0104] S1320, concatenating the multi-scale features of the input image of the current frame and the BEV features of the point cloud data with the BEV features of the motion-compensated historical frames along the channel dimension to obtain a multi-channel joint BEV feature;
[0105] S1330. Use a convolutional neural network to fuse multi-channel joint BEV features to obtain fused BEV features.
[0106] Exemplarily, a fusion module can be used to obtain a fusion feature map. The input of the fusion module is the output of the image feature extraction module and the output of the point cloud feature extraction module. The visual semantic information and spatial geometric information can be effectively integrated to generate a unified BEV representation that integrates rich semantic and structural features.
[0107] As an example, during the fusion process, the BEV features of the historical frames are first motion compensated. Based on the motion estimation between the current frame and the historical frames, the historical BEV features are transformed to the current frame coordinate system to achieve spatiotemporal alignment, which is specifically expressed as:
[0108]
[0109] in, = is the compensated historical frame BEV feature, is the BEV feature of the historical frame, T curr←hist is the motion transformation matrix from the historical frame to the current frame, and T() represents the motion compensation transformation operation.
[0110] The BEV features of the current frame image features and point cloud are concatenated with the historical BEV features after motion compensation along the channel dimension to form a multi-channel joint feature representation:
[0111]
[0112] in, is the image feature of the current frame, is the BEV feature of the point cloud.
[0113] The convolutional neural network is used to deeply fuse the spliced features. The convolution operation is used to enhance the interaction and coordination between different modalities and temporal information, and to promote the mutual complementation and reinforcement of semantic details, spatial geometry and temporal consistency, which can be expressed as:
[0114]
[0115] This fusion strategy not only preserves the rich semantics and precise geometry of the current frame image and point cloud features, but also incorporates temporal context from historical frames, significantly improving the stability and continuity of BEV features. The resulting unified BEV features combine detailed texture information, precise spatial structure, and good temporal consistency, effectively improving the performance and robustness of subsequent multi-task perception modules.
[0116] On this basis, a cross-task spatiotemporal consistency constraint mechanism is introduced. The BEV features of the historical frame and the current frame are aligned through the motion compensation algorithm, the stability loss of the static map is optimized, and cross-task temporal consistency modeling is achieved. This solves the problem of target ID jump rate and lane line cross-frame break rate caused by the asynchronous update of dynamic-static task timing.
[0117] In one embodiment, extracting road topology and semantic information based on the fused BEV features, and generating a map based on the road topology and the semantic information, includes:
[0118] S1410, using an encoder to downsample the BEV feature map level by level to obtain features at each level;
[0119] S1420, using a decoder to perform upsampling step by step through transposed convolution, and performing jump connections with features of corresponding levels of the encoder to obtain road topology and semantic information;
[0120] S1430: Modeling lane key points as graph nodes, and learning the connection relationship between nodes through a graph attention network;
[0121] S1440: Output a segmentation mask based on the road topology, the semantic information, and the connection relationship, where the segmentation mask includes lane labels, drivable area labels, and road boundary labels;
[0122] S1450: Generate a map according to the segmentation mask.
[0123] For example, a map segmentation module can be used to generate a map. The input of the map segmentation module is the output of the point cloud feature extraction module. The road topology semantic information can be extracted through a multi-level convolutional encoding and decoding architecture to achieve high-precision segmentation of lane lines, drivable areas, and road boundaries.
[0124] As an example, the map segmentation module can adopt an improved U-Net network structure. Its encoder part consists of four layers of convolution, each followed by batch normalization and ReLU activation, gradually downsampling the feature map to low resolution, capturing macro-semantic features such as global road topology and lane curvature; its decoder part gradually upsamples through transposed convolution and performs skip connections with the features of the corresponding encoder layer, effectively fusing shallow high-resolution details and restoring the sub-pixel accuracy of lane edges, thereby improving the spatial continuity and edge accuracy of the segmentation mask.
[0125] Optionally, to enhance the topological continuity of the segmentation results, a graph structure constraint loss function is designed. Lane key points are modeled as graph nodes, and the connection relationship between nodes is learned through the Graph Attention Network (GAT):
[0126]
[0127] Among them, L outIt is a constraint loss function based on the graph structure. The subsequent network is optimized by reducing this loss during training. ε is the set of all edges in the graph attention network, that is, the index set of all adjacent node pairs; h i and h j are the feature vectors of adjacent nodes i and j respectively, is the edge weight prediction function based on multi-layer perceptron (MLP), Indicates whether there is a real connection between nodes i and j. The loss promotes the continuity and accurate expression of the lane line structure by minimizing the error between the predicted connection weight and the actual state.
[0128] In addition, for pixel-level classification tasks, a weighted combination of Dice loss and cross entropy loss is used for optimization:
[0129] L seg =λ1·L Dice +λ2·L CE
[0130] Among them, L seg is the total loss of the map segmentation task, L Dice It is a loss for segmentation tasks, responsible for improving the segmentation recall rate of small target areas and reducing the impact of category imbalance on training; L CE is the cross entropy loss to ensure the accuracy of classification; λ1 and λ2 are hyperparameters that weigh the effects of the two, and are usually adjusted through experiments to obtain the best performance.
[0131] To ensure the stability of map segmentation between consecutive time frames, a static map stability loss is introduced. This loss aligns the segmentation results of the historical frames to the current frame coordinate system through a motion compensation algorithm, and calculates the pixel-level difference between the aligned results and the current frame segmentation results, which can be formalized as:
[0132] L stable =‖M t -T(M t-1 ,T t←t-1 )‖1
[0133] Among them, M t and M t-1 are the segmentation masks of the current and previous time frames, T t←t-1 is the motion transformation matrix, and T() is the motion compensation alignment operation. This loss effectively reduces the temporal abruptness and discontinuity of lane lines and road boundaries, improving the temporal robustness of the multi-task perception system.
[0134] Finally, the segmentation mask M∈R output by the module 200×200×3It contains three types of semantic labels: lane lines, drivable areas, and road boundaries, providing detailed scene semantic information and providing a reliable basis for path planning and decision-making of autonomous driving systems.
[0135] The map segmentation module effectively improves segmentation accuracy and topological coherence through multi-level convolutional feature extraction, graph structure topological constraints and comprehensive loss function design, meeting the performance requirements of the high-precision map semantic segmentation task in the multi-task fusion optimization method of the present invention.
[0136] In one embodiment, determining a three-dimensional bounding box and a motion state of a traffic participant based on the fused BEV features and the predicted position of the traffic participant in the map includes:
[0137] S1510, predicting candidate center point locations of traffic participants in the map using a heat map to obtain a probability distribution of center point locations;
[0138] S1520, outputting the three-dimensional bounding box parameters of each candidate center point according to the center point position probability distribution;
[0139] S1530. Align and aggregate the parsing results of each level through deformable convolution according to the three-dimensional bounding box parameters to output the motion state of the traffic participant.
[0140] For example, the target detection module outputs the motion state of traffic participants. The input of the target detection module is the output of the fusion module, and the 3D positioning and state estimation of traffic participants can be achieved through a sparse detection head and a multi-scale regression network.
[0141] As an example, the target detection module uses the anchor-free detection framework, which first predicts the probability distribution of the target center position through the heat map:
[0142] H(x,y)=σ(W h *B decoded (x,y))
[0143] Among them, B decoded ∈R H×W×256 represents the BEV feature map after the detection head is decoded, W h ∈R 1×1×256 is the convolution kernel parameter, and σ() is the Sigmoid activation function. This heatmap is used to predict the probability of each location in space being the object center. Peaks in the heatmap correspond to candidate object center locations. The heatmap is used to filter out sparse keypoints, and the sparse detection head performs subsequent bounding box parameter regression only on these candidate points, reducing computational effort and improving detection efficiency.
[0144] To improve the performance of small target detection, a multi-scale feature fusion mechanism can be introduced to align and aggregate the BEV features of different levels output by the decoder through deformable convolution:
[0145]
[0146] Among them, s is the feature level index, Δx s,k , Δy s,k is the spatial offset of the deformable convolution, K is the number of sampling points, and w s,k is the learnable weight, F fuse is the weighted fusion feature vector of multi-scale and spatially aligned features at the same spatial position, F s () represents the sampling value of the feature map of the sth layer at the corresponding position.
[0147] For each candidate center point (x, y), the regression convolution kernel is used to parallelly predict its 3D bounding box parameters, including the center offset Δx, Δy, height z, size w, l, h, and heading angle θ:
[0148] [Δx, Δy, z, w, l, h, θ] = W r *B decoded (x,y)
[0149] Among them, W r ∈R 7×1×1×256 is the regression convolution kernel weight.
[0150] The regression loss uses the IoU-guided L1 loss, which is calculated as:
[0151]
[0152] Among them, b i and Represent the true and predicted i-th bounding box parameter vectors respectively. The classification task is optimized using FocalLoss to suppress the contribution of a large number of simple negative samples to the gradient, thereby increasing the attention to difficult-to-classify samples.
[0153] On this basis, the decoded BEV features are parsed through a sparse detection head, key candidate points are screened using heatmaps, and bounding box parameters are sparsely regressed, effectively reducing computational overhead and improving detection efficiency. Combining multi-scale deformable convolutional feature fusion with a weighted regression loss, this method achieves precise 3D positioning and state estimation for objects of varying sizes and shapes, meeting the requirements for efficient and accurate detection within a multi-task fusion framework.
[0154] The target detection method based on the bird's eye view of the embodiment of the application can solve the detection positioning error and segmentation connectivity loss caused by the imbalance of geometric-semantic features of a traditional BEV method by constructing a geometric-semantic double-flow enhanced BEV feature space, combining geometric flow features with semantic flow features, and fusing the features in time sequence; the decoding network of the heterogeneous task is designed, the detection and segmentation independent task head networks are deployed on the shared BEV backbone network, the detection task head network can be understood as a target detection module, mainly used for decoding three-dimensional bounding boxes based on a sparse query mechanism, the segmentation task head network can be understood as a map segmentation module, mainly used for generating a pixel-level mask through a U-Net encoding-decoding structure, and dynamically adjusting the learning weights of each adapter according to the task loss gradient, which can solve the problems of power redundancy and feature conflict caused by multiple independent task branches; by introducing a cross-task spatiotemporal consistency constraint mechanism, the BEV features of the historical frame and the current frame are aligned by using a motion compensation algorithm, the dynamic target trajectory smoothness loss and the static map stability loss are jointly optimized, the cross-task spatiotemporal consistency modeling is realized, and the problems of target ID jump rate and lane line cross-frame fracture rate caused by the asynchronous updating of dynamic and static tasks in time sequence are solved.
[0155] Figure 4 A structure schematic diagram of a target detection device based on a bird's eye view provided by the embodiment of the application. The target detection device based on the bird's eye view provided by the embodiment comprises:
[0156] The image feature extraction module 210 is configured to extract multi-scale features from the input images of the multi-view surround-view camera.
[0157] The point cloud feature extraction module 220 is configured to convert the point cloud data of the laser radar into a sparse voxel structure and generate BEV features for the voxel features.
[0158] The fusion module 230 is configured to fuse the multi-scale features and the BEV features to obtain fused BEV features.
[0159] The map segmentation module 240 is configured to extract road topology and semantic information according to the fused BEV features, and generate a map according to the road topology and the semantic information.
[0160] The target detection module 250 is configured to determine a three-dimensional bounding box and a motion state of a traffic participant according to the fused BEV features and a prediction result of the position of the traffic participant in the map.
[0161] The device fuses multi-scale geometric features, BEV features and semantic information, and synchronizes the time sequence features of dynamic target detection and static map segmentation, thereby improving the reliability of target detection based on the bird's eye view, and being suitable for multi-task BEV perception of three-dimensional object detection, high-precision map segmentation and motion prediction.
[0162] Based on any of the above embodiments, the image feature extraction module 210 includes:
[0163] A first generating unit is configured to generate, for each input image of the multi-view surround view camera, a feature map of corresponding resolution using a convolution operation through a deep residual network, and generate multi-scale features based on the feature map of corresponding resolution using stacked residual blocks;
[0164] An alignment unit, used to align the multi-scale features of each input image;
[0165] The second generation unit is used to fuse the multi-scale features of each input image through a feature map pyramid network.
[0166] Based on any of the above embodiments, the point cloud feature extraction module 220 includes:
[0167] A division unit is used to divide the point cloud data of the laser radar into a regular voxel grid within a set range according to the three-dimensional coordinates of the point cloud in space;
[0168] The sparse mapping unit is used to process non-empty voxel features based on sparse convolution to obtain sparse features, and map the sparse features to a dense BEV space to obtain two-dimensional BEV features.
[0169] Based on any of the above embodiments, the fusion module 230 includes:
[0170] a motion compensation unit, configured to perform motion compensation on the BEV features of the historical frames so as to transform the BEV features of the historical frames into a coordinate system of the current frame;
[0171] A splicing unit is used to splice the multi-scale features of the input image of the current frame and the BEV features of the point cloud data with the BEV features of the motion-compensated historical frames along the channel dimension to obtain a multi-channel joint BEV feature;
[0172] The fusion unit is used to fuse multi-channel joint BEV features using a convolutional neural network to obtain a fused BEV feature.
[0173] Based on any of the above embodiments, the map segmentation module 240 includes:
[0174] A downsampling unit, configured to downsample the BEV feature map level by level using an encoder to obtain features at each level;
[0175] An upsampling unit, configured to use the decoder to upsample step by step through transposed convolution, and perform jump connections with the features of the corresponding level of the encoder to obtain road topology and semantic information;
[0176] The modeling unit is used to model lane key points as graph nodes and learn the connection relationship between nodes through the graph attention network (GAT);
[0177] a mask output unit, configured to output a segmentation mask according to the road topology, the semantic information, and the connection relationship, wherein the segmentation mask includes a lane label, a drivable area label, and a road boundary label;
[0178] A map generating unit is configured to generate a map according to the segmentation mask.
[0179] Based on any of the above embodiments, the target detection module 250 includes:
[0180] a second prediction unit, configured to predict candidate center point positions of traffic participants in the map using a heat map to obtain a probability distribution of center point positions;
[0181] a parameter output unit, configured to output three-dimensional bounding box parameters of each candidate center point according to the center point position probability distribution;
[0182] A detection unit is configured to align and aggregate the parsing results of each level through deformable convolution according to the three-dimensional bounding box parameters to output the motion state of the traffic participant.
[0183] The bird's-eye view-based target detection device provided in the embodiment of the present application can be used to execute the bird's-eye view-based target detection method provided in any of the above embodiments, and has corresponding functions and beneficial effects.
[0184] Figure 5 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present application is shown. The electronic device 10 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device 10 can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, user equipment, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided for example only and are not intended to limit the implementation of the present application as described and / or claimed herein.
[0185] like Figure 5As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0186] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks and wireless networks.
[0187] The processor 11 may be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above.
[0188] In some embodiments, the methods of the above embodiments may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the method described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to perform any of the above-described methods in any suitable manner (e.g., by means of firmware).
[0189] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0190] Computer programs for implementing the methods of the present application may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0191] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. A computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0192] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device 10 having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device 10. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0193] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0194] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.
[0195] An embodiment of the present application also provides a computer program product, including a computer program and / or instructions, which, when executed by a processor, implements the bird's-eye view-based target detection method as described in any of the above embodiments.
[0196] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this application can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of this application can be achieved. This is not limited herein.
[0197] The above specific embodiments do not constitute a limitation on the scope of protection of this application. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application shall be included within the scope of protection of this application.
Claims
1. A target detection method based on a bird's-eye view, characterized in that: include: Extract multi-scale features from the input images of the multi-view surround camera; Convert the LiDAR point cloud data into a sparse voxel structure and generate BEV features based on the voxel features; Fusing the multi-scale features and the BEV features to obtain a fused BEV feature; extracting road topology and semantic information according to the fused BEV features, and generating a map according to the road topology and the semantic information; A three-dimensional bounding box and a motion state of the traffic participant are determined based on the fused BEV features and a prediction result of a position of the traffic participant in the map.
2. The method according to claim 1, characterized in that The multi-scale feature extraction for each input image of the multi-view surround view camera includes: For each input image of the multi-view surround camera, a deep residual network is used to generate a feature map of corresponding resolution using a convolution operation, and multi-scale features are generated based on the feature map of corresponding resolution using stacked residual blocks; Align the multi-scale features of each input image; The multi-scale features of each input image are fused through the feature map pyramid network.
3. The method according to claim 1, characterized in that The step of converting the laser radar point cloud data into a sparse voxel structure and generating BEV features based on the voxel features includes: Divide the laser radar point cloud data into a regular voxel grid within a set range according to the three-dimensional coordinates of the point cloud in space; Non-empty voxel features are processed based on sparse convolution to obtain sparse features, and the sparse features are mapped to a dense BEV space to obtain a two-dimensional BEV feature.
4. The method according to claim 1, wherein The fusing the multi-scale features and the BEV features to obtain a fused BEV feature includes: Performing motion compensation on the BEV features of the historical frame to transform the BEV features of the historical frame into the coordinate system of the current frame; The multi-scale features of the input image of the current frame and the BEV features of the point cloud data are spliced with the BEV features of the motion-compensated historical frames along the channel dimension to obtain the multi-channel joint BEV features; Convolutional neural network is used to fuse multi-channel joint BEV features to obtain fused BEV features.
5. The method according to claim 1, wherein The extracting road topology and semantic information according to the fused BEV features, and generating a map according to the road topology and the semantic information, includes: The encoder is used to downsample the BEV feature map level by level to obtain the features of each level; The decoder is used to upsample the image layer by layer through transposed convolution and perform jump connections with the features of the corresponding level of the encoder to obtain road topology and semantic information; The lane key points are modeled as graph nodes, and the connection relationship between nodes is learned through the Graph Attention Network (GAT); Outputting a segmentation mask according to the road topology, the semantic information, and the connection relationship, wherein the segmentation mask includes a lane line label, a drivable area label, and a road boundary label; A map is generated based on the segmentation mask.
6. The method according to claim 1, characterized in that The determining of a three-dimensional bounding box and a motion state of the traffic participant based on the fused BEV features and the predicted result of the position of the traffic participant in the map includes: Predicting candidate center point locations of traffic participants in the map using a heat map to obtain a probability distribution of the center point locations; Outputting the three-dimensional bounding box parameters of each candidate center point according to the center point position probability distribution; According to the three-dimensional bounding box parameters, deformable convolution is used to align and aggregate the parsing results of each level to output the motion state of the traffic participant.
7. A target detection device based on a bird's-eye view, characterized in that: include: An image feature extraction module is used to extract multi-scale features from the input image of the multi-view surround camera; Point cloud feature extraction module, used to convert the point cloud data of the lidar into a sparse voxel structure and generate BEV features based on the voxel features; A fusion module, configured to fuse the multi-scale features and the BEV features to obtain a fused BEV feature; a map segmentation module, configured to extract road topology and semantic information based on the fused BEV features, and generate a map based on the road topology and the semantic information; The target detection module is configured to determine a three-dimensional bounding box and a motion state of the traffic participant based on the fused BEV features and the predicted result of the position of the traffic participant in the map.
8. An electronic device, characterized in that: include: at least one processor; a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the bird's-eye view-based target detection method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the target detection method based on a bird's-eye view according to any one of claims 1 to 6 is implemented.
10. A computer program product comprising a computer program and / or instructions, characterized in that When the computer program and / or instructions are executed by a processor, the bird's-eye view-based target detection method according to any one of claims 1 to 6 is implemented.
Citation Information
Cited By
Method and system for encrypting road surface point cloud and medium
CN121438249A
Full waveform decomposition method based on multi-branch convolutional neural network
CN121561412A
BEVLane-based lightweight lane line detection method and system
CN122024211A
Pure vision three-dimensional target detection system and method based on bird's-eye view angle
CN122290076A