A Multi-Camera Based BEV Visual Perception Method

Through the BEV visual perception method of multi-camera, the Encoder, Decoder and Loss designs are used to solve the problem of feature conversion from 2D images to 3D scenes, and the accuracy of autonomous driving perception is improved.

CN115512326BActive Publication Date: 2025-07-25SHANGHAI XUNXU ARTIFICIAL INTELLIGENCE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211277046.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-18
Publication Date
2025-07-25
Estimated Expiration
2042-10-18

AI Technical Summary

Technical Problem

In the prior art, BEV visual perception methods have difficulties in learning feature conversion from 2D images to 3D scenes, resulting in poor fusion of multi-view image features, affecting the accuracy of subsequent autonomous driving perception tasks.

Method used

Using a multi-camera-based BEV visual perception method, the nuscenes data set is used to design the Encoder structure, Decoder structure and Loss design, image features are extracted through DenseNet and PANET networks, combined with transformer decoder layer and self-attention mechanism, and optimize feature matching using Hungarian algorithm and focal loss to achieve accurate conversion from 2D images to 3D scenes.

Benefits of technology

Improve the understanding of 2D images to 3D scenes and improve the accuracy of autonomous driving perception tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115512326B_ABST
    Figure CN115512326B_ABST
Patent Text Reader

Abstract

The present invention discloses a BEV visual perception method based on multiple cameras, including model design. The nuscenes dataset is used, and the input is 6 panoramic camera images. The model design includes an encoder structure, a decoder structure, and a final Loss design. Based on densnet, each image is encoded to extract image convolution features, and then a PANET network is used to output 3-layer multi-scale feature maps to strengthen information propagation. The detection head includes 6 layers of transformer decoder layers. 300 / 600 / 900 object queries are preset in advance, and each query is a 256-dimensional fusion feature. The object query predicts the 3D reference point coordinates in the BEV space by a fully connected network. After being normalized by the tanh function, the coordinates represent the relative position in the space. The Hungarian algorithm is used for bipartite graph matching between the detection boxes predicted by the object queries and all the ground truth boxes. The present invention proposes an improved multi-view feature extraction network, which can effectively solve the understanding ability from 2D images to 3D scenes, thereby effectively improving the accuracy of subsequent perception tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of artificial intelligence and autonomous driving, and particularly to a BEV visual perception method based on multiple cameras. Background Art

[0002] In the field of autonomous driving visual perception, a popular direction in the past two or three years has been the more direct visual perception from the BEV perspective. Different from depth estimation that first explicitly obtains the depth of each pixel point and then supports other related tasks, various tasks such as end-to-end object detection, semantic segmentation, and trajectory prediction can be achieved from the BEV perspective. Since this method is more straightforward and can be better used by downstream planning and control (in the same coordinate system), it has quickly become an important research direction for the implementation of future autonomous driving perception.

[0003] Since BEV features need to be obtained by fusing multi-view image features, it is necessary to first extract features from multi-view images, and one of the important difficulties lies in the feature conversion learning from 2D images to 3D scenes. Summary of the Invention

[0004] In order to overcome the deficiencies of the prior art, the present invention provides a BEV visual perception method based on multiple cameras to solve the problem that BEV features need to be obtained by fusing multi-view image features, so it is necessary to first extract features from multi-view images, and one of the important difficulties lies in the feature conversion learning from 2D images to 3D scenes as mentioned in the above background art.

[0005] To solve the above technical problems, the present invention provides the following technical solution: A BEV visual perception method based on multiple cameras, including model design, using the nuscenes dataset, with the input being 6 panoramic camera pictures. Among them, the model design mainly consists of three parts: including an encoder structure, a decoder structure, and a final loss design;

[0006] Encoder Structure:

[0007] Based on densnet, each picture is encoded to extract image convolution features, and then through the PANET network, 3 layers of multi-scale feature maps are output to strengthen information propagation;

[0008] Decoder Structure:

[0009] The detection head includes 6 layers of transformer decoder layers, and 300 / 600 / 900 object queries are preset. Each query is a 256-dimensional fused feature. The object query predicts the 3D reference point coordinates (x, y, z) in the BEV space by a fully connected network. After the coordinates are normalized by the tanh function, they represent the relative position in the space;

[0010] Loss Design:

[0011] The Hungarian algorithm is used for bipartite graph matching between the detection boxes predicted by the object queries and all the ground truth boxes to find the optimal matching that minimizes the loss. The focal loss is used for the loss calculation between categories to reduce the influence caused by sample imbalance, and the L2 regression loss is used to calculate the regression loss to facilitate the network to give a stable solution.

[0012] As a preferred technical solution of the present invention, in the Decoder structure, among each layer, all the object queries interact with each other through the self-attention mechanism to obtain global information and avoid multiple queries converging to the same object. Then, the object queries perform feature matching with the image features. The 3D coordinates of the real world corresponding to each query are projected onto the image coordinates through the internal and external parameters of the camera, and linear interpolation is used to sample the corresponding multi-scale image features. If the projected coordinates fall outside the image range, zeros are filled, and then the image feature sampling is used to update the object queries.

[0013] As a preferred technical solution of the present invention, the updated object query predicts the category of the corresponding object and the parameters of the bounding box through two fully connected neural networks respectively. To avoid data deviation, the prediction offset δ of the center point of each object is used to update the coordinates of the reference points. The updated object queries and reference points of each layer are used as the input of the next decoder layer for calculation and update again, and the iteration is performed 6 times in total.

[0014] As a preferred technical solution of the present invention, in the Decoder structure, since the value range of the tanh function is between [-1, +1], the output of the hidden layer is limited to between [-1, +1], which can be regarded as being distributed around the 0 value with a mean of 0. In this way, from the hidden layer to the output layer, the data has the effect of normalization (mean of 0).

[0015] Compared with the prior art, the beneficial effects that the present invention can achieve are as follows:

[0016] 1. An improved multi-view feature extraction network is proposed to make the projection features richer;

[0017] 2. It can effectively solve the understanding ability from 2D images to 3D scenes. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 It is a flowchart of a BEV visual perception method based on multiple cameras according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] In order to make the technical means, creative features, achieved purposes and functions of the present invention easy to understand, the present invention will be further described below in conjunction with specific embodiments. However, the following embodiments are only the preferred embodiments of the present invention, not all of them. Based on the embodiments in the embodiments, other embodiments obtained by those skilled in the art without creative efforts all belong to the protection scope of the present invention. The experimental methods in the following embodiments are all conventional methods unless otherwise specified. The materials, reagents, etc. used in the following embodiments can all be obtained from commercial channels unless otherwise specified.

[0020] Embodiment:

[0021] As Figure 1 shown, the present invention provides a BEV visual perception method based on multiple cameras, including model design. The nuscenes dataset is used, and the input is 6 panoramic camera pictures. Among them, the model design mainly consists of three parts: including an encoder structure, a decoder structure, and a final loss design;

[0022] Encoder structure:

[0023] Based on densnet, each picture is encoded to extract image convolution features, and then a PANET network is used to output 3 layers of multi-scale feature maps to strengthen information propagation. The reason for choosing Densnet is that DenseNet directly connects the feature maps from different layers, which can achieve feature reuse and thus improve efficiency;

[0024] Decoder structure:

[0025] The detection head includes 6 layers of transformer decoder layers. 300 / 600 / 900 object queries are preset. Each query is a 256-dimensional fused feature. The object query predicts the 3D reference point coordinates (x, y, z) in the BEV space by a fully connected network. After the coordinates are normalized by the tanh function, they represent the relative position in the space. Since the value range of the tanh function is between [-1, +1], the output of the hidden layer is limited between [-1, +1], which can be regarded as being distributed near the 0 value with a mean of 0. In this way, from the hidden layer to the output layer, the data has the effect of normalization (mean of 0);

[0026] Among each layer, all object queries interact with each other through the self-attention mechanism to obtain global information and prevent multiple queries from converging to the same object. Then, the object query performs feature matching with the image features. The 3D coordinates of the real world corresponding to each query are projected onto the image coordinates through the internal and external parameters of the camera, and linear interpolation is used to sample the corresponding multi-scale image features. If the projected coordinates fall outside the image range, zeros are filled, and then the image features are sampled to update the object queries;

[0027] The updated object query predicts the category of the corresponding object and the parameters of the bounding box through two fully connected neural networks respectively. To avoid data deviation, the prediction offset δ of the center point of each object is used to update the coordinates of the reference points. The updated object queries and reference points of each layer are used as the input of the next decoder layer for calculation and update again, with a total of 6 iterations;

[0028] Loss design:

[0029] The Hungarian algorithm is used for bipartite graph matching between the detection boxes predicted by the object queries and all the ground truth boxes to find the optimal matching that minimizes the loss. The focal loss is used for calculating the loss between categories to reduce the impact caused by sample imbalance, and the L2 regression loss is used for calculating the regression loss to facilitate the network to give a stable solution.

[0030] An improved multi-view feature extraction network is proposed in a BEV visual perception method based on multiple cameras provided by the present invention, which can effectively solve the understanding ability from 2D images to 3D scenes, thereby effectively improving the accuracy of subsequent perception tasks.

[0031] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A BEV visual perception method based on multiple cameras, characterized in that: It includes model design. Using the nuScenes dataset, the input is 6 panoramic camera images. Among them, the model design mainly consists of three parts: including the encoder structure, the decoder structure, and the final loss design; Encoder structure: Based on densnet, each image is encoded to extract image convolution features, and then through the PANET network, 3-layer multi-scale feature maps are output to strengthen information propagation; Decoder structure: The detection head includes 6 layers of transformer decoder layers. 300 / 600 / 900 object queries are preset. Each query is a 256-dimensional fused feature. The object query predicts the 3D reference point coordinates (x, y, z) in the BEV space by a fully connected network. After being normalized by the tanh function, the coordinates represent the relative position in the space; Loss design: The Hungarian algorithm is used for bipartite graph matching between the detection boxes predicted by the object queries and all the ground truth boxes to find the optimal matching that minimizes the loss. The focal loss is used for the loss calculation between categories to reduce the impact caused by sample imbalance. The L2 regression loss is used to calculate the regression loss to facilitate the network to give a stable solution.

2. The BEV visual perception method based on multiple cameras according to claim 1, characterized in that: In the decoder structure, within each layer, all object queries interact with each other through the self-attention mechanism to obtain global information and avoid multiple queries converging to the same object. Then, the object query and the image features perform feature matching. The 3D coordinates of the real world corresponding to each query are projected to the image coordinates through the internal and external parameters of the camera, and linear interpolation is used to sample the corresponding multi-scale image features. If the projected coordinates fall outside the image range, zeros are filled, and then the image features are sampled to update the object queries.

3. A BEV visual perception method based on multiple cameras according to claim 2, characterized in that: The updated object query predicts the category of the corresponding object and the parameters of the bounding box through two fully connected neural networks. To avoid data deviation, the prediction offset δ of the center point of each object is used to update the coordinates of the reference points. The updated object queries and reference points of each layer are used as the input of the next decoder layer for calculation and update again, and the iteration is performed 6 times in total.

4. A multi-camera-based BEV visual perception method according to claim 1, characterized in that: In the decoder structure, since the value range of the tanh function is between [-1, +1], the output of the hidden layer is limited to between [-1, +1], which can be regarded as being distributed near the 0 value with a mean of 0. In this way, from the hidden layer to the output layer, the data has the effect of normalization (mean of 0).

Citation Information

Patent Citations

  • Projection full-convolution network three-dimensional model segmentation method based on fusion of multi-view-angle features

    CN108389251A

  • Map generation device, recording medium and map generation method

    CN113220805A