BEV sensing method and device, electronic equipment and storage medium

By combining DINOv2 model and LoRA technology, multi-view images are processed and BEV feature maps are generated, which solves the problem of insufficient computing resource limitation and generalization capabilities of existing BEV perception technologies, and achieves a more efficient and reliable BEV perception effect.

CN119919906APending Publication Date: 2025-05-02BEIJING TRUNK TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510005978.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-05-02

AI Technical Summary

Technical Problem

The existing BEV perception technology has significantly reduced its effectiveness under problems such as limitations in computing resources and insufficient generalization capabilities, especially in the face of environmental changes, severe weather or camera failures.

Method used

Using a combination of DINOv2 model and LoRA technology, multi-view image is processed through self-supervised learning and low-rank adapter pre-training, and 2D features are obtained and mapped to 3D space for fusion to generate BEV feature maps.

Benefits of technology

It reduces the impact of environmental changes on BEV perception, reduces the demand for computing resources, reduces the dependence on a large amount of labeled data, enhances generalization capabilities, and improves the reliability and efficiency of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919906A_ABST
    Figure CN119919906A_ABST
Patent Text Reader

Abstract

The invention provides a BEV sensing method and device, electronic equipment and a storage medium. According to the BEV sensing method, a DINOv2 model is used for processing a multi-view image of the surrounding environment of a vehicle to obtain a multi-view 2D feature, the multi-view 2D feature is mapped to a 3D space and fused to obtain a 3D fusion feature capable of serving as a complete feature representation of the surrounding environment of the vehicle, and finally, the 3D fusion feature is used for obtaining a BEV feature map of the surrounding environment of the vehicle. Wherein the DINOv2 model is obtained by pre-training through a self-supervised learning mode and a low-rank adaptation adapter. According to the method and the device, the influence of environment change on BEV perception can be reduced, the computing resource demand of BEV perception is reduced, dependence on annotated data is not needed, and the method and the device are more reliable, efficient and practical.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of autonomous driving technology, and in particular to a BEV perception method, device, electronic device and storage medium. Background Art

[0002] Bird's Eye View (BEV) perception technology aims to extract a two-dimensional plane representation from the images captured by multiple on-board cameras from a bird's eye view, which is crucial for the navigation and decision-making of autonomous vehicles. At present, the BEV perception scheme in related technologies has a significant decline in effectiveness when faced with problems such as limited computing resources and insufficient generalization capabilities. Summary of the invention

[0003] In view of this, the present disclosure provides a BEV perception method, device, electronic device and storage medium.

[0004] According to a first aspect of the present disclosure, a BEV perception method is provided, the method comprising:

[0005] Acquire multi-view images of the vehicle's surroundings;

[0006] Processing the multi-view images using a DINOv2 model to obtain multi-view 2D features, wherein the DINOv2 model is pre-trained by self-supervised learning and a low-rank adaptive adapter;

[0007] Mapping the multi-view 2D features to 3D space and fusing them to obtain 3D fused features, where the 3D fused features are complete feature representations of the vehicle's surroundings;

[0008] The 3D fusion features are used to obtain a BEV feature map of the vehicle's surroundings.

[0009] In some implementations of the first aspect of the present disclosure, the method further includes: before processing the multi-view image using the DINOv2 model, performing dynamic image enhancement on the multi-view image.

[0010] In some embodiments of the first aspect of the present disclosure, the use of the DINOv2 model to process the multi-perspective images to obtain multi-perspective 2D features includes: performing the following processing on the image of the vehicle's surroundings at each perspective to obtain the 2D features at that perspective: dividing the image of the vehicle's surroundings at a first perspective into blocks of a predetermined size; processing the blocks through the DINOv2 model to obtain 2D features of the first perspective, the 2D features of the first perspective including semantic information of the vehicle's surroundings at the first perspective and related information of the 3D spatial structure around the vehicle, the first perspective being any one of the multiple perspectives.

[0011] In some embodiments of the first aspect of the present disclosure, the DINOv2 model includes multiple Transformer encoder layers, each of which includes a self-attention layer; during the training process of the DINOv2 model, the LoRA adapter is used to introduce a low-rank matrix in the self-attention layer for parameter update.

[0012] In some embodiments of the first aspect of the present disclosure, mapping the multi-view 2D features to the 3D space and fusing them to obtain 3D fused features includes: mapping the multi-view 2D features to the 3D space to obtain multi-view 3D features; and fusing the multi-view 3D features to obtain the 3D fused features.

[0013] In some embodiments of the first aspect of the present disclosure, multi-perspective 2D features are mapped to 3D space to obtain multi-perspective 3D features in the following manner: a depth estimation network is used to predict the depth information of each pixel in an image of the vehicle's surroundings at a first perspective; the depth information is used to map the 2D features of the first perspective to 3D space to obtain the 3D features of the first perspective; wherein the first perspective is any one of the multiple perspectives.

[0014] In some embodiments of the first aspect of the present disclosure, the using of the 3D fusion features to obtain a BEV feature map of the vehicle's surroundings includes: discretizing the 3D fusion features into voxels to obtain voxel features of the vehicle's surroundings; and using the voxel features of the vehicle's surroundings to obtain a BEV feature map of the vehicle's surroundings.

[0015] According to a second aspect of the present disclosure, a BEV sensing device is provided, comprising:

[0016] An acquisition unit, used to acquire multi-view images of the environment surrounding the vehicle;

[0017] A 2D feature extraction unit, configured to process the multi-view images using a DINOv2 model to obtain multi-view 2D features, wherein the DINOv2 model is obtained by pre-training with a low-rank adaptive adapter in a self-supervised learning manner;

[0018] A 3D space conversion and fusion unit, used for mapping the multi-view 2D features to a 3D space and fusing them to obtain a 3D fused feature, wherein the 3D fused feature is a complete feature representation of the vehicle's surrounding environment;

[0019] The BEV feature extraction unit is used to obtain a BEV feature map of the vehicle's surrounding environment using the 3D fusion features.

[0020] According to a third aspect of the present disclosure, an electronic device is provided, comprising: one or more processors and a memory storing a program, wherein the program comprises instructions, and when the instructions are executed by the processor, the processor executes the above method.

[0021] According to a fourth aspect of the present disclosure, a computer-readable storage medium storing a program is provided, wherein the program includes instructions, which, when executed by one or more processors of a computing device, cause the computing device to execute the above method.

[0022] It can be seen from the above technical solutions that the disclosed embodiment adopts the DINOv2 model in BEV perception to obtain the 2D features of multi-view images. The DINOv2 model is pre-trained by self-supervised learning and a low-rank adaptation (LoRA) adapter. The combination of the DINOv2 model and the LoRA technology can not only reduce the impact of environmental changes such as light brightness changes, bad weather or camera failures on BEV perception, but also reduce the demand for computing resources, reduce the dependence on a large amount of labeled data, and enhance the generalization ability, making it more reliable, efficient and practical. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0024] Figure 1 A schematic diagram of a process of a BEV sensing method provided by an embodiment of the present disclosure;

[0025] Figure 2 A schematic diagram of a specific implementation flow of step 103 in the BEV perception method provided in an embodiment of the present disclosure;

[0026] Figure 3 A schematic diagram of the structure of a BEV sensing device provided in an embodiment of the present disclosure;

[0027] Figure 4 A schematic structural block diagram of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0028] The following will be combined with the drawings in the embodiments of the present disclosure to clearly and completely describe the technical solutions in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.

[0029] The terms used in the embodiments of the present disclosure are only for the purpose of describing specific embodiments, and are not intended to limit the present disclosure. The singular forms "a", "said" and "the" used in the embodiments of the present disclosure and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings.

[0030] As used herein, the words "if," "if," and the like may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)," depending on the context.

[0031] The following is a brief description of the relevant technology.

[0032] In the field of autonomous driving, the BEV perception method usually includes the following steps: using a deep learning model with a convolutional neural network (CNN) as the backbone network to extract 2D features of images from each perspective, converting these 2D features into 3D features, fusing the 3D features of each perspective and decoding them into a 2D BEV feature map for subsequent driving tasks such as target detection, tracking, and motion prediction.

[0033] The BEV perception method has the following main problems:

[0034] 1) Sensitive to environmental changes: Performance will degrade significantly when faced with different environmental conditions, such as brightness changes, bad weather, or camera failures;

[0035] 2) High computing resource requirements: BEV perception methods in related technologies often require the use of complex deep learning models and perform high-complexity calculations, which have high computing resource requirements and may lead to inefficiency and increased hardware costs in vehicle deployment;

[0036] 3) Strong dependence on training data: The BEV perception methods of related technologies usually require a large amount of labeled training data to train deep learning models, and the collection and labeling of these training data is time-consuming and expensive. In addition, the deep learning model is very sensitive to the distribution of training data and often performs poorly in scenarios that are not covered by training data;

[0037] 4) Limited generalization ability: The BEV perception method of related technologies performs well in specific data sets or specific scenarios, but has limited generalization ability in new scenarios or scenarios not covered by their model training data, and the performance will drop significantly.

[0038] In view of this, the embodiments of the present disclosure provide the following BEV perception method, device, electronic device and storage medium, which adopt the DINOv2 model to obtain the 2D features of multi-view images, and the DINOv2 model is pre-trained by self-supervised learning and low-rank adaptation (LoRA) adapter. The embodiments of the present disclosure combine the DINOv2 model and LoRA technology to not only reduce the impact of environmental changes such as light brightness changes, bad weather or camera failures on BEV perception, but also reduce the demand for computing resources, reduce the dependence on a large amount of labeled data and enhance the generalization ability, making it more reliable, efficient and practical.

[0039] The embodiments of the present disclosure can be applied to ports, highways, mines, farms, closed parks, urban transportation and other scenarios, and can be applied to logistics distribution, unmanned transportation, terminal distribution, car travel, automated agricultural operations, automated sanitation and many other aspects. Of course, the embodiments of the present disclosure can also be applied to any other intelligent control scenarios involving equipment such as vehicles, and the present disclosure does not limit the application scenarios and applicable fields of the embodiments of the present disclosure.

[0040] The disclosed embodiments can be applied to the control of various types of equipment such as multiple wheeled mobile robots, wheeled mobile robots, mobile robots, vehicles, aircraft, ships, and intelligent rail rapid transit systems (ART, Autonomous rail Rapid Transit). The vehicles can be, but are not limited to, passenger cars, commercial vehicles (e.g., trucks, buses, vans, etc.), special-purpose vehicles (e.g., ambulances, fire trucks, engineering vehicles, rescue vehicles, etc.), agricultural and industrial vehicles (e.g., harvesters, forklifts, etc.), transportation and logistics vehicles (e.g., container trucks, refrigerated trucks, etc.), new energy vehicles (e.g., electric vehicles, hybrid electric vehicles), and special vehicles (e.g., garbage trucks, sprinkler trucks, etc.). In other words, the "vehicle" in the disclosed embodiments is equivalent to the aforementioned various equipment.

[0041] The specific implementation of the embodiment of the present disclosure is described in detail below.

[0042] Figure 1 The flowchart of the BEV perception method provided by the embodiment of the present disclosure is shown. The BEV perception method of the embodiment of the present disclosure can be performed by the electronic device below, which can be implemented as but not limited to a domain controller installed in a vehicle. Figure 1 , the BEV sensing method may include:

[0043] Step 101, acquiring multi-view images of the vehicle's surroundings;

[0044] Step 102, using the DINOv2 model to process the multi-view image to obtain multi-view 2D features, the DINOv2 model is obtained by self-supervised learning and pre-training with a low-rank adaptive adapter;

[0045] Step 103, mapping the multi-view 2D features to the 3D space and fusing them to obtain a 3D fused feature, where the 3D fused feature is a complete feature representation of the vehicle's surrounding environment;

[0046] Step 104 , using the 3D fusion features to obtain a BEV feature map of the vehicle's surroundings.

[0047] The DINOv2 model is pre-trained through self-supervised learning and can learn robust feature representations from a large number of images. This pre-training method allows the model to maintain high performance and robustness when facing different environmental conditions, such as brightness changes, bad weather, or camera failures. As a result, compared with models that rely on specific environmental conditions, the DINOv2 model can provide more stable and reliable feature extraction, reducing the impact of environmental changes on BEV perception.

[0048] LoRA technology introduces a low-rank matrix in the self-attention layer of the DINOv2 model to update parameters. This method only requires training a small number of parameters, thereby reducing the consumption of computing resources. Compared with traditional fine-tuning methods, LoRA technology can significantly reduce the computational burden of the model while maintaining model performance, making the model more suitable for deployment in resource-constrained vehicle environments.

[0049] The self-supervised learning characteristics of the DINOv2 model reduce the reliance on large amounts of labeled data. In addition, LoRA technology further reduces the need for training data by updating only a small portion of the model's parameters. The combination of the DINOv2 model and LoRA technology not only reduces the reliance on labeled data, but also improves the model's performance in unseen scenarios and enhances the model's generalization ability. It can be seen that the combination of the DINOv2 model and LoRA can better adapt to new scenarios and a variety of environmental conditions, has stronger generalization capabilities, and can be easily applied to a variety of different scenarios.

[0050] The specific implementation of each step in the BEV perception method of the embodiment of the present disclosure is described in detail below.

[0051] In the disclosed embodiment, the multi-perspective image of the vehicle's surrounding environment includes images of the vehicle's surrounding environment captured simultaneously from multiple perspectives. Specifically, the multi-perspective image of the vehicle's surrounding environment can be collected by multiple on-board cameras mounted on the vehicle, and these on-board cameras can capture images of the vehicle's surrounding environment from different angles and transmit them to the domain controller via the on-board network. After receiving the images collected by each on-board camera, the domain controller can align the timestamps of these images to form a multi-perspective image of the vehicle's surrounding environment.

[0052] In some examples, the "multi-perspectives" of the embodiments of the present disclosure may include six perspectives, namely, the perspective in front of the vehicle, the perspective behind the vehicle, the perspective to the left of the vehicle, the perspective to the right of the vehicle, the perspective above the vehicle, and the perspective below the vehicle. The first perspective mentioned below may be any one of the six perspectives. In specific applications, cameras may be evenly deployed around the vehicle so that these cameras cover the surroundings of the vehicle without blind spots, and each camera captures an image of the surrounding environment of the vehicle that includes a portion of the scene around the vehicle.

[0053] Furthermore, before step 102, the BEV perception method of the embodiment of the present disclosure may also include: an image preprocessing step. Specifically, preprocessing such as cleaning, denoising, white balance adjustment, size standardization, etc. is performed on the images of each perspective in the multi-perspective image of the vehicle's surrounding environment, so as to perform the processing of step 102. In some examples, an image processing library (such as OpenCV) is used to perform the aforementioned image preprocessing such as denoising, white balance adjustment, size standardization, etc. The embodiment of the present disclosure does not limit the specific implementation method of image preprocessing.

[0054] In some examples, image preprocessing may also include: performing color correction on the image to eliminate the impact of ambient light changes, so as to improve the accuracy and reliability of subsequent multi-view 2D features and reduce the impact of environmental conditions on BEV perception.

[0055] Furthermore, before step 102, that is, before processing the multi-view images using the pre-trained DINOv2 model, the BEV perception method of the disclosed embodiment may further include: performing dynamic image enhancement on the multi-view images. Specifically, dynamic image enhancement is performed on the images of each view in the multi-view images of the vehicle's surroundings. Dynamic image enhancement helps to improve image quality, improve the accuracy and reliability of subsequent multi-view 2D features, and further reduce the impact of environmental conditions on BEV perception.

[0056] In some embodiments, dynamic image enhancement of multi-view images may include, but is not limited to, one or more of the following: 1) real-time monitoring of the light intensity of the vehicle's surrounding environment, and dynamically adjusting the contrast and brightness of the vehicle's surrounding environment images at each viewing angle according to the light intensity of the vehicle's surrounding environment. 2) dynamic image enhancement of the vehicle's surrounding environment images at each viewing angle is performed separately through a dynamic range compression algorithm such as High Dynamic Range (HDR) to improve the image quality of multi-view images under high dynamic lighting conditions. As a result, not only can the BEV perception system maintain high performance and stable operation under various environmental conditions, different lighting, and various weather conditions, but also the image quality of the multi-view images can be significantly improved through dynamic image enhancement, making the extraction of multi-view 2D features in step 102 more accurate.

[0057] In the disclosed embodiment, the DINOv2 model is based on Transformer, which uses a vision transformer (ViT) as the backbone network.

[0058] The DINOv2 model is trained by self-supervised learning. Specifically, the DINOv2 model can be trained under the self-supervised learning framework. During the DINOv2 model training process, large-scale and diverse datasets can be used, multi-scale image cropping and knowledge distillation techniques can be used, and the teacher-student architecture can be used to improve its learning efficiency and performance.

[0059] The DINOv2 model can include multiple Transformer encoder layers, each of which includes a self-attention layer. During the training process of the DINOv2 model, the LoRA adapter is used to introduce a low-rank matrix in the self-attention layer for parameter update. By introducing a low-rank matrix in the self-attention layer of the DINOv2 model for parameter update, the LoRA technology only needs to train a small number of parameters, which can significantly reduce the computing resource consumption during the training of the DINOv2 model and enable the DINOv2 model to adapt to the needs of BEV perception more quickly.

[0060] Furthermore, the BEV perception method of the embodiment of the present disclosure may also include: using the LoRA adapter to dynamically adjust the parameters of the DINOv2 model using an adaptive learning algorithm. Specifically, the parameters of the DINOv2 model can be further optimized using an online learning mechanism based on system performance feedback. By automatically adjusting the parameters of the DINOv2 model using an adaptive learning algorithm through the LoRA adapter, it can better adapt to different operating environments, task requirements and emerging target types, further reduce dependence on manual intervention, and achieve automatic adjustment and optimization of BEV perception.

[0061] In step 102, the multi-view 2D features include 2D features obtained after the images of the vehicle's surroundings at each viewing angle are processed by the DINOv2 model. In other words, the multi-view 2D features include 2D features of multiple viewing angles. If the multi-view images of the vehicle's surroundings include: a front viewing angle image, a rear viewing angle image, a left viewing angle image, a right viewing angle image, an upper viewing angle image, and a lower viewing angle image, the multi-view 2D features include: 2D features of the front viewing angle image (i.e., 2D features of the front viewing angle), 2D features of the rear viewing angle image (i.e., 2D features of the rear viewing angle), 2D features of the left viewing angle image (i.e., 2D features of the left viewing angle), 2D features of the right viewing angle image (i.e., 2D features of the right viewing angle), 2D features of the upper viewing angle image (i.e., 2D features of the upper viewing angle), and 2D features of the lower viewing angle image (i.e., 2D features of the lower viewing angle). Hereinafter, the multi-view 3D features are similar, and include 3D features of multiple viewing angles, and the 3D features of each viewing angle are obtained by the 2D features of the corresponding viewing angle.

[0062] In some embodiments, step 102 may include: performing the following processing on the image of the vehicle's surroundings at each perspective to obtain 2D features at that perspective: dividing the image of the vehicle's surroundings at a first perspective into blocks of a predetermined size; processing the blocks through a DINOv2 model to obtain 2D features of a first perspective, where the first perspective is any one of the multiple perspectives.

[0063] The 2D features of each perspective can be regarded as a 2D feature map, which contains the semantic features of the image at that perspective. These 2D feature maps not only represent the characteristic information of the image content (for example, the shape, texture, color, and spatial relationship of objects in the image), but also retain information related to the 3D spatial structure around the vehicle.

[0064] Specifically, the image of the vehicle's surroundings from a certain perspective is divided into fixed-size patches, the size of which can be set to 14x14 pixels. These patches are embedded through a linear layer to map the image space to a high-dimensional feature space. The DINOv2 model uses multiple Transformer encoder layers to process these patches in the high-dimensional feature space. Each Transformer encoder layer includes a self-attention layer and a feed-forward network (FFN). The self-attention layer calculates the attention score of each patch with all other images to capture the long-distance dependencies within the image. The FFN layer further processes the features from the self-attention layer to extract higher-level semantic information. After being processed by a series of Transformer encoder layers in the DINOv2 model, a 2D feature containing a sequence of class tokens and patch tokens can be obtained. The class token can be used to represent the category information of the entire image, and the patch tokens are the vector representations obtained after processing the patches. Therefore, the 2D features obtained by the DINOv2 model not only contain rich semantic information in the image, but also contain relevant information about the 3D spatial structure around the vehicle, and the 2D features from different perspectives have better continuity.

[0065] The DINOv2 model improves the accuracy and robustness of BEV perception by extracting rich semantic features. The DINOv2 model can learn rich feature representations by pre-training on unlabeled data, reducing the dependence on large amounts of labeled data. By pre-training the DINOv2 model on a variety of tasks and data sets, it can be generalized to new or unseen scenarios. The LoRA adapter achieves task adaptation through a small number of parameter updates, reducing computing resource consumption and accelerating the convergence of the DINOv2 model. Therefore, through the strong generalization ability of the DINOv2 model and the parameter efficiency of LoRA, the BEV perception method of the embodiment of the present disclosure can provide more stable BEV perception and good robustness under various environmental conditions, while having low demands on computing resources.

[0066] Figure 2 FIG. 1 is a schematic diagram showing a specific implementation flow of step 103 in the BEV sensing method provided by an embodiment of the present disclosure. Figure 2 , step 103 may include the following steps:

[0067] Step 201, mapping the multi-view 2D features to the 3D space to obtain the multi-view 3D features;

[0068] In some examples, a geometric transformation (such as a perspective transformation) may be used to map the 2D features of each viewing angle into a predetermined 3D space coordinate system to obtain the 3D features of the corresponding viewing angle.

[0069] In some examples, geometric transformation and depth estimation may be used to map 2D features of each perspective into the same 3D space coordinate system to obtain 3D features of each perspective.

[0070] In some examples, multi-perspective 2D features can be mapped to 3D space to obtain multi-perspective 3D features in the following manner: a depth estimation network is used to predict the depth information of each pixel in an image of the vehicle's surroundings at a first perspective, and the depth information is used to map the 2D features of the first perspective to 3D space to obtain the 3D features of the first perspective, where the first perspective is any one of the multiple perspectives.

[0071] In some examples, the multi-view 2D features can be pulled up to the 3D space by the 3D Deformable Attention (DFA3D) operator to obtain the corresponding multi-view 3D features. DFA3D is able to combine the multi-layer optimization capabilities of the Transformer to solve the single-pass problem in Lift-Splat, while maintaining the perception of depth to solve the depth ambiguity problem in the 2D attention mechanism. Specifically, DFA3D first uses the depth obtained by its depth estimation subnetwork to expand the 2D features of each view to the 3D space, and then uses DFA3D to aggregate features from the expanded 3D feature map, thereby effectively alleviating the depth ambiguity problem. Due to the existence of the Transformer-like architecture, DFA3D can gradually optimize the pulled-up features layer by layer.

[0072] Step 202 , fusing multi-view 3D features to obtain 3D fused features, where the 3D fused features are complete feature representations of the vehicle's surroundings.

[0073] Specifically, existing fusion technologies (such as attention fusion or feature pooling) or predetermined fusion strategies can be used to integrate multi-view 3D features to obtain 3D fusion features, which are complete feature representations of the vehicle's surroundings. Since 3D features from different perspectives provide different parts of information about the vehicle's surroundings (for example, the forward camera may not be able to see objects on the side of the vehicle, while the side camera can), integrating 3D features from multiple perspectives can make up for the shortcomings of each perspective and reduce the impact caused by occlusion or perspective limitations. The 3D fusion features obtained contain information from all perspectives, forming a comprehensive understanding of the vehicle's surroundings and improving the robustness of the overall environmental representation.

[0074] In some examples, an attention mechanism can be used to fuse multi-view 3D features to form 3D fusion features. The attention mechanism can be, but is not limited to, spatial cross-attention, temporal self-attention, dual cross-view spatial attention mechanism (VISTA), etc. The attention mechanism can identify the importance of features from different perspectives and assign different weights accordingly, thereby achieving effective feature fusion. This method is particularly suitable for processing occluded or dynamic scenes because it can combine feature maps of multiple frames to infer occluded areas and identify the motion state of objects.

[0075] The fusion of multi-view 3D features can be achieved through a variety of fusion strategies. In some examples, it can be achieved through early fusion, late fusion, or other similar methods. The following is an exemplary description of different fusion methods:

[0076] 1) Early fusion: Fusion is performed at the feature level; specifically, feature maps from different perspectives are directly concatenated in the channel dimension and then input into a model such as DINOv2 for processing. The advantage of this method is that it can capture complementary information between different perspectives and improve the expressiveness of features.

[0077] 2) Late fusion: Fusion is performed at the decision level; specifically, the 3D features of each view are processed independently to obtain their own prediction results, and then these prediction results are integrated to obtain 3D fused features. When this method is used, features from multiple views can be processed in parallel, which can reduce the amount of calculation.

[0078] Through the designed fusion strategy, the feature maps of different perspectives are effectively integrated to obtain a comprehensive, multi-perspective feature representation. This fused feature representation can be used for subsequent BEV perception tasks such as object detection and semantic segmentation.

[0079] In the disclosed embodiment, the DINOv2 model and the aforementioned feature fusion strategy work together to provide a comprehensive, robust and complete feature representation for the autonomous driving system, so that the 3D fused features can accurately understand and predict the environment around the vehicle under various environmental conditions.

[0080] As can be seen from the above, step 103 converts 2D features of different perspectives into 3D spatial representations and fuses them through 3D spatial conversion and fusion, so as to obtain 3D fused features as complete feature representations of the vehicle's surroundings.

[0081] In specific applications, the fusion strategy can be optimized in real time based on application conditions and performance feedback to adapt to different driving scenarios and environmental changes.

[0082] There are many specific implementation methods for obtaining a BEV feature map through 3D fusion features. In some embodiments, step 104 may include: discretizing the 3D fusion features into voxels (grid) to obtain voxel features of the vehicle's surroundings, and using the voxel features of the vehicle's surroundings to obtain a BEV feature map of the vehicle's surroundings. Specifically, operations such as pooling and convolution can be used to compress 3D voxel features into 2D BEV feature maps. Alternatively, the voxel features are projected onto a 2D BEV plane to obtain a BEV feature map of the vehicle's surroundings.

[0083] Furthermore, step 104 may also include: enhancing the BEV feature map of the vehicle's surroundings through an attention mechanism. In some examples, the BEV feature map may be enhanced through any of the following attention mechanisms: 1) Channel Attention: This attention mechanism focuses on the importance of different feature map channels. In the above processing of step 104, the feature representation of important channels can be enhanced by assigning weights to each channel of the 3D fusion feature, so that the features related to important channels in the final BEV feature map can be strengthened. 2) Spatial Attention: This mechanism focuses on the importance of different spatial positions in 3D features. In the processing of step 104, weights can be assigned to each spatial position in the voxel features of the vehicle's surroundings to enhance the feature representation of important positions, so that the final BEV feature map can better reflect the features of key areas in the image. As a result, important information in the BEV feature map, such as complementary information between different cameras, can be enhanced according to the needs of specific application scenarios, thereby further improving the accuracy, reliability and flexibility of BEV perception.

[0084] It should be noted that the above description of the specific implementation methods of each step in the BEV perception method provided in the embodiment of the present disclosure is only for example and is not intended to limit the specific implementation method of the embodiment of the present disclosure.

[0085] The BEV perception method provided in the embodiments of the present disclosure can be applied to various specific tasks, such as target detection and segmentation based on BEV features, motion prediction based on BEV features, driving decision-making and path planning based on BEV features, and construction of a high-precision map around the vehicle using BEV features.

[0086] The BEV perception method provided by the embodiments of the present disclosure can conveniently perform online learning and adaptive adjustment so as to be continuously optimized and adapt to the needs of various scenarios and various specific tasks.

[0087] Figure 3 FIG. 1 is a schematic diagram showing the structure of a BEV sensing device provided by an embodiment of the present disclosure. Figure 3 , the BEV sensing device 300 of the embodiment of the present disclosure may include:

[0088] An acquisition unit 301 is used to acquire multi-view images of the environment around the vehicle;

[0089] A 2D feature extraction unit 302 is used to process the multi-view images using a DINOv2 model to obtain multi-view 2D features, where the DINOv2 model is obtained by self-supervised learning and pre-training with a low-rank adaptive adapter;

[0090] A 3D space conversion and fusion unit 303 is used to map the multi-view 2D features to the 3D space and fuse them to obtain a 3D fusion feature, where the 3D fusion feature is a complete feature representation of the vehicle's surrounding environment;

[0091] The BEV feature extraction unit 304 is used to obtain a BEV feature map of the vehicle surroundings using the 3D fusion features.

[0092] Furthermore, the BEV perception device 300 of the embodiment of the present disclosure may also include: an image enhancement unit 305 for dynamically enhancing the multi-view images.

[0093] Furthermore, the 2D feature extraction unit 302 can be specifically used to perform the following processing on the image of the vehicle's surroundings at each perspective to obtain the 2D features at that perspective: divide the image of the vehicle's surroundings at a first perspective into blocks of a predetermined size; process the blocks through the DINOv2 model to obtain the 2D features of the first perspective, and the 2D features of the first perspective include semantic information of the vehicle's surroundings at the first perspective and related information of the 3D spatial structure around the vehicle, and the first perspective is any one of the multiple perspectives.

[0094] Furthermore, the 3D space conversion and fusion unit 303 can be specifically used to map the multi-view 2D features to the 3D space to obtain the multi-view 3D features, and fuse the multi-view 3D features to obtain the 3D fusion features. The following method is used to map the multi-view 2D features to the 3D space to obtain the multi-view 3D features: using a depth estimation network to predict the depth information of each pixel in the image of the vehicle's surrounding environment at a first perspective; using the depth information to map the 2D features of the first perspective to the 3D space to obtain the 3D features of the first perspective; wherein the first perspective is any one of the multiple perspectives.

[0095] Furthermore, the BEV sensing device 300 of the embodiment of the present disclosure may also include: a voxelization unit 306, which is used to discretize the 3D fusion features into voxels to obtain voxel features of the vehicle surrounding environment. The BEV feature extraction unit 304 can be specifically used to obtain a BEV feature map of the vehicle surrounding environment using the voxel features of the vehicle surrounding environment.

[0096] In a specific application, the BEV sensing device 300 can be implemented by software, hardware, or a combination of the two. For example, the BEV sensing device 300 can be implemented as software running in the electronic device 400 described below.

[0097] In addition, an embodiment of the present disclosure further provides a computer-readable storage medium on which a computer program is stored. The program includes instructions. When the instructions are executed by one or more processors of a computing device, the steps of the aforementioned BEV perception method are executed.

[0098] Figure 4 FIG. 1 is a schematic diagram showing the structure of an electronic device provided by an embodiment of the present disclosure. Figure 4 The electronic device 400 may include: one or more processors 401, and a memory 402 storing one or more programs, which are executed by the one or more processors 401 to implement the method flow shown in the above embodiments of the present disclosure and / or program units corresponding to each unit in the device.

[0099] The various components are interconnected using different buses and can be mounted on a common motherboard or otherwise as needed. The processor 401 can process instructions executed within the electronic device, including instructions stored in or on the memory to display graphical information of a user interface on an external input / output device (such as a display device coupled to the interface). In other embodiments, multiple processors and / or multiple buses can be used with multiple memories and multiple memories if desired.

[0100] Processor 401 may include one or more single-core processors or multi-core processors. Processor 401 may include any combination of general-purpose processors or dedicated processors (such as image processors, application processors, baseband processors, etc.).

[0101] The memory 402 is a computer-readable storage medium provided by the present disclosure, which can be used to store non-transient software programs, non-transient computer executable programs and units, such as the following in the embodiments of the present disclosure: Figure 1 The processor 401 executes the non-transient software programs, instructions and units stored in the memory 402 to perform the above method embodiments. Figure 1 The programs, instructions and units corresponding to the BEV perception method shown.

[0102] The electronic device 400 may further include: an input device 403 and an output device 404. The processor 401, the memory 402, the input device 403 and the output device 404 may be connected via a bus or other means. Figure 4 The example of connecting through bus is taken in the following.

[0103] The input device 403 can receive input digital or character information, and generate signal input related to user settings and function control, such as a touch screen, a keypad, a mouse, a track pad, a touch pad, an indicator rod, one or more mouse buttons, a trackball, a joystick and other input devices. The output device 404 may include a display device, an auxiliary lighting device (e.g., an LED) and a tactile feedback device (e.g., a vibration motor), etc. The display device may include, but is not limited to, a liquid crystal display (LCD), a light emitting diode (LED) display and a plasma display. In some embodiments, the display device may be a touch screen.

[0104] The above-mentioned programs (also referred to as software, software applications, or codes) include machine instructions for programmable processors, and these computer programs can be implemented using object-oriented programming languages, assembly or machine languages.

[0105] With the development of time and technology, the meaning of medium is becoming more and more extensive, and the propagation path of computer programs is no longer limited to tangible media, and can also be downloaded directly from the network, etc. Any combination of one or more computer-readable storage media can be used. Computer-readable storage media can be used but not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or devices, or any combination of the above. More specific examples (non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this document, computer-readable storage media can be any tangible medium containing or storing programs, which can be used by or in combination with instruction execution systems, devices or devices.

[0106] In a specific application, the electronic device 400 may be implemented as, but not limited to, a domain controller or other similar devices.

[0107] The technical solution provided by the present disclosure is described in detail above. The principles and implementation methods of the present disclosure are described in detail using specific examples. The description of the above embodiments is only used to help understand the method and core idea of ​​the present disclosure. At the same time, for those skilled in the art, according to the idea of ​​the present disclosure, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present disclosure.

[0108] The above description is only a preferred embodiment of the present disclosure and is not intended to limit the present disclosure. Any modifications, equivalent substitutions, etc. made within the spirit and principles of the present disclosure should be included in the protection scope of the present disclosure.

Claims

1. A BEV perception method, characterized in that: The method comprises: Acquire multi-view images of the vehicle's surroundings; Processing the multi-view images using a DINOv2 model to obtain multi-view 2D features, wherein the DINOv2 model is pre-trained by self-supervised learning and a low-rank adaptive adapter; Mapping the multi-view 2D features to 3D space and fusing them to obtain 3D fused features, where the 3D fused features are complete feature representations of the vehicle's surroundings; The 3D fusion features are used to obtain a BEV feature map of the vehicle's surroundings.

2. The method according to claim 1, characterized in that: The method further includes: before processing the multi-view image using the DINOv2 model, performing dynamic image enhancement on the multi-view image.

3. The method according to claim 1, characterized in that: The method of using the DINOv2 model to process the multi-perspective images to obtain multi-perspective 2D features includes: performing the following processing on the image of the vehicle's surroundings at each perspective to obtain the 2D features at that perspective: dividing the image of the vehicle's surroundings at a first perspective into blocks of a predetermined size; processing the blocks through the DINOv2 model to obtain the 2D features of the first perspective, wherein the 2D features of the first perspective include semantic information of the vehicle's surroundings at the first perspective and related information of the 3D spatial structure around the vehicle, and the first perspective is any one of the multiple perspectives.

4. The method according to claim 1, characterized in that: The DINOv2 model includes multiple Transformer encoder layers, each of which includes a self-attention layer; during the training process of the DINOv2 model, the LoRA adapter is used to introduce a low-rank matrix in the self-attention layer to update parameters.

5. The method according to claim 1, characterized in that Mapping the multi-view 2D features to the 3D space and fusing them to obtain 3D fused features includes: mapping the multi-view 2D features to the 3D space to obtain multi-view 3D features; and fusing the multi-view 3D features to obtain the 3D fused features.

6. The method according to claim 5, characterized in that The multi-view 2D features are mapped to the 3D space to obtain the multi-view 3D features in the following manner: a depth estimation network is used to predict the depth information of each pixel in the image of the vehicle's surrounding environment at a first view; the 2D features of the first view are mapped to the 3D space using the depth information to obtain the 3D features of the first view; wherein the first view is any one of the multiple views.

7. The method according to claim 1, characterized in that The method of obtaining a BEV feature map of the vehicle surrounding environment by using the 3D fusion feature includes: Discretizing the 3D fusion features into voxels to obtain voxel features of the vehicle surroundings; The voxel features of the vehicle's surroundings are used to obtain a BEV feature map of the vehicle's surroundings.

8. A BEV sensing device, characterized in that: include: An acquisition unit, used for acquiring multi-view images of the environment around the vehicle; A 2D feature extraction unit, configured to process the multi-view images using a DINOv2 model to obtain multi-view 2D features, wherein the DINOv2 model is obtained by pre-training with a low-rank adaptive adapter in a self-supervised learning manner; A 3D space conversion and fusion unit, used for mapping the multi-view 2D features to a 3D space and fusing them to obtain a 3D fused feature, wherein the 3D fused feature is a complete feature representation of the vehicle's surrounding environment; The BEV feature extraction unit is used to obtain a BEV feature map of the vehicle's surrounding environment using the 3D fusion features.

9. An electronic device, characterized in that: include: One or more processors and a memory storing a program, wherein the program comprises instructions, and when the instructions are executed by the processor, the processor executes the method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a program, wherein the program comprises instructions, and when the instructions are executed by one or more processors of a computing device, the instructions cause the computing device to execute the method according to any one of claims 1 to 7.