Three-dimensional perception method and system under incomplete observation condition, medium and program product
Through a three-dimensional perception model composed of BEV feature encoder and timing enhancement module, the problem of incomplete observation of surround viewing cameras in dynamic traffic scenarios is solved, and the environmental perception ability and robustness of the autonomous driving system are improved.
Patent Information
- Application Number
- CN202510439682.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-25
AI Technical Summary
The existing three-dimensional perception method based on surround view cameras cannot effectively deal with data loss and error caused by incomplete observations in dynamic traffic scenarios or inclement weather conditions, affecting the perception ability of the autonomous driving system.
A three-dimensional perception model consisting of BEV feature encoder, feature decomposition submodule, dual-stage panoramic feature learning submodule, feature aggregation submodule and timing enhancement module is used to supplement the missing perspective features and improve environmental perception capabilities through feature decomposition, panoramic feature learning and timing feature fusion.
Under incomplete observation conditions, the spatial perception ability and robustness of the autonomous driving system are significantly improved, ensuring efficient and stable operation in a dynamic environment, and achieving complete driving environment state perception.
Smart Images

Figure CN120375307A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent driving environment perception, and particularly to a three-dimensional perception method, system, medium and program product under incomplete observation conditions. Background Art
[0002] Compared with 2D object detection, 3D perception can provide more dimensional information, such as object depth, three-dimensional bounding box size, and heading angle, etc. These additional degrees of freedom make the estimation of the driving state more complete. The environment perception based on 3D space can not only comprehensively capture the dynamic changes of the surrounding environment, but also provide a clearer information basis for the scenario prediction and reliable planning of the autonomous driving system, thus significantly improving the safety and reliability of the autonomous driving system. The purpose of 3D perception is to comprehensively understand the driving scenario by using the data obtained by a series of environment perception sensors (mainly including lidar, 4D millimeter-wave radar, and surround-view cameras) for subsequent planning and decision-making. In the past, the 3D perception field was mainly dominated by lidar-based models because lidar can provide accurate three-dimensional point cloud data. However, lidar systems are costly, vulnerable to bad weather, and inconvenient to deploy. In contrast, vision-based systems have the advantages of low cost, easy deployment, and good scalability. Therefore, in recent years, 3D perception based on surround-view cameras has attracted extensive attention from researchers.
[0003] In an intelligent driving system, as the core sensor for realizing global environment perception, the data integrity of the surround-view camera is directly related to the performance of the 3D perception module and driving safety. However, in practical applications, due to sensor failures or the influence of harsh environmental factors, the camera often produces incomplete observations, resulting in missing or inaccurate input data, thereby weakening the overall perception ability of the autonomous driving system. Although existing methods have improved the perception performance of the environmental state to a certain extent, they generally assume that the sensor data is complete and of high quality, and do not fully consider the frequently occurring incomplete observations and data missing problems in practical applications. Therefore, traditional end-to-end perception methods often struggle to cope with the challenges brought by incomplete data in dynamic traffic scenarios, especially in mixed traffic or adverse weather conditions. Summary of the Invention
[0004] To solve the technical problems in the background art, the present invention proposes a three-dimensional perception method, system, medium and program product under incomplete observation conditions.
[0005] A three-dimensional perception method under incomplete observation conditions proposed by the present invention includes:
[0006] Obtain surround-view camera images under incomplete observation conditions; wherein, the surround-view camera images include images from different perspectives;
[0007] Perform 3D perception on the surround-view camera image using a preset 3D perception model to obtain a 3D perception result.
[0008] Preferably, the 3D perception model includes a BEV feature encoder, a BEV feature construction module, a temporal enhancement module, and a 3D environment perception module connected in sequence; among them, the BEV feature construction module includes a feature decomposition sub-module, a two-stage panoramic feature learning sub-module, and a feature aggregation sub-module connected in sequence.
[0009] Preferably, the BEV feature encoder is used to obtain an initial BEV feature based on the surround-view camera image;
[0010] The feature decomposition sub-module is used to decompose the initial BEV feature to obtain low-dimensional local features;
[0011] The two-stage panoramic feature learning sub-module is used to input the low-dimensional local features into a pre-constructed learnable BEV feature space to obtain low-dimensional complete features that match the low-dimensional local features, and obtain low-dimensional panoramic features based on the low-dimensional complete features;
[0012] The feature aggregation sub-module is used to aggregate the low-dimensional panoramic features to obtain re-aggregated BEV features;
[0013] The temporal enhancement module is used to perform feature fusion on the re-aggregated BEV features and the historical re-aggregated BEV features in the time series to obtain fused BEV features;
[0014] The 3D environment perception module is used to perform 3D environment perception on the fused BEV features to obtain a 3D environment perception result.
[0015] Preferably, performing 3D perception on the surround-view camera image using a preset 3D perception model to obtain a 3D perception result specifically includes:
[0016] Use the BEV feature encoder to obtain an initial BEV feature based on the surround-view camera image;
[0017] Use the feature decomposition sub-module to decompose the initial BEV feature to obtain low-dimensional local features;
[0018] Use the two-stage panoramic feature learning sub-module to input the low-dimensional local features into a pre-constructed learnable BEV feature space to obtain low-dimensional complete features that match the low-dimensional local features, and obtain low-dimensional panoramic features based on the low-dimensional complete features;
[0019] Use the feature aggregation sub-module to aggregate the low-dimensional panoramic features to obtain re-aggregated BEV features;
[0020] The temporal enhancement module is used to perform feature fusion on the reunited BEV features and the historical reunited BEV features in the time series to obtain the fused BEV features;
[0021] The 3D environment perception module is used to perform 3D environment perception on the temporally fused BEV features to obtain the 3D environment perception results.
[0022] Preferably, the BEV feature encoder includes a two-dimensional feature extractor and a PV-to-BEV conversion module, and the two-dimensional feature extractor includes a backbone network and a neck sub-module;
[0023] Among them, the initial BEV features are obtained according to the surround-view camera images by using the BEV feature encoder, specifically including:
[0024] The backbone network is used to perform multi-scale feature extraction on the images of different perspectives respectively to obtain multi-scale semantic features of different perspectives; the neck sub-module is used to perform feature fusion on the multi-scale semantic features of different perspectives respectively to obtain perspective view features of different perspectives;
[0025] The PV-to-BEV conversion module is used to perform view conversion on the perspective view features of different perspectives to obtain the initial BEV features.
[0026] Preferably, the PV-to-BEV conversion module is used to perform view conversion on the perspective view features of different perspectives to obtain the initial BEV features, specifically including:
[0027] For each perspective, according to the intrinsic matrix of the camera, the pixels of the vertical scan lines in the perspective view features are mapped one-to-one to the BEV space to obtain the positions of the BEV rays;
[0028] A mapping relationship between query and key is established for each image column feature in the perspective view feature and the position of each BEV ray; among them, the mapping relationship between query and key includes a query vector and a key vector;
[0029] According to the mapping relationship between query and key, the alignment score between each image column feature and the position of the BEV ray is calculated; according to the alignment score, the normalized weight is calculated;
[0030] According to the normalized weight and the key vector, the image column feature is softly aligned with its corresponding BEV ray position to obtain the local context vector of each BEV ray;
[0031] The global aggregation of the local context vectors of each BEV ray is performed through the self-attention mechanism to obtain the context vector of each BEV ray;
[0032] According to the context vector of each BEV ray, context information is assigned to the pixel features of all vertical scan lines in the perspective view features to obtain BEV ray features;
[0033] The BEV ray features of each perspective are concatenated to obtain the BEV feature of each perspective;
[0034] According to the BEV features of each perspective, an initial BEV feature is obtained.
[0035] Preferably, the temporal enhancement module includes a spatio-temporal alignment sub-module and a feature fusion sub-module;
[0036] The spatio-temporal alignment sub-module is used to align the historical reaggregated BEV features with the current reaggregated BEV features by an interpolation method;
[0037] The feature fusion sub-module is used to perform feature fusion in time series on the aligned BEV features and the historical BEV features to obtain fused BEV features.
[0038] Preferably, before performing 3D perception on the surround-view camera image using a preset 3D perception model to obtain a 3D perception result, it further includes:
[0039] Construct a 3D perception model and a training set; wherein, the training set includes multiple sample data pairs, and each sample data pair includes a surround-view camera image under complete observation conditions and a surround-view camera image under incomplete observation conditions;
[0040] The training set is input into the BEV feature encoder to obtain complete BEV image features and incomplete BEV image features;
[0041] The complete BEV image features and the incomplete BEV image features are alternately input into the feature decomposition sub-module in sequence to obtain low-dimensional local features under complete observation and low-dimensional local features under incomplete observation;
[0042] Use the two-stage panoramic feature learning sub-module to construct a learnable BEV feature space, and use the BEV feature space for the low-dimensional local features under complete observation and the low-dimensional local features under incomplete observation to perform panoramic feature learning to obtain local features
[0043] Use the feature aggregation sub-module to aggregate the local features to obtain reaggregated BEV features;
[0044] Use the temporal enhancement module to align the BEV features and the historical BEV features, and perform feature fusion in time series on the aligned BEV features and the historical BEV features to obtain temporally fused BEV features;
[0045] Perform 3D environmental perception on the temporal fusion BEV features using a 3D environmental perception module to obtain the 3D environmental perception result;
[0046] Based on the 3D environmental perception result and the environmental camera images under the corresponding complete observation conditions in the training set, obtain the total loss function;
[0047] Use the total loss function to iteratively train the 3D perception model until the 3D perception model converges to obtain the trained 3D perception model.
[0048] Preferably, the total loss function is L = δ·L det +δ·L occ ; where L represents the total loss function, δ represents the control weight, and L det represents the detection loss function, and L occ represents the occupancy total loss function;
[0049] Among them,
[0050] In the formula, V min represents the initial value of the control weight, V max represents the maximum value of the control weight, N represents the number of iterations, and i = 1, 2, 3,... N;
[0051] Among them, L occ = λ1×L ce +λ2×L lovasz +λ3×L geo +λ4×L sem ;
[0052] In the formula, L occ represents the occupancy total loss function, λ1, λ2, λ3, and λ4 represent coefficients, and L ce represents the cross-entropy loss, L lovasz represents the Lovasz softmax loss, L geo represents the geometric affinity loss, and L sem represents the semantic affinity loss;
[0053] Among them, L det = λ5·L cls +λ6·L reg ;
[0054] In the formula, L det represents the detection loss function, λ5 and λ6 represent coefficients, and L cls represents the classification loss, and L reg represents the regression loss.
[0055] In a second aspect, the present invention further provides a three-dimensional perception system under incomplete observation conditions, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the three-dimensional perception method under incomplete observation conditions described in any one of the first aspects.
[0056] In a third aspect, the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the three-dimensional perception method under incomplete observation conditions described in any one of the first aspects.
[0057] In a fourth aspect, the present invention further provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the steps of the three-dimensional perception method under incomplete observation conditions described in any one of the first aspects.
[0058] In the present invention, the proposed three-dimensional perception method, system, medium, and program product under incomplete observation conditions can perform three-dimensional perception on the omnidirectional camera images under incomplete observation conditions by using a preset three-dimensional perception model to obtain three-dimensional perception results, solving the problems brought by incomplete data in dynamic traffic scenarios, especially in mixed traffic or adverse weather conditions. Description of the Drawings
[0059] Figure 1 It is a schematic flowchart of the three-dimensional perception method under incomplete observation conditions in an embodiment proposed by the present invention.
[0060] Figure 2 It is a schematic diagram of the three-dimensional perception model in an embodiment proposed by the present invention.
[0061] Figure 3 It is a schematic diagram of the two-stage panoramic feature learning sub-module in an embodiment proposed by the present invention. Detailed Embodiments
[0062] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the drawings and in combination with the embodiments.
[0063] In a first aspect, referring to Figures 1-3 , a three-dimensional perception method under incomplete observation conditions proposed by the present invention includes:
[0064] Obtain omnidirectional camera images under incomplete observation conditions; wherein, the omnidirectional camera images include images from different perspectives;
[0065] Perform three-dimensional perception on the omnidirectional camera images by using a preset three-dimensional perception model to obtain three-dimensional perception results.
[0066] The present invention can perform three-dimensional perception on the omnidirectional camera images under incomplete observation conditions by using a preset three-dimensional perception model, and obtain a three-dimensional perception result, which solves the problems caused by incomplete data in dynamic traffic scenarios, especially in mixed traffic or bad weather conditions.
[0067] In this embodiment, the three-dimensional perception model includes: a BEV feature encoder, a BEV feature construction module, a temporal enhancement module, and a 3D environment perception module connected in sequence;
[0068] Among them, the BEV feature construction module includes a feature decomposition sub-module, a two-stage panoramic feature learning sub-module, and a feature aggregation sub-module connected in sequence;
[0069] The BEV feature encoder is used to obtain an initial BEV feature according to the omnidirectional camera image;
[0070] The feature decomposition sub-module is used to decompose the initial BEV feature to obtain a low-dimensional local feature;
[0071] The two-stage panoramic feature learning sub-module is used to input the low-dimensional local feature into a pre-constructed learnable BEV feature space to obtain a low-dimensional complete feature that matches the low-dimensional local feature, and obtain a low-dimensional panoramic feature according to the low-dimensional complete feature;
[0072] The feature aggregation sub-module is used to aggregate the low-dimensional panoramic features to obtain a re-aggregated BEV feature;
[0073] The temporal enhancement module is used to perform feature fusion on the re-aggregated BEV feature and the historical re-aggregated BEV feature in the time series to obtain a fused BEV feature;
[0074] The 3D environment perception module is used to perform 3D environment perception on the fused BEV feature to obtain a 3D environment perception result.
[0075] The 3D perception model in this embodiment can, while efficiently processing semantic information of different granularities through the BEV feature encoder, express the generation of BEV features from the PV view as a set of 1D sequence-to-sequence translation problems in natural language processing, solve the problem of the lack of depth information of remote object images in the surround-view camera in effective observations, significantly improve the spatial perception ability in complex environments, and is applicable to efficient environmental perception for autonomous driving. Moreover, in this embodiment, the panoramic feature decomposition sub-module, the two-stage panoramic feature learning sub-module, and the feature aggregation sub-module in the BEV feature construction module are used to extract panoramic features from incomplete observations, thereby supplementing the missing perspective features; moreover, the high-dimensional panoramic features can be decomposed into low-dimensional local scene features, reducing the complexity of high-dimensional feature learning, and thus ensuring that the autonomous driving system can still operate efficiently and stably under incomplete observations, realizing reliable perception of the complete driving environment state. In addition, this embodiment fuses historical BEV features through the temporal enhancement module, and uses the temporal continuity to further compensate for the information loss caused by incomplete observations in the current frame. This fusion of spatio-temporal features makes the 3D perception model more robust in dynamic environments.
[0076] Therefore, in this embodiment, the surround-view camera images are 3D-perceived using the preset 3D perception model to obtain 3D perception results, which specifically include:
[0077] Using the BEV feature encoder to obtain initial BEV features based on the surround-view camera images;
[0078] Using the feature decomposition sub-module to perform feature decomposition on the initial BEV features to obtain low-dimensional local features;
[0079] Using the two-stage panoramic feature learning sub-module to obtain low-dimensional complete features matching the low-dimensional local features in the pre-constructed learnable BEV feature space, and obtaining low-dimensional panoramic features based on the low-dimensional complete features;
[0080] Using the feature aggregation sub-module to aggregate the low-dimensional panoramic features to obtain re-aggregated BEV features;
[0081] Using the temporal enhancement module to perform feature fusion on the re-aggregated BEV features and the historical re-aggregated BEV features in the time series to obtain fused BEV features;
[0082] Using the 3D environment perception module to perform 3D environment perception on the temporally fused BEV features to obtain 3D environment perception results.
[0083] In this embodiment, the BEV feature encoder includes a two-dimensional feature extractor and a PV-to-BEV conversion module. The two-dimensional feature extractor is used to extract features from the surround-view camera images to obtain perspective view features from different perspectives. The PV-to-BEV conversion module is used to perform view conversion on the perspective view features from different perspectives to obtain initial BEV features.
[0084] Therefore, in this embodiment, using the BEV feature encoder to obtain initial BEV features from the surround-view camera images specifically includes: using the two-dimensional feature extractor to extract perspective view features from different perspectives from the surround-view camera images; using the PV-to-BEV conversion module to convert the perspective view features from different perspectives into initial BEV features.
[0085] This embodiment can convert the surround-view camera images into rough initial BEV features through the two-dimensional feature extractor and the perspective view (PV) to bird's-eye view (BEV) conversion module in the BEV feature encoder. By means of the attention mechanism, a one-to-one mapping is established between the vertical scan lines of the image and the BEV rays, and the problem of missing depth information is cleverly transformed into a sequence-to-sequence translation problem. That is to say, this embodiment can simulate the frame missing situation in a continuous time series when randomly discarding a single camera frame at each timestamp.
[0086] Specifically, when implemented, the two-dimensional feature extractor takes the surround-view camera image Img = {I1, I2, …, I n} as input and outputs the PV features F pv = {F I1 , F I2 , …, F In} of the surround-view camera image.
[0087] Among them, the two-dimensional feature extractor includes a backbone network and a neck sub-module. The backbone network is used to extract multi-scale semantic features in the PV view, and the neck sub-module is used to perform feature fusion on the multi-scale semantic features to obtain perspective view features with different granularities.
[0088] Specifically, when implemented, in this embodiment, the surround-view camera image is input into the backbone network, and the backbone network performs multi-scale semantic feature extraction in the PV view to obtain multi-scale semantic features. These multi-scale semantic features are then input into a neck sub-module for fusion to obtain perspective view features with different granularities, so as to make full use of semantic information with different granularities.
[0089] In one specific embodiment, the backbone network is a lightweight ResNet or a powerful Swin Transformer.
[0090] In one specific embodiment, the neck sub-module is an FPN network to fuse fine-grained features with directly upsampled coarse-grained features.
[0091] It should be understood that close-range objects have dense pixels and rich depth information. However, remote objects have relatively fewer pixels and severely lack depth information. Therefore, it is not easy to construct a BEV from surround-view camera images, especially for remote objects. To achieve the PV view to BEV view conversion, this embodiment proposes a PV to BEV conversion module, which converts the perspective view feature Fpv of the surround-view camera image into a rough initial BEV feature F bev , and this PV to BEV conversion process is carried out through a single end-to-end network.
[0092] Assuming a one-to-one correspondence between the vertical scan lines in the PV view and the top view in the BEV view, the rays are passed through the camera position to the top view. This can formulate the problem of generating a BEV view from an image as a sequence-to-sequence translation problem in natural language processing, which is crucial for real-time environment perception in autonomous driving.
[0093] The PV to BEV conversion module proposed in this embodiment directly maps each pixel in the PV view to the corresponding position in the BEV, utilizes the camera's intrinsic matrix, and achieves the alignment between PV and BEV through an attention mechanism. This embodiment formulates the PV to BEV conversion process as a one-to-one mapping between each vertical scan line in the PV and the BEV ray. Let I be the input image, where H and W represent the height and width of the image respectively, C be the camera's intrinsic matrix, and the goal is to predict the semantic BEV map for each class m ∈ M (assuming there are M classes in total), denoted as BEV m ; where X and Z represent the feature space dimensions in the BEV.
[0094] Among them, the PV to BEV conversion process is expressed by the formula p(BEV m |I, C) = Φ(I, C); in the formula, Φ(·) is a neural network, and Φ(·) is trained to solve semantic and positional uncertainties and is responsible for converting the features of the image into its corresponding BEV view.
[0095] It should be noted that in this embodiment, an attention mechanism is proposed to learn the mapping relationship between each vertical scan line in the image and the BEV ray. Specifically, the attention mechanism determines the contribution degree of each image pixel to the target ray by calculating the similarity between the image column feature and the position of the BEV ray. First, a mapping relationship of query and key is introduced for each image column feature and the position of each BEV ray, and the query vector and key vector are defined respectively as:
[0096] Query(y i ) = W q y i ; Key(h j ) = W k h j ; where yi represents the position on the ray, h j represents the pixel feature of the image column, and W q and W k represent the weight matrices obtained through learning.
[0097] Then, calculate the alignment score s i,j between the image column feature and the BEV ray feature, and this score measures the matching degree between the pixels in the image and the BEV ray;
[0098] Among them, where D is the dimension of the feature vector.
[0099] Then, according to the alignment score s i,j , obtain the normalized weight α i,j ; where, This weight represents the contribution degree of the image pixel when generating the ray feature.
[0100] Then, the generation of each BEV ray needs to assign a context information according to the pixel features of all vertical scan lines in the perspective view feature. This context vector is calculated through the soft alignment relationship between the image column feature and the position of its corresponding BEV ray. Specifically, the formula for generating the context vector is:
[0101]
[0102] Although the above soft alignment process can generate local context information for each ray, in order to ensure the spatial consistency of the features in the ray, the present invention further globally aggregates the context vector l i through the self-attention mechanism to ensure that the features on the ray are consistent with the structure of the overall scene. The last step of this step is to perform operations on all context vectors l iPerform global operations using a non-linear function that reasons in the entire polar coordinate system to obtain BEV ray features. Repeat this operation on the multi-view surround camera images and perform simple stitching in the BEV feature space. Do not stitch the BEV features of incomplete observation views, and finally obtain the initialized BEV feature F bev 。
[0103] Therefore, in this embodiment, the PV to BEV conversion module is used to convert the PV feature map into BEV image features, which specifically includes:
[0104] For each view, according to the intrinsic matrix of the camera, map the pixels of the vertical scan lines in the perspective view features one-to-one into the BEV space to obtain the positions of the BEV rays;
[0105] Establish a mapping relationship between queries and keys for each image column feature in the perspective view feature and the position of each BEV ray; among them, the mapping relationship between queries and keys includes query vectors and key vectors;
[0106] According to the mapping relationship between queries and keys, calculate the alignment score between each image column feature and the position of the BEV ray; according to the alignment score, calculate the normalized weight;
[0107] According to the normalized weight and the key vector, perform soft alignment between the image column feature and the position of its corresponding BEV ray to obtain the local context vector of each BEV ray;
[0108] Perform global aggregation on the local context vectors of each BEV ray through the self-attention mechanism to obtain the context vector of each BEV ray;
[0109] According to the context vector of each BEV ray, assign a context information to the pixel features of all vertical scan lines in the perspective view feature to obtain the BEV ray feature;
[0110] Stitch the BEV ray features of each view to obtain the BEV feature of each view;
[0111] According to the BEV features of each view, obtain the BEV feature.
[0112] When the surround camera of the intelligent vehicle is damaged, the quality of the input information decreases, which in turn leads to a decrease in the performance of the autonomous driving perception algorithm. Therefore, optimizing the modeling of BEV features under incomplete observations affects the safety of intelligent driving.
[0113] The BEV feature construction module under incomplete observations in this embodiment includes a feature decomposition sub-module, a two-stage panoramic feature learning sub-module, and a feature aggregation sub-module;
[0114] Among them, the feature decomposition sub-module is used to decompose the rough bird's-eye view image BEV feature F extracted in Step 1 bev as the input, and decompose it into local features under complete observation and local features under incomplete observation
[0115] The two-stage panoramic feature learning sub-module is used to alternately receive low-dimensional local features from complete and incomplete observations during the training phase or and output the updated local features for constructing the reunited BEV feature to gradually optimize the BEV feature space;
[0116] Among them, the feature aggregation sub-module is used to aggregate local features back into the reunited BEV feature
[0117] To reduce the computational and memory burdens generated during the high-dimensional panoramic feature learning process, this embodiment designs a feature decomposition model, which is used to decompose high-dimensional features into low-dimensional features for efficient learning before complex matrix calculations. Specifically, the feature decomposition sub-module decomposes the high-dimensional panoramic feature into low-dimensional local scene features or That is: In the formula, F de () represents a dimensionality transformation function, and F bev is regarded as the unupdated panoramic feature. This process allows the limited panoramic feature space B′ to retain limited local scene features because the same semantic occupancy elements (such as trucks, pedestrians, etc.) exhibit similar semantic features in different scenes.
[0118] The two-stage panoramic feature learning sub-module proposes to learn the BEV feature through the complete observation mode to solve the problem of missing incomplete observation features caused by the damage of the surround-view camera of the intelligent vehicle. The network structure of this two-stage panoramic feature learning sub-module is as Figure 3 shown. Using the observation switching training strategy and the panoramic feature space, it proposes to extract features with the complete observation mode from incomplete observations, thus effectively solving the problem of missing panoramic features caused by incomplete observations.
[0119] During the two-stage panoramic feature learning process, by alternately switching the input of features under complete and incomplete observations, the two-stage panoramic feature training alternately uses complete features and incomplete features as inputs. The two training stages are respectively named the "update stage" and the "freeze stage". A parameterized BEV feature space is constructed, and its parameters are updated only in the update stage. Therefore, complete features can be stored in the BEV feature space, and incomplete features can be matched with the corresponding complete features in this space.
[0120] In the update stage, the two-stage panoramic feature learning sub-module receives complete features as input; initializes the BEV feature space as a matrix and performs matrix multiplication operations to match low-dimensional features with the BEV feature space B to obtain the Address feature; and performs a SoftMax transformation on the Address feature to obtain the address vector
[0121] where, in the formula, cos(·) represents the cosine similarity, and the SoftMax transformation is the normalization process for . Among them, the address vector V is used to match the corresponding local scene features in the panoramic feature space; the local features updated through the Address feature will then be merged to form the panoramic feature In this stage, the updated local features are used to construct the panoramic feature
[0122] In the freeze stage, the two-stage panoramic feature learning sub-module receives the input from incomplete features . Different from the update stage, at this time, the parameters of the panoramic feature space B are frozen, allowing the complete features to be retained in the panoramic feature space B. Other training steps are similar to the update stage. The overall model uses the complete features matched with the incomplete features to construct the high-definition map. These matched complete features are regarded as the input for constructing the low-dimensional panoramic feature F p .
[0123] Finally, the limited low-dimensional features are used to aggregate into diverse high-dimensional features. These high-dimensional features can be regarded as synthetic features extracted from images taken from multiple perspectives, namely panoramic features. The low-dimensional features can be regarded as features obtained by dividing the image into small blocks, and each small block forms a feature part of the scene, that is, a combination of local scene features. The decomposition process enables the BEV feature construction model under incomplete observations to learn low-dimensional features, thereby reducing the computational burden; while the aggregation process enables the limited local scene features to be combined into diverse panoramic features, thereby improving the representation ability of the panoramic feature space in the BEV feature construction model under incomplete observations.
[0124] Specifically, in the local aggregation process of the feature aggregation sub-module, the local features are re-aggregated into the re-aggregated BEV features That is: In the formula, F ag (·) is another dimensional transformation function, represents the local features. This local aggregation process enables the limited local scene features to represent more diverse panoramic features.
[0125] It should be understood that the purpose of using the temporal enhancement module in this embodiment is to enhance the perception ability of dynamic objects or attributes by integrating historical information, thereby improving the dynamic environment understanding and prediction ability of the autonomous driving system.
[0126] Among them, the temporal enhancement module in this embodiment includes a spatio-temporal alignment sub-module and a feature fusion sub-module. First, the spatio-temporal alignment sub-module uses the position information of the ego vehicle to align the historical BEV features with the current BEV features. Specifically, the spatio-temporal alignment sub-module ensures the spatio-temporal consistency of the historical features and the current perception data through interpolation methods, thereby achieving the precise synchronization of historical data. This alignment process is very crucial, as it ensures that the historical features can accurately reflect the perception state of the vehicle at different times, providing a reliable basis for subsequent feature fusion. After the alignment is completed, the aligned historical BEV features will be transmitted to the feature fusion sub-module. At this stage, the feature fusion sub-module combines the historical features and the current perception input, considering the information of the temporal context, to generate a comprehensive representation of dynamic objects or attributes. The feature fusion sub-module can provide more accurate dynamic object recognition and trajectory prediction by organically fusing historical information and real-time perception data, significantly improving the overall perception accuracy and reliability of the system.
[0127] Therefore, the temporal enhancement module in this embodiment can not only capture the information in static scenes, but also effectively cope with the changes in dynamic scenes, improving the decision-making ability and safety of the autonomous driving system in complex environments.
[0128] In this embodiment, before performing 3D perception on the surround-view camera images using a preset 3D perception model to obtain a 3D perception result, it further includes:
[0129] Construct a 3D perception model;
[0130] Construct a training set and a test set, and use the training set and the test set to train and test the 3D perception model.
[0131] It should be noted that the data in the training set and the test set constructed in this embodiment comes from the large outdoor public dataset nuScenes.
[0132] Among them, the training set includes multiple sample data pairs, and each sample data pair includes the surround-view camera image I under complete observation conditions c and the surround-view camera image I under incomplete observation conditions inc .
[0133] In order to obtain a 3D perception model that passes the training, in this embodiment, using the training set to train the 3D perception model specifically includes:
[0134] Input the training set into the BEV feature encoder to obtain the complete BEV image feature and the incomplete BEV image feature;
[0135] Use the feature decomposition sub-module to decompose the complete BEV image feature and the incomplete BEV image feature into local features under complete observation and local features under incomplete observation
[0136] Use the two-stage panoramic feature learning sub-module to construct a learnable BEV feature space, and use the BEV feature space to perform panoramic feature learning on the local features under complete observation and local features under incomplete observation to obtain local features
[0137] Use the feature aggregation sub-module to aggregate the local features back to the BEV feature
[0138] Use the temporal enhancement module to align the BEV feature and the historical BEV feature, and perform feature fusion on the aligned BEV feature and the historical BEV feature in time series to obtain the time series fusion BEV feature;
[0139] Use the 3D environment perception module to perform 3D environment perception on the time series fusion BEV feature to obtain a 3D environment perception result;
[0140] Train the parameters of the 3D perception model according to the D environment perception results until the 3D perception model converges.
[0141] Among them, the convergence condition is that the number of optimization times reaches the preset number.
[0142] During the training process, adjust the parameters of the 3D perception model based on a preset dynamic task weight adjustment strategy, and efficiently unify multiple 3D perception tasks in a single network framework to achieve end-to-end training of the 3D perception model.
[0143] In the training stage, the model input proposed in this embodiment consists of paired inputs of complete observation I c and incomplete observation I inc , and these inputs are presented in the form of panoramic camera perspective images. It shares the parameters between the two training stages to prevent the model from becoming too large. In the testing and inference stages, the 3D perception model only receives incomplete observation I inc as input.
[0144] This embodiment adds a specific task head for different 3D perception tasks. For the occupancy task, the present invention uses two fully connected layers to map the feature channels to the number of occupancy categories. This embodiment uses the total occupancy loss function for end-to-end training. Among them, the total occupancy loss function includes cross-entropy loss L ce , Lovasz softmax loss L lovasz , geometric affinity loss L geo and semantic affinity loss L sem .
[0145] Among them, L occ =λ1×L ce +λ2×L lovasz +λ3×L geo +λ4×L sem ; in the formula, L occ represents the total occupancy loss function, and λ1, λ2, λ3, and λ4 represent coefficients.
[0146] For the detection task, the present invention uses a center heat map head of a specific category to predict the center positions of all objects, and estimates the size, rotation, and speed of the objects through a regression head. The classification loss and regression loss are respectively expressed as L cls , L reg . Among them, L det =λ5·L cls +λ6·L reg ; in the formula, L det represents the detection loss function, and λ5 and λ6 represent coefficients.
[0147] However, unifying multiple 3D perception tasks within a single network framework often leads to a decline in the performance of individual tasks as the network makes trade-offs between different objectives, and may even cause the network to fail to converge, which poses challenges to the end-to-end training of 3D perception models.
[0148] In the early stage of training, the 3D perception model has not yet learned reasonable features. Therefore, the feature representation of voxels is random in space and does not have practical physical or semantic meanings. Since the voxel features have not been well learned, the occupancy prediction and detection prediction of the 3D perception model are still very inaccurate, resulting in smaller weights and less influence of the loss terms for these tasks throughout the training process. Additionally, due to the inaccurate prediction of categories by the model in the initial training stage, the classification loss value is relatively large. It dominates the overall loss, masking the contributions of other loss terms (such as the occupancy task), making the optimization more difficult. Although there are corresponding loss functions for the occupancy head and detection head, their influence is not significant in the early stage of training, mainly because the model has not learned meaningful features, so the optimization effect of these loss terms is poor. In the setting of multi-task learning, the larger loss term will dominate the optimization direction of the model and affect the learning of other tasks.
[0149] To this end, this embodiment proposes a progressive loss weight adjustment strategy to dynamically adjust the loss weights. Specifically, a control parameter δ is added to the non-image-level losses, namely the occupancy loss and the detection loss, in order to adjust the loss weights at different training stages. The control weight δ is set to the initial value V min , and gradually increases to the maximum value V max within N training epochs. Therefore, the total loss function in this embodiment is L = δ·L det + δ·L occ ;
[0150]
[0151] where L represents the total loss function, V min represents the initial value of the control weight, V max represents the maximum value of the control weight, N represents the number of iterations, and i = 1, 2, 3, … N.
[0152] In this embodiment, by introducing a control weight into the total loss function, non-image-level losses such as occupancy prediction and detection prediction are dynamically adjusted, thereby alleviating the problem of weight imbalance among loss terms caused by unstable features in the initial stage of training. Since the model features are random in the initial stage and the losses of some tasks account for a relatively large proportion, in this embodiment, the control parameter is initially set to a small value in the overall loss and gradually increased to a preset threshold within a predetermined training period, so that the larger loss terms will not dominate the overall optimization direction, thereby achieving balanced and collaborative optimization of the losses of each task and meeting the comprehensive requirements of accuracy, robustness, and real-time performance for 3D perception tasks in complex traffic scenarios.
[0153] In a second aspect, the present invention also provides a three-dimensional perception system under incomplete observation conditions, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the three-dimensional perception method under incomplete observation conditions described in any one of the first aspects.
[0154] In a third aspect, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the three-dimensional perception method under incomplete observation conditions described in any one of the first aspects.
[0155] In a fourth aspect, the present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the steps of the three-dimensional perception method under incomplete observation conditions described in any one of the first aspects.
[0156] As mentioned above, the above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and all should be covered within the protection scope of the present invention.
Claims
1. A three-dimensional perception method under incomplete observation conditions, characterized in that Including: Obtain panoramic camera images under incomplete observation conditions; wherein, the panoramic camera images include images from different perspectives; Perform 3D perception on the panoramic camera images using a preset 3D perception model to obtain a 3D perception result.
2. The three-dimensional perception method under incomplete observation conditions according to claim 1, characterized in that The 3D perception model includes a BEV feature encoder, a BEV feature construction module, a temporal enhancement module, and a 3D environment perception module connected in sequence; wherein, the BEV feature construction module includes a feature decomposition sub-module, a two-stage panoramic feature learning sub-module, and a feature aggregation sub-module connected in sequence; Preferably, the BEV feature encoder is used to obtain an initial BEV feature based on the panoramic camera images; The feature decomposition sub-module is used to decompose the initial BEV feature to obtain low-dimensional local features; The two-stage panoramic feature learning sub-module is used to input the low-dimensional local features into a pre-constructed learnable BEV feature space to obtain low-dimensional complete features matching the low-dimensional local features, and obtain low-dimensional panoramic features based on the low-dimensional complete features; The feature aggregation sub-module is used to aggregate the low-dimensional panoramic features to obtain re-aggregated BEV features; The temporal enhancement module is used to perform feature fusion on the re-aggregated BEV features and historical re-aggregated BEV features in the time series to obtain fused BEV features; The 3D environment perception module is used to perform 3D environment perception on the fused BEV features to obtain a 3D environment perception result.
3. The three-dimensional perception method under incomplete observation conditions according to claim 2, wherein Performing 3D perception on the panoramic camera images using a preset 3D perception model to obtain a 3D perception result specifically includes: Using the BEV feature encoder to obtain an initial BEV feature based on the panoramic camera images; Using the feature decomposition sub-module to decompose the initial BEV feature to obtain low-dimensional local features; Using the two-stage panoramic feature learning sub-module to input the low-dimensional local features into a pre-constructed learnable BEV feature space to obtain low-dimensional complete features matching the low-dimensional local features, and obtain low-dimensional panoramic features based on the low-dimensional complete features; Using the feature aggregation sub-module to aggregate the low-dimensional panoramic features to obtain re-aggregated BEV features; Using the temporal enhancement module to perform feature fusion on the re-aggregated BEV features and historical re-aggregated BEV features in the time series to obtain fused BEV features; Using the 3D environment perception module to perform 3D environment perception on the temporally fused BEV features to obtain a 3D environment perception result.
4. The three-dimensional perception method under incomplete observation conditions according to claim 3, characterized in that The BEV feature encoder includes a two-dimensional feature extractor and a PV to BEV conversion module, and the two-dimensional feature extractor includes a backbone network and a neck sub-module; Among them, using the BEV feature encoder to obtain an initial BEV feature based on the panoramic camera images specifically includes: Using the backbone network to perform multi-scale feature extraction on the images from different perspectives to obtain multi-scale semantic features from different perspectives; Using the neck sub-module to perform feature fusion on the multi-scale semantic features from different perspectives to obtain perspective view features from different perspectives; Using the PV to BEV conversion module to perform view conversion on the perspective view features from different perspectives to obtain an initial BEV feature; Preferably, the perspective view features from different perspectives are transformed by the PV to BEV conversion module to obtain the initial BEV features, specifically including: For each perspective, according to the internal parameter matrix of the camera, the pixels of the vertical scan lines in the perspective view features are mapped one-to-one into the BEV space to obtain the positions of the BEV rays; A mapping relationship between queries and keys is established for each image column feature in the perspective view features and the positions of each BEV ray; wherein, the mapping relationship between queries and keys includes query vectors and key vectors; According to the mapping relationship between queries and keys, the alignment scores between each image column feature and the positions of the BEV rays are calculated; According to the alignment scores, the normalized weights are calculated; According to the normalized weights and the key vectors, the image column features are softly aligned with the positions of their corresponding BEV rays to obtain the local context vectors of each BEV ray; The local context vectors of each BEV ray are globally aggregated through the self-attention mechanism to obtain the context vectors of each BEV ray; According to the context vectors of each BEV ray, a context information is assigned to the pixel features of all vertical scan lines in the perspective view features to obtain the BEV ray features; The BEV ray features of each perspective are concatenated to obtain the BEV features of each perspective; According to the BEV features of each perspective, the initial BEV features are obtained.
5. The three-dimensional perception method under incomplete observation conditions according to claim 2, wherein The temporal enhancement module includes a spatio-temporal alignment sub-module and a feature fusion sub-module; The spatio-temporal alignment sub-module is used to align the historical reaggregated BEV features with the current reaggregated BEV features by an interpolation method; The feature fusion sub-module is used to fuse the aligned BEV features and the historical BEV features in time series to obtain the fused BEV features.
6. The three-dimensional perception method under incomplete observation conditions according to claim 2, wherein Before performing 3D perception on the panoramic camera images using the preset 3D perception model to obtain the 3D perception results, it further includes: Constructing a 3D perception model and a training set; wherein, the training set includes multiple sample data pairs, and each sample data pair includes a panoramic camera image under complete observation conditions and a panoramic camera image under incomplete observation conditions; Inputting the training set into the BEV feature encoder to obtain the complete BEV image features and the incomplete BEV image features; Inputting the complete BEV image features and the incomplete BEV image features into the feature decomposition sub-module alternately in sequence to obtain the low-dimensional local features under complete observation and the low-dimensional local features under incomplete observation; Construct a learnable BEV feature space using the two-stage panoramic feature learning sub-module, and use the BEV feature space to perform panoramic feature learning on the low-dimensional local features under complete observations and the low-dimensional local features under incomplete observations to obtain local features Use the feature aggregation sub-module to aggregate the local features to obtain the re-aggregated BEV features; Using the temporal enhancement module to align the BEV features and the historical BEV features, and fusing the aligned BEV features and the historical BEV features in time series to obtain the temporally fused BEV features; Using the 3D environment perception module to perform 3D environment perception on the temporally fused BEV features to obtain the 3D environment perception results; According to the 3D environment perception results and the environmental camera images under the corresponding complete observation conditions in the training set, the total loss function is obtained; Using the total loss function to iteratively train the 3D perception model until the 3D perception model converges to obtain the trained 3D perception model.
7. The three-dimensional perception method under incomplete observation conditions according to claim 6, characterized in that The total loss function is \(L = \delta\cdot L\) det +\delta\cdot L occ ; where \(L\) represents the total loss function, \(\delta\) represents the control weight, \(L\) det represents the detection loss function, \(L\) occ represents the occupancy total loss function; Among them, In the formula, V min represents the initial value of the control weight, V max represents the maximum value of the control weight, N represents the number of iterations, and i = 1, 2, 3, … N; Among them, L occ = λ1 × L ce + λ2 × L lovasz + λ3 × L geo + λ4 × L sem ; where L occ represents the occupancy total loss function, λ1, λ2, λ3, and λ4 represent coefficients, and L ce represents the cross-entropy loss, L lovasz represents the Lovasz softmax loss, L geo represents the geometric affinity loss, L sem represents the semantic affinity loss; where L det = λ5·L cls + λ6·L reg ; where L det represents the detection loss function, λ5 and λ6 represent coefficients, and L cls represents the classification loss, and L reg represents the regression loss.
8. A three-dimensional perception system under incomplete observation conditions, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the steps of the three-dimensional perception method under incomplete observation conditions according to any one of claims 1-7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the three-dimensional perception method under incomplete observation conditions according to any one of claims 1-7.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the three-dimensional perception method under incomplete observation conditions according to any one of claims 1-7.