Depth Image Estimation Method, Apparatus, Electronic Device, and Storage Medium

By using a feature interaction layer based on convolutional neural network and attention mechanism to process multiple circumferential images on unmanned vehicles, the problems of high cost of binocular depth estimation and poor visual effect are solved, high-precision depth estimation is achieved and camera number requirements are reduced.

CN115100265BActive Publication Date: 2025-07-29BEIJING PHIGENT TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210590600.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-27
Publication Date
2025-07-29
Estimated Expiration
2042-05-27

AI Technical Summary

Technical Problem

In the prior art, the binocular depth estimation method has a high cost and affects the aesthetics of the vehicle, while the monocular visual depth estimation method has poor effect.

Method used

A feature extraction layer based on a convolutional neural network and a feature interaction layer based on an attention mechanism are used to feature interactions between the overlapping areas of multiple circumferential images in a high-dimensional space to generate a depth image.

Benefits of technology

Improves the accuracy of depth estimation, reduces the cost of achieving surround viewing, and does not affect the aesthetics of the vehicle.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115100265B_ABST
    Figure CN115100265B_ABST
Patent Text Reader

Abstract

The present application provides a depth image estimation method, apparatus, electronic device, and storage medium. The method includes: inputting multiple panoramic images of a target vehicle at a target moment into a depth estimation model; the depth estimation model includes a feature extraction layer based on a convolutional neural network and a feature interaction layer based on an attention mechanism, and each of the multiple panoramic images includes an overlapping region with an adjacent panoramic image; calling the feature extraction layer to process the image features of the multiple panoramic images to obtain target feature maps corresponding to the multiple panoramic images; calling the feature interaction layer to perform feature interaction processing on the target feature maps and two adjacent target feature maps corresponding to the target feature maps to obtain a depth image corresponding to the multiple panoramic images. The present application can improve the accuracy of depth estimation and achieve a relatively low cost for the panoramic view effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image processing, and in particular, to a depth image estimation method, apparatus, electronic device, and storage medium. Background Art

[0002] In the field of autonomous driving, how to obtain the depth information of targets such as vehicle pedestrians is an important technical point in many current researches, such as 3D reconstruction, obstacle detection, SLAM (Simultaneous Localization and Mapping), etc.

[0003] Currently, the methods for obtaining depth information usually adopt monocular vision depth estimation methods and binocular depth estimation methods. Among them, when performing depth estimation by the binocular depth estimation method, there is usually at least an 80% overlapping area between the left-eye camera and the right-eye camera. If multiple binocular cameras are set on the vehicle to achieve a panoramic view effect, the cost is high and the aesthetics of the vehicle is affected at the same time. The monocular vision depth estimation method is a method for estimating depth from a single RGB image, and the depth information obtained by estimating depth based on a single RGB image alone has poor effects. Summary of the Invention

[0004] Embodiments of the present application provide a depth image estimation method, apparatus, electronic device, and storage medium to solve the problems in the related art that the binocular depth estimation method has a high cost and affects the aesthetics of the vehicle, while the monocular vision depth estimation method has poor effects.

[0005] To solve the above technical problems, the embodiments of the present application are implemented as follows:

[0006] In a first aspect, an embodiment of the present application provides a depth image estimation method, including:

[0007] Inputting multiple panoramic images of a target vehicle at a target moment into a depth estimation model; the depth estimation model includes a feature extraction layer based on a convolutional neural network and a feature interaction layer based on an attention mechanism, and each of the multiple panoramic images includes an overlapping area with an adjacent panoramic image;

[0008] Invoking the feature extraction layer to process the image features of the multiple panoramic images to obtain target feature maps corresponding to the multiple panoramic images;

[0009] Invoking the feature interaction layer to perform feature interaction processing on the target feature maps and two adjacent target feature maps corresponding to the target feature maps to obtain a depth image corresponding to the multiple panoramic images.

[0010] Optionally, the step of calling the feature extraction layer to process the image features of the multi-frame panoramic images to obtain the target feature maps corresponding to the multi-frame panoramic images includes:

[0011] Calling the feature extraction layer to extract the image features of the first dimension of the multi-frame panoramic images, and performing dimension conversion processing on the image features to output a target feature map including the image features of the second dimension; the voice information included in the image features of the second dimension is more than that of the first dimension.

[0012] Optionally, the feature interaction layer includes: a self-attention mechanism layer, a cross-attention mechanism layer, and a third normalization layer,

[0013] The step of calling the feature interaction layer to perform feature interaction processing on the target feature map and two adjacent target feature maps corresponding to the target feature map to obtain the depth image corresponding to the multi-frame panoramic images includes:

[0014] Calling the self-attention mechanism layer to process three fully-connected feature maps corresponding to the target feature map to obtain an attention mechanism feature map corresponding to the target feature map;

[0015] Calling the cross-attention mechanism layer to perform feature fusion processing on two fully-connected feature maps corresponding to the attention mechanism feature map and one fully-connected feature map corresponding to two attention mechanism feature maps adjacent to the attention mechanism feature map to obtain a fusion feature map corresponding to the target feature map;

[0016] Calling the third normalization layer to perform normalization processing on the image features in the fusion feature map to obtain the depth image of the panoramic image corresponding to the target feature map.

[0017] Optionally, the feature interaction layer further includes: a first normalization layer and three first fully-connected layers, wherein the first normalization layer is respectively connected to the input ends of the three first fully-connected layers, and the output ends of the three first fully-connected layers are respectively connected to the self-attention mechanism layer,

[0018] Before calling the self-attention mechanism layer to process three fully-connected feature maps corresponding to the target feature map to obtain an attention mechanism feature map corresponding to the target feature map, it further includes:

[0019] Calling the first normalization layer to perform normalization processing on the image features of the target feature map to obtain a first normalized feature map corresponding to the target feature map;

[0020] Calling the three first fully-connected layers to process the first normalized feature map according to the corresponding matrix parameters respectively to obtain three fully-connected feature maps corresponding to the target feature map.

[0021] Optionally, the feature interaction layer further includes: a second normalization layer and three second fully-connected layers. The second normalization layer is respectively connected to the input ends of the three second fully-connected layers, and the output ends of the three second fully-connected layers are connected to the cross-attention mechanism layer.

[0022] Before calling the cross-attention mechanism layer to perform feature fusion processing on the two fully-connected feature maps corresponding to the attention mechanism feature map and one fully-connected feature map corresponding to the two attention mechanism feature maps adjacent to the attention mechanism feature map to obtain the fusion feature map corresponding to the target feature map, it further includes:

[0023] Call the second normalization layer to perform feature fusion and normalization processing on the attention mechanism feature map and the first normalized feature map to obtain a second normalized feature map;

[0024] Call the three second fully-connected layers to respectively process the second normalized feature map according to the corresponding matrix parameters to obtain three fully-connected feature maps corresponding to the attention mechanism feature map.

[0025] Optionally, the step of calling the third normalization layer to perform normalization processing on the image features in the fusion feature map to obtain the depth image of the panoramic view image corresponding to the target feature map includes:

[0026] Call the third normalization layer to perform feature fusion processing on the fusion feature map and the second normalized feature map to generate a target fusion feature map;

[0027] Call the third normalization layer to perform normalization processing on the image features of the target fusion feature map to obtain the depth image of the panoramic view image corresponding to the target feature map.

[0028] In a second aspect, an embodiment of the present application provides a depth image estimation device, including:

[0029] A panoramic view image input module, configured to input multiple panoramic view images of a target vehicle at a target moment collected into a depth estimation model; the depth estimation model includes: a feature extraction layer based on a convolutional neural network and a feature interaction layer based on an attention mechanism, and each of the multiple panoramic view images includes an overlapping area with an adjacent panoramic view image;

[0030] A target feature map acquisition module, configured to call the feature extraction layer to process the image features of the multiple panoramic view images to obtain target feature maps corresponding to the multiple panoramic view images;

[0031] A depth image acquisition module, configured to call the feature interaction layer to perform feature interaction processing on the target feature map and two adjacent target feature maps corresponding to the target feature map, so as to obtain depth images corresponding to the multiple panoramic images.

[0032] Optionally, the target feature map acquisition module includes:

[0033] A target feature map output unit, configured to call the feature extraction layer to extract image features of a first dimension of the multiple panoramic images, and perform dimension conversion processing on the image features, and output a target feature map including image features of a second dimension; the semantic information included in the image features of the second dimension is more than that of the first dimension.

[0034] Optionally, the feature interaction layer includes: a self-attention mechanism layer, a cross-attention mechanism layer, and a third normalization layer,

[0035] The depth image acquisition module includes:

[0036] An attention mechanism feature map acquisition unit, configured to call the self-attention mechanism layer to process three fully-connected feature maps corresponding to the target feature map, so as to obtain an attention mechanism feature map corresponding to the target feature map;

[0037] A fused feature map acquisition unit, configured to call the cross-attention mechanism layer to perform feature fusion processing on two fully-connected feature maps corresponding to the attention mechanism feature map and one fully-connected feature map corresponding to two attention mechanism feature maps adjacent to the attention mechanism feature map, so as to obtain a fused feature map corresponding to the target feature map;

[0038] A depth image acquisition unit, configured to call the third normalization layer to perform normalization processing on the image features in the fused feature map, so as to obtain a depth image of the panoramic image corresponding to the target feature map.

[0039] Optionally, the feature interaction layer further includes: a first normalization layer and three first fully-connected layers, wherein the first normalization layer is respectively connected to the input ends of the three first fully-connected layers, and the output ends of the three first fully-connected layers are respectively connected to the self-attention mechanism layer,

[0040] The apparatus further includes:

[0041] A first image acquisition module, configured to call the first normalization layer to perform normalization processing on the image features of the target feature map, so as to obtain a first normalized feature map corresponding to the target feature map;

[0042] The first fully-connected feature map acquisition module is configured to call the three first fully-connected layers to process the first normalized feature map according to corresponding matrix parameters respectively, so as to obtain three fully-connected feature maps corresponding to the target feature map.

[0043] Optionally, the feature interaction layer further includes: a second normalization layer and three second fully-connected layers. The second normalization layer is respectively connected to the input ends of the three second fully-connected layers, and the output ends of the three second fully-connected layers are connected to the cross-attention mechanism layer.

[0044] The apparatus further includes:

[0045] The second image acquisition module is configured to call the second normalization layer to perform feature fusion and normalization processing on the attention mechanism feature map and the first normalized feature map, so as to obtain a second normalized feature map.

[0046] The second fully-connected feature map acquisition module is configured to call the three second fully-connected layers to process the second normalized feature map according to corresponding matrix parameters respectively, so as to obtain three fully-connected feature maps corresponding to the attention mechanism feature map.

[0047] Optionally, the depth image acquisition unit includes:

[0048] The target fusion image generation sub-unit is configured to call the third normalization layer to perform feature fusion processing on the fusion feature map and the second normalized feature map, so as to generate a target fusion feature map.

[0049] The depth image acquisition sub-unit is configured to call the third normalization layer to perform normalization processing on the image features of the target fusion feature map, so as to obtain the depth image of the panoramic image corresponding to the target feature map.

[0050] In a third aspect, an embodiment of the present application provides an electronic device, including:

[0051] A memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, the depth image estimation method described in any one of the above is implemented.

[0052] In a fourth aspect, an embodiment of the present application provides a readable storage medium. When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device can execute the depth image estimation method described in any one of the above.

[0053] In an embodiment of the present application, multiple panoramic images of a target vehicle at a target moment are input into a depth estimation model. The depth estimation model includes a feature extraction layer based on a convolutional neural network and a feature interaction layer based on an attention mechanism. Each of the multiple panoramic images contains an overlapping area with an adjacent panoramic image. The feature extraction layer is called to process the image features of the multiple panoramic images to obtain target feature maps corresponding to the multiple panoramic images. The feature interaction layer is called to perform feature interaction processing on the target feature maps and two adjacent target feature maps corresponding to the target feature maps to obtain depth images corresponding to the multiple panoramic images. In the embodiment of the present application, the feature interaction layer can interact the features of the overlapping areas of two adjacent panoramic images of the vehicle in a high-dimensional space, which can improve the accuracy of depth estimation and effectively solve the problem that the depth information obtained by the monocular vision depth estimation method has a poor effect. At the same time, compared with the binocular depth estimation method, two adjacent panoramic images do not require a large overlapping area (only 5% - 10% is required), so the number of cameras arranged on the vehicle can be reduced, the cost of realizing the panoramic effect can be reduced, and the aesthetic degree of the vehicle is not affected.

[0054] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the following specific embodiments of the present application are specifically given. Brief Description of the Drawings

[0055] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required to be used in the description of the embodiments of the present application will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0056] Figure 1 It is a flowchart of the steps of a depth image estimation method provided by an embodiment of the present application;

[0057] Figure 2 It is a flowchart of the steps of another depth image estimation method provided by an embodiment of the present application;

[0058] Figure 3 It is a flowchart of the steps of a method for obtaining a fully connected feature map provided by an embodiment of the present application;

[0059] Figure 4 It is a flowchart of the steps of another method for obtaining a fully connected feature map provided by an embodiment of the present application;

[0060] Figure 5The flowchart of another depth image estimation method provided by the embodiment of the present application;

[0061] Figure 6 The structural schematic diagram of a depth estimation model provided by the embodiment of the present application;

[0062] Figure 7 The structural schematic diagram of a feature interaction layer provided by the embodiment of the present application;

[0063] Figure 8 The structural schematic diagram of a depth image estimation device provided by the embodiment of the present application;

[0064] Figure 9 The structural schematic diagram of an electronic device provided by the embodiment of the present application. Detailed implementation manners

[0065] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0066] Refer to Figure 1 , which shows the flowchart of a depth image estimation method provided by the embodiment of the present application. As Figure 1 shown, the depth image estimation method may include the following steps:

[0067] Step 101: Input multiple panoramic images of the target vehicle at the target moment collected into the depth estimation model; the depth estimation model includes: a feature extraction layer based on a convolutional neural network and a feature interaction layer based on an attention mechanism.

[0068] The embodiment of the present application can be applied to a scenario where the panoramic images with overlapping regions in multiple panoramic images of the target vehicle collected are interacted in a high-dimensional space to improve the depth estimation accuracy.

[0069] The target vehicle refers to an autonomous vehicle. In this example, the target vehicle may be an unmanned delivery vehicle, an unmanned operation vehicle, etc. Specifically, the specific type of the target vehicle can be determined according to the actual situation, and this embodiment does not limit it.

[0070] The target moment refers to the moment when the driving scenario of the target vehicle is estimated for depth.

[0071] The surround view image refers to an image of the surrounding scene of the target vehicle at the target moment collected by a camera pre - installed on the target vehicle. Each of the multiple surround view images contains an overlapping area with an adjacent surround view image.

[0072] The depth estimation model refers to a model used to predict the depth information of an image. In this example, the depth estimation model may include: a feature extraction layer based on a convolutional neural network and a feature interaction layer based on an attention mechanism. As Figure 7 shown, the depth estimation model contains a CNN (i.e., the feature extraction layer) and a Transformer (i.e., the feature interaction layer).

[0073] When estimating the depth information of the surrounding scene of the target vehicle, multiple surround view images of the target vehicle at the target moment can be collected by a camera pre - installed on the target vehicle. For example, N (N is a positive integer, such as 6, etc.) cameras can be pre - installed around the body of the target vehicle to collect images covering a 360 - degree view around the body of the target vehicle, that is, multiple surround view images corresponding to the target vehicle are obtained.

[0074] After collecting multiple surround view images of the target vehicle at the target moment, the multiple surround view images can be input into the depth estimation model. As Figure 7 shown, the multiple surround view images of the target vehicle collected are image1, image2,..., image6, and these 6 images can be input into the depth estimation model.

[0075] After inputting the multiple surround view images of the target vehicle collected at the target moment into the depth estimation model, step 102 is executed.

[0076] Step 102: Invoke the feature extraction layer to process the image features of the multiple surround view images to obtain target feature maps corresponding to the multiple surround view images.

[0077] The target feature map refers to an image obtained after the feature extraction layer processes the image features of multiple surround view images. The dimension of the target feature map is larger than that of the surround view image.

[0078] After inputting multiple surround view images into the depth estimation model, the feature extraction layer can be invoked to process the image features of the multiple surround view images to obtain target feature maps corresponding to the multiple surround view images respectively. As Figure 7As shown, after inputting six panoramic images into the depth estimation model, six parallel CNNS can be called to process the six panoramic images respectively to obtain corresponding target feature maps. For example, image1 can obtain the target feature map feature1 through the CNN, image2 can obtain the target feature map feature2 through the CNN, ..., image6 can obtain the target feature map feature6 through the CNN, and so on.

[0079] The process of obtaining the high-dimensional feature maps corresponding to the panoramic images can be as follows: After inputting multiple panoramic images into the depth estimation model, the feature extraction layer of the feature extraction layer can be called to perform feature extraction processing on the panoramic images to extract the first-dimensional image features of the multiple panoramic images.

[0080] After calling the feature extraction layer to extract the first-dimensional image features of the multiple panoramic images, dimensionality conversion processing can be performed on the first-dimensional image features to obtain second-dimensional image features and output the target feature maps containing the second-dimensional image features. Among them, the semantic information contained in the second-dimensional image features is more than that in the first dimension, thus realizing the conversion of high-dimensional features.

[0081] After calling the feature extraction layer to process the image features of the multiple panoramic images to obtain the target feature maps corresponding to the multiple panoramic images, step 103 is executed.

[0082] Step 103: Call the feature interaction layer to perform feature interaction processing on the target feature map and two adjacent target feature maps corresponding to the target feature map to obtain the depth images corresponding to the multiple panoramic images.

[0083] After obtaining the target feature maps corresponding to the multiple panoramic images respectively, the feature interaction layer can be called to perform feature interaction processing on the target feature map and two adjacent target feature maps corresponding to the target feature map to obtain the depth images corresponding to the multiple panoramic images respectively. For example, Figure 6 As shown, after obtaining the target feature maps (i.e., feature1, feature2, ..., feature6) corresponding to the multiple panoramic images (i.e., image1, image2, ..., image6) respectively, the Transformer can be called to perform feature interaction processing on the target feature maps feature1, feature2, ..., feature6 and two adjacent target feature maps corresponding to each target feature map, so as to obtain the depth images corresponding to the multiple panoramic images. For example, Figure 6As shown, after passing through the Transformer, the depth images depth1 of image1, depth2 of image2, ..., depth6 of image6, etc. can be obtained.

[0084] In the above solution of this embodiment, by interacting the features of the overlapping regions of two adjacent panoramic images of the vehicle in a high-dimensional space, the accuracy of depth estimation can be improved, effectively solving the problem of poor depth information obtained by the monocular vision depth estimation method. At the same time, compared with the binocular depth estimation method, a large overlapping area is not required (only 5% - 10% is needed), so the number of cameras arranged on the vehicle can be reduced, and the cost of realizing the panoramic view effect can be reduced.

[0085] Next, the process of feature interaction processing can be described in detail with reference to the accompanying drawings.

[0086] Refer to Figure 2 , which shows the flowchart of steps of another depth image estimation method provided by an embodiment of the present application. As Figure 2 shown, the depth image estimation method may include: step 201, step 202, and step 203.

[0087] Step 201: Invoke the self-attention mechanism layer to process three fully connected feature maps corresponding to the target feature map, and obtain the attention mechanism feature map corresponding to the target feature map.

[0088] In an embodiment of the present application, the depth information prediction network layer may include: a self-attention mechanism layer, a cross-attention mechanism layer, and a third normalization layer. As Figure 7 shown, the depth information prediction network layer Transformer may include: Self-Attention (i.e., the self-attention mechanism layer), Cross-Attention (i.e., the cross-attention mechanism layer), and LayerNorm located in Cross-Attention (i.e., the third normalization layer).

[0089] After obtaining the target feature maps corresponding to multiple panoramic images through the feature extraction layer, the self-attention mechanism layer can be invoked to process three fully connected feature maps corresponding to the target feature map to obtain the attention mechanism feature map corresponding to the target feature map.

[0090] Specifically, after obtaining the target feature map, the target feature map can be processed sequentially through the normalization layer and the fully connected layer to obtain the corresponding fully connected feature map. Specifically, it can be described in detail in combination with Figure 3 as follows.

[0091] Refer to Figure 3, which shows the step flow chart of a method for obtaining a fully-connected feature map provided by an embodiment of the present application. As Figure 3 shown, the method for obtaining a fully-connected feature map may include: step 301 and step 302.

[0092] Step 301: Invoke the first normalization layer to perform normalization processing on the image features of the target feature map to obtain the first normalized feature map corresponding to the target feature map.

[0093] In this embodiment, the feature interaction layer may further include: a first normalization layer and three first fully-connected layers. Among them, the first normalization layer is respectively connected to the input ends of the three first fully-connected layers, and the output ends of the three first fully-connected layers are respectively connected to the self-attention mechanism layer. As Figure 7 shown, the LayerNorm located in front of Self-Attention is the first normalization layer, that is Figure 7 shown, the lowermost LayerNorm is the first normalization layer. There are three parallel first fully-connected layers (not shown in the figure) between the first normalization layer and Self-Attention.

[0094] After obtaining the target feature map, the target feature map can be used as the input of the first normalization layer of the feature interaction layer. As Figure 7 shown, after obtaining the target feature maps corresponding to N panoramic images, the target feature maps Camera1 feature, Camera2feature,...CameraN feature corresponding to N panoramic images can be used as the inputs of six first normalization layers respectively.

[0095] Furthermore, the first normalization layer can be invoked to perform normalization processing on the image features of the target feature map to obtain the first normalized feature map corresponding to the target feature map.

[0096] After obtaining the first normalized feature map corresponding to the target feature map, step 302 is executed.

[0097] Step 302: Invoke the three first fully-connected layers to process the first normalized feature map according to the corresponding matrix parameters respectively to obtain three fully-connected feature maps corresponding to the target feature map.

[0098] After obtaining the first normalized feature map corresponding to the target feature map, the three first fully-connected layers can be invoked to process the first normalized feature map according to the corresponding matrix parameters respectively to obtain three fully-connected feature maps corresponding to the target feature map.

[0099] It can be understood that the above matrix parameters are obtained through learning, and the matrix parameters of the three fully-connected layers are different.

[0100] After obtaining three fully-connected feature maps corresponding to the target feature map, these three fully-connected feature maps can be used as the input of the self-attention mechanism layer.

[0101] After calling the self-attention mechanism layer to process the three fully-connected feature maps corresponding to the target feature map and obtaining the attention mechanism feature map corresponding to the target feature map, step 202 is executed.

[0102] Step 202: Call the cross-attention mechanism layer to perform feature fusion processing on two fully-connected feature maps corresponding to the attention mechanism feature map and one fully-connected feature map corresponding to two attention mechanism feature maps adjacent to the attention mechanism feature map, so as to obtain the fused feature map corresponding to the target feature map.

[0103] After calling the self-attention mechanism layer to process the three fully-connected feature maps corresponding to the target feature map and obtaining the attention mechanism feature map corresponding to the target feature map, the cross-attention mechanism feature map can be called to perform feature fusion processing on two fully-connected feature maps corresponding to the attention mechanism feature map and one fully-connected feature map corresponding to two attention mechanism feature maps adjacent to the attention mechanism feature map, so as to obtain the fused feature map corresponding to the target feature map.

[0104] In a specific implementation, after obtaining the attention mechanism feature map, the attention mechanism feature map can pass through a second normalization layer and three fully-connected layers to obtain three fully-connected feature maps corresponding to the attention mechanism feature map. For this process, it can be combined with Figure 4 to be described in detail as follows.

[0105] Referring to Figure 4 , a step flowchart of another method for obtaining a fully-connected feature map provided by an embodiment of the present application is shown. As Figure 4 shown, the method for obtaining a fully-connected feature map may include: step 401 and step 402.

[0106] Step 401: Call the second normalization layer to perform feature fusion and normalization processing on the attention mechanism feature map and the first normalization feature map to obtain a second normalization feature map.

[0107] In this embodiment, the feature interaction layer may further include: a second normalization layer and three second fully-connected layers. The second normalization layer is respectively connected to the input ends of the three second fully-connected layers, and the output ends of the three second fully-connected layers are connected to the cross-attention mechanism layer. As Figure 7As shown, the LayerNorm between Self-Attention and Cross-Attention is the second normalization layer, and there are three fully connected layers (i.e., the second fully connected layers, not shown in the figure) between this second normalization layer and Cross-Attention.

[0108] After obtaining the attention mechanism feature map, the attention mechanism feature map can be used as the input of the second normalization layer to call the second normalization layer to perform feature fusion processing and normalization processing on the attention mechanism feature map and the first normalized feature map output by the first normalization layer, so as to obtain the second normalized feature map.

[0109] After obtaining the second normalized feature map, step 402 is executed.

[0110] Step 402: Call the three second fully connected layers to process the second normalized feature map respectively according to the corresponding matrix parameters to obtain three fully connected feature maps corresponding to the attention mechanism feature map.

[0111] After obtaining the second normalized feature map, three second fully connected layers can be called to process the second normalized feature map respectively according to the corresponding matrix parameters to obtain three fully connected feature maps corresponding to the attention mechanism feature map.

[0112] After obtaining the three fully connected feature maps corresponding to the attention mechanism feature map, the cross-attention mechanism layer can be called to perform feature fusion processing on two of the fully connected feature maps and one fully connected feature map of the adjacent attention mechanism feature map, where the matrix parameters corresponding to these two fully connected feature maps are adjacent, so as to obtain the fused feature map. For example, the first panoramic image is adjacent to the second panoramic image and the third panoramic image. When calling the cross-attention mechanism layer for feature fusion, after obtaining the attention mechanism feature map of the first panoramic image, three fully connected feature maps of the attention mechanism feature map can be obtained and two of the fully connected feature maps can be selected from them. Then, one fully connected feature map within the three fully connected feature maps corresponding to the attention mechanism feature map of the second panoramic image is obtained, and the two fully connected feature maps corresponding to the first panoramic image and the one fully connected feature map corresponding to the second panoramic image are used as the input of the cross-attention mechanism layer for feature fusion to obtain the first fused image feature. For the third panoramic image which is similar to the second panoramic image, feature fusion is performed again to obtain the second fused image feature. Finally, the cross-attention mechanism layer performs feature fusion on the first fused feature and the second fused feature, so that the fused image feature can be obtained and the corresponding fused feature map is output.

[0113] Of course, in the above example, the three fully connected layers can be represented by fc1, fc2, and fc3 respectively. The three FC images input in the cross-attention mechanism layer can be fc2(feat12) and fc3(feat12) corresponding to the current panoramic image, and fc1(feat22) corresponding to the adjacent panoramic image. Among them, feat12 represents the feature obtained by passing the image of camera 1 through the second LayerNorm layer, and feat22 represents the feature obtained by passing the image of camera 2 through the second LayerNorm layer. When performing image feature fusion of adjacent images in Cross-Attention, the fully connected feature map output by the first fully connected layer of the adjacent image and the fully connected feature maps output by the second and third fully connected layers of the current image can be jointly used as the input to the cross-attention mechanism layer.

[0114] It can be understood that the above example is only an example listed for better understanding of the technical solution of the embodiment of the present application, and does not serve as the sole limitation of this embodiment.

[0115] After obtaining the fused feature map corresponding to each panoramic image, step 203 is executed.

[0116] Step 203: Invoke the third normalization layer to perform normalization processing on the image features in the fused feature map to obtain the depth image of the panoramic image corresponding to the target feature map.

[0117] After obtaining the fused feature map corresponding to each panoramic image, the third normalization layer can be invoked to perform normalization processing on the images in the fused feature map to obtain the depth image of the panoramic image corresponding to the target feature map. In this process, the output of the second normalization layer can also be used as the input to the third normalization layer to jointly perform feature fusion with the fused feature map. Specifically, it can be combined with Figure 5 for the following detailed description.

[0118] Referring to Figure 5 , a step flowchart of another depth image estimation method provided by an embodiment of the present application is shown. As Figure 5 shown, the depth image estimation method may include: step 501 and step 502.

[0119] Step 501: Invoke the third normalization layer to perform feature fusion processing on the fused feature map and the second normalized feature map to generate a target fused feature map.

[0120] In this embodiment, after obtaining the fused feature map, the fused feature map and the second normalized feature map output by the second normalization layer can be jointly used as the input of the third normalization layer to call the third normalization layer to perform a fusion process on the fused feature map and the second normalized feature map, thereby generating a target fused feature map. Specifically, the third normalization layer can be called to extract the image features of the fused feature map and the second normalized feature map respectively, and perform a fusion process on the image features, and then a target fused feature map can be generated.

[0121] After generating the target fused feature map, step 502 is executed.

[0122] Step 502: Call the third normalization layer to perform a normalization process on the image features of the target fused feature map to obtain the depth image of the surround view image corresponding to the target feature map.

[0123] After generating the target fused feature map, the third normalization layer can be called to perform a normalization process on the image features of the target fused feature map to obtain the depth image of the surround view image corresponding to the target feature map.

[0124] In the above solution, through the above model structure provided by this embodiment (such as Figure 6 and Figure 7 ) synchronous learning of multiple surround view images at the same moment can be realized, and the depth estimation accuracy can be improved while the depth estimation efficiency is improved.

[0125] The depth image estimation method provided by the embodiment of the present application inputs multiple surround view images of a target vehicle at a target moment collected into a depth estimation model. The depth estimation model includes a feature extraction layer based on a convolutional neural network and a feature interaction layer based on an attention mechanism. Each of the multiple surround view images includes an overlapping area with an adjacent surround view image. The feature extraction layer is called to process the image features of the multiple surround view images to obtain target feature maps corresponding to the multiple surround view images. The feature interaction layer is called to perform a feature interaction process on the target feature maps and two adjacent target feature maps corresponding to the target feature maps to obtain depth images corresponding to the multiple surround view images. In the embodiment of the present application, the feature interaction layer can interact the features of the overlapping areas of two adjacent surround view images of the vehicle in a high-dimensional space, which can improve the accuracy of depth estimation and effectively solve the problem that the depth information obtained by the monocular vision depth estimation method has a poor effect. At the same time, compared with the binocular depth estimation method, two adjacent surround view images do not require a large overlapping area (only 5% - 10% is required), so the number of cameras arranged on the vehicle can be reduced, the cost of realizing the surround view effect can be reduced, and the beauty of the vehicle is not affected.

[0126] Referring to Figure 8 , a schematic structural diagram of a depth image estimation device provided by the embodiment of the present application is shown, such asFigure 8 As shown in the figure, the depth image estimation device 800 may include the following modules:

[0127] A panoramic image input module 810, configured to input multiple panoramic images of a target vehicle at a target moment collected into a depth estimation model; the depth estimation model includes a feature extraction layer based on a convolutional neural network and a feature interaction layer based on an attention mechanism, and each of the multiple panoramic images includes an overlapping area with an adjacent panoramic image;

[0128] A target feature map acquisition module 820, configured to call the feature extraction layer to process the image features of the multiple panoramic images to obtain target feature maps corresponding to the multiple panoramic images;

[0129] A depth image acquisition module 830, configured to call the feature interaction layer to perform feature interaction processing on the target feature map and two adjacent target feature maps corresponding to the target feature map to obtain a depth image corresponding to the multiple panoramic images.

[0130] Optionally, the target feature map acquisition module 820 includes:

[0131] A target feature map output unit, configured to call the feature extraction layer to extract the image features of the first dimension of the multiple panoramic images, perform dimension conversion processing on the image features, and output a target feature map including the image features of the second dimension; the semantic information included in the image features of the second dimension is more than that of the first dimension.

[0132] Optionally, the feature interaction layer includes a self-attention mechanism layer, a cross-attention mechanism layer, and a third normalization layer,

[0133] The depth image acquisition module 830 includes:

[0134] An attention mechanism feature map acquisition unit, configured to call the self-attention mechanism layer to process three fully connected feature maps corresponding to the target feature map to obtain an attention mechanism feature map corresponding to the target feature map;

[0135] A fusion feature map acquisition unit, configured to call the cross-attention mechanism layer to perform feature fusion processing on two fully connected feature maps corresponding to the attention mechanism feature map and one fully connected feature map corresponding to two attention mechanism feature maps adjacent to the attention mechanism feature map to obtain a fusion feature map corresponding to the target feature map;

[0136] A depth image acquisition unit, configured to call the third normalization layer to perform normalization processing on the image features in the fusion feature map to obtain a depth image of the panoramic image corresponding to the target feature map.

[0137] Optionally, the feature interaction layer further includes: a first normalization layer and three first fully connected layers, wherein the first normalization layer is connected to the input ends of the three first fully connected layers respectively, and the output ends of the three first fully connected layers are connected to the self-attention mechanism layer respectively.

[0138] The device also includes:

[0139] A first image acquisition module, configured to call the first normalization layer to perform normalization processing on the image features of the target feature map to obtain a first normalized feature map corresponding to the target feature map;

[0140] The first fully connected feature map acquisition module is used to call the three first fully connected layers to process the first normalized feature map according to the corresponding matrix parameters to obtain three fully connected feature maps corresponding to the target feature map.

[0141] Optionally, the feature interaction layer further includes: a second normalization layer and three second fully connected layers, wherein the second normalization layer is connected to the input ends of the three second fully connected layers respectively, and the output ends of the three second fully connected layers are connected to the cross attention mechanism layer.

[0142] The device also includes:

[0143] a second image acquisition module, configured to call the second normalization layer to perform feature fusion and normalization processing on the attention mechanism feature map and the first normalized feature map to obtain a second normalized feature map;

[0144] The second fully connected feature map acquisition module is used to call the three second fully connected layers to process the second normalized feature map according to the corresponding matrix parameters to obtain three fully connected feature maps corresponding to the attention mechanism feature map.

[0145] Optionally, the depth image acquisition unit includes:

[0146] a target fused image generating subunit, configured to call the third normalization layer to perform feature fusion processing on the fused feature map and the second normalized feature map to generate a target fused feature map;

[0147] The depth image acquisition subunit is used to call the third normalization layer to normalize the image features of the target fusion feature map to obtain a depth image of the surround view image corresponding to the target feature map.

[0148] The depth image estimation device provided by the embodiment of the present application inputs multiple panoramic images of the target vehicle at the target moment into a depth estimation model. The depth estimation model includes a feature extraction layer based on a convolutional neural network and a feature interaction layer based on an attention mechanism. Each of the multiple panoramic images contains an overlapping area with an adjacent panoramic image. The feature extraction layer is called to process the image features of the multiple panoramic images to obtain target feature maps corresponding to the multiple panoramic images. The feature interaction layer is called to perform feature interaction processing on the target feature maps and two adjacent target feature maps corresponding to the target feature maps to obtain a depth image corresponding to the multiple panoramic images. By the feature interaction layer in the embodiment of the present application, the features of the overlapping areas of two adjacent panoramic images of the vehicle are interacted in a high-dimensional space, which can improve the accuracy of depth estimation and effectively solve the problem that the depth information obtained by the monocular vision depth estimation method has a poor effect. At the same time, compared with the binocular depth estimation method, two adjacent panoramic images do not require a large overlapping area (only 5% - 10% is required), so the number of cameras arranged on the vehicle can be reduced, the cost of realizing the panoramic view effect can be reduced, and the beauty of the vehicle is not affected.

[0149] The embodiment of the present application provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, the above-mentioned depth image estimation method is implemented.

[0150] Figure 9 FIG. shows a schematic structural diagram of an electronic device 900 according to an embodiment of the present invention. As Figure 9 shown, the electronic device 900 includes a central processing unit (CPU) 901, which can execute various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 902 or computer program instructions loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the electronic device 900 can also be stored. The CPU 901, ROM 902, and RAM 903 are connected to each other through a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0151] Multiple components in the electronic device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, a microphone, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the electronic device 900 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0152] Each of the processes and treatments described above can be executed by the processing unit 901. For example, the method of any of the above embodiments can be implemented as a computer software program, which is tangibly included in a computer-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the CPU 901, one or more actions in the method described above can be executed.

[0153] The embodiments of the present application provide a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, each process of the above-described depth image estimation method embodiment is implemented, and the same technical effects can be achieved. To avoid repetition, it will not be elaborated here. Among them, the computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.

[0154] It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover a non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements that are not explicitly listed, or further includes elements inherent to such a process, method, article, or device. Without more limitations, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article, or device including that element.

[0155] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment method can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions to enable a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in the various embodiments of the present application.

[0156] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them fall within the protection scope of the present application.

[0157] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in the embodiments of the present application can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0158] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0159] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or groups can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling, direct coupling, or communication connection can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.

[0160] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0161] In addition, the functional units in each embodiment of the present application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0162] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, ROMs, RAMs, magnetic disks, or optical discs that can store program codes.

[0163] As described above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, and all of them should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. A depth image estimation method, characterized in that: Including: Inputting multiple panoramic images of a target vehicle at a target moment collected into a depth estimation model; The depth estimation model includes a feature extraction layer based on a convolutional neural network and a feature interaction layer based on an attention mechanism. Each of the multiple panoramic images contains an overlapping area with an adjacent panoramic image; Invoking the feature extraction layer to process the image features of the multiple panoramic images to obtain target feature maps corresponding to the multiple panoramic images; Invoking the feature interaction layer to perform feature interaction processing on the target feature maps and two adjacent target feature maps corresponding to the target feature maps to obtain depth images corresponding to the multiple panoramic images; Wherein, the feature interaction layer includes a self-attention mechanism layer, a cross-attention mechanism layer, and a third normalization layer. The step of invoking the feature interaction layer to perform feature interaction processing on the target feature maps and two adjacent target feature maps corresponding to the target feature maps to obtain depth images corresponding to the multiple panoramic images includes: Invoking the self-attention mechanism layer to process three fully connected feature maps corresponding to the target feature map to obtain an attention mechanism feature map corresponding to the target feature map; Performing feature fusion on a first fused image feature and a second fused image feature through the cross-attention mechanism layer to obtain a fused feature map corresponding to the target feature map; the first fused image feature is obtained by invoking the cross-attention mechanism layer to perform feature fusion processing on two fully connected feature maps corresponding to the attention mechanism feature map and one fully connected feature map corresponding to an attention mechanism feature map adjacent to the attention mechanism feature map; the second fused image feature is obtained by invoking the cross-attention mechanism layer to perform feature fusion processing on two fully connected feature maps corresponding to the attention mechanism feature map and one fully connected feature map corresponding to another attention mechanism feature map adjacent to the attention mechanism feature map; Invoking the third normalization layer to perform normalization processing on the image features in the fused feature map to obtain a depth image of the panoramic image corresponding to the target feature map.

2. The method according to claim 1, characterized in that The step of invoking the feature extraction layer to process the image features of the multiple panoramic images to obtain target feature maps corresponding to the multiple panoramic images includes: Invoking the feature extraction layer to extract the image features of the first dimension of the multiple panoramic images and perform dimension conversion processing on the image features, and outputting a target feature map containing the image features of the second dimension; the semantic information contained in the image features of the second dimension is more than that of the first dimension.

3. The method according to claim 1, characterized in that The feature interaction layer further includes a first normalization layer and three first fully connected layers. The first normalization layer is respectively connected to the input ends of the three first fully connected layers, and the output ends of the three first fully connected layers are respectively connected to the self-attention mechanism layer. Before invoking the self-attention mechanism layer to process three fully connected feature maps corresponding to the target feature map to obtain an attention mechanism feature map corresponding to the target feature map, it further includes: Call the first normalization layer to normalize the image features of the target feature map, and obtain the first normalized feature map corresponding to the target feature map; Call the three first fully-connected layers to process the first normalized feature map according to the corresponding matrix parameters respectively, and obtain three fully-connected feature maps corresponding to the target feature map.

4. The method according to claim 3, characterized in that The feature interaction layer further includes: a second normalization layer and three second fully-connected layers. The second normalization layer is respectively connected to the input ends of the three second fully-connected layers, and the output ends of the three second fully-connected layers are connected to the cross-attention mechanism layer. Before the cross-attention mechanism layer performs feature fusion on the first fused image feature and the second fused image feature to obtain the fused feature map corresponding to the target feature map, it further includes: Call the second normalization layer to perform feature fusion and normalization processing on the attention mechanism feature map and the first normalized feature map, and obtain the second normalized feature map; Call the three second fully-connected layers to process the second normalized feature map according to the corresponding matrix parameters respectively, and obtain three fully-connected feature maps corresponding to the attention mechanism feature map.

5. The method according to claim 4, wherein The step of calling the third normalization layer to normalize the image features in the fused feature map to obtain the depth image of the surround view image corresponding to the target feature map includes: Call the third normalization layer to perform feature fusion processing on the fused feature map and the second normalized feature map to generate a target fused feature map; Call the third normalization layer to normalize the image features of the target fused feature map to obtain the depth image of the surround view image corresponding to the target feature map.

6. A depth image estimation device, characterized in that It includes: A surround view image input module, configured to input multiple surround view images of a target vehicle at a target moment collected into a depth estimation model; The depth estimation model includes: a feature extraction layer based on a convolutional neural network and a feature interaction layer based on an attention mechanism. Each of the multiple surround view images includes an overlapping area with an adjacent surround view image; A target feature map acquisition module, configured to call the feature extraction layer to process the image features of the multiple surround view images, and obtain the target feature maps corresponding to the multiple surround view images; A depth image acquisition module, configured to call the feature interaction layer to perform feature interaction processing on the target feature map and two adjacent target feature maps corresponding to the target feature map, and obtain the depth images corresponding to the multiple surround view images; Wherein, the feature interaction layer includes: a self-attention mechanism layer, a cross-attention mechanism layer and a third normalization layer, and the depth image acquisition module includes: An attention mechanism feature map acquisition unit, configured to call the self-attention mechanism layer to process the three fully-connected feature maps corresponding to the target feature map, and obtain the attention mechanism feature map corresponding to the target feature map; A fused feature map acquisition unit is configured to perform feature fusion on the first fused image feature and the second fused image feature through a cross-attention mechanism layer to obtain a fused feature map corresponding to the target feature map; the first fused image feature is obtained by calling the cross-attention mechanism layer to perform feature fusion processing on two fully connected feature maps corresponding to the attention mechanism feature map and a fully connected feature map corresponding to an attention mechanism feature map adjacent to the attention mechanism feature map; the second fused image feature is obtained by calling the cross-attention mechanism layer to perform feature fusion processing on two fully connected feature maps corresponding to the attention mechanism feature map and a fully connected feature map corresponding to another attention mechanism feature map adjacent to the attention mechanism feature map; A depth image acquisition unit is used to call the third normalization layer to normalize the image features in the fusion feature map to obtain a depth image of the surround view image corresponding to the target feature map.

7. The device according to claim 6, characterized in that The target feature map acquisition module includes: A target feature map output unit is configured to call the feature extraction layer to extract image features of a first dimension of the plurality of surround view images, perform dimensionality conversion on the image features, and output a target feature map containing image features of a second dimension; the image features of the second dimension contain more speech information than the image features of the first dimension.

8. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the depth image estimation method according to any one of claims 1 to 5.

9. A readable storage medium, characterized in that: When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the depth image estimation method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Image processing method and device for intelligent driving

    CN114067292A