Monocular depth estimation device and method and electronic equipment

By performing spatial and channel condition random processing of the decoder of the monocular depth estimation device, the problem of insufficient accuracy in the monocular depth estimation method is solved, and more accurate depth images are generated, artifacts are reduced, and the estimation accuracy of image structure and object are improved.

CN120388060APending Publication Date: 2025-07-29FUJITSU LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410119410.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-29
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing monocular depth estimation method still has shortcomings in the accuracy and processing of artifacts in depth images, and cannot effectively improve the structure and object estimation accuracy in the image.

Method used

The decoder is used to process spatially conditioned random and channel-conditioned random fields on the features extracted by the encoder, and generate depth images through fusion processing to improve accuracy.

Benefits of technology

Improve the accuracy of depth images, reduce the discretized artifacts of depth intervals, and enhance the estimation accuracy of image structure and object.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388060A_ABST
    Figure CN120388060A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a monocular depth estimation device and method and electronic equipment. The monocular depth estimation device comprises: an encoder for encoding an input image to obtain at least two features having different sizes; the decoder is used for carrying out decoding processing on the at least two features so as to obtain decoding information, the decoder is provided with a plurality of conditional random field processing devices, and each conditional random field processing device carries out at least one of space conditional random field processing and channel conditional random field processing on the at least two features; and a depth image generation unit that generates a depth image on the basis of the decoding information. The accuracy of the depth image can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of image processing technology. Background Art

[0002] A depth image, also known as a distance image, is an image that uses the distance (i.e., depth) value from an image collector to each point in a scene as pixel values. Generally, it can be directly obtained by devices such as lidar, stereo cameras, or time-of-flight (TOF) cameras. Depth information can also be obtained by processing RGB images or videos.

[0003] Traditional depth estimation methods, such as structure from motion and stereo vision matching, are based on feature correspondences at multiple viewpoints. The method of inferring depth information from a single image is called the monocular depth estimation method. With the rapid development of deep neural networks, monocular depth estimation methods based on deep learning have been widely studied in recent years and have achieved good results in terms of accuracy.

[0004] The methods of using neural network models to improve the accuracy of monocular depth estimation can be roughly divided into four types: The first method is to increase training data, mixing images of multiple scenes as the training data set to improve the generalization ability of the model; The second method is based on traditional depth estimation methods, integrating Markov random fields or conditional random fields into the deep network to improve the accuracy of the model; The third method is based on the geometric characteristics of depth images, using these known geometric rules to guide model training and constrain the model; The fourth method is to design the network to improve the model's ability to access global information and local information.

[0005] Figure 1 It is a schematic composition diagram of a monocular depth estimation device. As Figure 1 shown, the monocular depth estimation device 100 generally includes three main parts: a feature extraction device 101, a depth prediction device 102, and a loss function calculation device 103. The feature extraction unit 101 may include an encoder and a decoder. The encoder can use various models for feature extraction. The decoder can be a standard feature upsampling decoder, etc. The depth prediction unit 102 can use the output information of the decoder to predict the final depth value. The loss function calculation unit 103 can be divided into two types. One is to directly use the difference between the predicted depth value and the true depth value to set the loss function, and the other is to use the geometric features of the depth to set the loss function.

[0006] It should be noted that the above introduction to the technical background is only for the convenience of clearly and completely explaining the technical solution of the present application and facilitating the understanding of those skilled in the art, and it cannot be considered that the above technical solutions are well-known to those skilled in the art just because these solutions are described in the background art part of the present application. Summary of the Invention

[0007] Although researchers have made many improvements to monocular depth estimation devices to enhance the accuracy of monocular depth estimation, the accuracy of current monocular depth evaluation is still relatively low, and it is unable to more accurately estimate the structures and objects in images. For example, the depth images generated by monocular depth estimation devices are not accurate enough, and the depth images have artifacts introduced by the discretization of the depth interval.

[0008] To address at least one of the above technical problems, embodiments of the present application provide a monocular depth estimation device, a method, and an electronic device. In the monocular depth estimation device, the decoder performs at least one of spatial conditional random field (sw-CRFs) processing and channel conditional random field (cw-CRFs) processing on the features extracted by the encoder, thereby improving the accuracy of the depth image.

[0009] According to one aspect of the embodiments of the present application, a monocular depth estimation device is provided, and the device includes:

[0010] An encoder for encoding an input image to obtain at least two features having different sizes;

[0011] A decoder for decoding at least two of the features to obtain decoded information, wherein the decoder has a plurality of conditional random field (CRFs) processing devices, and each of the conditional random field (CRFs) processing devices performs at least one of spatial conditional random field (sw-CRFs) processing and channel conditional random field (cw-CRFs) processing on at least two of the features; and

[0012] A depth image generation unit for generating a depth image based on the decoded information.

[0013] According to another aspect of the embodiments of the present application, a monocular depth estimation method is provided, and the method includes:

[0014] Encoding an input image using an encoder to obtain at least two features having different sizes;

[0015] Use a decoder to perform decoding processing on at least two of the features to obtain decoded information, wherein the decoder has a plurality of conditional random field (CRF) processing devices, and each of the conditional random field (CRF) processing devices performs at least one of spatial conditional random field (sw-CRFs) processing and channel conditional random field (cw-CRFs) processing on at least two of the features; and

[0016] Use a depth image generation unit to generate a depth image based on the decoded information.

[0017] According to another aspect of the embodiments of the present application, there is provided an electronic device, including a memory and a processor, where the memory stores a computer program, and the processor is configured to execute the computer program to implement the monocular depth estimation method as described above.

[0018] One of the beneficial effects of the embodiments of the present application is that the decoder performs at least one of spatial conditional random field (sw-CRFs) processing and channel conditional random field (cw-CRFs) processing on the features extracted by the encoder, thereby improving the accuracy of the depth image.

[0019] Referring to the following description and the drawings, specific embodiments of the embodiments of the present application are disclosed in detail, indicating the ways in which the principles of the embodiments of the present application can be adopted. It should be understood that the embodiments of the present application are not limited in scope thereby. Within the spirit and terms of the appended claims, the embodiments of the present application include many changes, modifications, and equivalents. Description of the Drawings

[0020] The included drawings are used to provide a further understanding of the embodiments of the present application, which form a part of the specification, are used to illustrate the embodiments of the present application, and are used to explain the principles of the present application together with the written description. Obviously, the drawings in the following description are only some embodiments of the present application, and for those of ordinary skill in the art, other embodiments can be obtained based on these drawings without creative efforts. In the drawings:

[0021] Figure 1 is a schematic diagram of a component of a monocular depth estimation device;

[0022] Figure 2 is a schematic diagram of the monocular depth estimation device of the embodiments of the present application;

[0023] Figure 3 is a schematic diagram of a conditional random field processing device;

[0024] Figure 4 is a schematic diagram of a fully connected spatial conditional random field model and a window-connected spatial conditional random field model;

[0025] Figure 5 It is a schematic diagram of fusing the processing result of the spatial conditional random field processing device with the processing result of the channel conditional random field processing device 212;

[0026] Figure 6 It is a schematic diagram of a monocular depth estimation method;

[0027] Figure 7 It is a schematic diagram of the electronic device according to the embodiment of the present application. Detailed implementation manners

[0028] Referring to the accompanying drawings, through the following description, the foregoing and other features of the embodiments of the present application will become apparent. In the description and drawings, specific embodiments of the present application are specifically disclosed, which show some embodiments in which the principles of the embodiments of the present application can be adopted. It should be understood that the present application is not limited to the described embodiments. On the contrary, the embodiments of the present application include all modifications, variations, and equivalents falling within the scope of the appended claims.

[0029] In the embodiments of the present application, terms such as "first", "second", etc. are used to distinguish different elements in terms of name, but do not indicate the spatial arrangement or time sequence of these elements, and these elements should not be limited by these terms. The term "and / or" includes any one and all combinations of one or more of the related listed terms. Terms such as "include", "comprise", "have", etc. mean the presence of the stated features, elements, components or assemblies, but do not exclude the presence or addition of one or more other features, elements, components or assemblies.

[0030] In the embodiments of the present application, the singular forms "a", "the", etc. include the plural forms and should be broadly understood as "a kind" or "a class" rather than being limited to the meaning of "one"; in addition, the term "the" should be understood to include both the singular form and the plural form unless otherwise clearly specified in the context. In addition, the term "according to" should be understood as "at least partially according to...", and the term "based on" should be understood as "at least partially based on...", unless otherwise clearly specified in the context.

[0031] Features described and / or illustrated for one embodiment can be used in the same or similar manner in one or more other embodiments, combined with the features in other embodiments, or replace the features in other embodiments. The term "include / comprise" as used herein means the presence of features, whole things, steps or components, but does not exclude the presence or addition of one or more other features, whole things, steps or components.

[0032] Embodiments of the first aspect

[0033] An embodiment of the present application provides a monocular depth estimation device.

[0034] Figure 2 It is a schematic diagram of the monocular depth estimation device according to the embodiment of the present application. As Figure 2 shown, the monocular depth estimation device 200 includes: an encoder 1, a decoder 2, and a depth image generation unit 3.

[0035] In an embodiment of the present application, the encoder 1 performs encoding processing on the input image, extracts features from the input image, and thus obtains at least two features with different scales.

[0036] The decoder 2 performs decoding processing on at least two features obtained by the encoder 1 to obtain decoded information. Among them, the decoder 2 has a plurality of (for example, more than 2) conditional random fields (CRFs) processing devices 21, and each conditional random fields processing device 21 performs at least one of spatial-window CRFs (sw-CRFs) processing and channel-wise CRFs (cw-CRFs) processing on the features obtained by the encoder 1.

[0037] The depth image generation unit 3 generates a depth image based on the decoding output by the decoder 2.

[0038] According to an embodiment of the first aspect of the present application, the decoder 2 performs at least one of spatial-window CRFs (sw-CRFs) processing and channel-wise CRFs (cw-CRFs) processing on the features extracted by the encoder 1, thereby improving the accuracy of the depth image.

[0039] In at least one embodiment, the input image may be an image captured by a camera in real time or an image stored in a storage device. The input image may be an RGB image or a grayscale image, etc. For example, if the input image is an RGB image, the size and number of channels of the input image can be expressed as H*W*3, where H represents the number of pixels included in the height of the input image, W represents the number of pixels included in the width of the input image, 3 represents the number of channels of the input image, and H*W represents the size (i.e., resolution) of the input image.

[0040] In at least one embodiment, the encoder 1 may employ an appropriate backbone network to generate at least two features with different scales. For example, the encoder 1 may extract 4 features from the input image, and the scale and number of channels of each feature are as follows: Feature 1, H / 4*W / 4*C; Feature 2, H / 8*W / 8*2C; Feature 3, H / 16*W / 16*4C; Feature 4, H / 32*W / 32*8C. In the scale and number of channels of each feature, the result of multiplying the first two numbers (e.g., H / 16 and W / 16) represents the scale (i.e., resolution) of the feature, and the third number (e.g., 4C) represents the number of channels of the feature. H and W represent the size of the input image, and C represents the number of channels of the feature with the highest resolution (i.e., Feature 1). In this application, the method by which the encoder 1 generates at least two features with different scales may refer to the related art, and the description in the specification of this application will not elaborate further.

[0041] In at least one embodiment, the decoder 2 has a plurality of conditional random field (CRF) processing devices 21. Each conditional random field (CRF) processing device 21 may perform at least one of spatial conditional random field (sw-CRFs) processing and channel conditional random field (cw-CRFs) processing on the features obtained by the encoder 1. For example, the conditional random field (CRF) processing device 21 may perform spatial conditional random field (sw-CRFs) processing on the features, or perform channel conditional random field (cw-CRFs) processing, or perform both spatial conditional random field (sw-CRFs) processing and channel conditional random field (cw-CRFs) processing.

[0042] In the following description, taking the case where the conditional random field (CRF) processing device 21 performs both spatial conditional random field (sw-CRFs) processing and channel conditional random field (cw-CRFs) processing on the features as an example, the working principle of the decoder 2 will be described.

[0043] Figure 3 is a schematic diagram of the conditional random field (CRF) processing device 21. As Figure 3 shown, each conditional random field (CRF) processing device 21 may include: a spatial conditional random field (sw-CRFs) processing device 211, a channel conditional random field (cw-CRFs) processing device 212, and a fusion device 213.

[0044] Among them, the spatio-temporal conditional random field (sw-CRFs) processing device 211 performs spatio-temporal conditional random field (sw-CRFs) processing on the features; the channel conditional random field (cw-CRFs) processing device 212 performs channel conditional random field (cw-CRFs) processing on the features; the fusion device 213 fuses the results of the spatio-temporal conditional random field (sw-CRFs) processing and the results of the channel conditional random field (cw-CRFs) processing to generate fusion information. The fusion information output by the fusion device 213 can be used as the information output by the conditional random field (CRFs) processing device 21.

[0045] As Figure 3 shown, each conditional random field (CRFs) processing device 21 may further include: a depthwise separable convolution (DW-Conv) device 214. The depthwise separable convolution (DW-Conv) device 214 performs depthwise separable convolution processing on the received information to generate a convolution result.

[0046] In addition, when the depthwise separable convolution (DW-Conv) device 214 generates a convolution result, the fusion device 213 may fuse the convolution result (e.g., denoted as Yw) generated by the depthwise separable convolution (DW-Conv) device 214 with the result of the spatio-temporal conditional random field (sw-CRFs) processing (e.g., denoted as Ys) and the result of the channel conditional random field (cw-CRFs) processing (e.g., denoted as Yc) to generate fusion information. This application is not limited thereto. When the conditional random field (CRFs) processing device 21 does not include the depthwise separable convolution (DW-Conv) device 214, the fusion device 213 may only fuse the result of the spatio-temporal conditional random field (sw-CRFs) processing and the result of the channel conditional random field (cw-CRFs) processing to generate fusion information.

[0047] As Figure 2 shown, the decoder 2 may further have: a pyramid pooling module (PPM) 22. The pyramid pooling module (PPM) 22 may perform pooling processing on at least two features to obtain pooled information. For example, all the features generated by the encoder 1 are input into the pyramid pooling module (PPM) 22 to generate pooled information.

[0048] In this application, as Figure 2 shown, the number of conditional random field (CRFs) processing devices 21 may be equal to or less than the number of features of different sizes generated by the encoder 1. For example, Figure 2Four features generated by the encoder 1 are shown. The number of conditional random field (CRFs) processing devices 21 can be four or less than four. When the number of conditional random field (CRFs) processing devices 21 is equal to the number of features, all features generated by the encoder 1 can be processed by the conditional random field. When the number of conditional random field (CRFs) processing devices 21 is less than the number of features, the computational load can be reduced.

[0049] Next, the case where the number of conditional random field (CRFs) processing devices 21 is equal to the number of features will be described.

[0050] One of the multiple conditional random field (CRFs) processing devices 21 can be connected to the pyramid pooling module (PPM) 22; the other conditional random field (CRFs) processing devices 21 can be connected to other conditional random field (CRFs) processing devices 21. That is, the information received by one conditional random field (CRFs) processing device 21 among the multiple conditional random field (CRFs) processing devices 21 is the pooled information generated by the pyramid pooling module (PPM) 22 and one feature, and the information received by at least one other conditional random field (CRFs) processing device 21 among the multiple conditional random field (CRFs) processing devices 21 is another feature and the fusion information generated by the fusion device 213 of other conditional random field (CRFs) processing devices 21 different from the other conditional random field (CRFs) processing device 21.

[0051] For example, in Figure 2 the four conditional random field (CRFs) processing devices 21 can be represented as conditional random field (CRFs) processing devices 21a, 21b, 21c, and 21d. The information received by the conditional random field (CRFs) processing device 21a is the pooled information generated by the pyramid pooling module (PPM) 22 and feature 1; the information received by the conditional random field (CRFs) processing device 21b is feature 2 and the fusion information generated by the fusion device 213 of the conditional random field (CRFs) processing device 21a; the information received by the conditional random field (CRFs) processing device 21c is feature 3 and the fusion information generated by the fusion device 213 of the conditional random field (CRFs) processing device 21b; the information received by the conditional random field (CRFs) processing device 21d is feature 4 and the fusion information generated by the fusion device 213 of the conditional random field (CRFs) processing device 21c.

[0052] As Figure 2As shown, in at least some embodiments, the decoder 2 further includes: a Multi-scale Deformable Attention (MSDA) device 24 and an Adaptive Feature Partitioning (AFP) device 25.

[0053] Among them, the Multi-scale Deformable Attention (MSDA) device 24 performs multi-scale deformable attention processing on at least two features generated by the encoder 1 to obtain a processing result. For example, all the features generated by the encoder 1 can be input into the Multi-scale Deformable Attention (MSDA) device 24, whereby the Multi-scale Deformable Attention (MSDA) device 24 performs multi-scale deformable attention processing on all the features generated by the encoder 1.

[0054] The Adaptive Feature Partitioning (AFP) device 25 can receive the fusion information respectively generated by multiple Conditional Random Field (CRF) processing devices 21 and perform adaptive feature partitioning processing.

[0055] As Figure 2 shown, the processing result of the Multi-scale Deformable Attention (MSDA) device 24 and the processing result of the Adaptive Feature Partitioning (AFP) device 25 are input into the depth image generation unit 3 as decoding information. Among them, the depth image generation unit 3 can be, for example, an Internal Scene Discretization (IDR) device. In addition, the present application is not limited thereto, and the depth image generation unit 3 can also be other devices.

[0056] In the present application, regarding the working principles of the Pyramid Pooling Module (PPM) 22, the Multi-scale Deformable Attention (MSDA) device 24, the Adaptive Feature Partitioning (AFP) device 25, and the Internal Scene Discretization (IDR) device, reference can be made to related technologies.

[0057] Next, the principle of the Conditional Random Field (CRF) processing device 21 will be described.

[0058] Figure 4 are schematic diagrams of a fully connected spatial conditional random field model and a window-connected spatial conditional random field model.

[0059] Figure 4 In (a) of , it is a schematic diagram of a graphical model representing a fully connected spatial conditional random field. Among them, for a certain node 41, this node 41 is connected to all other nodes in the image 40.

[0060] Figure 4(b) is a schematic diagram of a graphical model representing a window-connected spatial conditional random field, where nodes within a certain window are connected to other nodes within that window and not to nodes outside the window.

[0061] For example, for a certain node 41a in window 40a, this node 41a is connected to all other nodes in window 40a, but this node 41a is not connected to nodes in windows 40b, 40c, and 40d.

[0062] In Figure 4 , the image 40 can be, for example, an image corresponding to the feature space or channel space of the input image input to the encoder 1, and windows 40a, 40b, 40c, and 40d can be, for example, partial regions in the image 40; nodes 41, 41a, etc. can correspond to pixel units composed of more than one pixel in the input image.

[0063] For the graphical model of the fully connected spatial conditional random field shown in (a) of Figure 4 , the energy function can be defined as Equation (1) below.

[0064] E(x) = ∑ i ψ u (x i ) + ∑ ij ψ p (x i , x j ) (1)

[0065] where x i is the predicted value of node i, and j represents all other nodes in the model. The unary potential function ψ u (unary potential function) is calculated for each node by the predictor based on the image features.

[0066] The pairwise potential function ψ p used to connect node pairs can be expressed as Equation (2).

[0067] ψ p = μ(x i , x j ) f(x i , x j ) g(I i , I j ) h(p i , p j ) (2)

[0068] where, when i = j, μ(x i , x j ) = 1, otherwise μ(xi , x j ) = 0; I i represents the color of node i; p i represents the position of node i.

[0069] The pairwise potential function ψ p can impose constraints based on color and position information, making the predicted value x i , x j more reasonable.

[0070] For the graphical model of the fully connected spatial conditional random field shown in (a) as Figure 4 , the unary potential function ψ u is usually related to the distribution of the predicted value. For example, the unary potential function ψ u can be expressed as Equation (3) below.

[0071] ψ u (x i ) = -log P(x i |I) (3)

[0072] where I represents the input color image and P represents the probability distribution of the predicted value.

[0073] The pairwise potential function ψ p is usually calculated based on the colors and positions of node pairs (e.g., pixel pairs). For example, the pairwise potential function ψ p can be calculated by Equation (4) below.

[0074]

[0075] This pairwise potential function ψ p encourages nodes with different colors and far distances to obtain different predicted values, and penalizes nodes with similar colors and close distances for obtaining different predicted values.

[0076] For the graphical model of the window-connected spatial conditional random field shown in (b) as Figure 4 , the unary potential function ψ u can be calculated based on the features of the image. The features of the image can come from the network. For example, the encoder 1 uses the network to extract the features of the input image. Therefore, the unary potential function ψ u can be obtained through the network, as shown in Equation (5) for example.

[0077] ψ u (x i ) = θ u (I, x i ) (5)

[0078] Among them, θ is a parameter of the network, and the network is, for example, a unary network.

[0079] The pairwise potential can be composed of the predicted values of the current node and other nodes and the weights calculated based on the color and position information of the node pairs. For example, the pairwise potential can be expressed as Equation (6).

[0080]

[0081] Among them, is the feature map, and ω is the weight function. The pairwise potential is calculated node by node. Then, for node i, all its pairwise potentials are calculated to obtain the pairwise potential function ψ of node i p which is expressed as Equation (7).

[0082]

[0083] Among them, α and β are weight functions and can be calculated by the network. The query vector q and the key vector k are calculated according to the feature map of each region (patch) in the window, and all the query vectors q of all regions in the window are combined into a matrix Q, and all the key vectors k of all regions in the window are combined into a matrix K.

[0084] Then, the dot product of matrices Q and K is calculated to obtain the potential weight between any node pair; then, the predicted value matrix X is multiplied by the potential weight to obtain the final pairwise potential function ψ p Therefore, Equation (7) can be written as Equation (8):

[0085]

[0086] Among them, the predicted value matrix X is a matrix formed by combining the predicted values of each node (for example, each pixel of the input image). For example, the predicted value can be the depth value.

[0087] Thus, based on the window-connected spatial conditional random field model, the processing result of the spatial conditional random field for features can be expressed as the following Equation (9).

[0088]

[0089] In Equations (8) and (9), softmax can refer to the normalization function.

[0090] In this application, the spatial conditional random field (sw-CRFs) processing device 211 can obtain the spatial conditional random field (sw-CRFs) processing result based on Equation (9). Among them, the space of features can be divided into multiple windows, and Q and K in each windowT , X can be represented as Q s , X s , whereby, for each window of the feature space, Equation (9) is written as Equation (10).

[0091]

[0092] In the present application, the channel conditional random field (cw-CRFs) processing device 212 can perform random field processing along the channel dimension. For example, the channel can be divided into multiple heads, and attention processing can be applied to each head. Herein, the channel can refer to the channel of the features obtained by the encoder 1.

[0093] For channel conditional random field processing, the pairwise potential function ψ p can be calculated by Equation (11) as follows.

[0094]

[0095] Furthermore, the processing result of the channel conditional random field on the features can be expressed as Equation (12) below.

[0096]

[0097] In the present application, the channel conditional random field (cw-CRFs) processing device 212 can obtain the channel conditional random field (sw-CRFs) processing result based on Equation (12). Herein, Q, K T , X is obtained based on the channel information of the features. Among them, the channel dimension of the features can be divided into multiple blocks, and Q, K in each block T , X can be represented as Q c , X c , whereby, for each block of the channel dimension of the features, Equation (12) is written as Equation (13).

[0098]

[0099] Figure 5 is a schematic diagram of the fusion of the processing results of the spatial conditional random field (sw-CRFs) processing device 211 and the channel conditional random field (cw-CRFs) processing device 212.

[0100] As Figure 5 shown, the feature 1 extracted by the encoder 1 and the pooled information obtained by the pyramid pooling module (PPM) 22 are input into the conditional random field processing device 21a, and Q is generated based on the input information s, X s , Q c , X c , X. Among them, X s , X c , X can come from the same data: The data is processed by the spatio - conditional random field (sw - CRFs) processing device 211, and then it is called X s ; The data is input into the channel - conditional random field (cw - CRFs) processing device 212, and then it is called X c ; The data is input into the fusion device 213, and then it is called X.

[0101] The spatio - conditional random field (sw - CRFs) processing device 211 generates a processing result Y according to Q s , X s , generating a processing result Y s ; The channel - conditional random field (cw - CRFs) processing device 212 generates a processing result Y according to Q c , X c , generating a processing result Y c ; The depth - separable convolution (DW - Conv) device 214 generates a convolution result Y based on X w .

[0102] The fusion device 213 fuses the processing result Y s , the processing result Y c and the convolution result Y generated by the depth - separable convolution (DW - Conv) device 214 w to perform fusion processing, generating fusion information, and using it as the output information of the conditional random field processing device 21a.

[0103] Among them, the fusion processing of the fusion device 213 includes, for example, the following operations:

[0104] Operation 1: Y s interacts spatially with Y w . For example, Y s passes through two consecutive 1x1 convolutional layers, then through normalization and activation processing, and then generates attention weights in the spatial dimension through the sigmoid function. Finally, the attention weights are multiplied element - by - element with Y w to obtain the spatial interaction result.

[0105] Operation 2: Y c interacts channel - wise with Y w . For example, Y cAfter passing through a global average pooling layer, two consecutive 1x1 convolutional layers, followed by normalization and activation processing, a sigmoid function is then used to generate attention weights in the channel dimension. Finally, these attention weights are multiplied element-wise with Y w to obtain the channel interaction result through element-wise multiplication.

[0106] Operation 3: After concatenating the spatial interaction result of Operation 1 and the channel interaction result of Operation 2, it passes through a feed-forward neural network (FFN) module to generate the final fusion result as the above-mentioned fusion information.

[0107] It should be noted that Figure 5 shows the conditional random field processing unit device 21a. For the conditional random field processing unit devices 21b, 21c, 21c, it is necessary to Figure 5 replace the pooled information of the pyramid pooling module (PPM) 22 in with the fusion information output by the conditional random field processing unit device 21a, the fusion information output by the conditional random field processing unit device 21b, and the fusion information output by the conditional random field processing unit device 21c, respectively.

[0108] The monocular depth estimation device 200 according to the embodiment of the first aspect of the present application can improve the accuracy of the depth image. For example, tests using datasets of indoor scenes show that in terms of absolute relative error (Abs_Rel) and average error (e.g., expressed in log 10 ), the monocular depth estimation device 200 of the present application is superior to the prior art; tests using datasets of outdoor scenes show that in terms of threshold accuracy δ1, threshold accuracy δ2, threshold accuracy δ3, absolute relative error (Abs_Rel), and squared relative difference (Sq_Rel), the monocular depth estimation device 200 of the present application is superior to the prior art. Among them, the dataset of the indoor scene is, for example, the NYU Depth v2 dataset, and the dataset of the outdoor scene is, for example, the KITTI dataset. The threshold accuracy refers to: for each pixel of the image, calculate the ratio of the true depth value corresponding to the pixel to the depth value predicted by the monocular depth estimation device, count the number of pixels whose ratio is less than the threshold, calculate the percentage of the counted pixel number in the total pixel number of the image, and use this percentage as the threshold accuracy. Threshold accuracies δ1, δ2, and δ3 correspond to different thresholds. For example, the threshold corresponding to δ1 is 1.25, the threshold corresponding to δ2 is 1.25 2 , and the threshold corresponding to δ3 is 1.25 3 .

[0109] It should be noted that only the components or modules related to the present application are described above, but the present application is not limited thereto. The monocular depth estimation device 200 may further include other components or modules. For the specific content of these components or modules, reference may be made to the related art.

[0110] For simplicity, Figure 1 only the connection relationships or signal directions between the various components or modules are exemplarily shown in [the figure], but those skilled in the art should clearly understand that various related technologies such as bus connection can be adopted. The above-mentioned various components or modules can be implemented by hardware facilities such as a processor and a memory; the embodiments of the present application do not limit this.

[0111] The above-mentioned various embodiments only exemplarily illustrate the embodiments of the present application, but the present application is not limited thereto, and appropriate modifications can also be made on the basis of the above-mentioned various embodiments. For example, the above-mentioned various embodiments can be used alone, or one or more of the above-mentioned various embodiments can be combined.

[0112] Embodiments of the second aspect

[0113] The embodiments of the present application provide a monocular depth estimation method, corresponding to the monocular depth estimation device of the embodiments of the first aspect. The same content of the embodiments of the second aspect and the embodiments of the first aspect will not be repeated.

[0114] Figure 6 is a schematic diagram of a monocular depth estimation method. As Figure 6 shown, the method includes:

[0115] Operation 601: Use an encoder to encode the input image to obtain at least two features with different sizes;

[0116] Operation 602: Use a decoder to decode at least two of the features to obtain decoded information, wherein the decoder has a plurality of conditional random field (CRF) processing devices, and each of the conditional random field (CRF) processing devices performs at least one of spatial conditional random field (sw-CRFs) processing and channel conditional random field (cw-CRFs) processing on at least two of the features; and

[0117] Operation 603: Use a depth image generation unit to generate a depth image based on the decoded information.

[0118] In operation 602, the processing of each conditional random field (CRF) processing device includes:

[0119] Process the feature using a spatial conditional random field (sw-CRFs) processing device for spatial conditional random field (sw-CRFs) processing;

[0120] Process the feature using a channel conditional random field (cw-CRFs) processing device for channel conditional random field (cw-CRFs) processing; and

[0121] Use a fusion device to fuse the result of the spatial conditional random field (sw-CRFs) processing with the result of the channel conditional random field (cw-CRFs) processing.

[0122] In addition, in some embodiments, the processing of each of the conditional random field (CRFs) processing devices further includes:

[0123] Perform depthwise separable convolution processing on the received information using a depthwise separable convolution (DW-Conv) device to generate a convolution result.

[0124] Wherein, the fusion device also fuses the convolution result (Yw) generated by the depthwise separable convolution (DW-Conv) device with the result (Ys) of the spatial conditional random field (sw-CRFs) processing and the result (Yc) of the channel conditional random field (cw-CRFs) processing to generate fused information.

[0125] In some embodiments of operation 602, the decoding process of the decoder further includes:

[0126] Perform pooling processing on the at least two features using a pyramid pooling module (PPM) to obtain pooled information.

[0127] Wherein, the information received by one of the multiple conditional random field (CRFs) processing devices is the pooled information generated by the pyramid pooling module (PPM) and one of the features, and the information received by at least another one of the multiple conditional random field (CRFs) processing devices is another one of the features and the fused information generated by the fusion device of other conditional random field (CRFs) processing devices different from the other conditional random field (CRFs) processing device.

[0128] In some embodiments of operation 602, the decoding process of the decoder further includes:

[0129] Perform multi-scale deformable attention processing on the at least two features using a multi-scale deformable attention device (Multi-scale Deformable Attention, MSDA) to obtain a processing result; and

[0130] The Adaptive Feature Partitioning (AFP) device receives the fusion information respectively generated by the multiple Conditional Random Fields (CRFs) processing devices, and performs adaptive feature partitioning processing.

[0131] Wherein, the processing result of the Multi-Scale Deformable Attention (MSDA) device and the processing result of the Adaptive Feature Partitioning (AFP) device are input into the depth image generation unit as the decoding information.

[0132] Only the steps or processes related to the present application are described above, but the present application is not limited thereto. The method for monocular depth estimation may further include other steps or processes. For the specific content of these steps or processes, reference may be made to the prior art. In addition, only some structures of the model used in the method for monocular depth estimation are taken as examples to exemplarily illustrate the embodiments of the present application, but the present application is not limited to these structures, and appropriate modifications can also be made to these structures. The implementation manners of these modifications should all be included within the scope of the embodiments of the present application.

[0133] Each of the above embodiments only exemplarily illustrates the embodiments of the present application, but the present application is not limited thereto, and appropriate modifications can also be made on the basis of each of the above embodiments. For example, each of the above embodiments can be used alone, or one or more of the above embodiments can be combined.

[0134] As can be seen from the above embodiments, the depth image generated by the monocular depth estimation device based on the present application is clearer and more accurate.

[0135] Embodiments of the third aspect

[0136] The embodiments of the present application provide an electronic device, including the monocular depth estimation device 200 as described in the embodiments of the first aspect, the content of which is incorporated herein. The electronic device may be, for example, a computer, a server, a workstation, a laptop computer, a smart phone, etc.; however, the embodiments of the present application are not limited thereto.

[0137] Figure 7 is a schematic diagram of the electronic device of the embodiments of the present application. As Figure 7 shown, the electronic device 700 may include: a processor (such as a central processing unit CPU) 710 and a memory 720; the memory 720 is coupled to the central processing unit 710. The memory 720 can store various data; in addition, a program 721 for information processing is also stored, and the program 721 is executed under the control of the processor 710.

[0138] In some embodiments, the functions of the monocular depth estimation device 100 are integrated into the processor 710. Among them, the processor 710 is configured to implement the monocular depth estimation method as described in the embodiments of the second aspect.

[0139] In some embodiments, the monocular depth estimation device 100 is separately configured from the processor 710. For example, the monocular depth estimation device can be configured as a chip connected to the processor 710, and the function of depth estimation based on a monocular image is implemented through the control of the processor 710.

[0140] For example, the processor 710 is configured to perform the following controls: encoding the input image using an encoder to obtain at least two features with different sizes; decoding the at least two features using a decoder to obtain decoded information, where the decoder has a plurality of conditional random field (CRF) processing devices, and each of the conditional random field (CRF) processing devices performs at least one of spatial conditional random field (sw-CRFs) processing and channel conditional random field (cw-CRFs) processing on the at least two features; and generating a depth image using a depth image generation unit based on the decoded information.

[0141] In some embodiments, the processing of each of the conditional random field (CRF) processing devices includes:

[0142] Performing spatial conditional random field (sw-CRFs) processing on the feature using a spatial conditional random field (sw-CRFs) processing device;

[0143] Performing channel conditional random field (cw-CRFs) processing on the feature using a channel conditional random field (cw-CRFs) processing device; and

[0144] Using a fusion device to fuse the result of the spatial conditional random field (sw-CRFs) processing and the result of the channel conditional random field (cw-CRFs) processing.

[0145] In some embodiments, the processing of each of the conditional random field (CRF) processing devices further includes:

[0146] Performing depthwise separable convolution processing on the received information using a depthwise separable convolution (DW-Conv) device to generate a convolution result.

[0147] Among them, the fusion device also fuses the convolution result (Yw) generated by the depthwise separable convolution (DW-Conv) device with the result (Ys) of the spatial conditional random field (sw-CRFs) processing and the result (Yc) of the channel conditional random field (cw-CRFs) processing to generate fused information.

[0148] In some embodiments, the decoding process of the decoder further includes:

[0149] Using a Pyramid Pooling Module (PPM) to perform pooling processing on the at least two features to obtain pooled information.

[0150] In some embodiments, the information received by one Conditional Random Field (CRF) processing device among the multiple CRF processing devices is the pooled information generated by the Pyramid Pooling Module (PPM) and one of the features, and the information received by at least another CRF processing device among the multiple CRF processing devices is another one of the features and the fusion information generated by the fusion device of other CRF processing devices different from the other CRF processing device.

[0151] In some embodiments, the decoding process of the decoder further includes:

[0152] Using a Multi-scale Deformable Attention (MSDA) device to perform multi-scale deformable attention processing on the at least two features to obtain a processing result; and

[0153] Using an Adaptive Feature Partitioning (AFP) device to receive the fusion information generated by each of the multiple Conditional Random Field (CRF) processing devices and perform adaptive feature partitioning processing,

[0154] The processing result of the Multi-scale Deformable Attention (MSDA) device and the processing result of the Adaptive Feature Partitioning (AFP) device are input as the decoding information into the depth image generation unit.

[0155] In addition, as Figure 7 shown, the electronic device 700 may further include: an input / output (I / O) device 730, a display 740, etc.; among them, the functions of the above components are similar to those in the prior art and will not be elaborated here. It should be noted that the electronic device 700 does not necessarily have to include Figure 7 all the components shown in Figure 7 ; in addition, the electronic device 700 may further include components not shown in

[0156] This application embodiment also provides a computer-readable program, wherein when the program is executed in an electronic device, the program causes the computer to execute the monocular depth estimation method as described in the embodiments of the second aspect in the electronic device.

[0157] An embodiment of the present application also provides a storage medium storing a computer-readable program, wherein the computer-readable program causes a computer to execute the monocular depth estimation method as described in the embodiment of the second aspect in an electronic device.

[0158] The above devices and methods of the present application can be implemented by hardware or by a combination of hardware and software. The present application relates to such a computer-readable program that, when executed by a logic component, can cause the logic component to implement the above-described device or component, or cause the logic component to implement the above-described various methods or steps. The present application also relates to a storage medium for storing the above program, such as a hard disk, a magnetic disk, an optical disk, a DVD, a flash memory, etc.

[0159] The method / device described in combination with the embodiments of the present application can be directly embodied as hardware, a software module executed by a processor, or a combination of the two. For example, one or more of the functional block diagrams shown in the figure and / or a combination of one or more of the functional block diagrams can correspond to each software module of the computer program flow, and can also correspond to each hardware module. These software modules can respectively correspond to the various steps shown in the figure. These hardware modules can be implemented by solidifying these software modules using a field programmable gate array (FPGA).

[0160] The software module can be located in a RAM memory, a flash memory, a ROM memory, an EPROM memory, an EEPROM memory, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. A storage medium can be coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium; or the storage medium can be a component of the processor. The processor and the storage medium can be located in an ASIC. The software module can be stored in the memory of the mobile terminal or in a memory card insertable into the mobile terminal. For example, if the device (such as a mobile terminal) uses a larger capacity MEGA-SIM card or a large-capacity flash device, the software module can be stored in the MEGA-SIM card or the large-capacity flash device.

[0161] One or more of the functional blocks described in the accompanying drawings and / or one or more combinations of functional blocks can be implemented as a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, or any suitable combination thereof for performing the functions described in the present application. One or more of the functional blocks described in the accompanying drawings and / or one or more combinations of functional blocks can also be implemented as a combination of computing devices, for example, a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in communication combination with a DSP, or any other such configuration.

[0162] The present application has been described in conjunction with specific embodiments, but those skilled in the art should understand that these descriptions are exemplary and not a limitation on the scope of protection of the present application. Those skilled in the art can make various variations and modifications to the present application based on the principles of the present application, and these variations and modifications are also within the scope of the present application.

[0163] The present application also provides the following remarks:

[0164] 1. A monocular depth estimation method, the method comprising:

[0165] Encoding an input image using an encoder to obtain at least two features having different sizes;

[0166] Decoding at least two of the features using a decoder to obtain decoded information, wherein the decoder has a plurality of conditional random field (CRFs) processing devices, and each of the conditional random field (CRFs) processing devices performs at least one of spatial conditional random field (sw-CRFs) processing and channel conditional random field (cw-CRFs) processing on the at least two features; and

[0167] Generating a depth image using a depth image generation unit based on the decoded information.

[0168] 2. The method according to Remark 1, wherein

[0169] The processing of each of the conditional random field (CRFs) processing devices includes:

[0170] Performing spatial conditional random field (sw-CRFs) processing on the features using a spatial conditional random field (sw-CRFs) processing device;

[0171] Performing channel conditional random field (cw-CRFs) processing on the features using a channel conditional random field (cw-CRFs) processing device; and

[0172] Use a fusion device to fuse the result processed by the spatial conditional random field (sw-CRFs) with the result processed by the channel conditional random field (cw-CRFs).

[0173] 3. The method according to Note 2, wherein,

[0174] The processing of each of the conditional random field (CRFs) processing devices further includes:

[0175] Use a depthwise separable convolution (DW-Conv) device to perform depthwise separable convolution processing on the received information to generate a convolution result,

[0176] wherein, the fusion device also fuses the convolution result (Yw) generated by the depthwise separable convolution (DW-Conv) device with the result (Ys) processed by the spatial conditional random field (sw-CRFs) and the result (Yc) processed by the channel conditional random field (cw-CRFs) to generate fusion information.

[0177] 4. The method according to Note 3, wherein,

[0178] The decoding process of the decoder further includes:

[0179] Use a pyramid pooling module (PPM) to perform pooling processing on the at least two features to obtain pooled information.

[0180] 5. The method according to Note 4, wherein,

[0181] The information received by one of the multiple conditional random field (CRFs) processing devices is the pooled information generated by the pyramid pooling module (PPM) and one of the features,

[0182] The information received by at least another one of the multiple conditional random field (CRFs) processing devices is another one of the features and the fusion information generated by the fusion device of other conditional random field (CRFs) processing devices different from the other conditional random field (CRFs) processing device.

[0183] 6. The method according to Note 3, wherein,

[0184] The decoding process of the decoder further includes:

[0185] Use a multi-scale deformable attention device (Multi-scale Deformable Attention, MSDA) to perform multi-scale deformable attention processing on the at least two features to obtain a processing result; and

[0186] Receive the fusion information respectively generated by the plurality of Conditional Random Field (CRF) processing devices using an Adaptive Feature Partitioning (AFP) device, and perform adaptive feature partitioning processing.

[0187] The processing result of the multi-scale deformable attention device (MSDA) and the processing result of the adaptive feature partitioning (AFP) device are input into the depth image generation unit as the decoding information.

[0188] 7. A storage medium storing a computer-readable program, wherein the computer-readable program causes a processor coupled to the storage medium to execute the following method:

[0189] Use an encoder to perform encoding processing on an input image to obtain at least two features having different sizes;

[0190] Use a decoder to perform decoding processing on at least two of the features to obtain decoding information, wherein the decoder has a plurality of Conditional Random Field (CRF) processing devices, and each of the Conditional Random Field (CRF) processing devices performs at least one of spatial conditional random field (sw-CRFs) processing and channel conditional random field (cw-CRFs) processing on at least two of the features; and

[0191] Use a depth image generation unit to generate a depth image based on the decoding information.

[0192] 8. The storage medium according to note 7, wherein

[0193] The processing of each of the Conditional Random Field (CRF) processing devices includes:

[0194] Use a spatial conditional random field (sw-CRFs) processing device to perform spatial conditional random field (sw-CRFs) processing on the features;

[0195] Use a channel conditional random field (cw-CRFs) processing device to perform channel conditional random field (cw-CRFs) processing on the features; and

[0196] Use a fusion device to fuse the result of the spatial conditional random field (sw-CRFs) processing with the result of the channel conditional random field (cw-CRFs) processing.

[0197] 9. The storage medium according to note 7, wherein the method further includes:

[0198] The processing of each of the Conditional Random Field (CRF) processing devices further includes:

[0199] The received information is processed by a depthwise separable convolution (DW-Conv) device to generate a convolution result.

[0200] Among them, the fusion device also fuses the convolution result (Yw) generated by the depthwise separable convolution (DW-Conv) device with the result (Ys) processed by the spatial conditional random field (sw-CRFs) and the result (Yc) processed by the channel conditional random field (cw-CRFs) to generate fusion information.

[0201] 10. The method according to Note 7, wherein

[0202] The decoding process of the decoder further includes:

[0203] Pooling the at least two features using a pyramid pooling module (PPM) to obtain pooled information.

Claims

1. A monocular depth estimation device, characterized in that The device includes: An encoder for encoding an input image to obtain at least two features having different sizes; A decoder for decoding the at least two features to obtain decoded information, wherein the decoder has a plurality of conditional random field processing devices, and each of the conditional random field processing devices performs at least one of spatial conditional random field processing and channel conditional random field processing on the at least two features; and A depth image generation unit for generating a depth image based on the decoded information.

2. The device according to claim 1, wherein Each of the conditional random field processing devices includes: A spatial conditional random field processing device for performing spatial conditional random field processing on the feature; A channel conditional random field processing device for performing channel conditional random field processing on the feature; and A fusion device for fusing the result of the spatial conditional random field processing and the result of the channel conditional random field processing.

3. The device according to claim 2, wherein Each of the conditional random field processing devices further includes: A depthwise separable convolution device for performing depthwise separable convolution processing on the received information to generate a convolution result, The fusion device further fuses the convolution result generated by the depthwise separable convolution device with the result of the spatial conditional random field processing and the result of the channel conditional random field processing to generate fusion information.

4. The device according to claim 3, wherein The decoder further has: A pyramid pooling module for performing pooling processing on the at least two features to obtain pooled information.

5. The device according to claim 4, wherein The information received by one of the plurality of conditional random field processing devices is the pooled information generated by the pyramid pooling module and one of the features, The information received by at least another one of the plurality of conditional random field processing devices is another one of the features and the fusion information generated by the fusion device of other conditional random field processing devices different from the other conditional random field processing device.

6. The device according to claim 3, wherein The decoder further includes: A multi-scale deformable attention device for performing multi-scale deformable attention processing on the at least two features to obtain a processing result; and An adaptive feature partitioning device for receiving the fusion information generated by each of the plurality of conditional random field processing devices and performing adaptive feature partitioning processing, The processing result of the multi-scale deformable attention device and the processing result of the adaptive feature partitioning device are input as the decoded information into the depth image generation unit.

7. An electronic device including the monocular depth estimation device according to any one of claims 1 to 6.

8. A monocular depth estimation method, characterized in that, The method includes: Using an encoder to encode an input image to obtain at least two features having different sizes; Use a decoder to perform decoding processing on at least two of the said features to obtain decoded information, wherein the decoder has a plurality of conditional random field processing devices, and each of the conditional random field processing devices performs at least one of spatial conditional random field processing and channel conditional random field processing on at least two of the said features; and Use a depth image generation unit to generate a depth image based on the decoded information.

9. The method according to claim 8, wherein The processing of each of the conditional random field processing devices includes: Use a spatial conditional random field processing device to perform spatial conditional random field processing on the feature; Use a channel conditional random field processing device to perform channel conditional random field processing on the feature; and Use a fusion device to fuse the result of the spatial conditional random field processing with the result of the channel conditional random field processing.

10. The method according to claim 9, wherein The processing of each of the conditional random field processing devices further includes: Use a depthwise separable convolution device to perform depthwise separable convolution processing on the received information to generate a convolution result, wherein the fusion device also fuses the convolution result generated by the depthwise separable convolution device with the result of the spatial conditional random field processing and the result of the channel conditional random field processing to generate fusion information.