A method, apparatus, and electronic device for scene depth inference based on historical information

Through the scene depth inference method based on historical information, the network weights are updated using the time attention and space-time correlation modules, the data dependence and pose inaccurate problems of scene depth recovery in two-dimensional images are solved, and high-precision scene depth recovery and generalization capabilities are achieved.

CN114627176BActive Publication Date: 2025-07-18SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210139037.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-15
Publication Date
2025-07-18
Estimated Expiration
2042-02-15

AI Technical Summary

Technical Problem

The prior art requires a large amount of manual labeling data to restore the depth of scene from two-dimensional images, and inaccurate camera posture in the scheme based on unsupervised learning results in erroneous affine transformation, affecting the quality of the scene depth.

Method used

Using a scene depth inference method based on historical information, the camera posture accuracy and generalization ability are improved by obtaining image frames and pre-constructed depth estimation networks and camera motion networks, the error calculation module is used to update the weights, and combining the time attention module and the time-space correlation module to improve the camera position accuracy and generalization ability.

Benefits of technology

The impact of erroneous affine transformation caused by inaccurate poses is reduced, the algorithm's generalization ability to unknown scenes is improved, and the scene depth recovery is achieved without supervision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114627176B_ABST
    Figure CN114627176B_ABST
Patent Text Reader

Abstract

The present application provides a method, apparatus, and electronic device for scene depth inference based on historical information. The method includes: obtaining a first image frame and a second image frame; obtaining a first depth weight of a depth estimation network and a first motion weight of a camera motion network; using the first image frame, the second image frame, the first depth weight, and the first motion weight, and an error calculation module to calculate a first error; using the first error as a guiding signal to jointly update the first depth weight of the depth estimation network and the first motion weight of the camera motion network to obtain a second depth weight and a second motion weight; using the error calculation module to calculate a second error; and determining the scene depth of the second image frame and the relative pose between the first image frame and the second image frame according to the first error and the second error. This solution can reduce the influence of incorrect affine transformation caused by inaccurate pose.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of computer vision and image processing, and particularly relates to a method, apparatus, and electronic device for scene depth inference based on historical information. Background Art

[0002] Accurately recovering the scene depth from a two-dimensional image helps to better understand the three-dimensional structure of the scene, thereby better completing various vision tasks. However, ordinary cameras only capture two-dimensional images during shooting, losing the depth information of the scene. Therefore, how to recover the scene depth from two-dimensional images or video sequences has become a fundamental and challenging task in the field of computer vision. Although it is currently possible to recover a competitive scene depth from two-dimensional images, a large amount of manually labeled data is required to train the neural network, which is time-consuming and laborious. Moreover, once the model training is completed, the weights of the model are frozen, reducing the generalization ability of the algorithm for unknown scenes. In addition, for the scheme based on fully unsupervised learning to recover the scene depth from two-dimensional images, it is necessary to simultaneously predict the camera pose from adjacent frames, and inaccurate poses will produce incorrect affine transformation results, directly affecting the quality of the synthesized image and thus the quality of the recovered scene depth. Summary of the Invention

[0003] The objective of the embodiments of this specification is to provide a method, apparatus, and electronic device for scene depth inference based on historical information.

[0004] To solve the above technical problems, the embodiments of this application are implemented as follows:

[0005] In a first aspect, this application provides a method for scene depth inference based on historical information, the method including:

[0006] Obtain a first image frame and a second image frame of the image to be measured, where the first image frame is the image frame at the previous moment of the second image frame;

[0007] Obtain a first depth weight of a pre-constructed depth estimation network and a first motion weight of a pre-constructed camera motion network;

[0008] Using an error calculation module, calculate a first error for the first image frame, the second image frame, the first depth weight, and the first motion weight;

[0009] Use the first error as a guidance signal to jointly update the first depth weight of the depth estimation network and the first motion weight of the camera motion network to obtain a second depth weight and a second motion weight;

[0010] Using an error calculation module, calculate a second error for the first image frame, the second image frame, the second depth weight, and the second motion weight;

[0011] Determine the scene depth of the second image frame and the relative pose between the first image frame and the second image frame according to the first error and the second error.

[0012] In a second aspect, the present application provides a scene depth inference device based on historical information. The device includes:

[0013] A first acquisition module, configured to acquire a first image frame and a second image frame of a to-be-detected image, where the first image frame is the image frame at the previous moment of the second image frame;

[0014] A second acquisition module, configured to acquire a first depth weight of a pre-constructed depth estimation network and a first motion weight of a pre-constructed camera motion network;

[0015] A first processing module, configured to use an error calculation module to calculate a first error based on the first image frame, the second image frame, the first depth weight, and the first motion weight;

[0016] An update module, configured to use the first error as a guidance signal to jointly update the first depth weight of the depth estimation network and the first motion weight of the camera motion network to obtain a second depth weight and a second motion weight;

[0017] A second processing module, configured to use an error calculation module to calculate a second error based on the first image frame, the second image frame, the second depth weight, and the second motion weight;

[0018] A determination module, configured to determine the scene depth of the second image frame and the relative pose between the first image frame and the second image frame according to the first error and the second error.

[0019] In a third aspect, the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the scene depth inference method based on historical information as in the first aspect.

[0020] As can be seen from the technical solutions provided in the embodiments of the present specification above, the solution: recovers the scene depth from two-dimensional images in a completely unsupervised manner, injects the historical frame information in the memory unit into the current input unit through a temporal attention module, and models the spatial correlation of the spatio-temporal feature map to improve the accuracy of the camera pose and reduce the influence of incorrect affine transformation caused by inaccurate pose; during inference, online decision-making inference is used to improve the generalization ability of the algorithm for unknown scenes. Description of the Drawings

[0021] To more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0022] Figure 1 Flow diagram of the method for scene depth inference based on historical information provided by this application;

[0023] Figure 2 Joint training block diagram of the depth estimation network and the camera motion network provided by the embodiments of this application;

[0024] Figure 3 Principle diagram of the time attention module provided by the embodiments of this application;

[0025] Figure 4 Principle diagram of the spatio-temporal correlation module provided by the embodiments of this application;

[0026] Figure 5 Schematic diagram of the scene depth inference process provided by the embodiments of this application;

[0027] Figure 6 Structural diagram of the device for scene depth inference based on historical information provided by this application;

[0028] Figure 7 Structural diagram of the electronic device provided by this application. Detailed implementation manners

[0029] In order to enable those skilled in the art to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this specification.

[0030] In the following description, specific details such as specific system structures and technologies are proposed for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of this application. However, those skilled in the art should clearly understand that this application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of this application.

[0031] Without departing from the scope or spirit of the present application, various improvements and changes can be made to the specific embodiments of the description of the present application, which are obvious to those skilled in the art. Other embodiments obtained from the description of the present application are obvious to those skilled in the art. The description and examples of the present application are merely exemplary.

[0032] Regarding the use of "comprising", "including", "having", "containing", etc. in this article, they are all open-ended terms, meaning including but not limited to.

[0033] Unless otherwise specified, the "parts" in this application are all calculated by mass parts.

[0034] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0035] Referring to Figure 1 , which shows a schematic flowchart of a method for scene depth inference based on historical information provided by an embodiment of the present application.

[0036] As Figure 1 shown, the method for scene depth inference based on historical information may include:

[0037] S110. Obtain a first image frame and a second image frame of the image to be measured, where the first image frame is the image frame at the previous moment of the second image frame.

[0038] Among them, the image to be measured is any image for which the scene depth needs to be inferred and predicted, and the image to be measured may be a two-dimensional image.

[0039] The image to be measured is clipped into several adjacent-frame image frames at equal time intervals. Among them, the first image frame and the second image frame of the image to be measured are two image frames at adjacent moments, and the first image frame is the image frame at the previous moment of the second image frame. For example, the first image frame is the image frame at time t - 1, the second image frame is the image frame at time t, or the first image frame is the image frame at time t, and the second image frame is the image frame at time t + 1, etc.

[0040] S120. Obtain the first depth weight of the pre-constructed depth estimation network and the first motion weight of the pre-constructed camera motion network.

[0041] Among them, the depth estimation network is used to estimate the scene depth of a two-dimensional image. The depth estimation network may adopt a neural network with an encoder-decoder structure. The type and network structure of the neural network adopted by the depth estimation network are not limited.

[0042] The camera motion network is used to predict the relative pose between adjacent image frames.

[0043] Both the depth estimation network and the camera motion network are pre-constructed and pre-trained.

[0044] Referring to Figure 2 , which shows the joint training block diagram of the depth estimation network and the camera motion network provided by the embodiments of the present application ( Figure 2 In the leftmost original image, scene depth picture, and synthesized view in, all are color pictures, which are processed as grayscale pictures here). It can be understood that to train the depth estimation network and the camera motion network, first obtain the training set data. In this application, the training set data is a group of image frames at several adjacent three moments in the image. For example, the image frame at time t-1, the image frame at time t, and the image frame at time t+1 are one data in the training data set, and the image frame at time t-2, the image frame at time t-1, and the image frame at time t are also one data in the training data set, and so on. According to Figure 2 As shown in the training block diagram, taking the image frame at time t-1, the image frame at time t, and the image frame at time t+1 as the inputs of the camera motion network and the depth estimation network, the depth estimation network and the camera motion network are jointly trained using the objective function, that is, to guide the update of the depth weights of the depth estimation network and the motion weights of the camera motion network.

[0045] It can be understood that before the image frames at time t-1, time t, and time t+1 are input into the depth estimation network and the camera motion network, image preprocessing is first performed, for example, including randomly flipping, randomly cropping, and normalizing the image data, and converting the processed data into tensor data with a dimension of C×H×W. Here, the batch dimension is omitted. Among them, C represents the channel dimension size of the sample. During training, C = 3 for the depth estimation network, and C = 9 in the camera motion network (during testing, the input of the camera motion network is 3 image frames at adjacent moments). During the test inference period (that is, when implementing the scene depth inference method based on historical information of the present application), C = 6 in the camera motion network (during inference, the input of the camera motion network is 2 image frames at adjacent moments). H represents the height of the input sample image. Exemplarily, H = 256, and W represents the width of the input sample image. Exemplarily, W = 832.

[0046] Continuing to refer to Figure 2 , the camera motion network may include an encoder, a temporal attention module, and a spatio-temporal correlation module.

[0047] Among them, the encoder is used to extract the features of the stacked image frames to obtain a stacked feature map; the stacked image frames are obtained by stacking the first image frame and the second image frame along the channel dimension.

[0048] The temporal attention module is used to establish a global dependence relationship between the information of the historical memory unit and the information of the current input unit, and through the update unit, inject the information in the globally relevant historical memory unit into the current input unit, and at the same time store the globally relevant information in the current input unit into the historical memory unit as the historical memory unit for the next moment; the current input unit includes stacked feature maps, the stacked feature maps are updated to updated feature maps through the update unit, the historical memory unit includes a first memory feature map and a first temporal feature map, and the historical memory unit for the next moment includes a second memory feature map and a second temporal feature map.

[0049] The spatio-temporal correlation module is used to model the updated feature map / second memory feature map into a first / second spatio-temporal feature map with spatial correlation respectively.

[0050] In one embodiment, the temporal attention module uses shared temporal attention weights to establish a global dependence relationship between the information of the historical memory unit and the information of the current input unit, and through the update unit, inject the information in the globally relevant historical memory unit into the current input unit, and at the same time store the globally relevant information in the current input unit into the historical memory unit as the historical memory unit for the next moment, including:

[0051] Inject the feature information in the stacked feature maps and the feature information of the first memory feature map into the first temporal feature map to obtain a third temporal feature map;

[0052] Determine a temporal attention feature vector according to the third temporal feature map;

[0053] Determine a first feature vector according to the stacked feature maps;

[0054] Determine a second feature vector according to the first memory feature map;

[0055] Determine an input feature vector based on temporal attention according to the first feature vector and the temporal attention feature vector;

[0056] Determine a memory feature vector based on temporal attention according to the second feature vector and the temporal attention feature vector;

[0057] Adjust the input feature vector based on temporal attention and the memory feature vector based on temporal attention into corresponding first feature maps and second feature maps respectively;

[0058] Determine the updated feature map and the second memory feature map according to the first feature map and the second feature map;

[0059] Update the third temporal feature map to a second temporal feature map according to the updated feature map and the second memory feature map.

[0060] Exemplarily, refer toFigure 3 , which shows a schematic diagram of the principle of the time attention module provided in the embodiments of the present application. For convenience of description, assume that the input feature map (i.e., the stacked feature map) at time t is represented as The time feature map at time t-1 (i.e., the first time feature map) is The time memory feature map at time t-1 (i.e., the first memory feature map) is Its calculation process is as follows:

[0061] 1) Inject the feature information in the input feature map X at time t t and the feature information in the memory feature map X at time t-1 (m,t-1) into the time feature map X at time t-1 (time,t-1) to obtain the third time feature map

[0062]

[0063] Among them, represents the corresponding feature space, C represents the number of channels of the feature map, H represents the height of the feature map, W represents the width of the feature map, δ gelu (·) represents the activation function, W (i,t) , W (time,t-1) , W (m,t-1) represents the learned corresponding weight, b (i,t) , b (time,t-1) , b (m,t-1) represents the corresponding bias term.

[0064] 2) Calculate the time attention feature vector x (qk_time,t-1) according to the formula group (2):

[0065]

[0066]

[0067]

[0068]

[0069]

[0070] Among them, s represents the scalar scaling factor, "*" represents the product of corresponding elements, the function F split (·) represents slicing the feature map according to the channel dimension, the function F reshape (·) is used to adjust the feature map or feature vector into a preset shape, "T" represents transpose, represents the corresponding weight, represents the corresponding bias term.

[0071] 3) Adjust the input feature map X at time t according to formula group (3) (i,t) into a feature vector (i.e., the first feature vector) Adjust the memory feature map X at time t-1 (m,t-1) into a feature vector (i.e., the second feature vector)

[0072] x (i,t) = F reshape (X (i,t) )

[0073] x (m,t-1) = F reshape (X (m,t-1) ) (3)

[0074] 4) Calculate the input feature vector based on temporal attention and the memory feature vector based on temporal attention

[0075]

[0076]

[0077] 5) Adjust the feature vectors x (qk_i,t) and x (qk_m,t-1) into the corresponding feature maps (i.e., the first feature map) and (i.e., the second feature map):

[0078]

[0079]

[0080] 6) Calculate the information selection gate for selectively injecting the information in the memory feature map X at time t-1 (qk_m,t-1) into the input feature map X at time t (qk_i,t) :

[0081] G s = δ sig (W (qk_ms,t-1) X (qk_m,t-1) + b (qk_ms,t-1) + W (qk_is,t) X (qk_i,t) + b (qk_is,t) ) (6)

[0082] where the function δ sig (·) represents the sigmoid activation function, W(qk_ms,t-1) and W (qk_is,t) represent the corresponding weights, and b (qk_ms,t-1) and b (qk_is,t) represent the corresponding bias terms.

[0083] 7) Calculate the new feature map containing the memory feature map information according to formula (7)

[0084] X (im,t) = δ tanh (W (qk_imi,t) X (qk_i,t) + b (qk_imi,t) + G s *(W (qk_imm,t-1) X (qk_m,t-1) + b (qk_imm,t-1) )) (7)

[0085] where the function δ tanh (·) represents the tanh activation function, W (qk_imi,t) and W (qk_imm,t-1) represent the corresponding weights, and b (qk_imi,t) and b (qk_imm,t-1) represent the corresponding bias terms.

[0086] 8) Calculate the memory gate according to formula group (8) for updating the memory feature map X (qk_m,t-1) information at time t-1 to the memory feature map (the second memory feature map) at time t

[0087] G r = δ sig (W (qk_ir,t) X (qk_i,t) + b (qk_ir,t) + W (qk_mr,t-1) X (qk_m,t-1) + b (qk_mr,t-1) )

[0088] X (m,t) = (1 - G r ) * X (im,t) + G r * X (qk_m,t-1) (8)

[0089] where W (qk_ir,t) and W (qk_mr,t-1) represent the corresponding weights, and b (qk_ir,t) and b (qk_mr,t-1) represent the corresponding bias terms.

[0090] 9) Calculate the output gate according to formula (9) for updating the input feature map X (qk_i,t), obtain the updated new feature map (i.e., the updated feature map) as the input feature map for the next moment:

[0091] G o = δ sig (W (qk_io,t) X (qk_i,t) + b (qk_io,t) + W (qk_mo,t-1) X (qk_m,t-1) + b (qk_mo,t-1) )

[0092] X (io,t) = G o * X (im,t) (9)

[0093] where W (qk_io,t) and W (qk_mo,t-1) represent the corresponding weights, and b (qk_io,t) and b (qk_mo,t-1) represent the corresponding bias terms.

[0094] 10) Update the temporal feature map at time t - 1 to the temporal feature map at time t (i.e., the second temporal feature map) X (time,t) :

[0095]

[0096] where W (time,t-1) , W (io,t) and W (m,t) represent the corresponding weights, and b (time,t-1) , b (io,t) and b (m,t) represent the corresponding bias terms.

[0097] In order to utilize the global spatial structure information of the feature map and the dependency relationship between spatial structures to infer camera motion, a spatio-temporal correlation module as shown in Figure 4 is constructed to model the spatial context information using the global spatial correlation weights and to model the dependency relationship between the feature map channels of stacked frames to constrain the temporal information between stacked frames.

[0098] In one embodiment, the spatio-temporal correlation module is used to model the updated feature map / second memory feature map into the first / second spatio-temporal feature maps with spatial correlation respectively.

[0099] Among them, modeling the updated feature map into the first spatio-temporal feature map with spatial correlation includes:

[0100] Slice the updated feature map in the channel dimension to obtain the first sub-feature map, the second sub-feature map, and the third sub-feature map;

[0101] The updated feature map, the first sub-feature map, the second sub-feature map, and the third sub-feature map are respectively adjusted to the third feature vector, the first sub-feature vector, the second sub-feature vector, and the third sub-feature vector;

[0102] According to the first sub-feature vector and the second sub-feature vector, calculate the first spatial correlation matrix between the first sub-feature map and the second sub-feature map;

[0103] Use the first spatial correlation matrix to weight the third sub-feature vector to obtain the first spatially correlated feature vector;

[0104] According to the first spatially correlated feature vector and the third feature vector, determine the first spatio-temporal feature vector;

[0105] Adjust the first spatio-temporal feature vector to the first spatio-temporal feature map with spatial correlation.

[0106] Among them, modeling the second memory feature map into the second spatio-temporal feature map with spatial correlation includes:

[0107] Slice the second memory feature map in the channel dimension to obtain the fourth sub-feature map, the fifth sub-feature map, and the sixth sub-feature map;

[0108] The second memory feature map, the fourth sub-feature map, the fifth sub-feature map, and the sixth sub-feature map are respectively adjusted to the fourth feature vector, the fourth sub-feature vector, the fifth sub-feature vector, and the sixth sub-feature vector;

[0109] According to the fourth sub-feature vector and the fifth sub-feature vector, calculate the second spatial correlation matrix between the fourth sub-feature map and the fifth sub-feature map;

[0110] Use the second spatial correlation matrix to weight the sixth sub-feature vector to obtain the second spatially correlated feature vector;

[0111] According to the second spatially correlated feature vector and the fourth feature vector, determine the second spatio-temporal feature vector;

[0112] Adjust the second spatio-temporal feature vector to the second spatio-temporal feature map with spatial correlation.

[0113] Exemplarily, referring to Figure 4 , the calculation is as follows:

[0114] 1) Transform the input feature map (i.e., the updated feature map or the second memory feature map) according to the formula group (11) to obtain the feature map and evenly divide it into three sub-feature maps (i.e., the first / fourth sub-feature map, the second / fifth sub-feature map, and the third / sixth sub-feature map) respectively Among them, the function F split (·) represents slicing the feature map in the channel dimension.

[0115]

[0116] Among them, W mid is the corresponding weight, and b mid represents the corresponding bias term.

[0117] 2) Adjust the corresponding feature map into a feature vector according to the formula group (12) to obtain the corresponding feature vector Among them, the function F reshape (·) is used to adjust the feature map or feature vector into a preset shape.

[0118] x mid = F reshape (X mid )

[0119]

[0120]

[0121]

[0122] 3) Calculate the spatial correlation matrix between the first / fourth sub-feature maps and the second / fifth sub-feature maps according to formula (13) Among them, s represents the scalar scaling factor

[0123]

[0124] 4) Weight the third / sixth sub-feature quantities using the calculated spatial correlation matrix to obtain spatial correlation feature vectors (including the first spatial correlation feature vector and the second spatial correlation feature vector) as shown in formula (14)

[0125]

[0126] 5) Model the dependence relationship between the spatial structures of the feature maps according to formula (15), and calculate the first / second spatio-temporal feature vectors x time_corr :

[0127]

[0128] Among them, F C(·) consists of two layers of one-dimensional convolution with a kernel size of 3 and a stride of 1 and an activation function.

[0129] 6) Finally, the first / second spatio-temporal feature vector x with spatial correlation time_corr is adjusted into the first / second spatio-temporal feature map with spatial correlation

[0130] S130, the first image frame, the second image frame, the first depth weight, and the first motion weight, use an error calculation module to calculate the first error, including:

[0131] Input the first image frame and the second image frame into a pre-constructed depth estimation network. According to the first image frame and the first depth weight, obtain the first scene depth of the first image frame and the first encoder feature map of the first image frame. According to the second image frame and the first depth weight, obtain the second scene depth of the second image frame and the second encoder feature map of the second image frame;

[0132] Input the first image frame and the second image frame into a pre-constructed camera motion network. According to the first image frame, the second image frame, and the first motion weight, obtain the first relative pose between the first image frame and the second image frame;

[0133] The first scene depth, the second scene depth, and the first relative pose use an error calculation module to calculate the first error.

[0134] In one embodiment, the total error is determined according to the image synthesis error, the scene depth structure consistency error, the feature perception loss error, and the smoothness loss error. Among them, the total error includes the first error and the second error.

[0135] Specifically, the total error is determined according to the image synthesis error, the scene depth structure consistency error, the feature perception loss error, and the smoothness loss error, including:

[0136] Obtain the first image coordinates of the first image frame and the second image coordinates of the second image frame;

[0137] According to the first image coordinates, the camera internal parameters, and the first scene depth, determine the first world coordinates of the first image frame;

[0138] According to the second image coordinates, the camera internal parameters, and the second scene depth, determine the second world coordinates of the second image frame;

[0139] Affinely transform the first world coordinates of the first image frame to the second image frame panel to determine the third world coordinates after the affine transformation;

[0140] Affinely transform the second world coordinates of the second image frame to the first image frame panel to determine the fourth world coordinates after the affine transformation;

[0141] Project the third world coordinates and the fourth world coordinates onto a two-dimensional plane respectively to obtain the first post-affine transformation scene depth, the second post-affine transformation scene depth, and the corresponding first post-affine transformation image coordinates and second post-affine transformation image coordinates;

[0142] Determine the scene depth structure consistency error, the first depth structure inconsistency weight, and the second depth structure inconsistency weight according to the first scene depth, the second scene depth, the first post-affine transformation image coordinates, and the second post-affine transformation image coordinates;

[0143] Determine the first camera flow consistency occlusion mask and the second camera flow consistency occlusion mask according to the first image coordinates of the first image frame, the second post-affine transformation image coordinates, the second image coordinates of the second image frame, and the first post-affine transformation image coordinates;

[0144] Determine the image synthesis error according to the first depth structure inconsistency weight, the second depth structure inconsistency weight, the first camera flow consistency occlusion mask, and the second camera flow consistency occlusion mask;

[0145] Determine the feature perception loss error according to the first image frame, the second image frame, the first post-affine transformation image coordinates, and the second post-affine transformation image coordinates;

[0146] Determine the smoothness loss error according to the first scene depth, the second scene depth, the first image frame, and the second image frame;

[0147] Determine the total error according to the image synthesis error, the scene depth structure consistency error, the feature perception loss error, and the smoothness loss error.

[0148] Exemplarily, for convenience of description, assume that the trained scene depth fitting function is D t = F D (I t |W D ), where I t represents the two-dimensional image for which the scene depth at time t needs to be restored, W D represents the well-learned weight parameters in the scene depth fitting function F D (·), D t represents the scene depth of the two-dimensional image I t restored at time t, X (enc,t) represents the feature map of the image frame I t at time t output by the encoder, represents the pose from the image frame I t-1 at time t-1 to the image frame I t at time t, W T represents the pose transformation function F T(·) The weight parameters for good learning in, the camera internal parameters are represented as K, and P (xy,t-1) represents the image coordinates of the image frame I t-1 P (xyz,t-1) image frame I t-1 world coordinates of, P (xy,t) represents the image coordinates of the image frame I t P (xyz,t) represents the image frame I t world coordinates of. The total error calculation process is as follows:

[0149] 1) Calculate the world coordinates (i.e., the first world coordinates) P t-1 of the image frame I (xyz,t-1) image frame I t world coordinates of (i.e., the second world coordinates) P (xyz,t) :

[0150] P (xyz,t-1) = D t-1 * K - P (xy,t-1)

[0151] P (xyz,t) = D t * K - P (xy,t) (16)

[0152] where the "*" sign represents the product of the corresponding elements of the matrix.

[0153] 2) Calculate the affine transformation of the world coordinates P t-1 of the image frame I (xyz,t-1) to the image frame I t panel to obtain the world coordinates (i.e., the third world coordinates) P (proj_xyz,t) after affine transformation, and the world coordinates P t of the image frame I (xyz,t) affine transformation to the image frame I t-1 panel to obtain the world coordinates (i.e., the fourth world coordinates) P (proj_xyz,t-1) :

[0154]

[0155]

[0156] 3) Project the world coordinates P (proj_xyz,t) and P (proj_xyz,t-1) calculated by affine transformation onto the two-dimensional plane to obtain the scene depth D (proj,t) (i.e., the first scene depth after affine transformation) and D (proj,t-1)(i.e., the scene depth after the second affine transformation), and the corresponding image coordinates P after the affine transformation (proj_xy,t) (i.e., the image coordinates after the first affine transformation), P (proj_xy,t-1) (i.e., the image coordinates after the second affine transformation).

[0157] 4) Synthesize the image I t-1 and P (proj_xy,t-1) ; Synthesize the feature map X (syn,t) according to the encoder feature map X (enc,t-1) and P (proj_xy,t-1) ; Synthesize the depth map D (syn_enc,t) according to the estimated depth map D t-1 and P (projxy,t-1) ; Synthesize the depth map D (syn,t) according to the image frame I t and P (proj_xy,t) ; Synthesize the image I (syn,t-1) according to the encoder feature map X (enc,t) and P (proj_xy,t) ; Synthesize the feature map X (syn_enc,t-1) according to the estimated depth map D t and P (projxy,t) ; Synthesize the depth map D (syn,t-1) ; Calculate the forward camera flow U t-1 according to the image coordinates of I (proj_xy,t-1) and P forward ; Calculate the backward camera flow U t according to the image coordinates of I (proj_xy,t) and P backward ; Synthesize the forward camera flow U forward and P (proj_xy,t) ; Synthesize the backward camera flow U syn_forward according to U backward and P (projxy,t-1) ; syn_backward .

[0158] 5) Calculate the first camera flow consistency occlusion mask M (occ,t-1) and the second camera flow consistency occlusion mask M (occ,t) respectively according to the formula group (18):

[0159] M (occ,t-1) = Γ(‖U syn_forward + U backward ‖ 2 , α1(‖U syn_forward ‖ 2 + ‖U backward ‖ 2 ) + α2)

[0160] M (occ,t) = Γ(‖U syn_backward + U forward ‖2 , α1(‖U syn_backward ‖ 2 +‖U forward ‖ 2 )) + α2)

[0161]

[0162] 6) Calculate the scene depth structure consistency error E according to formula group (19) D and the first depth structure inconsistency weight M (D,t-1) , the second depth structure inconsistency weight M (D,t) .

[0163]

[0164]

[0165]

[0166] 7) Calculate the image synthesis error E according to formula (20) I :[[]]

[0167]

[0168] Among them,

[0169] 8) Calculate the feature perception loss error E according to formula (21) X :[[]]

[0170] E X = ERF(X (enc,t) , X (syn_enc,t) ) + ERF(X (enc,t-1) , X (syn_enc,t-1) ) (21)

[0171] 9) Calculate the smoothing loss error E according to formula (22) S :[[]]

[0172]

[0173] 10) Calculate the total error E according to formula (23):

[0174] E = λ I E I + λ D E D + λ X E X + λ S E S (23)

[0175] S140. Use the first error as a guidance signal to jointly update the first depth weight of the depth estimation network and the first motion weight of the camera motion network, obtaining the second depth weight and the second motion weight.

[0176] S150. For the first image frame, the second image frame, the second depth weight, and the second motion weight, use the error calculation module to calculate the second error, which may include:

[0177] Input the first image frame and the second image frame into the pre-constructed depth estimation network. According to the first image frame and the second depth weight, obtain the third scene depth of the first image frame and the third encoder feature map of the first image frame. According to the second image frame and the second depth weight, obtain the fourth scene depth of the second image frame and the fourth encoder feature map of the second image frame;

[0178] Input the first image frame and the second image frame into the pre-constructed camera motion network. According to the first image frame, the second image frame, and the second motion weight, obtain the second relative pose between the first image frame and the second image frame;

[0179] For the first image frame, the second image frame, the third scene depth, the fourth scene depth, the third encoder feature map, the fourth encoder feature map, and the second relative pose, use the error calculation module to calculate the second error.

[0180] This step refers to the specific calculation process of step S130, except that the first depth weight and the first motion weight in S130 are replaced with the second depth weight and the second motion weight, which will not be elaborated here.

[0181] S160. Determine the scene depth of the second image frame and the relative pose between the first image frame and the second image frame according to the first error and the second error.

[0182] Specifically, if the first error is greater than the second error, use the second scene depth as the scene depth of the second image frame, and use the first relative pose as the relative pose between the first image frame and the second image frame;

[0183] If the first error is less than or equal to the second error, use the fourth scene depth as the scene depth of the second image frame, and use the second relative pose as the relative pose between the first image frame and the second image frame.

[0184] Refer to Figure 5 , which shows a schematic diagram of the scene depth inference process. The process of inferring the scene depth is as follows:

[0185] 1) Using the weight W D of the depth estimation network obtained by training, and the weight W T of the camera motion networkAs the model weights of the depth estimation network and the camera motion network during inference, according to S130, the total error E and the pose transformation matrix from the historical frame to the current frame are calculated The scene depth D of the current frame t .

[0186] 2) Use the total error calculated in 1) above as a guiding signal to update the weights of the depth estimation network and the camera motion network, obtaining new model weights and

[0187] 3) According to the model weights obtained in 2) above and and S150, calculate the total error at this time The pose transformation matrix from the historical frame to the current frame The scene depth of the current frame

[0188] 4) Determine the scene depth of the current frame of the final output by comparing the magnitudes of the total error E and .

[0189] In the embodiments of the present application, a completely unsupervised form is adopted to recover the scene depth from two-dimensional images. Through the temporal attention module, the historical frame information in the memory unit is injected into the current input unit, and the spatial correlation of the spatio-temporal feature map is modeled to improve the accuracy of the camera pose and reduce the influence of incorrect affine transformations caused by inaccurate poses. During inference, online decision-making inference is used to improve the generalization ability of the algorithm for unknown scenes.

[0190] Referring to Figure 6 , which shows a schematic structural diagram of a scene depth inference device based on historical information described according to an embodiment of the present application.

[0191] As Figure 6 shown, the scene depth inference device 600 based on historical information may include:

[0192] A first acquisition module 610, configured to acquire a first image frame and a second image frame of the image to be measured, where the first image frame is the image frame at the previous moment of the second image frame;

[0193] A second acquisition module 620, configured to acquire the first depth weight of the pre-constructed depth estimation network and the first motion weight of the pre-constructed camera motion network;

[0194] A first processing module 630, configured to use the first image frame, the second image frame, the first depth weight, and the first motion weight, and adopt an error calculation module to calculate a first error;

[0195] An update module 640, configured to jointly update a first depth weight of the depth estimation network and a first motion weight of the camera motion network by using the first error as a guidance signal, so as to obtain a second depth weight and a second motion weight;

[0196] A second processing module 650, configured to use the first image frame, the second image frame, the second depth weight, and the second motion weight, and adopt an error calculation module to calculate a second error;

[0197] A determination module 660, configured to determine a scene depth of the second image frame and a relative pose between the first image frame and the second image frame according to the first error and the second error.

[0198] Optionally, the camera motion network includes an encoder, a temporal attention module, and a spatio-temporal correlation module;

[0199] The encoder is configured to extract features of the stacked image frames to obtain a stacked feature map; the stacked image frames are obtained by stacking the first image frame and the second image frame along the channel dimension;

[0200] The temporal attention module is configured to establish a global dependence relationship between the information of the historical memory unit and the information of the current input unit, and through an update unit, inject the information in the globally relevant historical memory unit into the current input unit, and at the same time store the globally relevant information in the current input unit into the historical memory unit as the historical memory unit for the next moment; the current input unit includes the stacked feature map, the stacked feature map is updated to an updated feature map through the update unit, the historical memory unit includes a first memory feature map and a first temporal feature map, and the historical memory unit for the next moment includes a second memory feature map and a second temporal feature map;

[0201] The spatio-temporal correlation module is configured to respectively model the updated feature map / second memory feature map into a first / second spatio-temporal feature map with spatial correlation.

[0202] Optionally, the scene depth inference device 600 based on historical information is further configured to:

[0203] Inject the feature information in the stacked feature map and the feature information in the first memory feature map into the first temporal feature map to obtain a third temporal feature map;

[0204] Determine a temporal attention feature vector according to the third temporal feature map;

[0205] Determine a first feature vector according to the stacked feature map;

[0206] Determine a second feature vector according to the first memory feature map;

[0207] Determine an input feature vector based on temporal attention according to the first feature vector and the temporal attention feature vector;

[0208] Determine a memory feature vector based on time attention according to the second feature vector and the time attention feature vector;

[0209] Respectively adjust the input feature vector based on time attention and the memory feature vector based on time attention into corresponding first feature maps and second feature maps;

[0210] Determine the updated feature map and the second memory feature map according to the first feature map and the second feature map;

[0211] Update the third time feature map to the second time feature map according to the updated feature map and the second memory feature map.

[0212] Optionally, the scene depth inference device 600 based on historical information is further configured to:

[0213] Slice the updated feature map in the channel dimension to obtain a first sub-feature map, a second sub-feature map, and a third sub-feature map;

[0214] Respectively adjust the updated feature map, the first sub-feature map, the second sub-feature map, and the third sub-feature map into a third feature vector, a first sub-feature vector, a second sub-feature vector, and a third sub-feature vector;

[0215] Calculate a first spatial correlation matrix between the first sub-feature map and the second sub-feature map according to the first sub-feature vector and the second sub-feature vector;

[0216] Perform weighted processing on the third sub-feature vector by using the first spatial correlation matrix to obtain a first spatially correlated feature vector;

[0217] Determine a first spatio-temporal feature vector according to the first spatially correlated feature vector and the third feature vector;

[0218] Adjust the first spatio-temporal feature vector into a first spatio-temporal feature map with spatial correlation;

[0219] Slice the second memory feature map in the channel dimension to obtain a fourth sub-feature map, a fifth sub-feature map, and a sixth sub-feature map;

[0220] Respectively adjust the second memory feature map, the fourth sub-feature map, the fifth sub-feature map, and the sixth sub-feature map into a fourth feature vector, a fourth sub-feature vector, a fifth sub-feature vector, and a sixth sub-feature vector;

[0221] Calculate a second spatial correlation matrix between the fourth sub-feature map and the fifth sub-feature map according to the fourth sub-feature vector and the fifth sub-feature vector;

[0222] Perform weighted processing on the sixth sub-feature vector by using the second spatial correlation matrix to obtain a second spatially correlated feature vector;

[0223] Determine the second spatio-temporal feature vector based on the second spatial correlation feature vector and the fourth feature vector;

[0224] Adjust the second spatio-temporal feature vector to the second spatio-temporal feature map with spatial correlation.

[0225] Optionally, the first processing module 630 is further configured to:

[0226] Input the first image frame and the second image frame into a pre-constructed depth estimation network, and obtain the first scene depth of the first image frame and the first encoder feature map of the first image frame according to the first image frame and the first depth weight, and obtain the second scene depth of the second image frame and the second encoder feature map of the second image frame according to the second image frame and the first depth weight;

[0227] Input the first image frame and the second image frame into a pre-constructed camera motion network, and obtain the first relative pose between the first image frame and the second image frame according to the first image frame, the second image frame and the first motion weight;

[0228] For the first image frame, the second image frame, the first scene depth, the second scene depth, the first encoder feature map, the second encoder feature map and the first relative pose, use an error calculation module to calculate the first error;

[0229] Optionally, the second processing module 650 is further configured to:

[0230] Input the first image frame and the second image frame into a pre-constructed depth estimation network, and obtain the third scene depth of the first image frame and the third encoder feature map of the first image frame according to the first image frame and the second depth weight, and obtain the fourth scene depth of the second image frame and the fourth encoder feature map of the second image frame according to the second image frame and the second depth weight;

[0231] Input the first image frame and the second image frame into a pre-constructed camera motion network, and obtain the second relative pose between the first image frame and the second image frame according to the first image frame, the second image frame and the second motion weight;

[0232] For the first image frame, the second image frame, the third scene depth, the fourth scene depth, the third encoder feature map, the fourth encoder feature map and the second relative pose, use an error calculation module to calculate the second error.

[0233] Optionally, the determination module 660 is further configured to:

[0234] If the first error is greater than the second error, use the second scene depth as the scene depth of the second image frame, and use the first relative pose as the relative pose between the first image frame and the second image frame;

[0235] If the first error is less than or equal to the second error, use the fourth scene depth as the scene depth of the second image frame, and use the second relative pose as the relative pose between the first image frame and the second image frame.

[0236] Optionally, the total error includes the first error and the second error; the total error is determined according to the image synthesis error, the scene depth structure consistency error, the feature perception loss error, and the smooth loss error.

[0237] Optionally, the first processing module 630 or the second processing module 650 is further configured to:

[0238] Obtain the first image coordinates of the first image frame and the second image coordinates of the second image frame;

[0239] Determine the first world coordinates of the first image frame according to the first image coordinates, the camera internal parameters, and the first scene depth;

[0240] Determine the second world coordinates of the second image frame according to the second image coordinates, the camera internal parameters, and the second scene depth;

[0241] Affinely transform the first world coordinates of the first image frame to the second image frame panel to determine the third world coordinates after affine transformation;

[0242] Affinely transform the second world coordinates of the second image frame to the first image frame panel to determine the fourth world coordinates after affine transformation;

[0243] Project the third world coordinates and the fourth world coordinates onto a two-dimensional plane respectively to obtain the first scene depth after affine transformation, the second scene depth after affine transformation, and the corresponding first image coordinates after affine transformation and the second image coordinates after affine transformation;

[0244] Determine the scene depth structure consistency error, the first depth structure inconsistency weight, and the second depth structure inconsistency weight according to the first scene depth, the second scene depth, the first image coordinates after affine transformation, and the second image coordinates after affine transformation;

[0245] Determine the first camera flow consistency occlusion mask and the second camera flow consistency occlusion mask according to the first image coordinates of the first image frame, the second image coordinates after affine transformation, the second image coordinates of the second image frame, and the first image coordinates after affine transformation;

[0246] Determine the image synthesis error according to the first depth structure inconsistency weight, the second depth structure inconsistency weight, the first camera flow consistency occlusion mask, and the second camera flow consistency occlusion mask;

[0247] Determine the feature perception loss error according to the first image frame, the second image frame, the first image coordinates after affine transformation, and the second image coordinates after affine transformation;

[0248] Determine the smooth loss error according to the first scene depth, the second scene depth, the first image frame, and the second image frame;

[0249] Determine the total error according to the image synthesis error, the scene depth structure consistency error, the feature perception loss error, and the smooth loss error.

[0250] A scene depth inference device based on historical information provided in this embodiment can execute the embodiments of the above method, and its implementation principle and technical effects are similar, which will not be elaborated here.

[0251] Figure 7 This is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. As Figure 7 shown, a schematic structural diagram of an electronic device 300 suitable for implementing the embodiments of the present application is shown.

[0252] As Figure 7 shown, the electronic device 300 includes a central processing unit (CPU) 301, which can execute various appropriate actions and processes according to the program stored in the read-only memory (ROM) 302 or the program loaded from the storage section 308 into the random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the device 300 are also stored. The CPU 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. The input / output (I / O) interface 305 is also connected to the bus 304.

[0253] The following components are connected to the I / O interface 305: an input section 306 including a keyboard, a mouse, etc.; an output section 307 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker; a storage section 308 including a hard disk, etc.; and a communication section 309 including a network interface card such as a LAN card, a modem, etc. The communication section 309 performs communication processing via a network such as the Internet. The drive 310 is also connected to the I / O interface 305 as needed. A removable medium 311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 310 as needed, so that the computer program read from it can be installed into the storage section 308 as needed.

[0254] Specifically, according to the embodiments of the present disclosure, the above reference Figure 1The described process can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product that includes a computer program tangibly embodied on a machine-readable medium, the computer program including program code for performing the above-described method for scenario depth inference based on historical information. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 309, and / or installed from the removable medium 311.

[0255] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than that noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0256] The units or modules involved in the embodiments described in this application can be implemented in software or in hardware. The described units or modules can also be provided in a processor. The names of these units or modules do not, in some cases, constitute a limitation to the units or modules themselves.

[0257] The systems, apparatuses, modules, or units illustrated in the above embodiments can be specifically implemented by a computer chip or entity, or by a product with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a mobile phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0258] As another aspect, the present application also provides a storage medium. The storage medium can be the storage medium included in the aforementioned device in the above embodiments; or it can exist separately and be not assembled into the device. The storage medium stores one or more programs, and the aforementioned programs are used by one or more processors to perform the method for scenario depth inference based on historical information described in this application.

[0259] A storage medium includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media do not include transitory computer-readable media such as modulated data signals and carrier waves.

[0260] It should be noted that the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, commodity or device including the element.

[0261] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiment.

Claims

1. A method for scene depth inference based on historical information, characterized in that, The method includes: Obtaining a first image frame and a second image frame of the image to be measured, where the first image frame is the image frame at the previous moment of the second image frame; Obtaining a first depth weight of a pre-constructed depth estimation network and a first motion weight of a pre-constructed camera motion network; Based on the first image frame, the second image frame, the first depth weight, and the first motion weight, using an error calculation module to calculate a first error; Using the first error as a guiding signal to jointly update the first depth weight of the depth estimation network and the first motion weight of the camera motion network to obtain a second depth weight and a second motion weight; Based on the first image frame, the second image frame, the second depth weight, and the second motion weight, using the error calculation module to calculate a second error; Determining the scene depth of the second image frame and the relative pose between the first image frame and the second image frame according to the first error and the second error.

2. The method according to claim 1, wherein The camera motion network includes an encoder, a temporal attention module, and a spatio-temporal correlation module; The encoder is used to extract the features of the stacked image frames to obtain a stacked feature map; the stacked image frames are obtained by stacking the first image frame and the second image frame along the channel dimension; The temporal attention module is used to establish a global dependence relationship between the information of the historical memory unit and the information of the current input unit, and through an update unit, inject the information in the globally relevant historical memory unit into the current input unit, and at the same time store the globally relevant information in the current input unit into the historical memory unit as the historical memory unit at the next moment; the current input unit includes the stacked feature map, the stacked feature map is updated to an updated feature map through the update unit, the historical memory unit includes a first memory feature map and a first temporal feature map, and the historical memory unit at the next moment includes a second memory feature map and a second temporal feature map; The spatio-temporal correlation module is used to model the updated feature map / the second memory feature map into a first / second spatio-temporal feature map with spatial correlation respectively.

3. The method according to claim 2, wherein The step of establishing a global dependence relationship between the information of the historical memory unit and the information of the current input unit, and through an update unit, injecting the information in the globally relevant historical memory unit into the current input unit, and at the same time storing the globally relevant information in the current input unit into the historical memory unit as the historical memory unit at the next moment includes: Injecting the feature information in the stacked feature map and the feature information in the first memory feature map into the first temporal feature map to obtain a third temporal feature map; Determining a temporal attention feature vector according to the third temporal feature map; Determining a first feature vector according to the stacked feature map; Determining a second feature vector according to the first memory feature map; Determining an input feature vector based on temporal attention according to the first feature vector and the temporal attention feature vector; Determining a memory feature vector based on temporal attention according to the second feature vector and the temporal attention feature vector; Respectively adjust the input feature vector based on temporal attention and the memory feature vector based on temporal attention into corresponding first feature maps and second feature maps; Determine the updated feature map and the second memory feature map according to the first feature map and the second feature map; Update the third temporal feature map to the second temporal feature map according to the updated feature map and the second memory feature map.

4. The method according to claim 2, wherein Model the updated feature map into a first spatio-temporal feature map with spatial correlation, including: Slice the updated feature map in the channel dimension to obtain a first sub-feature map, a second sub-feature map, and a third sub-feature map; Respectively adjust the updated feature map, the first sub-feature map, the second sub-feature map, and the third sub-feature map into a third feature vector, a first sub-feature vector, a second sub-feature vector, and a third sub-feature vector; Calculate a first spatial correlation matrix between the first sub-feature map and the second sub-feature map according to the first sub-feature vector and the second sub-feature vector; Perform weighted processing on the third sub-feature vector using the first spatial correlation matrix to obtain a first spatially correlated feature vector; Determine a first spatio-temporal feature vector according to the first spatially correlated feature vector and the third feature vector; Adjust the first spatio-temporal feature vector into the first spatio-temporal feature map with spatial correlation; Model the second memory feature map into a second spatio-temporal feature map with spatial correlation, including: Slice the second memory feature map in the channel dimension to obtain a fourth sub-feature map, a fifth sub-feature map, and a sixth sub-feature map; Respectively adjust the second memory feature map, the fourth sub-feature map, the fifth sub-feature map, and the sixth sub-feature map into a fourth feature vector, a fourth sub-feature vector, a fifth sub-feature vector, and a sixth sub-feature vector; Calculate a second spatial correlation matrix between the fourth sub-feature map and the fifth sub-feature map according to the fourth sub-feature vector and the fifth sub-feature vector; Perform weighted processing on the sixth sub-feature vector using the second spatial correlation matrix to obtain a second spatially correlated feature vector; Determine a second spatio-temporal feature vector according to the second spatially correlated feature vector and the fourth feature vector; Adjust the second spatio-temporal feature vector into the second spatio-temporal feature map with spatial correlation.

5. The method according to claim 1, wherein Based on the first image frame, the second image frame, the first depth weight, and the first motion weight, use an error calculation module to calculate a first error, including: Input the first image frame and the second image frame into a pre-constructed depth estimation network. According to the first image frame and the first depth weight, obtain the first scene depth of the first image frame and the first encoder feature map of the first image frame. According to the second image frame and the first depth weight, obtain the second scene depth of the second image frame and the second encoder feature map of the second image frame; Input the first image frame and the second image frame into a pre-constructed camera motion network, and obtain the first relative pose between the first image frame and the second image frame according to the first image frame, the second image frame, and the first motion weight; Based on the first image frame, the second image frame, the first scene depth, the second scene depth, the first encoder feature map, the second encoder feature map, and the first relative pose, use the error calculation module to calculate the first error; Based on the first image frame, the second image frame, the second depth weight, and the second motion weight, use the error calculation module to calculate the second error, including: Input the first image frame and the second image frame into a pre-constructed depth estimation network. According to the first image frame and the second depth weight, obtain the third scene depth of the first image frame and the third encoder feature map of the first image frame. According to the second image frame and the second depth weight, obtain the fourth scene depth of the second image frame and the fourth encoder feature map of the second image frame; Input the first image frame and the second image frame into a pre-constructed camera motion network, and obtain the second relative pose between the first image frame and the second image frame according to the first image frame, the second image frame, and the second motion weight; Based on the first image frame, the second image frame, the third scene depth, the fourth scene depth, the third encoder feature map, the fourth encoder feature map, and the second relative pose, use the error calculation module to calculate the second error.

6. The method according to claim 5, wherein The step of determining the scene depth of the second image frame and the relative pose between the first image frame and the second image frame according to the first error and the second error includes: If the first error is greater than the second error, use the second scene depth as the scene depth of the second image frame, and use the first relative pose as the relative pose between the first image frame and the second image frame; If the first error is less than or equal to the second error, use the fourth scene depth as the scene depth of the second image frame, and use the second relative pose as the relative pose between the first image frame and the second image frame.

7. The method according to claim 5 or 6, characterized in that, The total error includes the first error and the second error; The total error is determined according to the image synthesis error, the scene depth structure consistency error, the feature perception loss error, and the smoothness loss error.

8. The method according to claim 7, wherein The step that the total error is determined according to the image synthesis error, the scene depth structure consistency error, the feature perception loss error, and the smoothness loss error includes: Obtain the first image coordinates of the first image frame and the second image coordinates of the second image frame; According to the first image coordinates, the camera internal parameters, and the first scene depth, determine the first world coordinates of the first image frame; According to the second image coordinates, the camera internal parameters, and the second scene depth, determine the second world coordinates of the second image frame; Affinely transform the first world coordinates of the first image frame to the second image frame panel to determine the third world coordinates after the affine transformation; Affinely transform the second world coordinates of the second image frame to the first image frame panel to determine the fourth world coordinates after the affine transformation; Project the third world coordinates and the fourth world coordinates onto a two-dimensional plane respectively to obtain the first scene depth after the affine transformation, the second scene depth after the affine transformation, and the corresponding first image coordinates after the affine transformation and the second image coordinates after the affine transformation; Determine the scene depth structure consistency error, the first depth structure inconsistency weight, and the second depth structure inconsistency weight according to the first scene depth, the second scene depth, the first image coordinates after the affine transformation, and the second image coordinates after the affine transformation; Determine the first camera flow consistency occlusion mask and the second camera flow consistency occlusion mask according to the first image coordinates of the first image frame, the second image coordinates after the affine transformation, the second image coordinates of the second image frame, and the first image coordinates after the affine transformation; Determine the image synthesis error according to the first depth structure inconsistency weight, the second depth structure inconsistency weight, the first camera flow consistency occlusion mask, and the second camera flow consistency occlusion mask; Determine the feature perception loss error according to the first image frame, the second image frame, the first image coordinates after the affine transformation, and the second image coordinates after the affine transformation; Determine the smooth loss error according to the first scene depth, the second scene depth, the first image frame, and the second image frame; Determine the total error according to the image synthesis error, the scene depth structure consistency error, the feature perception loss error, and the smooth loss error; 9. A scene depth inference device based on historical information, characterized in that The device includes: A first acquisition module, configured to acquire a first image frame and a second image frame of a to-be-detected image, where the first image frame is the image frame at the previous moment of the second image frame; A second acquisition module, configured to acquire a first depth weight of a pre-constructed depth estimation network and a first motion weight of a pre-constructed camera motion network; A first processing module, configured to calculate a first error based on the first image frame, the second image frame, the first depth weight, and the first motion weight by using an error calculation module; An update module, configured to use the first error as a guiding signal to jointly update the first depth weight of the depth estimation network and the first motion weight of the camera motion network to obtain a second depth weight and a second motion weight; A second processing module, configured to calculate a second error based on the first image frame, the second image frame, the second depth weight, and the second motion weight by using the error calculation module; A determination module, configured to determine the scene depth of the second image frame and the relative pose between the first image frame and the second image frame according to the first error and the second error; 10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method for scene depth inference based on historical information according to any one of claims 1-8.

Citation Information

Patent Citations

  • Pose estimation method and device and computer readable storage medium

    CN112330589A

  • Image-based pose determination method and device, storage medium and electronic equipment

    CN112509047A