Video Depth Estimation Based on Temporal Attention
By utilizing the time consistency between video frames in video depth estimation, calculating the time focus graph and applying it to feature graphs, the problem of insufficient estimation accuracy in the prior art is solved, and more accurate depth estimation is achieved, suitable for Bokeh effect simulation and other applications.
Patent Information
- Application Number
- CN202010698819.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-04-06
- Filing Date
- 2020-07-20
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2040-07-20
AI Technical Summary
The prior art fails to effectively utilize the time consistency between video frames in video depth estimation, resulting in insufficient estimation accuracy.
Through a time-focus-based method, using the temporal consistency between video frames, the time-focus map is calculated and applied to the feature map to generate a feature map with time-focus, combining motion compensation and depth estimation to improve the accuracy of depth estimation.
Improves the accuracy of video depth estimation, especially when processing images taken by non-professional photographers or devices with smaller lenses, can simulate Bokeh effects and be applied in fields such as 3D object reconstruction, virtual reality, automotive automation and surveillance cameras.
Smart Images

Figure CN112288790B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims priority and the benefit of U.S. Provisional Application No. 62 / 877,246, filed Jul. 22, 2019, entitled “Video Depth Estimation Based on Temporal Attention,” the entire content of which is incorporated herein by reference. Technical Field
[0003] Aspects of embodiments of the present disclosure generally relate to image depth estimation. Background Art
[0004] Recently, there has been a great deal of interest in estimating the true depth of elements in a captured scene. Accurate depth estimation allows the separation of foreground (near) and background (far) objects in the scene. Accurate foreground - background separation allows the processing of captured images to simulate effects such as the Bokeh effect, which refers to the soft out - of - focus blur of the background. The Bokeh effect can be created by using the correct settings in an expensive camera equipped with a fast lens and a large aperture, or it can be simulated by adjusting the camera to be closer to the object and the object to be farther from the background to create a shallow depth of field. Thus, accurate depth estimation can allow the processing of images from non - professional photographers or cameras with small lenses (such as mobile phone cameras) to obtain more aesthetically pleasing images with the Bokeh effect focused on the object. Other applications of accurate depth estimation can include 3D object reconstruction and virtual reality applications, where it is desired to change the background or objects and present them according to the desired perceived virtual reality. Other applications of accurate depth estimation from a captured scene can be automotive automation, surveillance cameras, and autonomous driving applications, as well as enhancing security by improving object detection accuracy and estimating its distance from the camera.
[0005] The above information disclosed in the background art section is only for enhancing the understanding of the present disclosure, and thus it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0006] Aspects of embodiments of the present disclosure are directed to a video depth estimation system for video depth estimation based on temporal attention that utilizes the temporal consistency between frames of a video sequence and a method of using the system.
[0007] According to some embodiments of the present disclosure, a method for depth detection based on multiple video frames is provided. The method includes: receiving a plurality of input frames, the plurality of input frames including a first input frame, a second input frame, and a third input frame corresponding to different capture times respectively; convolving the first to third input frames to generate a first feature map, a second feature map, and a third feature map corresponding to different capture times; calculating a temporal attention map based on the first to third feature maps, the temporal attention map including a plurality of weights corresponding to different pairs of feature maps among the first to third feature maps, each weight in the plurality of weights indicating the similarity of the corresponding pair of feature maps; and applying the temporal attention map to the first to third feature maps to generate a feature map with temporal attention.
[0008] In some embodiments, the plurality of weights are based on learnable values.
[0009] In some embodiments, each weight A of the plurality of weights of the temporal attention map ij is represented as:
[0010]
[0011] where i and j are index values greater than zero, s is a learnable scaling factor, M r is a reshaped combined feature map based on the first to third feature maps, and c represents the number of channels in each of the first to third feature maps.
[0012] In some embodiments, applying the attention map includes calculating an element Y of the feature map with temporal attention i , as:
[0013]
[0014] where i is an index value greater than 0.
[0015] In some embodiments, the plurality of input frames are video frames of an input video sequence.
[0016] In some embodiments, the plurality of input frames are warped frames based on motion compensation of video frames.
[0017] In some embodiments, the method further includes: receiving a plurality of warped frames, including a first warped frame, a second warped frame, and a third warped frame; and spatially dividing each of the first to third warped frames into a plurality of patches, where the first input frame is a patch among the plurality of patches of the first warped frame, where the second input frame is a patch among the plurality of patches of the second warped frame, and where the third input frame is a patch among the plurality of patches of the third warped frame.
[0018] In some embodiments, the method further comprises: receiving a first video frame, a second video frame, and a third video frame, the first to third video frames being consecutive frames of a video sequence; compensating for motion between the first to third video frames based on optical flow to generate first to third input frames; and generating a depth map based on a feature map with temporal attention, the depth map including depth values of pixels of the second video frame.
[0019] In some embodiments, compensating for the motion comprises: determining optical flow of pixels of the second video frame based on pixels of the first and third video frames; and performing image warping on the first to third input frames based on the determined optical flow.
[0020] In some embodiments, the method further comprises: receiving a first video frame, a second video frame, and a third video frame, the first to third video frames being consecutive frames of a video sequence; generating a first depth map, a second depth map, and a third depth map based on the first to third video frames; compensating for motion between the first to third depth maps based on optical flow to generate first to third input frames; and convolving a feature map with temporal attention to generate a depth map, the depth map including depth values of pixels of the second video frame.
[0021] In some embodiments, the first to third input frames are warped depth maps corresponding to the first to third depth maps.
[0022] In some embodiments, generating the first to third depth maps comprises: generating a first depth map based on the first video frame; generating a second depth map based on the second video frame; and generating a third depth map based on the third video frame.
[0023] According to some embodiments of the present disclosure, there is provided a method for depth detection based on multiple video frames, the method comprising: receiving a plurality of warped frames, including a first warped frame, a second warped frame, and a third warped frame corresponding to different capture times; dividing each of the first to third warped frames into a plurality of patches including a first patch; receiving a plurality of input frames including a first input frame, a second input frame, and a third input frame; convolving the first patch of the first warped frame, the first patch of the second warped frame, and the first patch of the third warped frame to generate a first feature map, a second feature map, and a third feature map corresponding to different capture times; calculating a temporal attention map based on the first to third feature maps, the temporal attention map including a plurality of weights corresponding to different pairs of feature maps among the first to third feature maps, each weight in the plurality of weights indicating the similarity of the corresponding pair of feature maps; and applying the temporal attention map to the first to third feature maps to generate a feature map with temporal attention.
[0024] In some embodiments, the plurality of warped frames are motion-compensated video frames.
[0025] In some embodiments, the multiple warped frames are depth maps that are motion-compensated for multiple input video frames of a video sequence.
[0026] In some embodiments, the method further includes: receiving a first video frame, a second video frame, and a third video frame, the first through third video frames being consecutive frames of a video sequence; compensating for motion between the first through third video frames based on optical flow to generate first through third warped frames; and generating a depth map based on a feature map with temporal attention, the depth map including depth values of pixels of the second video frame.
[0027] In some embodiments, compensating for the motion includes: determining an optical flow of pixels of the second video frame based on pixels of the first and third input frames; and performing image warping on the first through third video frames based on the determined optical flow.
[0028] In some embodiments, the method further includes: receiving a first video frame, a second video frame, and a third video frame, the first through third video frames being consecutive frames of a video sequence; generating a first depth map, a second depth map, and a third depth map based on the first through third video frames; compensating for motion between the first through third depth maps based on optical flow to generate first through third input frames; and convolving a feature map with temporal attention to generate a depth map, the depth map including depth values of pixels of the second video frame.
[0029] In some embodiments, the first through third input frames are warped depth maps corresponding to the first through third depth maps.
[0030] According to some embodiments of the present disclosure, there is provided a system for depth detection based on multiple video frames, the system including: a processor; and a processor-local processor memory, wherein instructions are stored on the processor memory, and when executed by the processor, the instructions cause the processor to perform: receiving a plurality of input frames, the plurality of input frames including a first input frame, a second input frame, and a third input frame respectively corresponding to different capture times; convolving the first through third input frames to generate first, second, and third feature maps corresponding to different capture times; calculating a temporal attention map based on the first through third feature maps, the temporal attention map including a plurality of weights corresponding to different pairs of the first through third feature maps, each weight in the plurality of weights indicating a similarity of a corresponding pair of feature maps; and applying the temporal attention map to the first through third feature maps to generate a feature map with temporal attention. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] When considered in conjunction with the accompanying drawings, the present disclosure and many of its attendant features and aspects will become more readily understood, where like reference numerals denote like components, where:
[0032] Figure 1Shows a subsystem of a depth estimation system according to some embodiments of the present disclosure;
[0033] Figure 2 A-2D provides an RGB visualization of the operation of the temporal attention subsystem with respect to a reference video frame of an input video sequence according to some example embodiments of the present disclosure;
[0034] Figure 2 E-2H provides an RGB visualization of the operation of the temporal attention subsystem associated with different reference video frames of an input video sequence according to some example embodiments of the present disclosure;
[0035] Figure 3 Shows a subsystem of a depth estimation system according to some other embodiments of the present disclosure;
[0036] Figures 4A - 4B Shows two different methods for implementing the temporal attention subsystem according to some embodiments of the present disclosure; and
[0037] Figure 5 Is a block diagram illustration of a temporal attention scaler according to some embodiments of the present disclosure. Detailed Description
[0038] The detailed description set forth below is intended as a description of example embodiments of systems and methods for video depth estimation provided in accordance with the present disclosure, and is not intended to represent the only form in which the present disclosure may be constructed or utilized. The description sets forth the features of the present disclosure in connection with the illustrated embodiments. However, it is to be understood that the same or equivalent functions and structures may be achieved by different embodiments, which are also intended to be encompassed within the scope of the present disclosure. As elsewhere shown herein, like element numbers are intended to represent like elements or features.
[0039] Some embodiments of the present disclosure are directed to a video depth estimation system and a method of using the system for video depth estimation based on temporal attention that utilizes the temporal consistency between frames of a video sequence. Currently, depth estimation methods using an input video do not consider temporal consistency when estimating depth. Although some methods of the related art can utilize a video sequence during the training process, the prediction process is based on a single frame. That is, when estimating the depth of frame t, the information of frame t-1 or frame t+1 is not used. Since the temporal consistency between frames is ignored, this limits the accuracy of such methods of the related art.
[0040] According to some embodiments, a video depth estimation system (also referred to as a depth estimation system) is capable of estimating the real-world depth of elements in a video sequence captured by a single camera. In some embodiments, the depth estimation system includes three subsystems, a motion compensator, a temporal attention subsystem, and a depth estimator. By arranging these three subsystems in different orders, the depth estimation system according to some embodiments exploits temporal consistency in the RGB (Red, Green, and Blue) domain, or exploits temporal consistency in the depth domain according to some other embodiments.
[0041] Figure 1 Shown are the subsystems of a depth estimation system 1 according to some embodiments of the present disclosure.
[0042] Reference Figure 1 , a depth estimation system 1 according to some embodiments includes a motion compensator 100, a temporal attention subsystem 200, and a depth estimator 300. The motion compensator 100 receives a plurality of video frames 10, including a first video frame 11, a second video frame 12 (also referred to as a reference video frame), and a third video frame 13 that represent consecutive frames of a video sequence (e.g., successive frames).
[0043] In some embodiments, the motion compensator 100 is configured to compensate for pixel motion between the first through third video frames 11 - 13 based on optical flow and generate first through third input frames 121 - 123 (e.g., first through third warped frames). The motion compensator 100 can align the temporal consistency between consecutive frames (e.g., adjacent frames). The motion compensator 100 can include a spatio-temporal converter network 110 and an image warper 120. In some examples, the spatio-temporal converter network 110 can determine the optical flow (e.g., motion vectors) of pixels of consecutive frames, and generate a first optical flow map 111 indicating the optical flow of pixels from the first video frame 11 to the second video frame 12, and generate a second optical flow map 112 indicating the optical flow of pixels from the third video frame 13 to the second video frame 12. The image warper 120 utilizes the first and second optical flow maps 111 and 112 to warp the input frames 11 and 13 and generate first and third warped frames 121 and 123 (e.g., first and third RGB frames) that attempt to compensate for the motion of regions (i.e., pixels) of the input frames 11 and 13. The warped frame 122 can be the same as the second video frame 12 (e.g., reference frame). Changes in camera angle or perspective, occlusion, objects moving out of the frame, etc. may result in inconsistencies in the warped frames 121 - 123. If the warped frames 121 - 123 are directly fed to the depth estimator 300, such inconsistencies will confound the depth estimation. However, the temporal attention subsystem 200 can address this issue by extracting and emphasizing the consistent information between the motion-compensated warped frames 121 - 123.
[0044] As used herein, consistent information means that the characteristics (e.g., appearance, structure) of the same object are the same in consecutive (e.g., adjacent) frames. For example, when the motion compensator 100 correctly estimates the motion of a moving car in consecutive frames, the shape and color of the car appearing in consecutive (e.g., adjacent) warped frames may be similar. Consistency can be measured by the difference between the input feature map of the temporal attention subsystem 200 and the output feature map 292 of the temporal attention subsystem 200.
[0045] In some embodiments, the temporal attention subsystem 200 identifies which regions of the reference frame (e.g., the second / central video frame 12) are more important and should be given greater attention. In some examples, the temporal attention subsystem 200 identifies the differences between its input frames (e.g., warped frames 121 - 123) and assigns weights / confidence values to each pixel of the frames based on temporal consistency. For example, when a region changes from one frame to the next, the confidence level of the pixels in that region may become lower. The weights / confidence values of the pixels together form a temporal attention map, and the temporal attention subsystem 200 uses this temporal attention map to re - weight the frames it receives (e.g., warped frames 121 - 123).
[0046] According to some embodiments, the depth estimator 300 extracts the depth (depth map 20) of the reference frame (e.g., the second / central video frame 12) based on the output feature map 292 of the temporal attention subsystem 200.
[0047] Figure 2 A - 2D provides an RGB visualization of the operation of the temporal attention subsystem 200 with respect to a reference video frame of an input video sequence according to some example embodiments of the present disclosure. Figure 2 E - 2H provides an RGB visualization of the operation of the temporal attention subsystem 200 with respect to different reference video frames of an input video sequence according to some example embodiments of the present disclosure.
[0048] Figure 2 A and 2E show the reference frame 30 of the input video sequence of the temporal attention subsystem 200, Figure 2 B - 2D and 2E - 2H show the corresponding attention maps visualized in the B channel, G channel, and R channel. The temporal attention weight map is shown as the difference between the input and output of the temporal attention subsystem 200. In Figure 2 B - 2D, brighter colors indicate greater differences, corresponding to motion inconsistency. For example, if a pixel in the output of the temporal attention subsystem 200 is the same as the input, the difference map for that pixel will be 0 (shown as black). As Figure 2 shown in B - 2D, the attention is concentrated on the car because the car is the most important moving object. Due to the difficulty in estimating the motion of the leaves, little attention is given to the leaves. InFigure 2 In E-2H, the focus is on all the main areas with movement. Compared with Figure 2 A, Figure 2 the illumination of the reference frame in E is more complex (e.g., see the shadows), and the objects are closer to the camera. Therefore, Figure 2 the temporal consistency shown in F-2H is more complex. In addition, compared with Figure 2 the G channel and the R channel in G and 2H, Figure 2 the attention map of the B channel in F has a higher value in the sky. The reason is that in Figure 2 the reference frame of E, the B channel prefers to consider blue moving objects, and the sky is the largest moving "object".
[0049] Figure 3 Subsystems of the depth estimation system 1-1 according to some other embodiments of the present disclosure are shown. Except for the arrangement order of the motion compensator 100-1, the temporal attention subsystem 200-1, and the depth estimator 300-1, Figure 3 the depth estimation system 1-1 is substantially the same as Figure 1 the depth estimation system.
[0050] Referring to Figure 3 , according to some embodiments, the depth estimator 300-1 receives a plurality of video frames including consecutive video frames 11-13 from a video sequence, and uses a frame-by-frame depth estimation method (e.g., single image depth estimation, SIDE), and generates first to third depth maps 311, 312, and 313 corresponding to the first to third video frames 11-13 respectively.
[0051] In some embodiments, the motion compensator 100-1 receives the first to third depth maps 311-313 from the depth estimator 300-1. Therefore, the motion compensator 100-1 is applied to the depth domain, rather than being applied to the temporal domain like Figure 1 the motion compensator 100. Otherwise, the motion compensator 100-1 can be the same as Figure 1The motion compensator 100 is the same or substantially similar to the motion compensator 100 of . In some embodiments, the spatial-temporal converter network 110 generates optical flow maps 111-1 and 112-1 based on the first to third depth maps 311-313, and the image warper 120 uses these optical flow maps to generate warped estimated depth maps 121-1, 122-1 and 123-1. According to some embodiments, the temporal attention subsystem 200-1 is then applied to extract consistent information from the warped estimated depth maps 121-1, 122-1 and 123-1, followed by a convolution layer 400 to obtain a final output, which is a depth map 20-1 corresponding to a reference frame (e.g., the second video frame 12). The convolution layer 400 can be used to convert the output feature map 292 from the temporal attention subsystem 200-1 into a depth map 20-1.
[0052] Based on the trade-off between the motion compensator 100 / 100-1 and the depth estimator 300 / 300-1, it is possible to use Figure 1 Depth estimation system 1 or Figure 3 Depth estimation system 1-1. The processing bottleneck of depth estimation system 1 may be motion compensator 100 in the RGB domain, which may be relatively difficult to perform because the appearance of an object changes with changes in illumination and color distortion between different video frames. On the other hand, the processing bottleneck of depth estimation system 1-1 may be depth estimator 300-1. Motion compensation in the depth domain may be easier than in the RGB domain because changes in illumination and color distortion can be ignored. Therefore, when motion compensator 100 is very accurate (for example, when the accuracy of optical flow estimation is above a set threshold), depth estimation system 1 can be used. When depth estimator 300-1 is very accurate (for example, when its accuracy is greater than a set threshold), depth estimation system 1-1 can be used. According to some examples, a device that relies on depth estimation (such as driver assistance or an autonomous vehicle) may include Figure 1 The depth estimation system 1 and Figure 3 The depth estimation system 1-1 uses both, and appropriately switches between the two systems based on the accuracy of optical flow estimation and depth estimation.
[0053] Figures 4A - 4B Two different methods for implementing the time-focused subsystem 200 / 200-1 according to some embodiments of the present disclosure are shown. Figures 4A - 4B In the example, for ease of explanation, the input frames 201-203 of the temporal attention subsystem 200 / 200-1 are illustrated as RGB video frames; however, the embodiments of the present description are not limited thereto, and the input frames 201-203 may be warped frames 121-123 (e.g., Figure 1 ) or distorted depth maps 121-1 to 123-1 (as shown Figure 3 shown).
[0054] refer toFigure 4A According to some embodiments, the temporal attention subsystem 200 includes a feature map extractor 210 configured to convert the input frames 201-203 into feature maps 211-213, and the feature maps 211-213 are processed by a temporal attention scaler 220 for reweighting based on temporal attention consistency. The feature map extractor 210 can be a convolutional layer that applies convolutional filters with learnable weights to the elements of the input frames 201-203. Here, the temporal attention subsystem 200 receives and processes the entire input frames 201-203. Adding the feature map extractor 210 before the temporal attention scaler 220 allows the temporal attention scaler 220 to more easily cooperate with deep learning frameworks in the related art. However, the embodiments of the present disclosure are not limited to using the feature map extractor 210 before the temporal attention scaler 220, and in some embodiments, the input frames 201-203 can be directly fed into the temporal attention scaler 220.
[0055] Reference Figure 4B In some embodiments, the temporal attention subsystem 200-1 further includes a patch extractor 230 that divides each of the input frames 201-203 into a plurality of patches or sub-parts. Each patch of the input frame is processed separately from the other patches of the input frame. For example, the patch extractor 230 can divide the input frames 201-203 into four patches, thus generating four sets of patches / sub-parts. The first set of patches (i.e., 201-1, 202-1, and 203-1) can include the first patch of each of the input frames 201-203, and the fourth set of patches (i.e., 201-4, 202-4, and 203-4) can include the fourth patch of each of the input frames 201-203. Each set of patches is processed by the feature map extractor 210 and the temporal attention scaler 220 respectively. Different sets of patches can be processed in parallel, as Figure 4B shown, or can be processed serially. The patch feature maps generated based on each set of patches (e.g., 211-1, 212-1, and 213-1) can be combined together to form a single feature map with temporal attention (i.e., 292-1, 292-2, 292-3, and 292-4).
[0056] Although Figure 4B four sets of patches are shown, the embodiments of the present disclosure are not limited thereto. For example, the patch extractor 230 can divide each input frame into any suitable number of patches. Figure 4BThe temporal attention subsystem 200-1 can improve depth estimation accuracy because each processed patch group contains visually more spatially correlated information than the visual information of the entire frame. For example, in a frame where the sky in the background occupies the top of the frame and includes a car moving on the road, the sky only complicates the depth estimation of the moving car and may introduce inaccuracies. However, separating the sky and the car into different patches can allow the depth estimation system 1 / 1-1 to provide a more accurate estimate of the depth of the car in the reference frame.
[0057] Figure 5 is a block diagram illustration of a temporal attention scaler 220 according to some embodiments of the present disclosure.
[0058] According to some embodiments, the temporal attention scaler 220 includes a concatenation block 250, a reshaping and transposing block 260, a temporal attention map generator 270, a multiplier 280, and a reshaping block 290.
[0059] The temporal attention scaler 220 receives first to third feature maps 211, 212, and 213 and concatenates them into a combined feature map 252. Each of the feature maps 211-213 may have the same size CxWxH, where C represents the number of channels (e.g., it may correspond to the color channels of red, green, and blue), and W and H represent the width and height of the feature maps 211-213, which are the same as the width and height dimensions of the input video frames 201-203 (e.g., see Figure 4A and 4B ). The combined feature map 252 may have a size of 3CxWxH. As described above, the feature maps can be generated from the warped frames 121-123 or from the warped depth maps 121-1 to 123-1.
[0060] The reshaping and transposing block 260 can reshape the combined feature map 252 from three-dimensional (3D) to two-dimensional (2D) to compute a first reshaped map 262 of size (3C)x(WH), and can transpose the first reshaped map 262 to compute a second reshaped map 264 of size (WH)x(3C). The temporal attention map generator 270 generates a temporal attention map 272 of size (3C)x(3C) based on the first reshaped map 262 and the second reshaped map 264. The temporal attention map 272 may be referred to as a similarity map and includes a plurality of weights A ij (where i and j are indices less than or equal to C, i.e., the number of channels), where each weight indicates the similarity level of the corresponding pair of feature maps. In other words, each weight A ij indicates the similarity between the frames generating channels i and j. When i and j are from the same frame, the weight A ijMeasure a kind of self - attention. For example, if C = 3, the size of the time - attention map is 9x9 (e.g., channels 1 - 3 belong to feature map 211, channels 4 - 6 belong to feature map 212, channels 7 - 9 belong to feature map 213). The weight A in the time - attention map 272 14 (i = 1, j = 4) represents the similarity between feature map 211 and feature map 212. A higher weight value can indicate a higher similarity between the corresponding feature maps. Each weight A of the multiple weights in the time - attention map 272 ij can be represented by Equation 1:
[0061]
[0062] where and are one - dimensional vectors of the first reshaping map 262. is the dot - product between two vectors, s is a learnable scaling factor, and i and j are index values greater than 0 and less than or equal to C.
[0063] The multiplier 280 performs matrix multiplication between the time - attention map 272 and the first reshaping map 262 to generate a second reshaping map 282 with dimensions (3C)x(WH), which is reshaped from 2D to 3D by the reshaping block 290 to generate a feature map of the time - attention 292 with size 3CxWxH. The element Y of the output feature map 292 with time - attention i can be represented by Equation 2:
[0064]
[0065] where Y i can represent a single - channel feature map with size WxH.
[0066] According to some examples, multiple components of the depth - estimation system 1 / 1 - 1, such as the motion compensator, the time - attention subsystem, and the depth estimator, can correspond to neural networks and / or deep neural networks (a deep neural network is a neural network with more than one hidden layer for deep - learning techniques), and the process of generating these components can include using training data and algorithms, such as the back - propagation algorithm, to train the deep neural network. The training can include providing a large number of input video frames and depth maps of the input video frames with measured depth values. Then, the neural network is trained based on this data to set the above - mentioned learnable values.
[0067] The operations performed by the depth - estimation system according to some embodiments can be executed by a processor that executes instructions stored in the processor memory. When executed by the processor, these instructions cause the processor to perform the above - mentioned operations regarding the depth - estimation system 1 / 1 - 1.
[0068] While embodiments of the depth estimation system 1 / 1-1 are disclosed as operating on a set of three input frames with the second frame as the reference frame, embodiments of the present disclosure are not limited thereto. For example, embodiments of the present disclosure may employ a set of an odd number of input frames (e.g., 5 or 7 input frames), where the center frame serves as the reference frame for which the depth estimation system generates a depth map. Additionally, such input frames may represent a sliding window of frames of a video sequence. In some examples, increasing the number of input frames (e.g., from 3 to 5) may improve depth estimation accuracy.
[0069] It should be understood that although the terms "first", "second", "third", etc. may be used herein to describe various elements, components, regions, layers, and / or parts, these elements, components, regions, layers, and / or parts should not be limited by these terms. These terms are used to distinguish one element, component, region, layer, or part from another element, component, region, layer, or part. Thus, the first element, component, region, layer, or part discussed above may be referred to as the second element, component, region, layer, or part without departing from the scope of the inventive concept.
[0070] The terms used herein are for the purpose of describing particular embodiments and are not intended to limit the inventive concept. As used herein, the singular forms "a" and "an" are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that when used in this specification, the terms "comprises" and / or "comprising" specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items. Additionally, when describing embodiments of the inventive concept, the use of "may" means "one or more embodiments of the inventive concept". Further, the term "exemplary" is intended to indicate an example or illustration.
[0071] As used herein, the terms "using", "being used", and "having used" may be considered synonymous with the terms "utilizing", "being utilized", and "having utilized", respectively.
[0072] The depth estimation system and / or any other relevant devices or components according to embodiments of the present disclosure described herein can be implemented by utilizing any suitable hardware, firmware (e.g., application specific integrated circuit), software, or any suitable combination of software, firmware, and hardware. For example, the various components of the depth estimation system can be formed on one integrated circuit (IC) chip or on separate IC chips. Additionally, the various components of the depth estimation system can be implemented on a flexible printed circuit film, tape carrier package (TCP), printed circuit board (PCB), or formed on the same substrate. Further, the various components of the depth estimation system can be processes or threads running on one or more processors in one or more computing devices, executing computer program instructions and interacting with other system components to perform the various functions described herein. The computer program instructions are stored in a memory, which can be implemented in a computing device using a standard storage device, such as random access memory (RAM). The computer program instructions can also be stored in other non-transitory computer-readable media, such as optical discs, flash drives, etc. Moreover, those skilled in the art should recognize that, without departing from the scope of the exemplary embodiments of the present disclosure, the functions of various computing devices can be combined or integrated into a single computing device, or the functions of a particular computing device can be distributed among one or more other computing devices.
[0073] Although the present disclosure has been described in detail with specific reference to its illustrative embodiments, the embodiments described herein are not intended to be exhaustive or to limit the scope of the present disclosure to the exact forms disclosed. Those skilled in the art of the fields and technologies related to the present disclosure will understand that changes and alterations can be made to the structures and methods of the described assemblies and operations without significantly departing from the principles and scope of the present disclosure, as set forth in the following claims and their equivalents.
Claims
1. A method for depth detection based on multiple video frames, the method comprising: Receiving a plurality of input frames, the plurality of input frames including a first input frame, a second input frame, and a third input frame respectively corresponding to different capture times; Convolving the first to third input frames to generate a first feature map, a second feature map, and a third feature map corresponding to different capture times; Calculating a temporal attention map based on the first to third feature maps, the temporal attention map including a plurality of weights corresponding to different pairs of feature maps among the first to third feature maps, each weight in the plurality of weights indicating the similarity of the corresponding pair of feature maps; And Applying the temporal attention map to the first to third feature maps to generate a feature map with temporal attention.
2. The method according to claim 1, wherein the plurality of weights are based on learnable values.
3. The method according to claim 1, wherein each weight A among the plurality of weights of the time attention graph ij is represented as: where i and j are index values greater than zero, s is a learnable scale factor, M r is a reshaped combined feature map based on the first to third feature maps, and c represents the number of channels in each of the first to third feature maps.
4. The method according to claim 3, wherein, The application time attention graph includes calculating an element Y of a feature graph with time attention i , as follows: Where i is an index value greater than 0.
5. The method according to claim 1, wherein, The plurality of input frames are video frames of an input video sequence.
6. The method according to claim 1, wherein the plurality of input frames are motion-compensated warped frames based on video frames.
7. The method according to claim 1, further comprising: Receiving a plurality of warped frames including a first warped frame, a second warped frame, and a third warped frame; and Spatially dividing each of the first to third warped frames into a plurality of patches, Among them, The first input frame is a patch among the plurality of patches of the first warped frame, Wherein, the second input frame is a patch among the plurality of patches of the second warped frame, and Wherein, the third input frame is a patch among the plurality of patches of the third warped frame.
8. The method according to claim 1, further comprising: Receiving a first video frame, a second video frame, and a third video frame, the first to third video frames being consecutive frames of a video sequence; Compensating for the motion between the first to third video frames based on optical flow to generate the first to third input frames; and Generating a depth map based on the feature map with temporal attention, the depth map including depth values of pixels of the second video frame.
9. The method according to claim 8, wherein, The compensation for the motion includes: Determining the optical flow of pixels of the second video frame based on the pixels of the first and third video frames; And Performing image warping on the first to third input frames based on the determined optical flow.
10. The method according to claim 1, further comprising: Receiving a first video frame, a second video frame, and a third video frame, the first to third video frames being consecutive frames of a video sequence; Generating a first depth map, a second depth map, and a third depth map based on the first to third video frames; Compensating for the motion between the first to third depth maps based on optical flow to generate the first to third input frames; and Convolving the feature map with temporal attention to generate a depth map, the depth map including depth values of pixels of the second video frame.
11. The method according to claim 10, wherein, The first to third input frames are warped depth maps corresponding to the first to third depth maps.
12. The method according to claim 10, wherein, Generating the first to third depth maps includes: Generating a first depth map based on the first video frame; Generating a second depth map based on the second video frame; And Generating a third depth map based on the third video frame.
13. A method for depth detection based on multiple video frames, the method comprising: Receive a plurality of warped frames, the plurality of warped frames including a first warped frame, a second warped frame, and a third warped frame corresponding to different capture times; Divide each of the first to third warped frames into a plurality of patches including a first patch; Convolve the first patch of the first warped frame, the first patch of the second warped frame, and the first patch of the third warped frame to generate a first feature map, a second feature map, and a third feature map corresponding to different capture times; Calculate a temporal attention map based on the first to third feature maps, the temporal attention map including a plurality of weights corresponding to different pairs of feature maps among the first to third feature maps, each weight among the plurality of weights indicating the similarity of the corresponding pair of feature maps; and Apply the temporal attention map to the first to third feature maps to generate a feature map with temporal attention.
14. The method according to claim 13, wherein The plurality of warped frames are motion-compensated video frames.
15. The method according to claim 13, wherein The plurality of warped frames are motion-compensated depth maps corresponding to a plurality of input video frames of a video sequence.
16. The method according to claim 13, further comprising: Receive a first video frame, a second video frame, and a third video frame, the first to third video frames being consecutive frames of a video sequence; Compensate for the motion between the first to third video frames based on optical flow to generate the first to third warped frames; and Generate a depth map based on the feature map with temporal attention, the depth map including depth values of pixels of the second video frame.
17. The method according to claim 16, wherein The compensation for the motion includes: Determine the optical flow of the pixels of the second video frame based on the pixels of the first and third video frames; and Perform image warping on the first to third video frames based on the determined optical flow.
18. The method according to claim 13, further comprising: Receive a first video frame, a second video frame, and a third video frame, the first to third video frames being consecutive frames of a video sequence; Generate a first depth map, a second depth map, and a third depth map based on the first to third video frames; Compensate for the motion between the first to third depth maps based on optical flow to generate the first to third warped frames; and Convolve the feature map with temporal attention to generate a depth map, the depth map including depth values of pixels of the second video frame.
19. The method according to claim 18, wherein the first to third warped frames are warped depth maps corresponding to the first to third depth maps.
20. A system for depth detection based on a plurality of video frames, the system comprising: A processor; and Processor-local processor memory, wherein, Instructions stored in the processor memory, which when executed by the processor, cause the processor to perform: Receive a plurality of input frames, the plurality of input frames including a first input frame, a second input frame, and a third input frame corresponding to different capture times respectively; Convolve the first to third input frames to generate a first feature map, a second feature map, and a third feature map corresponding to different capture times; Calculate a temporal attention map based on the first to third feature maps, the temporal attention map including a plurality of weights corresponding to different pairs of feature maps among the first to third feature maps, each weight among the plurality of weights indicating the similarity of the corresponding pair of feature maps; and Apply the temporal attention map to the first to third feature maps to generate a feature map with temporal attention.