Method and apparatus for depth estimation of an image

A neural network-based method estimates depth information from monocular images by outputting multiple statistical values and reliability indicators, enhancing accuracy and efficiency in depth estimation.

JP7687612B2Active Publication Date: 2025-06-03SAMSUNG ELECTRONICS CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2021086601
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-07-29
Filing Date
2021-05-24
Publication Date
2025-06-03
Estimated Expiration
2041-05-24

AI Technical Summary

Technical Problem

Existing depth estimation technologies require multiple images and complex alignment processes, limiting their efficiency in estimating depth information from a single monocular image.

Method used

A neural network is designed to output multiple statistical values related to depth for each pixel in an image, allowing for depth estimation based on these values and their reliability indicators.

Benefits of technology

This approach improves the accuracy and efficiency of depth estimation by providing reliable depth information for each pixel in a monocular image, even in complex scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007687612000008
    Figure 0007687612000008
  • Figure 0007687612000009
    Figure 0007687612000009
  • Figure 0007687612000010
    Figure 0007687612000010
Patent Text Reader

Abstract

To provide a method and a device for estimating a depth of a video.SOLUTION: A method for estimating a depth of a video comprises the steps of: acquiring a first statistic value related to a depth for each pixel in an input video, on the basis of a first channel of output data which is acquired by applying the input video to a neural network; acquiring a second statistic value related to the depth for each pixel in the input video, on the basis of a second channel of the output data; and estimating depth information of each pixel in the input video, on the basis of the first statistic value and the second statistic value. The neural network in one embodiment learns depth information in learning data, using probability distribution related to the depth of each pixel in the video based on the first statistic value and the second statistic value which are acquired corresponding to a known video.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] It relates to a method and an apparatus for depth estimation of an image.

Background Art

[0002] Depth information of an image is information regarding the distance between a camera and an object, and there is an increasing interest in a method of inferring a geometric structure (e.g., the position of a vanishing point, a horizon boundary, etc.) within a scene using the depth information of the image. Understanding such a geometric structure is used for understanding a scene such as grasping the exact position of an object or the three-dimensional relationship between objects, and is utilized in various fields such as the production of 3D images, robot vision, and human computer interface.

[0003] A person can identify whether an object is at a long distance or a short distance through parallax, which is the displacement of an object viewed with both eyes. Based on such a principle, depth information of an object within an image can be estimated through the parallax of two two-dimensional images captured by two cameras with different positions on the x-axis (horizontal axis). That is, depth information of an image can be obtained based on the geometric relationship between two images captured by two cameras at different positions. However, the depth estimation technology using stereo images requires two images and requires operations for aligning the two images. Here, various technologies for estimating depth information from an image have been developed, and technologies for estimating depth information from a two-dimensional monocular image captured by one camera have been studied.

Summary of the Invention

Problems to be Solved by the Invention

[0004] An embodiment provides a technique for estimating the depth of an image based on a neural network learned to output two or more statistical values regarding depth for each pixel within the image.

[0005] Embodiments provide a technique for designing an objective function such that two or more output data of a neural network have statistical characteristics, and training a neural network to output two or more statistical values related to depth for each pixel in a video.

Means for Solving the Problem

[0006] The neural network device according to one aspect includes a camera that captures an input video, applies the input video to a neural network, obtains a first statistical value related to the depth of a pixel in the input video, obtains another first statistical value related to the depth of another pixel in the input video, obtains a second statistical value related to the depth of the pixel in the input video, obtains another second statistical value related to the depth of the other pixel in the input video, estimates the depth information of the pixel in the input video based on the first statistical value and the second statistical value, and estimates the depth information of the other pixel in the input video based on the other first statistical value and the other second statistical value, and at least one processor.

[0007] The first statistical value and the other first statistical value may correspond to a first type of statistical value.

[0008] The second statistical value and the other second statistical value may correspond to a second type of statistical value.

[0009] The second statistical value may indicate the reliability with respect to the first statistical value, and the other second statistical value may indicate the reliability with respect to the other first statistical value.

[0010] The second statistical value may correspond to the standard deviation or variance with respect to the depth value of the pixel, and the other second statistical value may correspond to the standard deviation or variance with respect to the depth value of the other pixel.

[0011] The processor can selectively correct the first statistical value according to the reliability indicated by the second statistical value, and selectively correct the other first statistical value according to the reliability indicated by the other second statistical value.

[0012] The processor can generate an input video reconstructed from the reference video based on the first statistical value, the second statistical value, the other first statistical value, and the other second statistical value.

[0013] Based on the comparison result between the generated and reconstructed input video and the input video, the neural network can be trained.

[0014] The first statistical value and the other first statistical value are obtained from the output video corresponding to the first channel among the output channels of the neural network, and the second statistical value and the other second statistical value can be obtained from the output video corresponding to the second channel among the output channels of the neural network.

[0015] A neural network method according to an embodiment includes: obtaining a first statistical value related to depth for each pixel in the input video based on a first channel of output data obtained by applying the input video to a neural network; obtaining a second statistical value related to depth for each pixel in the input video based on a second channel of the output data; and estimating depth information of each pixel in the input video based on the first statistical value and the second statistical value.

[0016] The first statistical value can include an average of depth values based on the probability of the depth value of each pixel in the input video, and the second statistical value can include a variance or standard deviation of the depth value based on the probability of the depth value of each pixel in the input video.

[0017] The step of estimating the depth information may include a step of determining the reliability of the first statistical value based on the second statistical value, and a step of determining whether to adopt the first statistical value of the corresponding pixel as depth information based on the reliability of each pixel.

[0018] The step of determining the reliability may include a step of reducing the reliability as the second statistical value increases, and a step of increasing the reliability as the second statistical value decreases.

[0019] The step of estimating the depth information may include a step of obtaining a probability distribution based on the first statistical value and the second statistical value, and a step of performing random sampling for correcting the first statistical value based on the probability distribution.

[0020] The step of performing the random sampling may include a step of selectively performing random sampling based on the reliability of the first statistical value indicated by the second statistical value.

[0021] The neural network method according to an embodiment may further include a step of capturing the input video using a camera of a terminal, and a step of generating 3D position information corresponding to the terminal and peripheral objects included in the input video using the depth information.

[0022] The input video may include at least one of a monocular video and a stereo video.

[0023] A neural network method according to an embodiment includes: applying an input video to a neural network, and obtaining a first statistical value and a second statistical value related to depth for each pixel in the input video; obtaining a probability distribution related to the depth of each pixel in the input video based on the first statistical value and the second statistical value; and training the neural network using an objective function based on the probability distribution related to the depth of each pixel in the input video and predetermined depth information of each pixel in the input video.

[0024] The first statistical value includes an average of depth values based on the probability of the depth value of each pixel in the input video, and the second statistical value can include a variance or a standard deviation of the depth values based on the probability of the depth value of each pixel in the input video.

[0025] The step of obtaining a first statistical value and a second statistical value related to depth for each pixel in the input video can include: obtaining a first statistical value related to depth for each pixel in the input video based on a first channel of output data obtained by applying the input video to a neural network; and obtaining a second statistical value related to depth for each pixel in the input video based on a second channel of the output data.

[0026] The step of obtaining a first statistical value and a second statistical value related to depth for each pixel in the input video can include: applying a feature map extracted from the input video to a first network in the neural network to obtain the first statistical value; and applying the feature map to a second network in the neural network to obtain the second statistical value.

[0027] The step of training the neural network may include: estimating the depth information of each pixel in the input video based on the probability distribution of the depth of each pixel in the input video and the reference video corresponding to the input video; and training the neural network using an objective function based on the estimated depth information of each pixel in the input video and the predetermined depth information of each pixel in the input video.

[0028] The reference video may include at least one of a stereo image, a monocular video, and pose information of the camera that captured the monocular video.

[0029] The step of training the neural network may further include: obtaining the predetermined depth information of each pixel in the input video based on the reference video corresponding to the input video and the input video.

[0030] The step of training the neural network may include: synthesizing an output video based on the reference video corresponding to the input video and the probability distribution of the depth of each pixel in the input video; and training the neural network using an objective function based on the input video and the output video.

[0031] The step of synthesizing the output video may include: generating a second pixel in the output video corresponding to the first pixel based on the probability distribution of the depth of the first pixel in the input video and the pixel corresponding to the depth of the first pixel in the reference video corresponding to the input video.

[0032] The step of generating a second pixel in the output video corresponding to the first pixel may include: generating a second pixel in the output video corresponding to the first pixel by weighted summing the values of the pixels corresponding to the depth of the first pixel in the reference video corresponding to the input video based on the probability distribution of the depth of the first pixel.

[0033] The step of training the neural network can include the step of training the neural network based on an objective function that minimizes the difference between a first pixel in the input video and a second pixel in the output video corresponding to the first pixel.

[0034] The step of training the neural network can include the step of generating a mask corresponding to each pixel in the input video based on the first statistic and the second statistic, and the step of training the neural network using the mask and the objective function.

[0035] The step of training the neural network can include the step of optimizing parameters for each layer of the neural network using the objective function.

[0036] The input video can include at least one of a monocular video and a stereo video.

[0037] A neural network device according to an embodiment includes at least one processor that obtains a first statistic regarding depth for each pixel in the input video based on a first channel of output data obtained by applying the input video to a neural network, obtains a second statistic regarding depth for each pixel in the input video based on a second channel of the output data, and estimates depth information for each pixel in the input video based on the first statistic and the second statistic.

[0038] The first statistic can include an average of depth values based on the probability of the depth value of each pixel in the input video, and the second statistic can include a variance or a standard deviation of depth values based on the probability of the depth value of each pixel in the input video.

[0039] In estimating the depth information, the processor determines the reliability of the first statistical value based on the second statistical value, and can determine whether to adopt the first statistical value of the corresponding pixel as depth information based on the reliability of each pixel.

[0040] The neural network device further includes a camera that captures the input video, and the processor can generate 3D position information corresponding to at least one of the device of the camera and the surrounding objects included in the input video by using the depth information.

[0041] The input video can include at least one of a captured monocular video and a captured stereo video.

[0042] The processor trains the neural network based on the probability distribution regarding the depth of each pixel in the first video within the learning data, and the probability distribution can be obtained based on the first statistical value and the second statistical value regarding the depth of each pixel in the first video.

[0043] In obtaining the first statistical value, the processor obtains the first statistical value for each pixel in the input video from the first channel, and in obtaining the second statistical value, the processor can obtain the second statistical value for each pixel in the input video from the second channel.

Advantages of the Invention

[0044] By using the depth estimation of the video using the estimation model for two or more statistical values output from the neural network according to the embodiment and the learning process of the neural network in the estimation model, it is possible to have the effects of improving the accuracy of depth estimation and the learning efficiency.

Brief Description of the Drawings

[0045]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

DETAILED DESCRIPTION OF THE INVENTION

[0046] The embodiments described below can be variously modified. The scope of the patent application is not limited or restricted by such embodiments. The same reference numerals presented in each drawing indicate the same members. The specific structural or functional descriptions disclosed in this specification are merely exemplified for the purpose of explaining the embodiments, and the embodiments can be implemented in various different forms, and the present invention is not limited to the embodiments described in this specification.

[0047] The terms used in this specification are only for the purpose of describing specific embodiments and are not intended to limit the present invention. Singular expressions include plural expressions unless the context clearly dictates otherwise. In this specification, terms such as "comprising" or "having" indicate the presence of the features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and should not be construed as precluding the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0048] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Commonly used predefined terms should be interpreted as having a meaning consistent with their meaning in the context of the related art, and should not be interpreted in an idealized or overly formal sense unless clearly defined herein.

[0049] Also, when describing with reference to the drawings, the same components are given the same reference numerals regardless of the drawing reference signs, and duplicate descriptions thereof are omitted. When it is determined that a detailed description of a known technology related to the description of an embodiment would unnecessarily obscure the gist of the present invention, the detailed description thereof is omitted.

[0050] Also, in the description of the components of the embodiments, terms such as first, second, A, B, (a), (b), etc. may be used. Such terms are for distinguishing the components from other components, and the essence, order, or sequence of the corresponding components is not limited by such terms. When it is described that any component is "connected", "coupled", or "joined" to another component, it should be understood that the component may be directly connected or joined to the other component, but additional components may be "connected", "coupled", or "joined" between each component. (Depth Estimation Model) FIG. 1 is a diagram for explaining a model for estimating the depth of a video according to an embodiment.

[0051] Referring to FIG. 1, a model 110 for estimating the depth of a video according to an embodiment (hereinafter referred to as an estimation model) uses a pre-trained machine learning model such as a neural network to obtain from the input video 101 multiple types of information regarding the depth of the pixels included in the video, for example, the first statistical value 102 of the first type and the second statistical value 103 of the second type. In the following embodiments, as an example of the learning model, a neural network will be described. This is for convenience of explanation and the configuration of the learning model is not limited to a neural network.

[0052] The input video 101 according to an embodiment is a two-dimensional image (or video) captured by a camera and may include a monocular video captured by a monocular camera. A monocular video means a video captured by one camera. That is, the estimation model 110 according to an embodiment is a model that outputs depth information corresponding to each pixel in the video when a monocular video is input.

[0053] The depth of an image is a value based on the distance between the camera that captured the image and the object within the image, and the depth of the image can be indicated in terms of image pixels. For example, the depth of an image can be indicated in a depth map that includes a value related to depth for each pixel in the image. A depth map means one image or one channel of an image that contains information regarding the distance from the observation viewpoint to the object surface. For example, a depth map is a data structure of a size corresponding to the number of pixels in the image that stores values related to the depth corresponding to each pixel in the image, and is an image in which the values related to the depth corresponding to each pixel in the image are imaged in grayscale or the like. In a stereo system, the depth of an image is estimated based on the disparity, which is the positional difference of corresponding points between the image captured by the right camera (hereinafter referred to as the right image) and the image captured by the left camera (hereinafter referred to as the left image). That is, in a stereo system, the depth information of an image may include disparity information, and the depth value of a pixel in the image may be expressed like the corresponding disparity value. Also, in a stereo system, the probability regarding the depth of a pixel in the image is the probability regarding the corresponding disparity.

[0054] Referring back to FIG. 1, the information 102, 103 regarding the depth of the pixels included in the video, which is the output of the neural network according to one embodiment, includes statistical values based on the probabilities of the depth values that each pixel in the video has. For example, the first type of statistical value may include the average regarding the depth value of each pixel, and the first type of statistical value may include the variance (or, standard deviation). The first statistical value 102, which is one of the outputs of the neural network according to one embodiment, indicates information regarding the estimated depth value of each pixel, and the second statistical value 103, which is a further output with the first statistical value, indicates information regarding the reliability of the estimation regarding the depth of each pixel. For example, when the first statistical value output corresponding to a specific pixel is the average of the depth values based on the probability of the depth value of the corresponding pixel, the first statistical value can indicate the estimated depth value of the corresponding pixel. For example, when the second statistical value output corresponding to a specific pixel is the variance of the depth values based on the probability of the depth value of the corresponding pixel, the second statistical value can indicate information regarding the reliability of the estimation regarding the depth value of the corresponding pixel. In this case, the larger the variance, the lower the reliability is indicated, and the smaller the variance, the higher the reliability is indicated.

[0055] According to one embodiment, each output of the neural network is a video in which the statistical values corresponding to each pixel in the video are imaged in grayscale or the like. In other words, a grayscale image corresponding to the first statistical value of the pixels in the video and a grayscale image corresponding to the second statistical value of the pixels in the video can be generated in the output video of the neural network.

[0056] The estimation model 110 according to an embodiment can be implemented as a neural network that estimates the probabilistic characteristics of the depth corresponding to each pixel in the input video. According to an embodiment, the neural network may include an input layer, at least one hidden layer, and an output layer. Each layer of the neural network includes at least one node, and the relationship between the plurality of nodes is defined non-linearly. The neural network according to an embodiment is designed as a DNN (deep neural network), a CNN (convolutional neural network), an RNN (recurrent neural network), etc., and may include various artificial neural network structures. The neural network according to an embodiment outputs two or more probabilistic characteristics, and thus may include two or more output channels that output the respective probabilistic characteristics.

[0057] Hereinafter, the neural network according to an embodiment will be described by taking a DNN as an example.

[0058] FIG. 2 is a diagram illustrating the structure of a neural network according to an embodiment. FIG. 2 shows the structure of a neural network designed on a U-net base in an hourglass form. The neural network shown in FIG. 2 is merely an example of the structure of the neural network according to an embodiment, and the structure of the neural network according to an embodiment is not limited thereto, and may be designed in various structures.

[0059] Referring to FIG. 2, a neural network according to an embodiment is composed of at least one connected layer for extracting a feature map 201 from an input video, and includes a first network 210 for obtaining a first statistical value 202 from the feature map 201 and a second network 220 for obtaining a second statistical value 203. For example, the first network 210 and the second network 220 including the extracted feature map 201 and an input connection may operate in parallel with each other. FIG. 2 shows a case where the first network and the second network are included to output the first statistical value and the second statistical value in the neural network of the estimation model. However, the neural network according to an embodiment may be designed such that the output channel includes two channels. That is, the output channel according to an embodiment can be designed to include a first channel related to the first statistical value and a second channel related to the second statistical value. For example, any one of the illustrated first network 210 and second network 220 may be configured to output a multi-channel having a channel for outputting the first statistical value 202 and another channel for outputting the second statistical value 203.

[0060] Referring to FIG. 2, the layers included in the neural network according to an embodiment relate to specific functions and may be classified according to the functions. For example, the DS layer of the neural network is classified as a feature extractor, the Res-block (residual blocks or ResNet architecture) is classified as a feature analyzer, and the US layer is classified as a depth generator. However, classifying the layers according to the functions is for the convenience of explanation and does not limit the structure of the neural network according to an embodiment.

[0061] The neural network according to one embodiment is trained so that each of two or more output channels estimates a value having an appropriate meaning. In the objective function to be optimized for the neural network according to one embodiment, each output channel can be trained so that the meaning of each channel is reflected. For example, the neural network is trained based on the objective function until the result by the objective function of the neural network becomes equal to or higher than a threshold related to accuracy.

[0062] Hereinafter, an example will be described in the case where the output of the neural network according to one embodiment includes two channels regarding the first statistical value and the second statistical value, the first statistical value is the average, and the second statistical value is the variance.

[0063] A learning method for depth estimation of a video according to one embodiment includes a step of applying an input video to a neural network and obtaining a first statistical value and a second statistical value regarding depth for each pixel in the input video, a step of obtaining a probability distribution regarding the depth of each pixel in the input video based on the first statistical value and the second statistical value, and a step of training the neural network using an objective function based on the probability distribution regarding the depth of each pixel in the input video and the actual depth information (or labeled depth information) of each pixel in the input video.

[0064] The learning data for training the neural network according to one embodiment may include an input video and the actual depth information of the input video. According to one embodiment, the actual depth information of the input video is obtained from a reference video corresponding to the input video. The reference video according to one embodiment is a video corresponding to the input video in which the actual depth information of the input video is known. For example, when the input video is a video captured by a left camera in a stereo system, the video captured by the right camera becomes the reference video, and a video including camera shooting information (for example, camera pose) becomes the reference video. In other words, the learning data according to one embodiment may include an input video and the actual depth information of the input video, or an input video and a reference video in which the actual depth information of the input video is known.

[0065] A method for training a neural network according to an embodiment can calculate a gradient that minimizes the difference between the depth information estimated by the neural network and the actual depth information of the input video by utilizing a stereo video database or a video database storing training data, and perform an optimization process. According to an embodiment, based on the actual depth information and the estimated depth information of the input video, a loss of the neural network occurs using a pre-designed objective function, and a gradient is generated based on the generated loss. The generated gradient is adopted to optimize and train the parameters of the neural network, and the neural network can be trained by a method such as the steepest descent method.

[0066] According to an embodiment, the difference between the estimated depth information and the actual depth information of the input video may include the difference between the output video and the input video generated based on the depth information estimated by the neural network and the reference video. That is, a loss of the neural network can be generated based on the reference video and the output video. According to an embodiment, the loss based on the reference video and the output video can include the loss generated based on the features extracted from the input video and the features extracted from the reference video.

[0067] FIG. 3 is a diagram showing a process of generating an output video for training a neural network according to an embodiment.

[0068] Referring to FIG. 3, in a stereo image database storing training data, an image captured by the left camera can be used as an input left image, and an image captured by the right camera can be used as a right image to train a neural network (DNN). According to an embodiment, when the image captured by the left camera is used as the input video, the image captured by the right camera may be used as the reference video.

[0069] According to one embodiment, an input video (input left image) is applied to a neural network (DNN), and a first statistical value (mean) and a second statistical value (std.) related to depth are obtained for each pixel in the input video. An image generator according to one embodiment can generate an output video based on the probability distribution related to the depth of each pixel in the input video and a reference video. Here, the probability distribution related to the depth of each pixel in the input video is determined based on the first statistical value (mean) and the second statistical value (std.) as the distribution of the probability of what value the random variable (for example, the depth value of the corresponding pixel) has.

[0070] FIG. 4 is a diagram showing the probability distribution related to the depth of pixels in an input video according to one embodiment. More specifically, FIG. 4 shows a Gaussian probability distribution by the mean μ and variance σ of the depth values of the (x, y) pixels in the input video. Here, the mean μ and variance σ of the depth values of the x, y pixels are statistical values output by the neural network. That is, the probability distribution according to one embodiment is obtained based on the statistical values related to the depth of the pixels output by the neural network. The statistical values output by the neural network according to one embodiment are various, and the probability distributions obtained based on the statistical values according to one embodiment are also various. For example, the probability distribution according to one embodiment may include a Gaussian distribution, a Laplace distribution, etc. obtained based on the mean and variance of the depth.

[0071] Referring to FIG. 3 again, an output video (reconstructed left image) according to one embodiment is generated by calculating the mean of the pixel values in the reference video corresponding to each depth based on the probability distribution related to the depth of each pixel in the input video. In other words, the second pixel in the output video is generated by weighted-summing the values of the pixels corresponding to the depth of the first pixel in the reference video corresponding to the input video based on the probability distribution related to the depth of the first pixel in the input video.

[0072] For example, the output video is synthesized by the following mathematical formula (1).

[0073] [Number] Here,

[0074] [Number] is the pixel value of the pixel (x, y) in the output video, J(x - d, y) is the pixel value of the pixel in the reference video corresponding to each depth of the pixel (x, y) in the input video, and P x、y (d) means the probability that the disparity of the pixel (x, y) in the input video is d. As described above, the disparity of the pixel (x, y) corresponds to the depth of the pixel (x, y). That is, according to one embodiment, based on the average disparity (or depth) estimated at a specific pixel in the input video, the pixel values of the neighboring disparities are weighted-summed by a Gaussian distribution of the probability regarding the disparity, thereby generating the pixel in the corresponding output video. According to one embodiment, P x、y (d) can be calculated as in the mathematical formula (2) by a Gaussian distribution regarding the disparity.

[0075] [Number] Here, μ(x, y) is the average of the disparities of the pixel (x, y) in the input video, which is one of the outputs of the neural network, and σ(x, y) means the standard deviation of the disparities of the pixel (x, y) in the input video, which is one of the outputs of the neural network.

[0076] The loss of the neural network according to one embodiment is generated based on a predefined objective function. According to one embodiment, the predefined objective function is designed in various ways and may include, for example, the L2 loss function, the cross entropy objective function, the softmax activation function, and the like.

[0077] Referring to FIG. 3 again, the neural network according to one embodiment can be trained using an objective function designed based on the output video (reconstructed left image) and the input video (input left image). For example, when generating an output video by expressing a Gaussian distribution with the mean and variance, which are statistical characteristics of the estimated depth, the objective function to be optimized for the neural network learning according to one embodiment can be designed as shown in the following formula (3).

[0078]

Equation

[0079]

Equation

[0080] That is, the neural network according to one embodiment can be trained based on an objective function that reflects all of the two or more output data indicating the statistical characteristics related to the depth in terms of pixels.

[0081] The two or more outputs of the neural network according to one embodiment are used to estimate the depth information of the video. According to one embodiment, the second statistical value indicating the reliability of the estimation serves as a criterion for determining whether to use the first statistical value. More specifically, the estimation model according to one embodiment determines, based on a predetermined criterion, the reliability information indicated by the second statistical value, and determines whether to use the first statistical value for the estimated depth in terms of pixels. Also, according to one embodiment, the second statistical value indicating the reliability of the estimation serves as a criterion for determining whether to correct the first statistical value. For example, according to the second statistical value, if it is determined that the reliability is low, the estimation model according to one embodiment may perform processing in addition to the first statistical value in terms of pixels. In this case, the additional processing of the first statistical value may include processing for correcting the first statistical value based on the second statistical value. For example, the estimation model according to one embodiment can obtain a probability distribution based on the mean as the first statistical value and the variance as the second statistical value, and correct the first statistical value by performing random sampling according to the probability distribution.

[0082] FIG. 5 is a diagram showing the configuration of a model for estimating the depth information of a video based on the output data of a neural network according to one embodiment.

[0083] Referring to FIG. 5, an estimation model 510 according to one embodiment may include a processing model 530 that processes the output data of a neural network 520 including a first statistical value 502 and a second statistical value 503 to estimate the depth information 504 of the video. According to one embodiment, the processing model 530 may include a model that outputs the depth information 504 of the input video 501 based on the first statistical value 502 indicating information regarding the depth value estimated in terms of pixels and the second statistical value 503 indicating the reliability of the estimation in terms of pixels. According to one embodiment, the processing model 530 serves as a filter that determines the reliability of the first statistical value based on the second statistical value and selects the first statistical value in terms of pixels based on the determined reliability. According to a further embodiment, the processing model 530 may perform processing for correcting the first statistical value in terms of pixels based on the reliability.

[0084] When the second statistical value is the variance of the depth value for each pixel, if the variance of a specific pixel is greater than a predetermined reference value, the estimated reliability of the corresponding pixel is evaluated as low. That is, the processing model 530 according to one embodiment can determine the reliability of the first statistical value by decreasing the reliability as the second statistical value increases and increasing the reliability as the second statistical value decreases. According to one embodiment, when the reliability of the first statistical value corresponding to a specific pixel is evaluated as low, in estimating the depth of the video, the depth information estimated corresponding to the specific pixel may be excluded or corrected. The processing model 530 according to one embodiment can obtain a probability distribution based on the average as the first statistical value and the variance as the second statistical value, and correct the first statistical value by performing random sampling according to the obtained probability distribution. Or, when the reliability of a specific pixel is evaluated as low, it may be corrected by referring to the depth values of the surrounding pixels of the corresponding pixel.

[0085] The two or more output data of the neural network according to one embodiment are used to improve the learning efficiency and accuracy in the learning process of the estimation model. In the learning process of the depth estimation model according to one embodiment, by not using the first statistical value corresponding to a specific pixel as learning data based on the second statistical value, or by reaching the learning data after selectively applying post-processing, it is possible to prevent a chain of learning failures due to depth estimation errors. For example, according to the output result of the neural network, by processing the probability information of the depth for each pixel, the difference between the output video and the input video can be selectively back-propagated for each pixel. The difference between the output video and the input video may include, for example, the difference between the aforementioned left video and the reconstructed left image (where reconstruction means reconstruction based on the right video in the reference video during the learning process of the neural network).

[0086] Alternatively, in the learning process of the neural network according to an embodiment, a mask can be generated to perform learning on a pixel-by-pixel basis based on the second statistical value. According to one embodiment, based on the second statistical value, a binary mask corresponding to each pixel is generated, the mask is applied to the kernels in the neural network, and learning epochs may be performed by inactivating the kernels corresponding to some pixels. Alternatively, according to one embodiment, a weighted value mask corresponding to each pixel is generated based on the second statistical value, the weighted value mask is applied to the kernels in the neural network, and learning epochs may be performed by assigning weights to the kernels corresponding to each pixel.

[0087] In addition, in the learning process of the neural network according to an embodiment, based on the first statistical value and the second statistical value, the probability that an object in the input video is occluded by other objects can be calculated for each pixel of the input video. According to one embodiment, based on the probability of occlusion by other objects calculated for each pixel, a binary mask corresponding to each pixel or a weighted value mask having a value between 0 and 1 for each pixel may be generated. According to one embodiment, by applying the generated mask to the loss of the neural network, the accuracy of learning can be improved.

[0088] For example, in order to measure the reconstruction loss of a neural network, in the process of reconstructing (or synthesizing) the object corresponding to the k-th pixel in the input video captured by the left camera in the stereo video database stored in the training data, the pixel value of the reference video captured by the left camera separated only by disparity may be used with reference to the position of the k-th pixel. Here, the right camera and the left camera are separated, and there is a disparity between the video captured by the right camera and the video captured by the left camera. Here, if a plurality of objects overlap in the direction in which the camera capturing the reference video is facing, the object located behind among the plurality of objects may be blocked by the object located in front, causing incorrect loss propagation. Here, by calculating the probability that the object corresponding to each pixel is blocked by another object to generate a mask and applying the generated mask to the loss of the neural network, the accuracy of learning can be improved.

[0089] Using the distributions obtained from the first statistical value and the second statistical value of each pixel of the input video, the probability that the object corresponding to each pixel is blocked by another object can be calculated through the following formula (4).

[0090]

Equation

[0091]

Equation

[0092] That is, Equation (4) calculates the probability that the k-th pixel and the (k + j)-th pixel do not refer to the values at the same position in the reference video, and v k If the value is high, it can be said that the k-th pixel value is a condition that can be sufficiently reconstructed. A mask in binary form may be generated using the probability calculated through Equation (4), or a probability value having a value between 0 and 1 may be used as a weighted value.

[0093] That is, a plurality of data output by the neural network according to one embodiment are used to improve the learning efficiency and performance of the neural network and can contribute to the accurate estimation of the first statistical value.

[0094] In addition, the output of the neural network including the first statistical value and the second statistical value according to one embodiment may be variously applied to the technology of estimating and using the depth information of the video. For example, when the output of the neural network includes a first statistical value indicating information regarding the estimated depth value of a pixel and a second statistical value indicating information regarding the reliability of the estimation for the first statistical value of each pixel, by transmitting both the depth value and the reliability information to various algorithms that use the depth information of the video as an input, it is possible to contribute to the overall performance improvement of the algorithms.

[0095] FIG. 6 is a diagram showing an output video obtained by inputting an input video according to one embodiment into an estimation model.

[0096] Referring to FIG. 6, the output video according to one embodiment includes a video that images the average of the depth values (or disparities) of each pixel in the input video and a video that images the standard deviation of the depth values (or disparities). Through the video that images the standard deviation of the depth values, the reliability of the average of the estimated depth values can be evaluated. The brighter the output video regarding the standard deviation of the depth values, the greater the standard deviation, which means the lower the reliability. Also, the darker the output video regarding the standard deviation of the depth values, the smaller the standard deviation of the depth values, which means the higher the reliability.

[0097] Referring to FIG. 6, the area displayed in the video is usually an area where depth estimation is difficult, and it is confirmed that the area displayed in the output video regarding the standard deviation of the depth values is brightly shown. For example, usually in FIG. 6, the bright area of the video regarding the standard deviation means that the reliability of the estimated depth value is low. In other words, it means that it is difficult to estimate the depth value of the corresponding area by the normal approach method.

[0098] On the other hand, referring to the video (Estimated disparity mean) regarding the average of the depth values of scenes #1 to #3 shown in FIG. 6, it can be seen that the depth values estimated in the bright area of the video (Estimated disparity std.) regarding the standard deviation of the depth values, that is, the area where depth estimation is difficult (for example, complex region, fine-grained texture region, object with unknown bottom-line), well reflect the actual depth values.

[0099] That is, referring to the area shown in FIG. 6, it can be seen that the second statistical value output by the neural network is used as reliability information regarding the estimation. Therefore, by using the estimation model of the depth of the video with two or more statistical values in the neural network and in the learning process of the neural network in the estimation model, effects such as improvement in the accuracy and learning efficiency in depth estimation can be achieved.

[0100] FIG. 7 is a diagram showing an inference process of a model for estimating the depth of a video according to an embodiment.

[0101] Referring to FIG. 7, the estimation model according to an embodiment may include a neural network (DNN), and the input video may include a monocular input image. Although not shown in FIG. 7, the estimation model according to an embodiment may take as an input video a stereo video which is two or more videos that capture the same object at different positions. In this case, the estimation model according to an embodiment may be used to obtain depth information for at least one pixel in the stereo video where stereo matching is not performed, or for at least one pixel where stereo matching is not accurately performed.

[0102] Referring to FIG. 7, an estimation model according to an embodiment can extract a feature map by applying an input video to a feature extraction unit including at least one layer in a DNN. The extracted feature map passes through at least one layer constituting a feature analyzer in the DNN and is converted and processed into a form suitable for probability estimation regarding depth. The extracted feature map and / or the converted and processed feature map are applied to a depth probability generator including at least one layer in the DNN, and finally, a first statistical value (Mean of depth probability distribution) and a second statistical value (Variance of depth probability distribution) regarding the depth of each pixel in the video are output. FIG. 7 illustrates the feature extraction unit, the feature analyzer, and the depth probability generator as the configuration of the DNN. However, the feature extraction unit, the feature analyzer, and the depth probability generator may each be configured by different networks, or may be configured in a structure in which a plurality of network structures are combined.

[0103] An estimation model according to an embodiment can estimate the depth information of a video based on the first statistical value and the second statistical value. The depth information of a video according to an embodiment may be variously used in the field related to the 3D information of the video. For example, the depth information of the video may be used to generate a 3D video or to generate 3D position information corresponding to at least one of a terminal that captured the input video and peripheral objects included in the input video.

[0104] FIG. 8 is a diagram showing a learning process of a model for estimating the depth of a video according to an embodiment.

[0105] Referring to FIG. 8, the training data for training the estimation model according to one embodiment may include a stereo image. Although not shown in FIG. 8, the training data according to one embodiment may include a monocular video. When the reference video according to one embodiment includes a monocular video, it may further include camera pose information such as the shooting position and shooting angle of the camera that captured the video. In other words, when the training data includes a monocular video, the training data may further include information for obtaining the depth information of the frame images constituting the monocular video.

[0106] Referring to FIG. 8, the estimation model according to one embodiment can be trained to output a first statistical value and a second statistical value related to depth from the input video based on the training data.

[0107] FIG. 9 is a diagram showing the structure of an apparatus for estimating the depth of a video according to one embodiment.

[0108] Referring to FIG. 9, the apparatus for estimating the depth of a video according to one embodiment can output depth information regarding the input video using the above-described estimation model. The estimation apparatus according to one embodiment may include at least one processor that estimates the depth of the input video using a depth estimation model. The depth estimation model according to one embodiment may include a model that estimates the depth information of the input video using a neural network as described above.

[0109] That is, the model for estimating the depth of an image according to one embodiment is embodied in the form of a chip and installed in a device that uses 3D information of the image to estimate the depth of the image. For example, the estimation model embodied in an advanced form is installed in a monocular depth sensor for an automobile and used to provide services related to automobile driving, such as generating 3D position information with objects around the automobile based on the depth information estimated from a monocular image taken by the automobile. As another example, the estimation model embodied in the form of a chip is installed in AR glasses and used to provide AR-related services, such as generating 3D position information using the depth information estimated from a 2D input image.

[0110] As shown in FIG. 2, the depth estimation model according to one embodiment may be designed as a DNN. The multiple layers included in the DNN are involved in functions such as feature extraction, feature analysis, and depth probability generation executed by the DNN. The configurations of the DNN shown in FIGS. 7 to 9 show examples in which the multiple layers included in the DNN are classified according to functions, and the configuration of the DNN is not necessarily limited to this.

[0111] The method according to the embodiment is embodied in the form of program instructions implemented via various computer means and recorded on a computer-readable recording medium. The recording medium includes program instructions, data files, data structures, etc. alone or in combination. The recording medium and the program instructions may be specially designed and configured for the purpose of the present invention, or may be known and usable by those skilled in the art of computer software technology. Examples of computer-readable recording media include magnetic media such as hard disks, floppy (registered trademark) disks, and magnetic tapes, optical recording media such as CD-ROMs, DVDs, magneto-optical media such as floptical disks, and hardware devices specially configured to store and execute program instructions such as ROMs, RAMs, flash memories, etc. Examples of program instructions include not only machine language codes generated by compilers but also high-level language codes executed by a computer using an interpreter or the like. The hardware device may be configured to operate as one or more software modules to execute the operations shown in the present invention, and vice versa.

[0112] Software includes a computer program, code, instruction, or a combination of one or more of them, and can configure a processing device to operate as desired or command the processing device independently or in combination. Software and / or data can be permanently or temporarily embodied in any type of machine, component, physical device, virtual device, computer storage medium or device, or signal wave being transmitted in order to be interpreted by the processing device or provide instructions or data to the processing device. Software can be distributed on a computer system connected to a network and stored or executed in a distributed manner. Software and data can be stored in one or more computer-readable recording media.

[0113] As described above, the embodiments have been described with reference to the limited drawings. However, those with ordinary knowledge in the art can apply various technical modifications and variations based on the above description. For example, the described technology may be executed in an order different from the described method, and / or the components such as the described system, structure, device, circuit, etc. may be combined or assembled in a form different from the described method, and appropriate results can be achieved even if they are replaced or substituted by other components or equivalents.

[0114] Therefore, other embodiments, other implementations, and those equivalent to the claims also fall within the scope of the claims described below.

Description of Reference Numerals

[0115] 101: Input video 102: First statistical value 103: Second statistical value 110: Estimation model 201: Feature map 202: First statistical value 203: Second statistical value 210: First network 220: Second network 501: Input video 502: First statistical value 503: Second statistical value 504: Depth information 510: Estimation model 520: Neural network 530: Processing model

Claims

1. A camera for capturing an input video, applying the input video to a neural network, obtaining a first statistical value regarding the depth of pixels in the input video, obtaining another first statistical value regarding the depth of other pixels in the input video, obtaining a second statistical value regarding the depth of the pixels in the input video, obtaining another second statistical value regarding the depth of the other pixels in the input video, estimating the depth information of the pixels in the input video based on the first statistical value and the second statistical value, at least one processor for estimating the depth information of the other pixels in the input video based on the other first statistical value and the other second statistical value, comprising, wherein the first statistical value and the other first statistical value correspond to a first type of statistical value, and the second statistical value and the other second statistical value correspond to a second type of statistical value, a neural network device.

2. The neural network device according to claim 1, wherein when estimating the depth information, the processor determines the reliability of the first statistical value based on the second statistical value and determines the reliability of the other first statistical value based on the other second statistical value.

3. The neural network device according to claim 2, wherein the second statistical value corresponds to the standard deviation or variance with respect to the depth value of the pixels, and the other second statistical value corresponds to the standard deviation or variance with respect to the depth value of the other pixels.

4. The processor of the neural network device according to claim 2, selectively corrects the first statistical value according to the reliability indicated by the second statistical value, and selectively corrects the other first statistical value according to the reliability indicated by the other second statistical value.

5. The processor of the neural network device according to claim 4, generates a reconstructed input video from a reference video based on the first statistical value, the second statistical value, the other first statistical value, and the other second statistical value, and trains the neural network based on the comparison result between the generated and reconstructed input video and the input video.

6. The neural network device according to claim 1, wherein the first statistical value and the other first statistical value are obtained from an output video corresponding to a first channel among the output channels of the neural network, and the second statistical value and the other second statistical value are obtained from an output video corresponding to a second channel among the output channels of the neural network.

7. A step of obtaining a first statistical value related to depth for each pixel in the input video based on a first channel of output data obtained by applying the input video to a neural network; A step of obtaining a second statistical value related to depth for each pixel in the input video based on a second channel of the output data; A step of estimating depth information of each pixel in the input video based on the first statistical value and the second statistical value; A neural network method including:

8. The first statistical value includes an average of depth values based on probabilities of depth values of each pixel in the input video; The neural network method according to claim 7, wherein the second statistical value includes a variance or a standard deviation of depth values based on probabilities of depth values of each pixel in the input video.

9. The step of estimating the depth information includes: A step of determining a reliability of the first statistical value based on the second statistical value; A step of determining whether to adopt the first statistical value of a corresponding pixel as depth information based on the reliability of each pixel; The neural network method according to claim 7, including:

10. The step of determining the reliability includes: A step of reducing the reliability as the second statistical value is larger; A step of increasing the reliability as the second statistical value is smaller; The neural network method according to claim 9, including:

11. The step of estimating the depth information includes: A step of obtaining a probability distribution based on the first statistical value and the second statistical value; A step of performing random sampling for correction of the first statistical value based on the probability distribution; The neural network method according to claim 7, including:

12. The step of performing the random sampling includes a step of determining the reliability of the first statistical value based on the second statistical value and selectively performing random sampling based on the determined reliability of the first statistical value. The neural network method according to claim 11.

13. A step of capturing the input video using a camera of a terminal; A step of generating 3D position information corresponding to the terminal and peripheral objects included in the input video using the depth information; The neural network method according to claim 7, further including:

14. The neural network method according to claim 7, wherein the input video includes at least one of a monocular video and a stereo video.

15. A computer program stored in a medium for causing, in combination with hardware, the method according to any one of claims 7 to 14 to be executed.

16. Applying an input video to a neural network and obtaining a first statistical value and a second statistical value related to depth for each pixel in the input video; Obtaining a probability distribution related to the depth of each pixel in the input video based on the first statistical value and the second statistical value; Training the neural network using an objective function based on the probability distribution related to the depth of each pixel in the input video and the predetermined depth information of each pixel in the input video; A neural network method comprising:

17. The first statistical value includes an average of depth values based on the probability of the depth value of each pixel in the input video, The second statistical value includes the variance or standard deviation of the depth value based on the probability of the depth value of each pixel in the input video, The neural network method according to claim 16.

18. The step of obtaining a first statistical value and a second statistical value related to depth for each pixel in the input video is: Obtaining a first statistical value related to depth for each pixel in the input video based on a first channel of output data obtained by applying the input video to a neural network; Obtaining a second statistical value related to depth for each pixel in the input video based on a second channel of the output data; The neural network method according to claim 16, comprising:

19. The step of obtaining a first statistical value and a second statistical value related to depth for each pixel in the input video is: Applying a feature map extracted from the input video to a first network in the neural network and obtaining the first statistical value; Applying the feature map to a second network in the neural network and obtaining the second statistical value; The neural network method according to claim 16, comprising:

20. The step of training the neural network is: Estimating the depth information of each pixel in the input video based on the probability distribution related to the depth of each pixel in the input video and a reference video corresponding to the input video; Training the neural network using an objective function based on the estimated depth information of each pixel in the input video and the predetermined depth information of each pixel in the input video; The neural network method according to claim 16, including

21. The reference video is a stereo image, a monocular video, and pose information of the camera that captured the monocular video, The neural network method according to claim 20, including at least one of

22. The step of training the neural network further includes the step of obtaining the predetermined depth information of each pixel of the input video based on the reference video corresponding to the input video and the input video. The neural network method according to claim 16.

23. The step of training the neural network is synthesizing an output video based on the reference video corresponding to the input video and the probability distribution regarding the depth of each pixel in the input video, training the neural network using an objective function based on the input video and the output video. The neural network method according to claim 16, including

24. The step of synthesizing the output video includes generating a second pixel in the output video corresponding to the first pixel based on the probability distribution regarding the depth of the first pixel in the input video and the pixel corresponding to the depth of the first pixel in the reference video corresponding to the input video. The neural network method according to claim 23.

25. The step of generating the second pixel in the output video corresponding to the first pixel includes generating the second pixel in the output video corresponding to the first pixel by weighted summing the values of the pixels corresponding to the depth of the first pixel in the reference video corresponding to the input video based on the probability distribution regarding the depth of the first pixel. The neural network method according to claim 24.

26. The step of training the neural network includes training the neural network based on an objective function that minimizes the difference between the first pixel in the input video and the second pixel in the output video corresponding to the first pixel. The neural network method according to claim 23.

27. The step of training the neural network is generating a mask corresponding to each pixel in the input video based on the first statistical value and the second statistical value, training the neural network using the mask and the objective function. The neural network method according to claim 16, including

28. The step of training the neural network includes the step of optimizing parameters for each layer of the neural network using the objective function, according to the neural network method of claim 16.

29. The input video includes at least one of a monocular video and a stereo video, according to the neural network method of claim 16.

30. A computer program stored in a medium for causing hardware to execute the method according to any one of claims 16 to 29.

31. Based on a first channel of output data obtained by applying an input video to a neural network, a first statistical value regarding depth is obtained for each pixel in the input video, Based on a second channel of the output data, a second statistical value regarding depth is obtained for each pixel in the input video, At least one processor for estimating depth information of each pixel in the input video based on the first statistical value and the second statistical value; A neural network device including the above.

32. The first statistical value includes an average of depth values based on the probability of the depth value of each pixel in the input video, The second statistical value includes the variance or standard deviation of the depth value based on the probability of the depth value of each pixel in the input video, according to the neural network device of claim 31.

33. The processor is In estimating the depth information, Determine the reliability of the first statistical value based on the second statistical value, Based on the reliability of each pixel, determine whether to adopt the first statistical value of the corresponding pixel as depth information, according to the neural network device of claim 31.

34. Further includes a camera for capturing the input video, The processor generates 3D position information corresponding to at least one of the device of the camera and peripheral objects included in the input video using the depth information, according to the neural network device of claim 31.

35. The input video includes at least one of a captured monocular video and a captured stereo video, according to the neural network device of claim 31.

36. The processor is Train the neural network based on the probability distribution regarding the depth of each pixel in the first video in the training data, The probability distribution is obtained based on the first statistical value and the second statistical value regarding the depth of each pixel in the first video, according to the neural network device of claim 31.

37. The processor is In obtaining the first statistical value, the first statistical value is obtained for each pixel in the input video from the first channel. The neural network device according to claim 31, wherein in obtaining the second statistical value, the second statistical value is obtained for each pixel in the input video from the second channel.

Citation Information

Patent Citations

  • Depth prediction from image data using statistical models

    JP2019526878A