Apparatus and method for generating depth map
By generating a two-dimensional structural framework and utilizing disparity estimation, combined with artificial neural networks to optimize the depth map generation process, the problems of depth information accuracy and high computational resources are solved, achieving more efficient and accurate depth map generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-20
- Publication Date
- 2026-03-27
AI Technical Summary
Existing technologies lack the accuracy and reliability of depth information when generating depth maps, leading to a decline in the quality of 3D images and high computational resource requirements, making it difficult to achieve efficient and accurate depth estimation.
A receiver and a structural frame determiner are used to generate a two-dimensional structural frame. Combined with a depth determiner, a depth map is generated through an artificial neural network and disparity estimation. The depth value is then determined and optimized using the two-dimensional structural frame and disparity information.
It improves the accuracy and consistency of depth maps, reduces computational complexity and resource requirements, is suitable for depth estimation of stereo and multi-view images, and provides more detailed depth values.
Smart Images

Figure CN121753067A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an apparatus and method for generating depth maps, particularly but not limited to generating depth maps containing images of people or animals. Background Technology
[0002] Traditionally, the technical processing and use of images have been based on two-dimensional imaging, but there is an increasing emphasis on explicitly considering the third dimension in image processing.
[0003] For example, three-dimensional (3D) displays have been developed that add a third dimension to the viewing experience by providing different views of the scene being viewed to the viewer's eyes.
[0004] One example is the free-viewpoint use case, which allows (within limitations) spatial navigation of a scene captured by multiple cameras. This can be done, for example, on a smartphone or tablet and can provide a game-like experience. Alternatively, the data can be viewed on an augmented reality (AR) or virtual reality (VR) headset.
[0005] In many applications, it may be desirable to generate viewing images for new viewing directions. Although various algorithms are known for generating such new view images based on image and depth information, they tend to be highly dependent on the accuracy of the provided (or derived) depth information.
[0006] The quality of the 3D image presented from the new view depends on the quality of the received image and the depth data. In many practical applications and scenarios, the provided depth information is often not optimal. In fact, in many real-world applications and usage scenarios, depth information may not be as accurate as expected, which can introduce errors, artifacts, and / or noise into the processed and generated images.
[0007] For example, in many applications, a multi-camera system placed at different locations within a scene is used to capture 3D scenes. Then, multi-view matching between the cameras can be used to generate specific depth values. However, depth estimation is problematic and often results in non-ideal depth values. This can again lead to artifacts and a degraded 3D image quality in the newly synthesized view.
[0008] To improve depth information, various techniques have been proposed for post-processing and / or improving depth estimation and / or depth maps. However, these are often not optimal, often not the most accurate and reliable, and / or may be difficult to implement, for example, due to the required computational resources.
[0009] An example of post-processing for depth maps is disclosed in EP4013049A1. In this method, a scanning method is applied, where the depth map can be initialized and then iteratively updated using the scanning method. The depth of the current pixel is updated based on a candidate set of candidate depth values, typically the depth values of neighboring pixels. However, while this method can improve depth maps in many situations, it is often not optimal in all situations and does not always generate the most accurate depth map. It also tends to be computationally demanding.
[0010] Therefore, improved methods for generating / processing / modifying depth information would be advantageous, and in particular, methods for processing depth maps that allow for increased flexibility, simplified implementation, reduced complexity, reduced resource requirements, improved depth information, more reliable and / or more accurate depth information, improved 3D experience, and improved quality and / or performance of depth-based rendered images would be advantageous. Summary of the Invention
[0011] Therefore, the present invention seeks to mitigate, alleviate or eliminate one or more of the above-mentioned disadvantages, preferably alone or in any combination.
[0012] According to one aspect of the present invention, an apparatus for generating a depth map of a first image of a scene is provided, the apparatus comprising: a receiver arranged to receive at least the first image; a structural frame determiner arranged to process the first image to determine a two-dimensional structural frame of an object in the first image, the two-dimensional structural frame being defined by a set of points and interconnections between the points, the two-dimensional structural frame representing a projection of a three-dimensional structural frame of the object in the scene onto the image space of the image; and a depth determiner arranged to generate the depth map of the first image based on the two-dimensional structural frame.
[0013] This invention can improve depth maps, thereby improving, for example, the quality of 3D image processing and perceptual rendering. In particular, in many embodiments and scenarios, the method can provide more consistent and / or accurate depth maps. In many embodiments, the processing can provide improved depth maps while maintaining sufficiently low complexity and / or resource requirements.
[0014] An advantage of many embodiments is that this method is well-suited for use and integration with depth estimation techniques, such as disparity-based depth estimation using stereo or multi-view images. In many embodiments, this method can allow for refinement of depth values determined by other methods.
[0015] A depth map indicates the depth values of pixels in a first image. Depth values can be any value indicating depth, including, for example, disparity values, z-coordinates, or distances from the viewpoint.
[0016] The two-dimensional structural frame can be a projection of the three-dimensional structural frame onto the image space / plane of the first image, and in particular, each point of the two-dimensional structural frame can be a projection of the (corresponding) point of the three-dimensional structural frame onto the image plane of the first image.
[0017] The two-dimensional / three-dimensional structural framework can be a skeleton (framework / model). The object can specifically be a human or an animal.
[0018] In some embodiments, the two-dimensional structural frame includes only two-dimensional information and may only provide information about the structural frame within the image plane. For example, the two-dimensional structural frame may be defined solely by the positions of points in the image; for instance, it may include the image coordinates of each point in the two-dimensional structural frame.
[0019] In some embodiments, the two-dimensional model may also include information about third coordinates. For example, depth coordinates may also be provided for each point of the two-dimensional structural frame. For example, the points of the two-dimensional structural frame may be defined by image coordinates and may also include depth values. In fact, the two-dimensional structural frame may also provide information for defining a three-dimensional structural frame, and in many such cases, the two-dimensional structural frame can be directly converted to three-dimensional space, especially to a three-dimensional structural frame.
[0020] In some embodiments, the apparatus may include an image synthesizer arranged to synthesize an image based on / according to a depth map and a first image viewpoint (different from the viewpoint of the first image).
[0021] According to an optional feature of the invention, the structural frame determiner includes a trained artificial neural network arranged to receive the first image as input and generate points of the two-dimensional structural frame as output.
[0022] In many embodiments, this can provide improved performance and / or implementation. In particular, in many scenarios, it can provide an efficient method for determining accurate 2D structural frames.
[0023] In some embodiments, the structural frame determiner is arranged to generate probability maps of at least some points, the probability maps of which indicate the probability of the point being at different locations in the image space; and wherein the depth map generator is arranged to generate a depth map based on the probability maps.
[0024] According to an optional feature of the invention, the receiver is arranged to receive a plurality of images, the structural frame determiner is arranged to generate a two-dimensional structural frame for at least some of the plurality of images, and the depth determiner is arranged to determine a depth value of the depth map based on the disparity between points of the two-dimensional frame structure of the at least some images.
[0025] In many embodiments, this can provide improved performance and / or implementation.
[0026] According to an optional feature of the invention, the depth determiner is arranged to determine an estimated three-dimensional frame structure of the object based on the disparity between points of the two-dimensional frame structure in the at least some images, and to determine the depth values of the points of the two-dimensional frame structure in the at least some images by projecting the points of the estimated three-dimensional frame structure into the image space of the at least some images.
[0027] In many embodiments, this can provide improved performance and / or implementation.
[0028] According to an optional feature of the invention, the depth determiner is arranged to generate a depth value for the interconnection between the first point and the second point based on the depth values of the first point and the second point of the two-dimensional frame structure.
[0029] In many embodiments, this can provide improved performance and / or implementation.
[0030] In some embodiments, the depth map determiner is arranged to determine the depth value of the interconnection between the first point and the second point by interpolating the depth values between the first point and the second point.
[0031] In some embodiments, the depth determiner is arranged to determine the 3D position estimate of the interconnection between points of the estimated 3D frame structure based on the estimated position of the points of the estimated 3D frame structure, and to determine the depth value by projecting the interconnection onto the image space of the first image.
[0032] According to an optional feature of the invention, the depth determiner is arranged to determine an initial depth map and generate the depth map by performing the following steps on at least a first pixel of the depth map: determining a set of candidate depth values, the set of candidate depth values including depth values of other pixels in the depth map besides the first pixel and at least a first candidate depth value determined according to the two-dimensional structural frame; determining a cost value for each candidate depth value in the set of candidate depth values in response to a cost function; selecting a first depth value from the set of candidate depth values in response to the cost value; and determining an updated depth value for the first pixel in response to the first depth value; wherein the cost value of the first candidate depth value depends on the difference between the candidate depth value and the depth value of the two-dimensional structural frame.
[0033] In many embodiments, this can provide improved performance and / or implementation. The cost function can be implemented as an evaluation function, and the cost value can be indicated by an evaluation value. The selection of a first depth value in response to the cost value can be implemented as the selection of a first depth value in response to an evaluation value determined from the evaluation function. The incrementing evaluation value / function is the decrementing cost value / function. The selection of the first depth value can be the selection of the candidate depth value with the lowest cost value from a set of candidate depth values, corresponding to / equivalent to the selection of the candidate depth value with the highest evaluation value from a set of candidate depth values.
[0034] According to an optional feature of the invention, the cost value of the first candidate depth value depends on the distance between the position of the candidate depth value and the position of the depth value of the two-dimensional structural frame.
[0035] In many embodiments, this can provide improved performance and / or implementation.
[0036] According to an optional feature of the invention, the depth determiner is arranged to extract visual features from the first image and determine an image mask of the object in the first image, and wherein the depth determiner is further arranged to determine a depth estimate of the depth map based on the visual features, the mask, and the two-dimensional structural frame.
[0037] In many embodiments, this can provide improved performance and / or implementation.
[0038] According to an optional feature of the invention, the depth determiner includes a trained artificial neural network arranged to receive the visual features, the mask, and the two-dimensional structural frame as input and determine the depth map as output.
[0039] In many embodiments, this can provide improved performance and / or implementation.
[0040] According to an optional feature of the invention, the depth determiner includes a trained artificial neural network arranged to receive the first image as input and generate the mask as output.
[0041] In many embodiments, this can provide improved performance and / or implementation.
[0042] According to an optional feature of the invention, the depth determiner is arranged to receive a first set of depth values of the image and determine the depth map by adjusting the first set of depth values according to the two-dimensional frame structure.
[0043] In many embodiments, this can provide improved performance and / or implementation.
[0044] According to an optional feature of the invention, the at least first image comprises a plurality of images representing the scene from different viewpoints, and the apparatus includes a disparity estimator arranged to determine the first set of depth values by disparity estimation between the plurality of images.
[0045] In many embodiments, this can provide improved performance and / or implementation.
[0046] According to an optional feature of the invention, the depth determiner is arranged to determine the depth value of the first pixel of the depth map by reducing the difference between the depth value of the first depth pixel in the first set of depth values and the depth value of the two-dimensional frame structure, the reduction depending on the distance between the first depth pixel and the two-dimensional depth structure.
[0047] In many embodiments, this can provide improved performance and / or implementation.
[0048] According to another aspect of the present invention, a method for determining a depth map of a first image of a scene is provided, the method comprising: receiving at least the first image; processing the first image to determine a two-dimensional structural framework of an object in the first image, the two-dimensional structural framework being defined by a set of points and interconnections between the points, the two-dimensional structural framework representing a projection of a three-dimensional structural framework of the object in the scene onto the image space of the image; and generating the depth map of the first image based on the two-dimensional structural framework.
[0049] These and other aspects, features, and advantages of the invention will become apparent and will be elucidated with reference to one or more embodiments described below. Attached Figure Description
[0050] Embodiments of the invention will be described by way of example only with reference to the accompanying drawings, in which... Figure 1 Examples of apparatus for generating depth maps according to some embodiments of the present invention are shown; Figure 2 An example of the structure of an artificial neural network is shown; Figure 3 An example of a node in an artificial neural network is shown; Figure 4 Examples of methods for determining a depth map of an image according to some embodiments of the present invention are shown; Figure 5 Examples of methods for determining a depth map of an image according to some embodiments of the present invention are shown; Figure 6 Examples of methods for determining a depth map of an image according to some embodiments of the present invention are shown; and Figure 7 Some elements of a processor for implementing an apparatus are shown in some embodiments of the invention. Detailed Implementation
[0051] Images representing a scene are now sometimes supplemented with depth maps, which provide depth information for pixels in the image. This additional information can allow for novel view compositing or depth-related editing operations, thus providing a number of additional services. A depth map tends to provide a depth value for each of a plurality of pixels, typically arranged in an array with a first number of horizontal rows and a second number of vertical columns. The depth value provides depth information for pixels in the associated image. In many embodiments, the resolution of the depth map may be the same as the resolution of the image, so that each pixel of the image may have a one-to-one link to a depth value in the depth map. However, in many embodiments, the resolution of the depth map may be lower than the resolution of the image, and in some embodiments, the depth values of the depth map may be common to multiple pixels of the image (specifically, the depth map pixels may be larger than the image pixels).
[0052] The depth value can be any value that indicates depth, specifically including depth coordinate values (e.g., directly providing the z-value of a pixel) or disparity values. In many embodiments, the depth map can be a rectangular array of pixels (having rows and columns), where each pixel provides a depth ( / disparity) value.
[0053] The accuracy of scene depth representation is a key parameter for the quality of novel view images synthesized and perceived by the user. Therefore, generating accurate depth information is crucial. Achieving accurate values may be relatively easy for artificial scenes (e.g., computer games), but it can be very difficult for applications involving, for example, capturing real-world scenes.
[0054] Several different methods have been proposed for depth estimation. One approach is to use pattern matching to estimate the disparity between different images of a scene captured from different viewpoints. However, this disparity estimation is inherently imperfect. Another approach involves adding depth sensors, such as those using time-of-flight or structured light techniques. However, these depth sensors often contain noise and have limited measurement range and spatial resolution. A third approach utilizes pre-defined (hypothetical) depth information within the scene. For example, for outdoor scenes (and indeed for most typical indoor scenes), objects below the image tend to be closer than objects above the image (e.g., the distance from the floor or ground to the camera gradually increases with height, the sky tends to be further back than the ground below, etc.). Therefore, a predetermined depth profile can be used and fitted to the data to estimate appropriate depth map values.
[0055] However, most depth estimation techniques tend to fall short of achieving perfect depth estimates, and in many applications, improved depth values are often more reliable and / or accurate, which would be of great benefit.
[0056] The following section describes a method for generating depth maps. In many scenarios, this method can generate improved depth maps and provide more accurate depth information.
[0057] Figure 1 Some elements of an apparatus for generating a depth map of an image are shown. The apparatus includes a receiver 101 arranged to receive at least one (2D) image of a (3D) scene, and typically receives multiple images of the scene. Each image represents a scene from a given viewpoint, typically within a region called a capture point, and when multiple images are received, multiple images are typically provided for different capture points, i.e., each image is captured from a different viewpoint.
[0058] Receiver 101 is coupled to a first circuit, which will also be referred to as structural frame determiner 103, and is arranged to process an image to determine a two-dimensional structural frame (2D structural frame) of an object in the image. The 2D structural frame represents the three-dimensional structural frame (3D structural frame) of an object in the scene, and specifically, the 2D structural frame is the projection of the 3D structural frame onto the image space / plane of the image.
[0059] A 3D structural frame can be an internal framework, structure, skeleton, or framework that structurally supports an object. In practice, in many embodiments, the 3D structural frame can be the skeleton of an object, and in many embodiments, the object can be a person or animal, and the 3D structural frame can be the skeleton of a person or animal. For other objects, such as a car, the 3D structural frame can be a chassis or other supporting frame or basic structure.
[0060] In many embodiments, a 3D structural framework can be represented by a set of points connected by interconnections. The interconnections can specifically have a predetermined shape and are typically line segments. Therefore, a 3D structural framework can be given by a set of points connected by line segments. Each point is a point in the 3D space of the scene, and thus each point can be represented by a set of coordinates (specifically scene coordinates).
[0061] Therefore, the projection of a 3D structural framework onto a 2D image space / plane can also be represented by a set of points connected by interconnections, which can specifically be linear. Thus, a 2D structural framework can be given as a set of points with (typically linear) connections. Each point in a 2D structural framework can be represented by a set of coordinates (specifically image coordinates). Points in a 3D structural framework are typically represented by three coordinate values, while points in a 2D structural framework can be represented by two coordinate values.
[0062] A 2D structural frame specifically represents the projection of a 3D structural frame onto the image space / plane of an image. Specifically, each point of the 2D structural frame is the projection of a point of the 3D structural frame onto the image space / plane of the received image. The 2D structural frame can be considered to represent the 3D structural frame because it is projected onto the viewport of the image, and therefore it can represent how the 3D structural frame will be seen from the viewpoint of the image. The 2D structural frame correspondingly represents the 3D structural frame as seen in the image. The 2D structural frame thus provides information about the representation of objects in the image, specifically information about the representation of the structural frame of the objects.
[0063] The structural frame determiner 103 is arranged to determine the 2D structural frame of an object from the received image. In some embodiments, the structural frame determiner 103 may be arranged to first determine the 3D structural frame and then project it onto the image plane of the image, or in some embodiments it may be arranged to directly determine the 2D structural frame.
[0064] The determination of a 2D structural frame can be based on predetermined structural frame data that indicates the attributes of the object's structural frame. This data can specifically be constraint data, such as indications of relative attributes between points, such as maximum and minimum distances and / or angles between pairs of points. Such structural frame attribute data can reflect the object's restrictive properties. For example, for a structural frame serving as a human skeleton, the data could, for instance, indicate the maximum or minimum distance between two interconnected points, such as the maximum and / or minimum distance between points corresponding to the wrist and elbow points, reflecting how points can move relative to each other (e.g., for a 3D structural frame, the distance between the wrist and elbow points would be fixed, which would also provide constraints on the 2D structural frame (particularly in conjunction with constraint information about the movement of points relative to each other)). In some cases, the structural frame determiner 103 may include a predetermined model of the object's structural frame with multiple variable parameters that can be modified to adapt the model to a specific instantiation matching the structural frame.
[0065] In some embodiments, the structural frame determiner 103 may be arranged to determine a 2D structural frame by detecting specific image objects or regions / points in an image that may correspond to different parts of an object. For example, if the object is specifically a football player, the structural frame determiner 103 may be arranged to detect the face, hands, torso and arms, football boots, knees, etc. Detection may typically be based on unique colors (skin tone, the distinctive color of a football jersey, the color of football boots, etc.) and the relative distances and positions between these colors (e.g., the knee is below the football shorts and above the football boots, etc.). The position of each of these items may correspond to a point in the structural frame, and the parameters of the structural frame model can then be adjusted to provide an optimal fit to the detected points / image regions.
[0066] In some embodiments, the 2D structural frame may be defined only in the image plane and may include only 2D information, that is, it may include only information on how the structural frame is constructed in the image plane (and thus specifically may include only 2D information, such as providing only 2D image coordinates for the points of the structural frame).
[0067] However, in other embodiments, the 2D structural frame may also include some information related to the three-dimensional scene and location. For example, image plane information provided by point and / or interconnected image coordinates may be combined with depth data, for example, indicating distance in the z-direction. Thus, the 2D information can be supplemented by 3D data, in which case the 2D structural frame can provide not only information about the projection of the 3D frame onto the image plane, but also information about the transformation from the image plane to 3D coordinates. Therefore, in this case, the 2D structural frame can also provide information that directly allows for the determination of the 3D structural frame, and it can be practically considered that the 2D structural frame also directly represents / corresponds to the 3D structural frame.
[0068] In many embodiments, the structural frame determiner 103 includes a trained artificial neural network arranged to receive an image as input and generate a two-dimensional structural frame as output.
[0069] For example, artificial neural networks can be configured to directly provide an output corresponding to a 2D structural frame in response to a single image input. In fact, several algorithms based on a single image are known to provide not only the image coordinates of the points in a 2D structural frame but also the depth. Such algorithms can, for example, directly provide a 2D / 3D structural frame represented by the image coordinates and depth of all points in the structural frame (such as all points in a human skeleton).
[0070] A suitable artificial neural network can be a hierarchical network of nodes, where each node has a node value. Figure 2 An example of a portion of an artificial neural network is shown.
[0071] The node value of a given node can be calculated by incorporating contributions from some, or typically all, nodes in the previous layer of the artificial neural network. Specifically, the node value can be calculated as a weighted sum of the node values output by all nodes in the previous layer. Typically, biases can be added, and the results can be processed by activation functions. Activation functions typically provide the basic components of each neuron by introducing nonlinearity. This nonlinearity and activation function play a crucial role in the learning and tuning of the neural network. Therefore, the node value is generated based on the node values of the previous layer.
[0072] An artificial neural network may specifically include an input layer 201, which comprises multiple nodes that receive input data values from the artificial neural network. Therefore, the node values of the input layer nodes can typically be directly the input data values of the artificial neural network and cannot be calculated from other node values.
[0073] Artificial neural networks may also include zero, one or more hidden layers 203 or processing layers. For each such layer, node values are typically generated based on the node values of the previous layer, specifically by weighted combination and the addition of bias, followed by an activation function (e.g., a sigmoid, ReLU, or Tanh function may be applied).
[0074] Specifically, such as Figure 3 As shown, each node (also called a neuron) can receive input values (from nodes in the previous layer) and compute a node value based on these values. Typically, this involves first generating a linear combination of values as input values, each of which is weighted according to its own weight: Where w is the weight, x is the node of the previous layer, and n is the index of the different node in the previous layer.
[0075] Then, the activation function can be applied to the resulting combination. For example, the node value l can be determined as: This function can be, for example, the modified linear unit function described in Xavier Glorot, Antoine Bordes, and Yoshua Bengio's "Fourteenth International Conference on Artificial Intelligence and Statistics" (PMLR 15: pp. 315-323, 2011): Other commonly used functions include the sigmoid function or the tanh function. In many implementations, multiple functions can be used to compute the node output or value. For example, activation functions such as the following can be used to combine both the ReLU and sigmoid functions: Such operations can be performed by each node of an artificial neural network (usually in addition to the input node).
[0076] The artificial neural network also includes an output layer 205, which provides the output from the artificial neural network; that is, the output data of the artificial neural network is the node values of the output layer. For the hidden / processing layers, the output node values are generated as a function of the node values of the previous layer. However, unlike the hidden / processing layers, whose node values are generally inaccessible or not further usable, the node values of the output layer are accessible and provide the results of the artificial neural network's operations.
[0077] Many different network architectures and toolkits have been developed for artificial neural networks, and in many embodiments, artificial neural networks can be based on and customized from such networks. An example of a network architecture suitable for the aforementioned applications is WaveNet by van den Oord et al., described in Oord, Aaron van den, Sander Dieleman, HeigaZen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu, “Wavenet: A generative model for raw audio.” arXiv preprint arXiv:1609.03499 (2016).
[0078] WaveNet is an architecture for synthesizing temporal signals using dilated causal convolutions, and it has been successfully applied to audio signals. For WaveNet, the following activation function is typically used: Where * represents the convolution operator, ⊙ represents the element-wise multiplication operator, σ(∙) is the sigmoid function, k is the layer index, f and g represent the filter and gate, respectively, and W represents the weights of the learned artificial neural network. The filter product of the equation typically provides a filtering effect, while the gate product provides weighting of the result, which in many cases can effectively reduce the contribution of a node to near zero (i.e., it can allow or “cut off” a node from contributing to other nodes, thus providing a “gate” function). In different cases, the gate function may result in the node’s output being negligible, while in other cases it will contribute significantly to the output. Such functions can greatly help neural networks learn and train efficiently.
[0079] Many different network architectures and toolkits have been developed for artificial neural networks, and in many embodiments, artificial neural networks can be based on and customized from such networks. An example of a network architecture that can be applied to the above applications is Long Short-Term Memory (LSTM) [Sepp Hochsreiter’s, which is described in Hochreiter, Sepp and Jürgen Schmidhuber. “Long short-term memory” Neural computation 9.8 (1997): 1735-1780].
[0080] LSTM is an architecture used for classifying and regressing time-domain signals using cyclic causality or bidirectional evaluation, and it has been successfully applied to audio signals. For Where * denotes matrix multiplication, o denotes Hadamard product, x is the input vector, ht-1 represents the output vector at the previous time step, W, V, and U are network weights, and b is the bias vector.
[0081] In theory, classic (or "vanilla") artificial neural networks can track any long-term dependencies in an input sequence. The problem with vanilla artificial neural networks is inherently computational (or practical): when training a vanilla artificial neural network using backpropagation, long-term gradients from backpropagation can either "vanish" (i.e., tend to zero) or "explode" (i.e., tend to infinity) because the computations involved employ finite-precision numerical values. Artificial neural networks using LSTM units partially solve the vanishing gradient problem because LSTM units allow gradients to propagate unchanged. However, LSTM networks can still face the exploding gradient problem.
[0082] In some cases, artificial neural networks can also be configured to include additional contributions that allow the network to be dynamically tuned or tailored to meet specific desired properties or characteristics of the generated output. For example, a set of values can be provided to tune the artificial neural network. These values can be achieved by contributing to some nodes of the network. These nodes can be nodes of a specific input, but are typically nodes of hidden or processing layers. Such tuning values can, for example, be weighted and summed as a contribution to a weighted sum / correlation value for a given node.
[0083] The above description relates to neural network methods applicable to many embodiments and implementations. However, it should be understood that many other types and structures of neural networks can be used. In fact, many different neural network generation methods have been and are being developed, including neural networks using complex structures and processes different from those described above. This method is not limited to any particular neural network method, and any suitable method can be used without departing from the invention.
[0084] Artificial neural networks (ANNs) are adapted for a specific purpose through a training process that adjusts / tunes / modifies the weights and other parameters (e.g., biases) of the ANN. It should be understood that many different training processes and algorithms are known for training ANNs. Typically, training is based on a large training set, where a large number of input data examples are fed to the network. Furthermore, the output of the ANN is usually (directly or indirectly) compared to an expected or desired result. A cost function can be generated to reflect the desired outcome of the training process. In a typical scenario known as supervised learning, the cost function typically represents the distance between the prediction for a particular input data and the true baseline. Based on the cost function, the weights can be changed, and by repeating this process with the modified weights, the ANN can be tuned to achieve a state that minimizes the cost function.
[0085] More specifically, during the training phase, a neural network can have two distinct information flows: from input to output (forward propagation) and from output to input (backward propagation). In forward propagation, as described above, the data is processed by the neural network, while in backpropagation, the weights are updated to minimize the cost function. Typically, this backpropagation follows the gradient direction of the cost function landscape. In other words, by comparing the predicted output with a baseline of true values from a set of data inputs, the direction in which the cost function is minimized and backpropagated can be estimated by updating the weights accordingly. Other known methods for training artificial neural networks include, for example, the Levenberg-Marquardt algorithm, the conjugate gradient method, and Newton's method.
[0086] In the current context, training can specifically include a training set comprising a potentially large number of images of a scene and data representing 2D structural frames of suitable objects (e.g., people) within those images. For example, 2D structural frames can be manually determined for the training images, and these manually constructed 2D structural frames can then be used as a reference for evaluating 2D structural frames generated by an artificial neural network. In many embodiments, the training images can be synthetically generated and can be generated to include suitable objects. The synthesis of images of a scene can be based on 2D structural frames of objects in the scene. For example, an image of a landscape can include a person already synthesized based on a human skeleton model. Therefore, the artificial neural network can be trained by analyzing these synthetic images and comparing the resulting 2D structural frames with those used to synthesize the images. A large number of such images and training data can be generated and used to train the network.
[0087] Specific examples of artificial neural networks suitable for the method can be found in He, K., Gkioxari, G., Dollar, P., and Girshick, R., “Mask R-CNN”. arXiv e-prints, 2017. doi:10.48550 / arXiv.1703.06870. This paper describes a training method and an artificial neural network that is quite versatile and can be used for multiple tasks: object detection (bounding boxes), object instance segmentation (object masks), and keypoint detection.
[0088] Therefore, the structural frame determiner 103 is arranged to generate a structural frame defined by key points. The structural frame also includes interconnections between several (typically) pairs of points. Interconnections can constrain how points are positioned, particularly how points move / change position relative to each other. In some embodiments, the predetermined structural frame model may have, for example, multiple interconnections between a plurality of points and some of a set of points. The predetermined model may have multiple variable parameters. For example, the positions of points may be variable but constrained by the interconnections. The constraints provided by each interconnection may, for example, be a set of acceptable values for the distance and / or orientation between points connected by the interconnection. The structural frame of an image can then be determined by fitting the predetermined model to the image, for example, by varying the parameters to produce a best fit. In many embodiments, an artificial neural network can be trained to implicitly perform this fitting by receiving an image and providing the output of the positions of points for a given structural frame. For example, in an embodiment where the structural frame is a human skeleton, the artificial neural network can be trained to provide the positions of a predetermined set of points corresponding to selected points of the skeleton (e.g., skull, shoulder joint, hip joint, elbow, etc.).
[0089] Structural frames can also be created based on depth maps from depth sensors (e.g., using time-of-flight). For example, for a chair, a structural frame can be created in 3D space based on the depth map.
[0090] In some embodiments, the method described above can be performed in three-dimensional space, and then the 3D structural frame is projected onto the 2D structural frame. However, in many embodiments, the method can be applied directly to / performed in a two-dimensional image plane, thus directly providing a 2D structural frame. Neural networks used for keypoint detection of an object category (e.g., a person) directly from an input color image typically use the latter approach.
[0091] In some embodiments, the structural frame determiner 103 may provide a confidence score for each point, for example, with a value between 0 and 1. These confidence scores can provide valuable information about the extent to which we can rely on a given point or connecting line segment when generating candidate depth values for subsequent dense (per-pixel) depth map estimation steps.
[0092] The following sections will describe various methods for generating depth maps.
[0093] In some embodiments, receiver 101 is arranged to receive multiple images of a scene captured from different viewpoints, and structural frame determiner 103 is arranged to generate a two-dimensional structural frame for some images, and typically for all images. Thus, a 2D structural frame is generated for each of the multiple images of the scene. In such an embodiment, depth determiner 105 may be arranged to determine the depth of a depth map based on the disparity between corresponding points of the two-dimensional frame structure in at least some images.
[0094] The depth determiner 105 can, for example, use a disparity vector to determine the depth of corresponding points in different images representing different viewpoints. Therefore, for two images of a scene viewed from different viewpoints, the relative position of a given corresponding point (e.g., a point corresponding to a person's shoulder) within a 2D structural frame is used in the two images. The 3D position of this point in world space can be determined based on the difference in position between the two images.
[0095] In some embodiments, only two images (and two 2D structural frames) are considered, and for a pixel corresponding to a point of the 2D structural frame in that image, the depth value of the depth map of one of the images can be set to the determined depth value. In some cases, multiple pairs of images can be considered, and a depth value can be determined, for example, for each pair of images / 2D structural frames, and the final depth value can be determined based on these images, for example, by averaging.
[0096] This method can accordingly determine the depth values of points in a 2D structural frame. In many embodiments, the depth determiner 105 can be arranged to determine an estimated 3D frame structure of the object based on the disparity between points in the 2D frame structure of at least some images. In some embodiments, the 3D structural frame can be generated through geometric calculations based on an image, the viewpoint of the image, the location of the point in the image, and the determined depth. This can be repeated for all points in the 2D structural frame to generate 3D points for the 3D structural frame. The interconnections of the 3D structural frame can then be determined as interconnections between the same points in the 2D structural frame.
[0097] In other embodiments, more sophisticated methods can be used to determine the estimated 3D structural framework. For example, multiple determinations of 3D points can be made based on different image pairs, and these can be combined, for example by averaging, or by applying a predetermined 3D structural framework model.
[0098] In many embodiments, a 2D structural framework for one or more images can be generated by projecting a 3D structural framework onto the image plane of an image. In many embodiments and scenarios, this can provide an improved 2D structural framework to be determined, which can further lead to improved depth determination and an improved depth map to be generated.
[0099] In many embodiments, the depth determiner 105 can be arranged to generate depth values for the interconnections between points. Specifically, the depth values for the interconnections between the first and second points of the two-dimensional frame structure can be determined based on the depth values of the first and second points.
[0100] In some embodiments, this can be done directly in the image plane of the 2D structural frame. Specifically, the depth values along the interconnection of two points can be interpolated between two depth values of the two points. For example, for a linear interconnection, the interpolation can simply be a weighted linear combination of the depths of the two points, with the weights depending on the distance to each point.
[0101] In some embodiments, the depth values of the interconnects can be based on considerations in the 3D domain. For example, a depth determiner 105 can determine the 3D positions of two points and then determine the positions of different portions of the interconnect between those two points in 3D space. Similarly, in some embodiments, interpolation, such as linear interpolation, can be performed in 3D space. Thus, the 3D positions of the interconnects between two 3D points in a 3D structural frame can be determined by interpolating between the points of the two 3D points. The resulting positions can then be projected onto a 2D plane to generate a set of depth values for the corresponding interconnects in the 2D structural frame.
[0102] In some embodiments, the depth determiner 105 can be effectively arranged to construct connecting lines in 3D space, drawing these values by interpolating points on the lines and projecting these points onto image space. In some embodiments, the depth determiner 105 can be arranged to directly project / or generate points in 2D image space and interpolate depth values on the projected / corresponding lines.
[0103] In some embodiments, the depth determiner 105 may be arranged to first determine an initial depth map and then continue to consider the 2D structural framework to update the depth map.
[0104] The initial depth map can be determined in any suitable form, including, for example, through disparity estimation from multi-angle images, dedicated range / distance measurements, etc. Therefore, the initial depth map can be determined without considering a 3D or 2D structural frame.
[0105] Then, the depth determiner 105 can determine the output depth map by updating the depth values in the initial depth map according to the 2D structural frame / 3D structural frame. The depth determiner 105 can specifically achieve this by continuing to determine a set of candidate depth values for a given depth pixel / depth value, which includes or contains the depth values of other pixels in the depth map besides the first pixel and one or more depth values determined from the 2D structural frame, and is typically the depth value of the 2D structural frame.
[0106] The depth determiner 105 can use a suitable cost function to determine the cost of each candidate depth value. Then, one of the depth values can be selected based on the generated cost. Next, an updated depth value can be determined based on the selected depth value, for example, simply by setting the depth value of the depth map to the selected depth value. The cost of the depth value determined according to the 2D structural frame (e.g., the depth value of the 2D structural frame) is determined based on the difference between the candidate depth value and the depth value of the 2D structural frame, and is typically determined based on the depth value of the 2D structural frame that is closest to the pixel for which its depth value was determined. It will also typically depend on the distance between the 2D structural frame and the candidate pixel value.
[0107] More specifically, the depth determiner 105 can further determine a set of candidate depth values for a given pixel. This set of candidate depth values includes the depth values of a group of candidate pixels. This set of candidate depth values for a given pixel can consist of the depth values of neighboring pixels, either spatially or temporally. For example, the set of candidates could include the depth values of pixels in the neighborhood surrounding the current pixel, such as a group of pixels within a given distance of the current pixel or within a window / kernel surrounding the current pixel.
[0108] In many embodiments, the set of candidate pixels also includes the depth value of the current pixel itself, that is, the first depth value is itself one of the set of candidate depth values.
[0109] Furthermore, in many embodiments, the set of candidate depth values may also include depth values from other depth maps. For example, in many embodiments where the image is part of a video stream, one or more depth values from previous and / or subsequent frames / images may also be included in the set of candidate depth values, or depth values from other views from which depth maps are simultaneously estimated.
[0110] In some embodiments, the set of candidate depth values may also include values that are not direct depth values from the depth map. For example, in some embodiments, the set of candidate depth values may include one or more fixed depth values or, for example, relative offset depth values, such as depth values that are a fixed offset larger or smaller than the current initial depth value. Another example is that the set of candidate depth values may include one or more random or semi-random depth values.
[0111] Then, the depth determiner 105 can determine the cost of the group of candidate depth values, specifically, it can determine the cost of each candidate depth value in the group of candidate depth values.
[0112] The cost value can be determined based on a cost function, which can depend on several different parameters, as will be described in more detail later. In many embodiments, the cost function for candidate depth values of pixels in the current depth map depends on the differences between image values in the multi-view images, which are offset by the disparity corresponding to the depth value. Thus, for a first depth value, or perhaps each candidate depth value in a set of candidate depth values belonging to a set of candidate depth values, the cost function can monotonically decrease based on the differences between two view images in an image region of the multi-view images having disparity between two image views that match the candidate depth value. The image region can specifically be an image region that includes pixels of the current pixel and / or candidate depth values. The image region can typically be relatively small, for example, including no more than 1%, 2%, 5%, or 10% of the image and / or, for example, no more than 100, 1000, 2000, 5000, or 10000 pixels.
[0113] In some embodiments, the depth determiner 105 can determine the disparity between two images that match a given candidate depth value. It can then apply this disparity to identify regions in one of the two images that are offset to the other image by the disparity. A difference metric between image signal values (e.g., RGB values) in the two regions can be determined. Therefore, the difference metric between the two images / image regions can be determined based on the assumption that the candidate depth value is correct. The smaller the difference, the more likely the candidate depth value is an accurate reflection of the depth. Therefore, the smaller the difference metric, the smaller the cost function.
[0114] The image region can typically be a small area around the first / current pixel, and in some embodiments it can actually include only the first / current pixel.
[0115] For a candidate depth value corresponding to the current depth map, the cost function may accordingly include the cost contribution of matching between two multi-view images depending on the disparity corresponding to the candidate depth value.
[0116] In many embodiments, the cost of candidate depth values from other depth maps associated with the image (such as time-offset depth maps and images) may also include the corresponding image matching cost contribution.
[0117] In some embodiments, the cost of some candidate depth values may not include the contribution of image matching cost. For example, a fixed cost value may be assigned for a predetermined fixed depth offset that is not related to the depth map or image.
[0118] A cost function can usually be determined such that it indicates the likelihood that the depth value reflects the accurate or correct depth value of the current pixel.
[0119] It should be understood that determining the evaluation value based on the evaluation function and / or selecting candidate depth values based on the evaluation value is essentially the same as determining the cost value based on the cost function and / or selecting candidate depth values based on the cost value. By applying the function to the evaluation value, the evaluation value can be simply transformed into a cost function, where the function is any monotonically decreasing function. A higher evaluation value corresponds to a lower cost value; for example, selecting the candidate depth value with the highest evaluation value has exactly the same cost value as selecting the lowest value.
[0120] The depth determiner 105 can then select a depth value from the set of candidate depth values in response to the cost of that set. The selected candidate depth value will then be referred to as the selected depth value.
[0121] In many embodiments, this selection may be to choose a candidate depth value that determines the lowest cost value. In some embodiments, other parameters may also be considered to evaluate more complex criteria (equivalently, such considerations can generally be considered as part of the (modified) cost function).
[0122] Therefore, for the current pixel, the method can select a candidate depth value that is considered most likely to reflect the correct depth value of the current pixel determined by the cost function.
[0123] Then, an updated depth value is determined for the current pixel based on the selected depth value. The exact update will depend on the specific requirements and preferences of each embodiment. For example, in many embodiments, the previous depth value of the first pixel may simply be replaced by the selected depth value. In other embodiments, the update may take into account the initial depth value; for example, the updated depth value may be determined as a weighted combination of the initial depth value and the selected depth value, where, for example, the weight depends on the absolute cost of the selected depth value.
[0124] Therefore, an updated or modified depth value is determined for the current pixel. This process can iterate over some or virtually all pixels in the depth map. It should be understood that although the above description focuses on the application to a single pixel, a block procedure can be applied, where, for example, the determined updated depth value is applied to all depth values within a block that includes the first pixel.
[0125] The specific selection of candidate depth values can depend on the desired operation and performance of the application. Typically, a set of candidate depth values will include multiple pixels in the neighborhood of the current pixel. A kernel, region, or template can cover the current pixel, and pixels within the kernel / region / template can be included in this set of candidates. Additionally, a kernel is included, comprising the same pixels from different time frames (typically immediately before or after the current frame providing the depth map) and potential neighboring pixels; however, this kernel is typically much smaller than the kernel in the current depth map. At least two offset depth values are also typically included (corresponding to increases and decreases in depth, respectively).
[0126] However, to reduce computational complexity and resource requirements, the number of depth values included in this set of candidates is typically heavily limited. Specifically, since all candidate depth values are evaluated for each new pixel and typically for all pixels in the depth map for each iteration, each additional candidate value results in a large number of additional processing steps.
[0127] In many typical applications, it is generally preferred that there be no more than about 5-20 candidate depth values in a set of candidate depth values for each pixel. In many practical scenarios, to achieve real-time processing of video sequences, it is necessary to limit the number of candidate depth values to around 10. However, the relatively small number of candidate depth values makes determining / selecting which candidate depth values to include in the set crucial.
[0128] If the cost function of one pixel is less than or lower than the cost function of another pixel, it means that, for all other parameters (i.e., all parameters except that the two pixels are in the same position), the cost value determined by the cost function is smaller / lower. Similarly, if the cost function of one pixel is greater than or higher than the cost function of another pixel, it means that, for all other parameters considered, the cost value determined by the cost function is larger / higher. Furthermore, if the cost function of one pixel exceeds that of another pixel, it means that, for all other parameters considered, the cost value determined by the cost function exceeds that of the other pixel.
[0129] For example, the cost function typically considers multiple different parameters. For instance, the cost value can be determined as C = f(d, a, b, c, ...), where d refers to the position of the pixel relative to the current pixel (e.g., distance), and a, b, c, ... reflect other parameters considered, such as the image signal value of the associated image, the value of other depth values, smoothness parameters, etc.
[0130] If the cost function C = f(d, a, b, c, ...) of pixel A is lower than the cost function f(d, a, b, c, ...) of pixel B, and if the parameters a, b, c, ... are the same for both pixels (and similar for the other terms), then the cost function f(d, a, b, c, ...) of pixel A is lower than that of pixel B.
[0131] The exact cost function will depend on the specific implementation. In many embodiments, the cost function includes a cost contribution that depends on the difference between image values of a multi-view image of pixels offset by a disparity matching the depth value. As previously mentioned, a depth map can be a map of images (or a set of images) of a multi-view image set capturing a scene from different viewpoints. Therefore, a disparity will exist between the positions of the same object in different images, and this disparity depends on the depth of the object. Therefore, for a given depth value, the disparity between two or more images of the multi-view image can be calculated. Thus, in some embodiments, for a given candidate depth value, the disparity with other images can be determined, thereby determining the position of a first pixel location in other images under the assumption that the depth value is correct. Image values, such as color or brightness values, of one or more pixels at the corresponding positions can be compared, and an appropriate disparity metric can be determined. If the depth value is indeed correct, the image values are more likely to be the same compared to the case where the depth value is not correct, and the disparity metric is smaller. Therefore, the cost function can include consideration of the differences between image values; specifically, the cost function can reflect the increased cost of increasing the disparity metric.
[0132] The depth determiner 105 can be specifically arranged to add candidate depth values determined from the 2D structural frame. For example, for a structural frame representing a human skeleton, the depth determiner 105 can add one or more candidate depth values depending on the depth of the skeletal portion (e.g., arm / leg). In particular, it can add one or more depth values that are depth values directly determined for the structural frame, including, for example, depth values of points or interconnections of the 2D structural frame. For example, skeletal lines (e.g., the upper arm) will intersect with a pixel grid, thus providing candidate values for all intersecting pixels and their neighboring pixels. In many cases, candidate depth values can also include depth values close to the 2D structural frame.
[0133] Then, the depth determiner 105 can determine a cost function for these candidate depth values of the structural frame, to include a cost contribution that depends on the difference between the candidate depth value and the depth value of the 2D structural frame.
[0134] For example, a pixel spatially (in the projected image space) very close to a person's skeletal arm may have the same or very similar depth to the nearest point on the line corresponding to the skeletal arm. More specifically, during candidate-based multi-view depth estimation, the depth determiner 105 can evaluate the following cost terms for a given pixel: Where C match It depends on the cost of the matching error between the current view and one or more other views, and C smoothness Spatial smoothness is weighted and depth transitions within regions with constant color intensity are penalized. Cost function C match and C smoothness This is well known to those skilled in the art, and therefore will not be described further in this article.
[0135] Cost component C frame This indicates the deviation of the depth from the depth that can be predicted by a nearby 2D structural frame. An example of a suitable version could be: Where D frame It is the (encoded) depth value of the 2D structural frame portion closest to the pixel, and |xx frame | is the estimated Euclidean distance between the pixel and a portion of the 2D structural frame. It should be noted that for distances exceeding a given pixel distance w, the cost will be set to the maximum cost to ensure that structural frame depth values are not selected for pixels that are too far away; for example, this prevents the selection of a human body depth that is far greater than, say, the width of an arm or leg.
[0136] Figure 4The method is illustrated by first passing the source view image through frame detection phase 401. In the next step 403, multi-view frame depth estimation is performed to assign a depth value to each vertex in the frame, after which depth values can be interpolated on line interconnections (e.g., corresponding to the skeletal portion of a person). In the following step 405, dense depth estimation then takes into account proximity and estimated frame / skeleton depth, among other factors, during multi-view candidate depth generation and matching. This process evaluates a cost function that also depends on 2D structural frame information.
[0137] A significant advantage of this approach is that, in principle, a training phase is not required as part of the depth estimation.
[0138] In some embodiments, the depth determiner 105 may be arranged to determine depth values based on visual features from an image combined with a 2D structural framework and a mask that reflects the contours of an object.
[0139] In addition to the depth determiner 105 being arranged to generate a structural framework (e.g., using an artificial neural network as described above), in some embodiments, the depth determiner 105 may include a visual feature extractor arranged to extract visual features from an image.
[0140] Such features can specifically include, for example, vectors at each pixel that describe the appearance of a local region (such as edge orientation, color distribution, brightness distribution, etc.).
[0141] Visual features can be extracted, for example, using some form of artificial neural network. Currently, many artificial neural networks have been developed that can perform this operation and visual feature extraction, such as the developed Unet, HRNet, and Mask-RCNN neural networks. In some embodiments, visual feature extraction can be based on specially designed (handcrafted) features. Such feature extractors can include features such as SIFT (Scale Invariant Feature Transform), SURF (Speed-Up Robust Features), and ORB (Oriented Fast and Rotated BRIEF).
[0142] As a very concrete example, such features can specifically generate the output of, for example, object detection neural networks (such as Masked R-CNN) with N=10, 20, ..., 50 layers. The first N layers of a convolutional neural network, when applied to an image, typically result in a shape / size (C, H, W) tensor, where C is typically greater than the 8 channels representing the features, and H and W are typically smaller than the height and width of the original image, respectively. Therefore, the features often model different aspects of a spatial region.
[0143] In many scenarios, such visual features can characterize an image and provide relevant information that allows for depth estimation.
[0144] In many embodiments, the determination of the object mask can be based on using a trained artificial neural network. This artificial neural network can be arranged to receive a first image as input and generate an object mask as output.
[0145] For example, a variety of different images can be used to train an artificial neural network, where image regions corresponding to people (or other specific objects) have already been manually identified. The output of the artificial neural network for a given set of images can then be compared with a manually generated mask to provide a cost for updating the artificial neural network. In some embodiments, virtual images can be generated based on a 3D model of the scene, and masks can be determined to correspond to image segments reflecting people in the virtual images. Similarly, using such images as input can produce an output mask, which can be compared with a mask determined directly from the model to provide a cost for tuning / training.
[0146] In other embodiments, non-artificial neural network methods can be used, where, for example, image segmentation is performed, and then a mask can be generated by combining fragments that meet a given criterion that may represent a person. As another example, known techniques such as region growing or feature clustering can be used to generate, for example, a mask of a person.
[0147] The depth determiner 105 can be arranged to generate depth data from a mask, visual features, and structural frames, and it can specifically include a trained artificial neural network that can receive visual features, structural frames, and a mask as input data and can provide depth values as output for at least a portion of an image, for example, specifically for providing depth values for portions of an image within an object mask.
[0148] This artificial neural network can be trained using images whose depth values have been determined, for example, manually or otherwise. Such training images can be analyzed using the same feature extraction, frame determination, and mask determination used in the depth determiner 105 during normal operation. The resulting masks, frames, and visual features can be used as input to the artificial neural network, and the generated output can be compared with depth values determined manually or otherwise. As a concrete example of training data, a suitable model can be used to synthesize a synthetic / virtual scene. Images representing the virtual world can then be generated, and corresponding depth values can be determined. These images can then undergo feature extraction, frame determination, mask determination, and depth determination via the artificial neural network, where the results are compared with depth values, and the artificial neural network is adjusted in response.
[0149] This trained artificial neural network can provide very good depth estimation in many scenarios, and this can often be achieved without considering information other than visual features, masks, and the 2D structural frame. Specifically, for objects such as humans and animals, where the 2D structural frame represents the skeleton, depth estimation can often be accurate even without considering other information, given sufficient training data. Given enough training data, this method can provide accurate depth information without requiring multiple images representing the scene from different orientations or any dedicated depth measurement. However, for use cases where a new view needs to be synthesized from multiple reference views, it is still possible to use multiple images of the view... Figure 1 To the depths.
[0150] In some embodiments, depth determination can therefore be based on visual features / attributes of an image. The method can extract a 2D structural framework (e.g., representing a human pose) and a mask from a single, for example, RGB image, and combine this information with the extracted visual features to produce a depth map of an object (e.g., a human body).
[0151] Figure 5 An example of an element determined by the depth of this method is shown. In this example, the object is specifically a human body, and the structural framework is a human skeleton structure / model.
[0152] In this example, a single image of a single view is analyzed, and depth data is determined based on that single image. In this method, a structural framework determination is performed on the image, which in this specific example is the extraction of a human skeleton model, and can specifically be performed by an artificial neural network as described above. Visual feature extraction is also performed on the same image, generating visual features as described above. Furthermore, human mask segmentation is performed on the image, which extracts / determines the masks of detected human bodies in the image.
[0153] The resulting structural framework, visual features, and mask are then fed into a data fusion algorithm that generates depth data. Specifically, the information can be fed into an artificial neural network, which then continues to generate, for example, a depth map of an image.
[0154] It has been found that this method provides efficient and advantageous depth determination in many scenarios. In many embodiments, it can advantageously use multiple artificial neural networks and provide the specific advantage that different (smaller) artificial neural networks can focus on different aspects and then be combined to provide (computationally) more efficient and / or more accurate depth determination.
[0155] In some embodiments, the depth determiner 105 is arranged to receive a set of initial depth values of an image, and then it can determine a depth map by adjusting that set of initial depth values according to a two-dimensional frame structure. The depth determiner 105 may receive initial depth values, such as an initial depth map, and then continue to refine / modify these depth values based on a 2D structural frame determined for the image. The depth determiner 105 may specifically impose constraints on the depth values of the depth map.
[0156] Specifically, in many embodiments, the depth determiner 105 may be arranged to bias depth values to depth values determined for the 2D structural frame, wherein the bias increases as the distance from the depth pixel to the 2D structural frame decreases.
[0157] In many embodiments, receiver 101 may be arranged to receive multiple images representing a scene from different viewpoints. Figure 1 The apparatus may include a disparity estimator 107, which is arranged to receive multiple images and perform disparity-based depth estimation based on images from different viewpoints. Such disparity estimation may be based on detecting corresponding or matching image segments in different images and then determining the corresponding depth based on the disparity / positional offset of the image segments in the different images and the viewpoint differences between them, as is known in the art.
[0158] Therefore, in some embodiments, disparity estimation can be used to generate an initial depth map from, for example, a set of images. This depth map can provide a suitable depth estimate and a useful depth map. However, in Figure 1 In the example, the depth map is further improved by adjusting the depth based on information from the 2D structural framework. This additional information can provide more consistent and accurate improved depth values, and in particular, can provide improved depth value determination for, for example, people or animals in a scene / image.
[0159] Specifically, the depth determiner 105 can modify depth values that approximate a 2D structural frame, making these depth values more consistent with the 2D structural frame. The 2D structural frame introduces constraints on objects; for example, for a human skeletal structural frame representing a given human pose, the positions / depths of different parts of the body (such as arms, legs, head, torso, etc.) are constrained, particularly the positions / depths of different points / pixels of the body relative to each other. For example, depth variations along directions corresponding to the interconnections between, for example, points corresponding to the wrist and points corresponding to the elbow will be very limited and, for example, correspond to linear transitions. The depth determiner 105 can accordingly adjust the depth values of the depth map along corresponding lines in the image represented by the 2D structural frame. As a low-complexity example, a spatial low-pass filter can be applied along a portion of the image from the wrist point to the elbow point.
[0160] Figure 6An example of an element determined by the depth of this method is shown. In this example, the object is specifically a human body, and the structural framework is the human skeleton structure.
[0161] In this example, two images of a scene including people are received, captured from different angles; that is, two view images from different viewpoints are received. Based on these two images, depth estimation is performed, specifically, depth estimation based on disparity matching is performed to generate a depth map.
[0162] In this method, a structural frame determination is performed on one of the images. In this specific example, the structural frame determination is the extraction of a human skeleton model, and can be specifically performed by an artificial neural network, as previously described. Typically, a structural frame can be determined for a view image that is the same viewpoint from which the depth map is also determined.
[0163] The initial depth map's depth values are then processed and updated using a structural frame / human skeleton model to generate an output depth map. This depth refinement can vary in different embodiments, and as described above / below, several different methods can be used. Typically, the depth values of the depth map can be modified to constrain them to match the depth values of the structural frame; specifically, the depth values can be biased to the depth values determined by the 2D structural frame.
[0164] In many embodiments, the depth determiner 105 may be arranged to bias the depth values near the 2D structural frame to the depth values of the 2D structural frame.
[0165] For example, depth pixels consistent with the 2D structural frame can be identified first, and the depth values of these pixels can be determined based on initial values. For example, the depth transition between two interconnecting points of the 2D structural frame can be considered a linear depth transition. Therefore, a linear depth transition can be applied by fitting the depth values along the interconnect, or, for example, the depth values of the interconnecting points can be assumed to be correct, and the depth values along the intermediate depth values of the interconnect can be modified / set to provide a gradual linear transition from the depth of one endpoint to the other.
[0166] The depth determiner 105 can then continue aligning the depth values around pixels that coincide with the 2D structural frame with the depth values of the determined aligned pixels. As a concrete example, a low-pass spatial filter can be applied. Such a low-pass filter can, for example, have a kernel asymmetric to the depth pixels relative to the determined depth values, and this asymmetry causes depth pixels facing towards and typically encompassing the 2D structural frame to have higher weights than pixels farther from the 2D structural frame.
[0167] In many embodiments, the depth determiner 105 can be arranged to reduce the difference between the depth values of the depth map and the depth values of the 2D structural frame, wherein this reduction depends on the distance between the pixel of the depth value and the 2D structural frame. Therefore, the closer a pixel is to the 2D structural frame, the stronger the reduction, and thus the more the depth values are aligned with the depth values of the 2D structural frame. This method can be applied, for example, to each pixel within a given distance of the 2D structural frame, thereby producing a more accurate depth map.
[0168] For example, when a 2D structural framework represents the human skeleton, this method may result in the depth values of pixels reflecting the corresponding human body being more closely aligned with the depth values of the underlying skeleton and the current human pose. Information about the human pose and body in the image provided by the 2D structural framework is used to constrain / adjust the depth values to make them more accurate.
[0169] In some embodiments, a convolutional neural network can be trained, taking as input a depth mapped from the skeleton (structural framework), a mask mapping which pixels the skeleton intersects with, and an initial depth map. Given a baseline ground truth depth map, the coefficients of the convolutional depth refinement neural network are fitted.
[0170] In some embodiments, a 2D structural frame, such as a human skeleton frame indicating human posture, can be used as supplementary information for depth map adjustment / improvement. The initial raw depth data can be generated in various ways, including, for example, through multi-view matching and disparity estimation. The 2D structural frame determined from one (or more) views can be used as supporting information for post-processing of the initial depth map, thereby refining the original initial depth data.
[0171] For example, this method can reduce noise, artifacts, and inaccuracies in depth maps / estimations. This improvement can specifically provide more accurate and consistent depth maps for parts corresponding to objects (e.g., the human body). This is particularly advantageous because errors and noise in such parts can have a greater impact on the perceived image quality of any image generated based on the depth map. For example, a viewer might notice distortion or noise / inaccuracies in the rendering of the human body.
[0172] One or more devices can be specifically implemented in one or more appropriately programmed processors. For example, an artificial neural network can be implemented in one or more such appropriately programmed processors. Different functional blocks, in particular artificial neural networks, can be implemented in separate processors and / or can be implemented, for example, in the same processor. Examples of suitable processors are provided below.
[0173] Figure 7This is a block diagram illustrating an example processor 700 according to an embodiment of the present disclosure. Processor 700 can be used to implement one or more processors that implement the means or elements thereof as described above (particularly including one or more artificial neural networks). Processor 700 can be any suitable processor type, including but not limited to microprocessors, microcontrollers, digital signal processors (DSPs), field-programmable gate arrays (FPGAs) (wherein the FPGA is programmed to form a processor), graphics processing units (GPUs), application-specific integrated circuits (ASICs) (wherein the ASIC is designed to form a processor), or combinations thereof.
[0174] Processor 700 may include one or more cores 702. Core 702 may include one or more arithmetic logic units (ALUs) 704. In some embodiments, in addition to or in place of ALU 704, core 702 may include a floating-point logic unit (FPLU) 706 and / or a digital signal processing unit (DSPU) 708.
[0175] Processor 700 may include one or more registers 712 communicatively coupled to core 702. Registers 712 may be implemented using dedicated logic gates (e.g., flip-flops) and / or any memory technology. In some embodiments, registers 712 may be implemented using static memory. Registers may provide data, instructions, and addresses to core 702.
[0176] In some embodiments, processor 700 may include one or more levels of cache memory 710 communicatively coupled to core 702. Cache memory 710 may provide computer-readable instructions to core 702 for execution. Cache memory 710 may provide data for core 702 to process. In some embodiments, the computer-readable instructions may have already been provided to cache memory 710 by local memory (e.g., local memory attached to external bus 716). Cache memory 710 may be implemented using any suitable cache memory type, such as metal-oxide-semiconductor (MOS) memory, static random access memory (SRAM), dynamic random access memory (DRAM), and / or any other suitable memory technology.
[0177] Processor 700 may include controller 714, which controls inputs to processor 700 from other processors and / or components included in the system and / or outputs from processor 700 to other processors and / or components included in the system. Controller 714 may control data paths in ALU 704, FPLU 706, and / or DSPU 708. Controller 714 may be implemented as one or more state machines, data paths, and / or dedicated control logic. The gates of controller 714 may be implemented as stand-alone gates, FPGAs, ASICs, or any other suitable technology.
[0178] Register 712 and cache 710 can communicate with controller 714 and core 702 via internal connections 720A, 720B, 720C, and 720D. These internal connections can be implemented as buses, multiplexers, crossbar switches, and / or any other suitable connection technology.
[0179] Inputs and outputs for processor 700 may be provided via bus 716, which may include one or more conductive lines. Bus 716 may be communicatively coupled to one or more components of processor 700, such as controller 714, cache 710, and / or register 712. Bus 716 may be coupled to one or more components of the system.
[0180] Bus 716 may be coupled to one or more external memories. The external memory may include read-only memory (ROM) 732. ROM 732 may be a mask ROM, electrically programmable read-only memory (EPROM), or any other suitable technology. The external memory may include random access memory (RAM) 733. RAM 733 may be static RAM, battery-backed static RAM, dynamic RAM (DRAM), or any other suitable technology. The external memory may include electrically erasable programmable read-only memory (EEPROM) 735. The external memory may include flash memory 734. The external memory may include a magnetic storage device such as a disk 736. In some embodiments, the external memory may be included within the system.
[0181] This invention can be implemented in any suitable form, including hardware, software, firmware, or any combination thereof. The invention can optionally be implemented, at least in part, as computer software running on one or more data processors and / or digital signal processors. The elements and components of embodiments of the invention can be implemented physically, functionally, and logically in any suitable manner. In practice, functionality can be implemented in a single unit, multiple units, or as part of other functional units. Therefore, the invention can be implemented in a single unit or can be physically and functionally distributed among different units, circuits, and processors.
[0182] Although the invention has been described in conjunction with some embodiments, it is not intended to be limited to the specific forms set forth herein. Rather, the scope of the invention is defined only by the claims. Furthermore, while features may appear to have been described in conjunction with specific embodiments, those skilled in the art will recognize that various features of the described embodiments can be combined according to the invention. Throughout this application, the terms “2D” and “3D” are equivalent to “two-dimensional” and “three-dimensional,” respectively. In the claims, the term “comprising” does not exclude the presence of other elements or steps.
[0183] Furthermore, although listed separately, multiple devices, elements, circuits, or method steps can be implemented by, for example, a single circuit, unit, or processor. Additionally, although individual features may be included in different claims, these features can be advantageously combined together, and inclusion in different claims does not imply that the combination of features is infeasible and / or disadvantageous. Moreover, including a feature in one class of claims does not imply a limitation on that class, but rather indicates that the feature is equally applicable to other claim classes. Furthermore, the order of features in a claim does not imply that the features must operate in any particular order; in particular, the order of steps in a method claim does not imply that the steps must be performed in that order. Rather, these steps can be performed in any suitable order. Furthermore, singular references do not exclude plural. Therefore, references to “a,” “an,” “first,” “second,” etc., do not exclude plural. Reference numerals in the claims are provided only as clarifying examples and should not be construed as limiting the scope of the claims in any way.
Claims
1. An apparatus for generating a depth map of a first image of a scene, the apparatus comprising: A receiver (101) is arranged to receive at least the first image; A structural frame determiner (103) is arranged to process the first image to determine a two-dimensional structural frame of an object in the first image, the two-dimensional structural frame being defined by a set of points and the interconnections between the points, the two-dimensional structural frame representing a projection of the three-dimensional structural frame of the object in the scene onto the image space of the image; A depth determiner (105) is arranged to generate the depth map of the first image based on the two-dimensional structural framework.
2. The apparatus according to claim 1, wherein, The structural frame determiner (103) includes a trained artificial neural network configured to receive the first image as input and generate points of the two-dimensional structural frame as output.
3. The apparatus according to any of the preceding claims, wherein, The receiver (101) is arranged to receive a plurality of images, the structural frame determiner (103) is arranged to generate a two-dimensional structural frame for at least some of the plurality of images, and the depth determiner (105) is arranged to determine the depth value of the depth map based on the disparity between points of the two-dimensional frame structure of the at least some images.
4. The apparatus according to claim 3, wherein, The depth determiner (105) is arranged to determine an estimated three-dimensional frame structure of the object based on the disparity between points of the two-dimensional frame structure in the at least some images, and to determine the depth values of the points of the two-dimensional frame structure in the at least some images by projecting the points of the estimated three-dimensional frame structure into the image space of the at least some images.
5. The apparatus according to claim 3 or 4, wherein, The depth determiner (105) is arranged to generate a depth value for the interconnection between the first point and the second point based on the depth values of the first point and the second point of the two-dimensional frame structure.
6. The apparatus according to any of the preceding claims, wherein, The depth determiner (105) is configured to determine an initial depth map and generate the depth map by performing the following steps on at least a first pixel of the depth map: A set of candidate depth values is determined, the set of candidate depth values including the depth values of other pixels in the depth map besides the first pixel and at least a first candidate depth value determined according to the two-dimensional structural framework; In response to the cost function, the cost value of each candidate depth value in the set of candidate depth values is determined; In response to the cost value of the set of candidate depth values, a first depth value is selected from the set of candidate depth values; In response to the first depth value, an updated depth value for the first pixel is determined; The cost value of the first candidate depth value depends on the difference between the candidate depth value and the depth value of the two-dimensional structural frame.
7. The apparatus according to claim 6, wherein, The cost of the first candidate depth value depends on the distance between the position of the candidate depth value and the position of the depth value of the two-dimensional structural frame.
8. The apparatus according to any of the preceding claims, wherein, The depth determiner (105) is arranged to extract visual features from the first image and determine an image mask of the object in the first image, and wherein the depth determiner (105) is further arranged to determine a depth estimate of the depth map based on the visual features, the mask and the two-dimensional structural frame.
9. The apparatus according to claim 8, wherein, The depth determiner (105) includes a trained artificial neural network configured to receive the visual features, the mask, and the two-dimensional structural frame as input and determine the depth map as output.
10. The apparatus according to claim 8 or 9, wherein, The depth determiner includes a trained artificial neural network configured to receive the first image as input and generate the mask as output.
11. The apparatus according to any of the preceding claims, wherein, The depth determiner (105) is arranged to receive a first set of depth values of the image and determine the depth map by adjusting the first set of depth values according to the two-dimensional frame structure.
12. The apparatus according to claim 11, wherein, The at least first image comprises a plurality of images representing the scene from different viewpoints, and the apparatus includes a disparity estimator (107) arranged to determine the first set of depth values by disparity estimation between the plurality of images.
13. The apparatus according to claim 11 or 12, wherein, The depth determiner (105) is arranged to determine the depth value of the first pixel of the depth map by reducing the difference between the depth value of the first depth pixel in the first set of depth values and the depth value of the two-dimensional frame structure, the reduction depending on the distance between the first depth pixel and the two-dimensional depth structure.
14. A method for determining a depth map of a first image of a scene, the method comprising: Receive at least the first image; The first image is processed to determine a two-dimensional structural framework of an object in the first image. This two-dimensional structural framework is defined by a set of points and the interconnections between those points. The two-dimensional structural framework represents a projection of the three-dimensional structural framework of the object in the scene onto the image space of the image. The depth map of the first image is generated based on the two-dimensional structural framework.
15. A computer program product comprising a computer program code module, wherein when the program is run on a computer, the computer program code module is adapted to perform all the steps of claim 1.
Citation Information
Patent Citations
Apparatus and method for processing a depth map
EP4013049A1