Generates displacement maps for pairs of input datasets of image or audio data
Through the feature extractor and displacement unit based on neural network, a disparity map of the solid image pair is generated, which solves the problems of high computing cost and high complexity in the prior art, and realizes efficient and low-complexity disparity map generation.
Patent Information
- Application Number
- CN201980036652.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-05-30
- Filing Date
- 2019-05-29
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2039-05-29
AI Technical Summary
In the prior art, when generating a parallax map of a volume image pair, the calculation cost is high and the complexity is high, making it difficult to achieve efficient calculations.
A feature extractor branch based on neural network is used to generate feature maps of multiple scales through multi-layer convolution and pooling layers, and a displacement unit and a displacement refinement unit are used to generate displacement maps of different levels through the hierarchical level of the feature map.
Generate high-quality and high-resolution parallax maps at low complexity and low computing costs, suitable for autonomous vehicles and other computer vision applications.
Smart Images

Figure CN112219223B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method and apparatus for generating a displacement map of a first input data set and a second input data set in an input data set pair (eg, a disparity map of a left / first image and a right / second image in a stereoscopic image pair). Background Art
[0002] Accurately estimating depth from stereo images (in other words, generating disparity maps from stereo images) is a core problem in many computer vision applications such as autonomous (self-driving) vehicles, robotic vision, augmented reality. More generally, the study of the displacement between two (correlated) images is a widely used tool today.
[0003] Accordingly, several different approaches can be used to generate a disparity map based on stereo images.
[0004] In the methods disclosed in US 5,727,078, US 2008 / 0267494 A1, US 2011 / 0176722 A1, US 2012 / 0008857 A1, US 2014 / 0147031 A1 and US 9,030,530 B2, low-resolution images are generated from stereo images recorded from a scene. A disparity analysis is performed on these low-resolution images, and the disparity map obtained by the analysis is enlarged, or its accuracy is improved by gradually applied enlargements.
[0005] A separate layered series of images with increasingly lower resolutions is generated for each stereo image in WO 00 / 27131 A1 and WO 2016 / 007261 A1. In these approaches, a disparity map is generated based on the coarsest level of left and right images. In these documents, upscaling of low-resolution disparity maps is applied.
[0006] In CN 105956597 A, stereo images are processed with the aid of a neural network to obtain a disparity map. In the following papers (some of which are available as preprints in the arXiv open access database of the Cornell University Library), a neural network is used to generate feature maps and, for example, disparity or other types of outputs of stereo images:
[0007] A. Kendall et al.: End-to-end learning of geometry and context for deepstereo regression, 2017, arXiv:1703.04309 (hereinafter referred to as Kendall);
[0008] Y. Zhong et al.: Self-supervised learning for stereo matching with self-improving ability, 2017, arXiv:1709.00930 (hereinafter referred to as Zhong);
[0009] N. Mayer et al.: FlowNet: Learning optical flow with convolutional networks, 2015, arXiv:1504.06852 (hereinafter referred to as Mayer);
[0010] Ph. Fischer et al.: FlowNet: Learning optical flow with convolutional networks, 2015, arXiv:1504.06852 (hereinafter referred to as Fischer);
[0011] J. Pang et al.: Cascade residual learning: A two-stage convolutional neural network for stereo matching, 2017, arXiv:1708.09204 (hereinafter referred to as Pang);
[0012] C. Godard et al.: Unsupervised monocular depth estimation with left-right consistency, 2017, arXiv:1609.03677 (hereinafter referred to as Godard).
[0013] A disadvantage of many prior art approaches to applying neural networks is the high complexity of implementation. In addition, in most known approaches, the computational cost is high.
[0014] In view of the known methods, there is a need for a method and apparatus for generating a displacement map of a first input data set and a second input data set in an input data set pair (e.g., a disparity map for generating a stereoscopic image pair), whereby the displacement map (e.g., a disparity map) can be generated in a computationally cost-effective manner. Summary of the invention
[0015] The main object of the present invention is to provide a method and device for generating a displacement map of a first input data set and a second input data set in an input data set pair (for example, a disparity map of a stereo image pair), which method and device are free from the disadvantages of the prior art methods to the greatest possible extent.
[0016] It is a further object of the present invention to provide a method and apparatus for generating a displacement map, whereby a displacement map of good quality and resolution can be generated in a computationally cost-effective manner with as low complexity as possible.
[0017] The objects of the invention are achieved by a method according to claim 1 and by an apparatus according to claim 13. Preferred embodiments of the invention are defined in the dependent claims.
[0018] Throughout this document, "displacement map" shall refer to some n-dimensional generalized disparity, i.e., a generalization of the well-known stereo (left-right) disparity (disparity map is a special case of displacement map) or other kinds of displacement maps (such as flow maps in the case of studying optical flow). However, the embodiments described in detail with reference to the accompanying drawings are explained with the help of the example of the well-known stereo disparity (see Figure 2 ). A displacement map (eg, an n-dimensional generalized disparity map) is a mapping that assigns an n-dimensional displacement vector to each spatial location (such as a pixel on an image).
[0019] Assume that an input data set (e.g., an image) or a feature map (these are generalized images and feature maps processed by the method and apparatus according to the present invention) is given a multidimensional tensor having at least one spatial dimension and / or a temporal dimension (the dimensionality of the spatial dimension is otherwise unconstrained, it can be 1, 2, 3 or more), and a channel dimension (the channel corresponding to the dimension can be called a feature channel or simply a channel). If the channel dimension of the input or feature map is N, then the input or feature map has N channels, for example, an RGB image has 3 channels and a grayscale image has 1 channel. The meaning of the expression "spatial dimension" is not constrained here, it is only used to distinguish these dimensions from the channel dimension, and the temporal dimension can also be defined separately. Subsequently, the (generalized) displacement defined above can be represented as a tensor having the same spatial dimension as the input data set and having a (coordinate) channel dimension, the length of which is the same as the dimension of the spatial dimension of the input data set tensor, or less than the dimension of the spatial dimension when the task is restricted to a subset of the spatial dimension. For each spatial position (i.e., pixel or voxel, generally speaking, data element) of the input data set, the displacement map thus determines the direction and magnitude of the displacement, i.e., the displacement vector. The coordinates of these displacement vectors should form the coordinate channel dimensions of the tensor representing the displacement map (the channels of the displacement map may be referred to as coordinate channels to distinguish them from the feature channels of the feature map).
[0020] The concept of displacement in the sense of generalized disparity should also include those cases where the displacement vector is restricted to a subset of possible displacement vectors. For example, a typical stereo disparity of a two-dimensional left input data set tensor and a right input data set tensor (in this case, an image) can be described by constraining the generalized disparity so that only displacements in the horizontal direction are allowed. Thus, the displacement vectors can be described by one coordinate instead of two coordinates because their vertical (y) coordinate is always zero (this is the case in the example illustrated in the accompanying drawings). Therefore, in this case, the generalized disparity can be represented by a tensor whose coordinate channel dimension is 1 instead of 2, and whose values represent only horizontal displacements. That is, the number of coordinate channels of the disparity tensor can be equal to the actual dimension of the subspace of possible displacement vectors, that is, 1 instead of 2 in this case.
[0021] In three spatial dimensions, image elements are often referred to as voxels, in two dimensions they may be referred to as pixels, and in one dimension they may be referred to as samples (e.g., in the case of speech recordings) or sample values. In this document, the spatial locations of the input dataset tensors (and feature maps) are referred to as "data elements", regardless of the number of spatial dimensions.
[0022] Thus, the input data sets need not be RGB images, they can be any kind of multidimensional data, including but not limited to: time series (such as audio samples), ECG, 2D images encoded in any color space (such as RGB, grayscale, YCbCr, CMYK), 2D thermal imager images, 2D depth maps from depth cameras, 2D depth maps generated by sparse LIDAR or RADAR scans, or 3D medical images (such as MRI, CT). The technical representations of the displacement map are not limited to those described above, and other representations can also be conceived. In general, the input data set (which can be simply referred to as input, input data or input data set) typically has data elements in which images or other information are stored, that is, the input data set typically includes digitized information (for example, it is a digitized image), and in addition, the input data set is typically a record (that is, an image of a camera).
[0023] There are several special cases of application of displacement maps (generalized disparity) that are worth noting. In the case of stereo disparity, the two input images (a special case of the input dataset) correspond to images of the left and right cameras (i.e., members of a stereo pair), and the generalized disparity is actually a 2D vector field constrained to the horizontal dimension, so it can be (and usually is) represented by a 1-coordinate channel tensor with two spatial dimensions.
[0024] In another case of displacement maps (generalized disparity), i.e. in the case of 2D optical flow, two input images (input data sets) correspond to frames taken by the same camera at different times (e.g. a previous frame and a current frame), i.e. these frames constitute an input image pair. In the case of 3D image registration (e.g. another image matching process for medical purposes), one image is, for example, a diffusion MRI brain scan and the other image is a reference brain scan (i.e. in this process, the recording obtained by the study is compared with the reference), these two images constitute an input data set pair; and the generalized disparity is a 3D vector field, i.e. a tensor with three spatial dimensions and three coordinate channels in each tensor position of the three spatial coordinates. As mentioned above, the method and apparatus according to the present invention can be generalized to any spatial dimension in which an appropriate configuration of the feature extractor is applied.
[0025] Another possible use case is to match two audio recordings in time, where the dimensionality of the spatial dimension is one (it is actually the time dimension), the two input data sets correspond to an audio recording and another reference recording, and the generalized disparity (displacement) has one spatial dimension (the time dimension) and only one coordinate channel (this is because the input data sets have one spatial dimension). In this case, the feature extractor is applied to the time-dependent function of the audio amplitude (with a function value for each time instance). In addition to these examples, other applications can be envisioned, which may have even more than three spatial dimensions. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Preferred embodiments of the present invention are described below by way of example with reference to the accompanying drawings, in which:
[0027] Figure 1 An exemplary neural network-based feature extractor branch is explained,
[0028] Figure 2 A block diagram showing an embodiment of the method and apparatus according to the present invention is shown.
[0029] Figure 3 shows a block diagram of a parallax unit in one embodiment,
[0030] Figure 4 is a schematic illustration of a shifting step in one embodiment,
[0031] Figure 5A A block diagram illustrating an embodiment of a comparator unit applied in a disparity (refinement) unit,
[0032] Figure 5B Possible variations for calculating the disparity in the comparator unit are explained, and
[0033] Figure 6 A block diagram of an embodiment of a disparity refinement unit is shown. DETAILED DESCRIPTION
[0034] Figure 1 An exemplary neural network based feature extractor branch 15 (ie, a single branch feature extractor) is shown applied to an image (as an example of an input dataset). Figure 1 The details are for illustration only. The parameters shown in the figure can be selected in several different ways, see below for details. The illustrated feature extractor branch 15 can be used exemplarily in the method and apparatus according to the present invention. Figure 2 , feature extractor branches 27 and 29 are shown (which are part of feature extractor 25); these feature extractor branches are connected to Figure 1 The feature extractor branch 15 is very similar to , in order to simplify the Figure 1 and Figure 2 Description of this aspect of the invention.
[0035] Also like Figure 2 As shown in the figure, the pixel group selected at a level is the convolution input (see the convolution input 22a, 32a exemplarily indicated in the best resolution level (the highest level in the figure)) of the corresponding convolution (see convolution 24a, 34a) between the levels. The output of the convolution is taken as the convolution output (see the convolution output 26a, 36a exemplarily indicated in the highest possible level) by summing each pixel of the convolution input weighted by the corresponding element of the convolution kernel (which may also be called a kernel or filter). If the number of output channels in the convolution kernel is greater than 1, multiple weighted sums (one weighted sum for each output channel) will be calculated and stacked together to form the convolution output. Figure 2 The grid size in illustrates the size of the convolution kernel (see convolution inputs 22a, 32a). Specific convolution inputs, kernels, convolutions and convolution outputs are shown for illustrative purposes only. In the illustrated example, a 3×3 kernel (i.e., a convolution with a kernel size of 3 pixels in each spatial dimension) is applied. In the 2D case, the kernels are simply labeled 3×3, 1×1, etc., but in general, it is relevant to define the kernel size in each spatial dimension. Kernels of other sizes can also be applied, i.e., kernels with unequal sides to each other. For example, a 3×1 kernel and a 1×3 kernel can be used in sequence. The number of feature channels can also be tuned by parameters of each feature map.
[0036] As emphasized above, Figure 1is only an illustration of a neural network based feature extractor (branch). A learnable convolution (which may also be referred to simply as convolution or convolution operator) is applied to an input image 10a, thereby successively generating a series of feature maps 10b, 10c, 10d and 10e. The convolution applied in embodiments of the present invention (i.e., in the neural network feature extractor and in the comparator unit) is a convolution with a learned filter (kernel) rather than a convolution with a predefined filter as applied in some known approaches mentioned above. The convolution sweeps across the image by moving the convolution kernel across the image to which it is applied. For simplicity, in Figure 1 and Figure 2 In , only exemplary locations of convolution kernels are shown for each level of the feature extractor.
[0037] Preferably, each convolution is followed by a nonlinearity layer (usually a ReLU (short for Rectified Linear Unit), where f(x)=max(0,x)).
[0038] The convolution can reduce the spatial dimension of the input layer if its stride is > 1. Alternatively, a separate pooling (usually average pooling or max pooling) layer can be used to reduce the spatial dimension (usually by 2).
[0039] When the stride is 1, the kernel moves one pixel at a time during the sweep mentioned above. When the stride is 2, the kernel jumps 2 pixels at a time when moving. The stride value is usually an integer. With higher strides, the kernels overlap less and the resulting output has a smaller spatial dimension (see also the following for Figure 1 Analysis of the illustrated examples).
[0040] As mentioned above, the spatial dimension can also be reduced with the help of pooling layers. Commonly applied average pooling layers and maximum pooling layers are applied as follows. Maximum pooling is applied to groups of pixels, selecting the pixel with the maximum value for each feature channel. For example, if one unit of the maximum pooling layer considers a 2×2 pixel group, where each pixel has a different intensity in each channel, the pixel with the maximum intensity is selected for each channel, and the 2×2 pixel group is thus reduced to a 1×1 pixel group (i.e., to a single pixel) by combining the channel-by-channel maximum values. The average pooling layer is applied similarly, but it outputs the average of the pixel values in the pixel group studied for each channel.
[0041] exist Figure 1In the example shown in , the width and height parameters of each level (image and feature map) are indicated. Image 10a has a height H and a width W. These parameters for the first feature map 10b (the first feature map in a row of feature maps) are H / 4 and W / 4, i.e. the spatial parameters are reduced by a factor of 4 between the starting image 10a and the first feature map 10b. The dimensions of feature map 10c are H / 8 and W / 8. Correspondingly, in the next step, the spatial parameters are reduced by a factor of 2, which is similar for subsequent steps (the dimensions of feature map 10d are H / 16 and W / 16, and for feature map 10e the dimensions are H / 32 and W / 32). This means that the feature map hierarchy (see also Figure 2 The number of pixels of the feature maps in the feature map level 50 of the feature extractor 10a is successively reduced, that is, the resolution is gradually reduced from image 10a to the last feature map 10e. In other words, feature maps with lower and lower resolutions are obtained in the sequence (at deeper and deeper levels of the feature extractor). Since these feature maps are at deeper and deeper stages of the feature extractor, it is expected that these feature maps have fewer and fewer image-like properties and more and more feature-like properties in the sequence. Since the size of the feature maps is getting smaller and smaller, the kernel covers an increasingly larger area in the initial image. As a result, the recognition of increasingly larger objects in the image can be achieved in the feature maps of lower resolution. As an example, pedestrians, pets or cars can be recognized in lower resolution levels (these are feature maps that usually include higher-level features, so these feature maps can also be called high-level feature maps), but only parts or details of these objects (for example, heads or ears) can be recognized at higher resolution levels (these feature maps at these levels usually include lower-level features, so these feature maps can also be called low-level feature maps).
[0042] The feature extractor is preferably pre-trained to advantageously achieve better results and avoid overfitting. For example, the (neural network-based) machine learning component for the device and corresponding method is taught (or pre-trained) with many different pets and other feature objects. With appropriate pre-training, higher efficiency can be achieved in identifying objects (see below for details explained with the example of disparity analysis).
[0043] During pre-training / training, a loss value (an error-like amount) is generated using an appropriate loss function. During training, a loss value is preferably generated on the output of the last displacement refinement unit (i.e., the output of the displacement refinement unit that outputs the final displacement map when the device is untrained and used to generate a displacement map; this is, for example, a disparity refinement unit). The loss function is preferably also applied to other layers of the neural network. The multiple loss functions can be considered as a single loss function because these loss functions are summed. The loss function can be placed on the output of all displacement units and at least one displacement refinement unit (refined to scales of 1 / 32, 1 / 16, 1 / 8, and 1 / 4 compared to the starting image in the exemplary refinement step). With the loss function applied to these units, reasonable displacement maps can be achieved at all levels.
[0044] Furthermore, pre-training and training are typically two separate phases, where the tasks and data to be performed, as well as the network architecture and the loss function applied, can be different. Furthermore, pre-training can include multiple phases, where the network is trained on successively more specific tasks (e.g., ImageNet->synthetic data->real data->real data on more specific domains; any of these phases can be skipped, and any task for which appropriate data is available, easier to teach, or from which better / more diverse feature capture can be expected can be applied for learning.
[0045] In one example, displacements are not taught in the pre-training phase, so a different loss function must be used. When pre-training is complete, all parts of the network that were used for pre-training but not for displacement map generation are omitted when starting the displacement map training phase (this may be the case for the last layer or two, if any). Layers that are necessary for displacement map generation but not necessary for pre-training are then placed on the network, and the weights of these new layers are initialized in some way. This allows the network to be taught partially or fully on a new task (in this case displacement map generation) or in a new way.
[0046] In one example, in the case of pre-training with ImageNet, the classification part, which is tasked with classifying images into appropriate categories (e.g., object classification: bulldog, St. Bernard, cat, table), which will be discarded later, is placed at the end of the network. Thus, in this variant, filters with better discriminative properties are implemented. The classification part (the end of the network) is then eliminated, since no classification is explicitly performed in the main task, only in pre-training.
[0047] In addition, such pre-training as follows can also be applied (alone or as a second stage after ImageNet learning), in which the same architecture is used but displacement is taught in a different way and / or on different data. This is advantageous because a large amount of synthetic data can be generated in a fairly cheap way in terms of computational cost. In addition, perfect ground truth displacement maps (e.g., disparity or optical flow) can be generated for synthetic data, whereby it is best to teach with these synthetic data. However, in the case of real images, the network taught on synthetic data cannot achieve the same quality. Therefore, training on real data is also necessary. In the case where, for example, a large number of stereo image pairs can be obtained from some different environment (these stereo image pairs are then applied to pre-training), training on real data can also be some type of pre-training, but may have fewer teaching images from the real target domain (these images are then retained for training). This is the case in examples where the goal is to achieve good results on the KITTI benchmark but only a small amount of teaching data is available from the environment.
[0048] The loss function is typically back-propagated in the device to fine-tune the machine learning (neural network) component, but any other training algorithm may be used instead of back-propagation. In supervised approaches to displacement (e.g., disparity) generation, the resulting displacement map (and calculated depth values) may be compared to, for example, LIDAR data; such data is not required in self-supervised or unsupervised approaches.
[0049] Now turn to the issue of channel dimension in conjunction with the feature map hierarchy. Performing convolutions usually increases the feature channel dimension (for RGB images, this dimension is initially 3, for example). The increase in channel dimension is Figure 1 This is illustrated in the figure by explaining the thickness of the objects in the starting image and feature map. The typical number of channels at 1 / 32 scale is 1024 to 2048.
[0050] It should be emphasized that Figure 1 This is just an illustration of a feature extractor branch based on a convolutional neural network. Any other layer configuration with any kind of learnable convolutional layers, pooling layers, and various non-linearities can be used. Layer configurations can also include multiple computation paths that can be combined later (such as GoogleNet variants, or skip connections in the case of ResNet). Applicable layer configurations have two things in common: 1. They can learn a transformation on the input, and 2. They have feature outputs (a base feature map at the bottom of the feature map hierarchy, and at least one (usually more) intermediate feature maps) typically at multiple scales. The feature extractor network can also be called the "base network".
[0051] In summary, the image 10a is used as input to the feature extractor branch 15. The feature map 10b is the first feature map in the illustrated exemplary feature map hierarchy, which includes further feature maps 10c, 1Od, 1Oe.
[0052] exist Figure 1 , a convolution 14a with an exemplary convolution input 12a is applied to a starting image 10a (which is one of the images in a stereo pair), and a convolution output 16a on a first feature map 10a is shown (which is the output of the convolution 14a). The exemplary convolution 14a is applied to a 3×3 pixel group (applied to all one or more channels of the starting image 10a), and its convolution output 16a is a single pixel (with a specific number of channels equal to the number of channels defined by the convolution kernel, see the convolution kernel depth for explaining the number of channels. The specific convolution 14a, and the corresponding convolution input 12a and convolution output 16a merely constitute illustrative examples, and the convolution kernel is moved across the entire image (for a sequence of feature maps, across the entire feature map) according to a prescribed rule. Naturally, the kernel size may be different, or other parameters may be different. For feature maps 10b-10e, Figure 1 Further exemplary convolutions 14b, 14c, 14d having convolution inputs 12b, 12c, 12d and convolution outputs 16b, 16c, 16d are illustrated in FIG.
[0053] The method and apparatus according to the present invention are suitable for generating a displacement map of a first input data set and a second input data set in an input data set pair (in Figure 2 In the example of , the displacement map is a disparity map of a left image 20a and a right image 30a in a stereo image pair; the members of the image pair under study are of course related in some way), and each input data set has at least one spatial dimension and / or a temporal dimension. According to the present invention, a displacement map (i.e., at least one displacement map) is generated. As illustrated in the figure, an embodiment of the present invention is a hierarchical feature matching method and apparatus for stereo depth estimation (based on disparity maps). Due to the relatively low computational cost and effectiveness, the method and apparatus according to the present invention are advantageously fast. Thus, the method and apparatus according to the present invention are suitable for using a neural network to predict a displacement map (e.g., a disparity map from a stereo image pair). The method and apparatus according to the present invention are robust to small errors in camera calibration and correction, and perform well on non-texture areas.
[0054] In one embodiment, the steps of the method according to the present invention are given as follows (by Figure 2 ).
[0055] In the first step, a first input data set and a second input data set (the concepts "first" and "second" in the names of the input data sets do not of course indicate the order of the input data sets; these concepts only serve to show the difference between the two input data sets; these input data sets are Figure 2 The example of the left image 20a and the right image 30a) is processed by a neural network-based feature extractor 25 (which may also be referred to as a neural (network) feature extractor or a feature extractor using a neural network; in this embodiment, it includes two branches 27 and 29) to generate a feature map hierarchy 50, which includes feature map basis pairs (in Figure 2 In the example, feature maps 20e and 30e; usually the coarsest feature maps) and feature map refinement pairs (in Figure 2 The feature maps are preferably generated successively by means of a feature extractor), each pair of feature maps forming a level of a feature map hierarchy 50, the resolution of the feature map refinement pair being less coarse than the resolution of the feature map base pair in all dimensions of at least one spatial dimension and / or temporal dimension thereof. Accordingly, the feature map hierarchy comprises at least two feature map pairs (a feature map base pair and at least one feature map refinement pair).
[0056] Thus, one or more feature map refinement pairs may be included in the feature map hierarchy. In the following, for a refinement pair (which may be the only one refinement pair, or more preferably, from Figure 2 The method according to the present invention is described with reference to a single (only or selected) feature map refinement pair until turning to the case with multiple (ie at least two) feature map refinement pairs.
[0057] Referring to these pairs as "base" and "refinement" is not limiting, and any other names may be used for these pairs (e.g., first, second, etc.). Figure 2 In the illustrated example, as the feature maps are progressively downscaled, increasingly coarser resolutions are obtained.Thus, each successive feature map pair is preferably downscaled relative to the feature map pair of the previous level in the feature map hierarchy 50.
[0058] Thus, in the above steps, a hierarchical series (sequence) of feature maps, called feature map levels, is generated. In other words, the neural network-based feature extractor generates feature maps at multiple scales. Accordingly, the number of feature map levels is at least two; that is, the device according to the present invention includes a displacement unit and at least one displacement refinement unit (e.g., a disparity unit and at least one disparity refinement unit). Throughout this description, disparity is the main example of displacement. The number of feature map levels is preferably between 2 and 10, in particular between 3 and 8, and more preferably between 3 and 5 (and in particular between 5 and 6). Figure 2 4 in the example of ). The feature map hierarchy includes feature map pairs, i.e., there is a feature map "sub-level" corresponding to each of the first input data set (e.g., left image) and the second input data set (e.g., right image), i.e., the hierarchy has two branches; the feature maps from the sub-levels that constitute a pair at a certain level of the feature map hierarchy have the same image size and resolution (number of pixels). Accordingly, both of these hierarchy branches contribute to the feature map pair in one level; in other words, one level extends into these two branches of the hierarchy. In summary, the feature map hierarchy is a set of feature maps (feature atlas) having multiple levels, with a pair of feature maps in each level; the wording "hierarchy" in the name indicates that the feature atlas has increasingly coarse feature maps at its levels.
[0059] exist Figure 2 Referring to the members of the input dataset pair (stereoscopic image pair) as left and right images in the example will indicate that these images are recorded from different perspectives, since the left and right images in a stereoscopic pair can always be defined. The left and right images may also be referred to simply as the first image and the second image, or, for example, as the first side image and the second side image, respectively (these considerations of left and right also apply to the distinction between feature maps, sub-units of disparity (refinement) units, etc.). Figure 1 A single feature extractor branch 15 is illustrated. Figure 2 In FIG. 1 , a block diagram of an embodiment of the method and apparatus according to the present invention is shown. Figure 2 In the embodiment of Figure 2 A neural network based feature extractor 25 with two feature extractor branches 27 and 29 on the left and right sides (a separate branch for each of the left and right images).
[0060] According to the above description, according to Figure 2 , the level referred to as the previous level is the higher level (higher in the diagram). Thus, the feature map of the "current" (or currently studied / processed) level is downscaled relative to the feature maps of the higher levels. The predetermined amount applied for downscaling can be a freely chosen (integer) number. The downscaling factor does not have to be the same throughout the hierarchy, e.g. Figure 2In the illustrative example of , factors of 2 and 4 are applied. In addition, other scale lifting factors may be used, in particular powers of 2. The upscaling factors applied to the refinement of the displacement (e.g., disparity) map may have to correspond to these scale lifting factors. All scale lifting factors can be 2, but this will result in much slower performance (longer runtime) and will not produce much greater improvement in accuracy. Applying a scale lifting factor of 2 to an image, the area of the scaled image becomes one-fourth of the original image. Thus, if the runtime is 100% in the original image, the runtime is 25% in the image scaled by a factor of 2 and 6.25% in the image scaled by a factor of 4. Accordingly, if a scale lifting factor of 4 is applied at the beginning of the procedure and a scale lifting factor of 2 is applied for the remaining scale lifting, rather than using a scale lifting factor of 2 throughout the entire scale lifting procedure, the runtime can be greatly reduced. Accordingly, different scale lifting factors can be applied between each level pair, but it is also conceivable to use the same scale lifting factor at all levels.
[0061] The feature map of the “current” level always has a coarser resolution (lower number of pixels) than the previous level, and as Figure 2 As explained in , it is preferred to have a higher number of feature map channels. In the figure, the number of channels of the feature map is explained by the thickness of the feature map. The preferred increase in the number of channels corresponds to the fact that in the feature map of the "current" level, the picture information of the original image (left image and right image) is processed more in most cases than in the previous level. The number of channels also corresponds to the different features to be studied. In some cases, it is not worth increasing the number of channels too much; therefore, the number of channels of two consecutive levels can even be the same. Accordingly, in a preferred case, the "current" level feature map contains less residual picture-like information compared to the feature map of the previous level, and their information structure becomes more similar to the final feature map (with the lowest resolution, that is, these are usually the coarsest feature maps) (the "current" level feature map is more "feature map-like"). In addition, at deeper and deeper levels in the neural network (with increasingly coarse feature maps), higher and higher level features are identified. For example, only some basic objects (e.g., lines, blocks) can be distinguished at the starting level, but in the deeper level of the hierarchy, many different objects (e.g., people, vehicles, etc.) can be identified. To maintain information storage capacity, the number of channels must be increased when reducing the feature map / image size.
[0062] Thereafter, in the next step, an initial displacement map (in the case of FIG. 1 ) for the feature map basis pair in the feature map level 50 is generated in a displacement (e.g., disparity) generation operation based on matching the first feature map in the feature map basis pair with the second feature map in the feature map basis pair. Figure 2In the embodiment of the present invention, the disparity map is calculated, namely the left disparity map 46e and the right disparity map 48e, that is, not only one initial disparity map, but two initial disparity maps). The feature map basis is Figure 2 In the illustration of FIG. 1 , there is a pair of feature maps 20 e and 30 e; these feature maps are the coarsest feature maps in the illustrated hierarchy because the hierarchy includes feature maps with increasingly coarse resolutions at levels starting with images 20 a and 30 a. Accordingly, an initial displacement map is generated for a base (typically the last, lowest resolution, coarsest) pair of feature maps. This initial displacement map (e.g., a disparity map) is a low-resolution displacement map whose resolution is increased to obtain a final displacement map (i.e., in FIG. 1 ). Figure 2 In the example of , a disparity map can be appropriately used as an output of a stereoscopic image).
[0063] Furthermore, in a step corresponding to the displacement refinement operation, the initial displacement map (e.g., disparity map) is upscaled in all dimensions of at least one spatial dimension and / or temporal dimension thereof to a feature map refinement pair (feature maps 20a, 30d in the figure, see for details) in the feature map level (feature map level 50 in the figure) using corresponding upscaling factors. Figure 6 ) of the feature map refinement pair (i.e., whereby the initial displacement map can be upscaled to the scale of the appropriate feature map refinement pair), and the values of the initial displacement map are multiplied by the corresponding upscaling factors to generate an upscaled initial displacement map (the feature map and the displacement [e.g., disparity] map have the same spatial and / or temporal dimensions as the input dataset, but these typically have different resolutions), and then a warping operation is performed by applying the upscaled initial displacement map to the first feature map in the feature map refinement pair (feature maps 20d, 30d in the illustration) in the feature map hierarchy 50, thereby generating a warped version of the first feature map in the feature map refinement pair. Accordingly, in the upscaling step, the corresponding size of the displacement map is enlarged (e.g., by upscaling using a factor of 2, in which the distance between two pixels becomes twice its previous distance), and the values of the displacement map are multiplied by the upscaling factors (this is because the distance will be larger in the upscaled disparity map), see below for details in an embodiment.
[0064] The following will combine Figure 6An exemplary implementation of the warping operation is described in detail. In short, a previous displacement (e.g., disparity) map is applied to one side feature map from a different level in the hierarchy to obtain a well-behaved input for comparison with the other side feature map. In other words, the warping operation shifts the feature map by the value of the displacement (e.g., disparity) map on a pixel-by-pixel basis. As a result, the warped version of the feature map will be closer to the other feature map to be compared with it (the displacement [e.g., disparity] calculated for a coarser level is itself an approximation of the displacement, whereby one feature map can be transformed into the other feature map in the pair at a level). Thus, by means of the warping, an approximation of the other feature map in a pair is obtained (which is referred to above as a well-behaved input). When this approximation is compared to the other feature map in the pair itself, a correction to the (approximate) previous displacement (e.g., disparity) can be obtained. The warping operation allows only a more limited set of shifts to be applied when seeking an appropriate displacement (e.g., disparity) refinement in a displacement refinement unit at a certain level (see for details). Figure 6 Description of ). The displacement map (e.g., disparity) typically has channels (referred to as coordinate channels in this document) corresponding to the spatial or temporal dimensions of the input or feature map (or may have fewer coordinate channels in limited cases). When the displacement map is applied to the feature map, the displacement values stored in the coordinate channels of the displacement map are used as shift offsets for each pixel of the feature map. The shift operation is performed independently for all feature channels of the feature map without blending the feature channels with each other.
[0065] In the context of this document, the definition defined above is used for warping. In other approaches, this step may be called pre-warping, and at the same time, the term warping may be used for other operations.
[0066] Next, the warped version of the first feature map in the feature map refinement pair is matched with the second feature map in the feature map refinement pair to obtain a corrected displacement map for the warped version of the first feature map in the feature map refinement pair and the second feature map in the feature map refinement pair (the displacement [e.g., disparity] map is simply a correction of the coarser displacement map due to adjustment using the warping operation, see below for details), which is added to the upscaled initial displacement map to obtain an updated displacement map (for the feature map refinement pair, i.e., typically for the next level, in the case of the first execution of this step in Figure 2 In the example, there is a left disparity image 46d and a right disparity image 48d).
[0067] Therefore, in the above steps, use Figure 2 The figure mark (discussing the disparity map as a special case of the displacement map), using the next level of feature map pairs to the existing disparity map at hand (this is for Figure 2In other words, in the at least one disparity refinement unit, the corresponding hierarchical structure of the comparator refines the previous detection. The purpose is to improve the resolution and accuracy of the disparity (and generally the displacement) map. This is accomplished by a displacement refinement unit (e.g. a disparity refinement unit) in the apparatus according to the present invention (and therefore preferably also in the method according to the present invention). The next coarsest feature map pair (as an exemplary refined pair) has a higher resolution (higher number of pixels) than the coarsest feature map pair (as a good candidate for the base pair). Accordingly, a higher resolution correction can be obtained for the disparity map based on this data. Therefore, before the initial disparity map and the disparity map correction obtained based on the feature map (referred to above as the corrected disparity map) are added (combined), the size of the initial disparity map is upscaled. In Figure 2 In the example of , the upscaling is performed by a factor of 2, since there is a 2-fold upscaling between the coarsest feature map and the next coarsest feature map. Thus, in general, for each level, the disparity map of the previous level is upscaled by the upscaling factor between the feature maps of the previous level and the current level before being added at each level. For details of upscaling in the illustrated embodiment, see also Figure 6 The addition of different disparity maps means adding them to each other on a "pixel-by-pixel" (more generally, data element-by-data element) basis; thus, the disparity (and generally the displacement) values in the pixels are added, or in the case of generalized disparity vectors, the addition of disparity vectors is performed in this step (generally these are performed on displacement vectors).
[0068] In the above steps, matching of two feature maps is performed. Matching is a comparison of two feature maps (one of which is distorted in a displacement / disparity refinement operation), based on which a displacement (e.g., disparity) map of the two feature maps can be obtained (generated). In other words, in the matching operation, the displacement (e.g., disparity) is calculated based on the two feature maps. In the illustrated embodiment, matching is performed by applying a shift value (for generating multiple shifted versions of the feature map), and the comparison is performed taking into account the multiple shifted versions. In particular, the following description will be made in conjunction with Figure 3 and Figure 6 Describe this matching method in detail.
[0069] The above describes in detail the basic building blocks of the present invention, i.e., the correspondence between feature map base pairs and feature map refinement pairs. In general, the feature map refinement pair mentioned above is the only refinement pair, or the refinement pair that is closest to the feature map base pair (see below for details). The following describes in detail the typical case of having more than one feature map refinement pair.
[0070] Thus, in a preferred embodiment (such as Figure 2 In the illustrated embodiment), the feature map level (in Figure 2 In the example with reference numeral 50) at least two (a plurality of) feature map refinement pairs (in Figure 2 In the example of 20b, 20b, 20c, 30c, 20d, 30d). In other words, the following describes a situation with multiple feature map refinement pairs. Having more than one feature map refinement pair is advantageous because in this embodiment, the refinement of the displacement map is performed in multiple stages and a displacement map with better resolution can be obtained.
[0071] In this embodiment, the feature map basis pair closest to the feature map hierarchy (in Figure 2 In the example of the feature map pair 20e, 30e), the first feature map refinement pair (in Figure 2 The resolution of the feature map pair 20d, 30d in the example is less coarse than the resolution of the feature map basis pair, and each successive feature map refinement pair (in the example of FIG. 1 ) is less close to the feature map basis pair in the feature map hierarchy than the first feature map refinement pair. Figure 2 The resolution of feature map pairs 20b, 30b, 20c, 30c) in the example of FIG. 1 is not as coarse as the resolution of adjacent feature map refinement pairs that are closer to feature map base pairs in the feature map hierarchy than to corresponding successive feature map refinement pairs.
[0072] Furthermore, in the present embodiment, a displacement (e.g., parallax) refinement operation is performed using the first feature map refinement pair (this is the first displacement refinement operation introduced above), and a corresponding further displacement refinement operation is performed for each successive feature map refinement pair, wherein in each further displacement refinement operation, an updated displacement map obtained for an adjacent feature map refinement pair that is closer to a feature map basis pair in the feature map hierarchy than the corresponding successive feature map refinement pair is used as an initial displacement map, which is upscaled to the scale of the corresponding successive feature map refinement pair using a corresponding upscaling factor during upscaling in the corresponding shift refinement operation (i.e., whereby the initial displacement map can be upscaled to the scale of an appropriate [next] feature map refinement pair), and the value of the updated initial displacement map is multiplied by the corresponding upscaling factor.
[0073] use Figure 2 , if there are one or more subsequent levels, a subsequent count is made for each subsequent level (from the lowest (i.e., coarsest) disparity level, in Figure 2The refinement of the disparity map is performed at higher and higher disparity levels (with higher and higher resolution in the image). As mentioned above, the present invention also covers the case where the feature extractor has two levels (i.e., there are two levels of feature maps in addition to the level of the input image, i.e., one level with feature map base pairs, and one level with feature map refinement pairs), to which the disparity units and the disparity refinement units correspond in the disparity hierarchy of the disparity map of the image pair suitable for obtaining a stereoscopic image. In fact, having further intermediate levels in the hierarchy (between the levels of the input data set and the base feature map) can result in better performance. If more levels are applied, the resulting disparity map is obtained in more refinement steps and with smaller jumps in the size of the disparity map. Accordingly, the corresponding feature map sequence has more feature map levels applied for refinement.
[0074] and Figure 2 Compared to the example of , a further disparity refinement unit may be applied, i.e. such a unit may be applied with the original images (left and right images) as input for the disparity refinement. However, the computational cost of such a further disparity refinement unit is higher (which also means a longer running time), so it may be preferable to skip this unit. The situation where such an additional disparity refinement unit is not applied (such as the one explained) does not result in a big disadvantage. Therefore, from the efficiency point of view, it is simple and straightforward to omit the disparity refinement unit from the level of the analyzed images. Therefore, depending on the structure of the disparity refinement unit, the method obtains the final disparity map after processing the coarsest feature map pair or, alternatively, after also processing the starting image itself.
[0075] In summary, in the method according to the present invention, a plurality of different levels of feature maps are generated by a neural network-based feature extractor. Furthermore, during displacement (e.g., disparity) analysis, different levels of displacement are generated based on hierarchical levels of feature maps by displacement units and displacement refinement units. As a result, the total number of feature map levels and starting image levels is equal to or greater than the total number of displacement units and the at least one displacement refinement unit (e.g., disparity unit and at least one disparity refinement unit; each unit operates on a corresponding single feature map level). Figure 2 In the example of , the first number is larger (it is five and the number of units is four). In this example, the number of feature map levels is equal to the number of units.
[0076] However, such feature map levels may also be arranged as follows: these feature map levels are not processed by the displacement (refinement) unit, i.e., not feature map basis pairs or refinement pairs, but feature map pairs that are not processed by the displacement (refinement) unit. Accordingly, additional levels with a coarser resolution than the feature map basis pairs may also be conceived. In other words, all disparity (refinement) units have corresponding feature map pairs that may be processed by the corresponding units as input.
[0077] According to the above details, in one embodiment, the input data set pair is an image pair of stereo images, the displacement map is a disparity map, the displacement generation operation is a disparity generation operation, and the displacement refinement operation is a disparity refinement operation.
[0078] Generally speaking, according to the present invention, there is at least one feature map (which may be referred to as an intermediate feature map) between the base feature map which is the last one in the hierarchy and the analyzed image.
[0079] As is clear from the above details, in one embodiment, the feature map has one or more feature channels, and matching is performed by taking into account the one or more feature channels of the feature map in the displacement (e.g., disparity) generation operation and / or in the displacement (e.g., disparity) refinement operation (see Figure 5A , where all C channels are taken into account in the convolution of the comparator (i.e., in the matching), and the initial displacement (e.g., disparity) map and the corrected displacement (e.g., disparity) map are generated using the same or a smaller number of coordinate channels than the dimensionality of at least one spatial dimension and / or temporal dimension of the input dataset, respectively (these displacement maps and the input dataset have the same dimensionality).
[0080] In addition, as explained, Figure 2 In the embodiment of the present invention, a pair of left initial disparity map 46e and right initial disparity map 48e are generated in the disparity generation operation, and based on the pair of left initial disparity map and right initial disparity map and the pair of left corrected disparity map and right corrected disparity map generated in the disparity refinement operation, a pair of left updated disparity map and right updated disparity map are generated in the disparity refinement operation (the updated disparity map is used as the initial disparity map of the next level). More generally, a pair of first initial displacement map and second initial displacement map are generated in the displacement generation operation, and based on the pair of first initial displacement map and second initial displacement map and the pair of first corrected displacement map and second corrected displacement map generated in the displacement refinement operation, a pair of first updated displacement map and second updated displacement map are generated in the displacement refinement operation.
[0081] Some embodiments of the present invention relate to an apparatus for generating a displacement map of a first input data set and a second input data set in an input data set pair (in Figure 2In an example, the apparatus is adapted to generate a disparity map of a left image and a right image in a stereoscopic image pair), each input data set having at least one spatial dimension and / or a temporal dimension. Figure 2 An embodiment of the device is also explained. The device is suitable for performing the steps of the method according to the invention. In one embodiment, the device according to the invention comprises (by means of Figure 2 )
[0082] - a neural network based feature extractor 25 adapted to process the first input data set and the second input data set (in Figure 2 In the example of the left image 20a and the right image 30a), a feature map hierarchy 50 is generated, which includes a feature map base pair and a feature map refinement pair (with respect to the feature map pair ( Figure 2 left feature map and right feature map at the same level in (see above), each pair of feature maps constitutes a level of feature map hierarchy 50, and the resolution of the feature map refinement pair in all dimensions of at least one spatial dimension and / or temporal dimension is not as coarse as the resolution of the feature map basis pair;
[0083] - comprising a first comparator unit (in one embodiment in Figure 3 The shift unit (in the first comparator unit 64, 74) is shown in Figure 2 In the example of FIG. 4 , specifically the disparity unit 40 e), the displacement unit is adapted to displace the first feature map in the feature map basis pair with the second feature map in the feature map basis pair ( Figure 2 The feature maps 20e and 30e are members of the basis pair) to generate an initial displacement map for the feature map basis pair in the feature map level 50 (in Figure 2 In the embodiment of the present invention, the initial disparity map, that is, the left disparity map 46e and the right disparity map 48e, as described in detail above);
[0084] -Displacement refinement element (in Figure 2 Specifically, in the example of , the disparity refinement unit; Figure 2 4 shows three disparity refinement units, namely, disparity refinement units 40b, 40c, 40d), comprising:
[0085] - Upscaling unit (in Figure 6 In the embodiment of the invention, the upscaling units 120 and 130 are used, for details, see below), which are adapted to upscale the initial displacement map in all dimensions of at least one spatial dimension and / or time dimension to the feature map refinement pair in the feature map level 50 (in Figure 2 20d, 30d), and is adapted to multiply the values of the initial displacement map by the corresponding upscaling factor to generate an upscaled initial displacement map,
[0086] - a twisting unit (in one embodiment, the twisting units 124, 134, see Figure 6 ), the warping unit is adapted to refine the first feature map in the feature map pair (in Figure 2 In the embodiment, the feature maps 20d and 30d are members of a refinement pair for scaling the initial disparity map up and down) to perform a warping operation, thereby generating a warped version of the first feature map in the feature map refinement pair,
[0087] - A second comparator unit (for details about the second comparator units 126 and 136, see Figure 6 Embodiment; The first and second comparator units may have the same structure, for example, Figure 5A The Figure 5B ), the second comparator unit being adapted to match the distorted version of the first feature map in the feature map refinement pair with the second feature map in the feature map refinement pair to obtain a corrected displacement map for the distorted version of the first feature map in the feature map refinement pair and the second feature map in the feature map refinement pair, and
[0088] - Addition unit (see Figure 6 The adding unit 128, 138 in the embodiment of the present invention is adapted to add the corrected displacement map and the upscaled initial displacement map to obtain an updated displacement map (for feature map refinement pair).
[0089] In a preferred embodiment (see, for example, Figure 2 ), as mentioned above, the device comprises at least one further displacement refinement unit (on top of the displacement refinement unit described above; for the special case of parallax, see Figure 2 The plurality of disparity refinement units 40b-40d in the feature map hierarchy include at least two feature map refinement pairs (in Figure 2 In the example of , the feature map refinement pair 20b, 30b, 20c, 30c, 20d, 30d), where the feature map base pair closest to the feature map level (in Figure 2 In the example of the feature map pair 20e, 30e), the first feature map refinement pair (in Figure 2 In the example of , the resolution of feature map pair 20d, 30d) is not as coarse as the resolution of the feature map base pair, and the resolution of each successive feature map refinement pair that is closer to the feature map base pair in the feature map hierarchy than the first feature map refinement pair is not as coarse as the resolution of an adjacent feature map refinement pair that is closer to the feature map base pair in the feature map hierarchy than the corresponding successive feature map refinement pair.
[0090] Furthermore, in this embodiment, a displacement refinement unit is applied to a first feature map refinement pair, and a respective further displacement refinement unit is applied to each successive feature map refinement pair, wherein in each further displacement refinement unit, an updated displacement map obtained for a neighbouring feature map refinement pair that is closer to a feature map base pair in the feature map hierarchy than the respective successive feature map refinement pair is used as an initial displacement map, which initial displacement map is upscaled to the scale of the respective successive feature map refinement pair using a respective upscaling factor during upscaling in the respective displacement refinement operation, and the values of the updated displacement map are multiplied by the respective upscaling factor.
[0091] Regarding prior art approaches, some prior art approaches apply neural network-based feature extraction (CN105956597 A, Kendall, Zhong, Mayer, Fischer, Pang, Goard), but the structure of the disparity (or displacement in general) refinement applied according to the present invention is not disclosed in any of the prior art approaches. In other words, the structure applied in the present invention in which an intermediate level feature map (one of which is distorted) is utilized in the sequence responsible for displacement (e.g., disparity) refinement is not disclosed in the above-cited prior art and cannot be derived therefrom.
[0092] In the above-mentioned article, Mayer applied Fischer's method to disparity estimation. In Fischer, a neural network-based method was applied to optical flow (the flow from time "t" to time "t+1"). Fischer applied correlation operations at relatively early [less coarse, relatively high resolution] levels of the feature hierarchy (this is the only comparison of feature maps in this method), i.e., only feature maps from these levels were calculated for the two images. In this method, hierarchical refinement using feature map pairs was not applied.
[0093] In Kendall, a neural network based feature extractor is used for stereo images. In this approach, downsampling and upsampling are applied sequentially. Unlike the present invention, in Kendall, the disparity values are deducted from the cost volume using a so-called "soft argmin" operation. In the Kendall approach, a cost volume is generated in which a reference and other feature maps transformed using all possible disparity values are cascaded together. Accordingly, in this approach, the reference is always the first level, and since they apply a 3×3×3 kernel, when the kernel (which has been cascaded) is on the reference, the reference only has an impact on two consecutive levels. Therefore, disadvantageously, all other levels cannot "see" the reference.
[0094] Zhong's approach is very similar to Kendall's. In this approach, feature quantities are generated. One feature map is selected as a reference, and different disparity values are applied to another feature map. In the feature quantity obtained by cascading, the transformed feature map is sandwiched between the two references. This approach is disadvantageously complex and computationally expensive.
[0095] Furthermore, in contrast to the approach of Kendall and Zhong, this approach is used in certain embodiments of the present invention (see Figure 5A and 5B ), where, according to the structure of the comparator, the result is equivalent to the case where the characteristic maps of different shifts are compared one by one with the reference. Additionally, due to the construction details of the comparator unit, the approach applied in this embodiment of the invention is very advantageous from the aspect of computational cost. For details, see the following Figure 5A and 5B Description.
[0096] Pang discloses a framework for disparity analysis that is significantly different from the present invention. In Pang's approach, a first neural network produces a coarse disparity map, which is refined by a second neural network in a manner different from the approach of the present invention. In contrast, in the framework invention, a high-resolution disparity map can be obtained with the aid of a hierarchically applied disparity refinement unit. In Pang's approach, temporary results are not used for prediction of the next level because the temporary results are only summed in the final stage. In contrast, in the approach of the present invention, temporary results are applied for warping, and accordingly, the computational cost of successive levels is reduced.
[0097] In Godard, hierarchical disparity refinement is not applied. In this approach, the loss is studied at multiple levels to control the procedure at several points in the training process. In this approach, the neural network processes only one camera image (e.g., the left image), i.e., a hierarchy of left and right feature maps is not generated.
[0098] In WO 00 / 27131 A1, no machine learning or neural network-based approach is applied. Therefore, no transformation of feature space is applied in this prior art document. Instead, convolution with predefined filters is applied in the approach of this document. In other words, no features are studied in WO 00 / 27131 A1 (this is unnecessary because disparity basically characterizes the image). Disadvantageously, in contrast to the present invention, the approach of WO 00 / 27131 A1 does not focus on relevant features. However, matching can be performed more effectively based on features. For example, in the case where more identical color blocks are located in the image in the stereoscopic image pair, it is more difficult to correspond between these color blocks in the image space (i.e., based on color) than in the feature space (i.e., based on semantic content). In summary, the feature-based matching applied in the present invention is more advantageous, because for those areas where relevant features are located, disparity must be accurate. It has been identified that, from the point of view of the features, the low resolution of the base feature map is not disadvantageous, since the presence of features can also be retrieved at low resolutions (the feature map has a much richer content at lower resolutions, with more and more feature-like properties, see above). This is in contrast to the approach of WO 00 / 27131 A1, where the low resolution levels include much less information. Furthermore, in the approach of WO 00 / 27131 A1, the same content is compared at different resolutions.
[0099] In summary, in the known approaches, the hierarchical refinement of displacement maps (eg, disparity maps) does not present a way like in the present invention, that is, at a given level, the previous coarser displacement map is refined with the help of the feature map of this level.
[0100] Figure 2 A high-level architecture of an embodiment of the method of the present invention is shown. In the illustrated exemplary embodiment, the corresponding feature map pairs consist of feature maps 20b and 30b, 20c and 30c, 20d and 30d, and 20e and 30e. In the illustrated exemplary embodiment, the same convolutional neural network (CNN) is applied to both the left and right images of the stereo pair (which can be processed as a 2-element batch, or using two separate CNNs with shared filters). In many cases, the neural network is implemented in a way that the neural network can process more images at the same time; these images are processed in parallel, and finally feature maps for each image can be obtained. This approach produces the same result as if the neural network would be copied, and one of the neural networks would process the left image, while the other neural network would process the right image.
[0101] Naturally, the members of a feature map pair have the same size (i.e., Figure 2The same spatial dimensions are indicated, H / 4×W / 4, H / 8×W / 8, etc.) and the same number of characteristic channels (illustrated by the thickness of each member of the pair).
[0102] The extracted features are fed to a unit that generates (disparity unit 40e) or refines (disparity refinement units 40d, 40c, 40b) a disparity (or generally a displacement) image. The units used in this document may also be referred to as modules. Figure 2 As explained in , disparity is first computed at the base (usually the coarsest) scale (for the feature map base pair, i.e., for feature maps 20e and 30e), and a refinement unit is used to upscale and improve these predictions. This results in Figure 2 The hierarchical structure shown in .
[0103] exist Figure 2 In the example illustrated in , the size of the output of the disparity unit 40e (i.e., the coarsest (lowest in the figure, base) left disparity map 46e (having the lowest resolution in the left disparity map) and the coarsest (lowest in the figure, base) right disparity map 48e (having the lowest resolution in the right disparity map)) is 1 / 32 of the input left image 20a and right image 30a. In the next two levels, the size of the disparity map is 1 / 16 and 1 / 8 of the input left image 20a and right image 30a, respectively.
[0104] Note that the scales of 1 / 32, 1 / 16, etc. are only for demonstration. Any sequence of scales with any increments can be used (not just 2×, see the detailed description of size reduction above in conjunction with the introduction of convolutional and pooling layers). Finally, the output of the disparity refinement unit (according to Figure 2 With reference numeral 40b) are the network outputs (disparity maps 46 and 48), i.e. the resulting disparity maps (in this example a pair of maps) obtained by various embodiments of the method and apparatus according to the present invention. Other scaling or convolutions may be applied to improve the result. The disparity (refinement) unit also includes a learnable convolution (in other words, a unit that performs convolution, or simply a convolution unit, see below). Figure 4-6 ), so these convolutions should be trained how to perform their tasks.
[0105] It should be noted that the system of units (feature extractors, displacement / disparity [refinement] units) used in the present invention is preferably fully distinguishable and the system of units is end-to-end learnable. All displacement (refinement) units can then pass gradients backward, so they behave like ordinary CNN layers. This helps to learn features that are just (i.e., appropriate) for the task. CNN can be pre-trained on the commonly used ImageNet (see Olga Russakovsky et al.: ImageNet Large Scale Visual Recognition Challenge, 2014, arXiv: 1409.0575) or on any other appropriate task, or can be used without pre-training (by applying only the training phase). The network used can be taught in a supervised or unsupervised manner. In the supervised approach, the input image and the output displacement (e.g., disparity) map are presented to the network. Since it is difficult to obtain dense displacement / disparity maps of the real world, these are usually simulated data. In the unsupervised approach, the network output is used to distort the left image to the right and the right image to the left. They are then compared to the real images and a loss is approximated based on how well they match. Additional loss components can be added to improve the results (see above for details).
[0106] exist Figure 3 , a flow chart of a disparity unit 40e in one embodiment is shown (in one approach, this is the first unit (module) applied to the feature map in the hierarchy and is at the lowest level in the diagram). It starts from the same hierarchical level ( Figure 2 The left feature map 20e and the right feature map 30e are received as input by the left shifter unit 62 and the right shifter unit 72 (the shifter unit may be simply referred to as a shifter, or alternatively referred to as a shifter module), respectively, and these shifted versions are fed to the corresponding comparator units 74 and 64 (which may also be simply referred to as comparators, or alternatively referred to as comparator modules; there are cross connections 66, 68 between the shifter units 62, 72 and the comparator units 64, 74), which are also learnable (see below).
[0107] Accordingly, the disparity unit 40e responsible for generating an initial disparity map (more precisely, a pair of left and right initial disparity maps in the illustrated example) operates as follows. Figure 3Starting the description of the operation from the left side of (the right side is equivalent), the left feature map 20e as the first input is fed into both the shifter unit 62 and the comparator unit 64. Due to the cross connections 66, 68, the comparator unit 64 will be used to compare the shifted right feature map 30e with the left feature map 20e fed into the comparator unit 64. The shifter unit 62 forwards the shifted version of the left feature map 20e to the comparator unit 74, where the right feature map 30e is fed into the comparator unit 74 on the right hand side. The shifter unit applies a global shift to the feature map, i.e. the same shift amount is applied to all pixels of the feature map (more shifts can be applied in the shifter unit at the same time, see the description of the following figures).
[0108] When the comparator unit 74 performs its task (i.e., performs a comparison to obtain a disparity map at its output), it compares the shifted version of the left feature map 20e with the unshifted version of the right feature map. In practice, the comparator unit selects (or interpolates, see below) the most appropriate shift (i.e., the shift whereby a given left feature map pixel best corresponds to the corresponding pixel of the right feature map) for each pixel on a pixel-by-pixel basis, and takes these shift values to the disparity map, again on a pixel-by-pixel basis for each pixel, e.g., for the pixel position from the right feature map 20e (i.e., for the unshifted feature map in the pair). Accordingly, the right disparity map may include disparity values from the perspective of the right feature map (and, therefore, the right image).
[0109] The comparator unit 64 performs the same process on the shifted versions of the left feature map 20e and the right feature map 30e. The comparator unit 64 outputs a left initial disparity map 46e, where the disparity values are given from the aspect of the left feature map (left image).
[0110] Therefore, in Figure 3 In an embodiment of the method illustrated in , in a displacement (eg, parallax) generation operation:
[0111] -Feature map basis pair ( Figure 2 A plurality of shifted feature maps of the first feature map in the feature maps 20e, 30e illustrated in ) are generated by applying a plurality of different shifts (shift / shift values in the case where the shift can be given by a single number, or shift / shift vectors in the case where the shift can be given by more than one coordinate) to the first feature map in the feature map basis pair, and
[0112] - Initial displacement map ( Figure 2The left disparity map 46e and the right disparity map 48e in the figure are obtained by generating a resulting shift for each data element (pixel) position of the second feature map in the feature map basis pair based on studying the matching between the multiple shifted feature maps of the first feature map in the feature map basis pair and the second feature map in the feature map basis pair (correspondingly, based on the matching, a resulting shift (i.e., a corresponding shift) is obtained for each pixel position, thereby giving a displacement [e.g., disparity] value corresponding to the position).
[0113] To refine the displacement (e.g., disparity) map, Figure 6 Similar steps are taken in the embodiment illustrated in (a similar scheme of applying displacement is performed). Accordingly, in a displacement (e.g., parallax) refinement operation:
[0114] -Feature map refinement pair ( Figure 2 The plurality of shifted feature maps of the first feature map in the feature map 20d, 30d) are generated by applying a plurality of different shifts to a distorted version of the first feature map in the feature map refinement pair,
[0115] - The correction displacement map is obtained by generating a resulting shift for each data element (pixel) position of the second feature map in the feature map refinement pair based on studying the matching between the plurality of shifted feature maps of the first feature map in the feature map refinement pair and the second feature map in the feature map refinement pair.
[0116] These embodiments (i.e., Figure 3 and Figure 6 The method steps and device units explained in the above) can be applied individually or in combination. Thus, in short, the shifting method can be applied to any unit of the displacement unit and the at least one displacement refinement unit (correspondingly, applied to the exemplary disparity unit and at least one disparity refinement unit).
[0117] Figure 4 An illustrative block diagram showing a shift operation ( Figure 4 For illustrative purposes; shifter units 62 and 72 and further shifter units below serve the same purpose). Figure 4 In FIG. 6 , the shifter unit 65 copies the input feature map 60 (labeled “input feature” in the figure) to multiple locations in the output feature map. For each offset, the shifter unit 65 produces a shifted version of the input feature map (in Figure 4 In the example of , the shifted versions 75, 76, 77 and 78 are obtained.
[0118] For each offset and position in the input feature map, a given position is translated by a given offset to produce a position in the output feature map. The corresponding pixel of the input feature map at the input position is then copied to the output position of the output feature map. Thus, geometrically speaking, the output feature map becomes a translated version of the input feature map. There will be pixels at the edges in the output feature map whose positions do not correspond to valid pixel positions in the input feature map. The features at these pixels will be initialized to a predefined value, which can be zero, another constant, or a learnable parameter, depending only on the channel index of the feature map (for example, for each channel, there is a constant fill value, and then if there is an output pixel without a corresponding input pixel, the feature vector at the pixel is initialized to a vector consisting of these fill values). In the case of parallax, for each different shift, the horizontal shift value is preferably increased / decreased by one pixel (that is, the entire feature map is globally shifted by a predetermined amount of pixels). Accordingly, the side of the image without input can be filled with zeros or a learnable vector.
[0119] The shifter unit has a parameter, which is the range of the shift, in Figure 4 In the case of - when the shift can be given by a number - the range is [-4, 0] (the definition of the range can be generalized to more dimensions). This describes which (integer) shifts are generated. On the last basic level, the shifters 62 and 72 of the disparity unit 40e (see Figure 3 ) generates an asymmetric shift (shift value) (since disparity can only be unidirectional, all features "advance" in the same direction between the left and right images). The shifter units in the disparity refinement unit (e.g., shifter units 122 and 132 in disparity refinement unit 40d, see Figure 6 ), the shifting is symmetric (this is because refinements are produced in these levels, which can be of either sign). The shifter unit has no learnable parameters. The extremes of the range are preferably chosen based on the maximum allowed disparity and based on the maximum allowed correction factor. The shifter unit multiplies the number of channels by n=max 范围 -min 范围 +1, i.e., in the above example, a total of five shifts are applied (including shift 0 and shift -4 in addition to shifts -1, -2 and -3).
[0120] In the general case, as mentioned above, the possible shifts can be any N-dimensional vector. The shift will be selected according to the specific task (e.g., adaptively). For example, in the case of 2D parallax described in detail in this document, only horizontal vectors are considered as possible shifts. In the case of optical flow, the shift vector can be in any direction, for example, such shifts can be considered where the 2D coordinates (e.g., x and y) can vary bidirectionally in any direction. This means 5×5=25 different shifts, because both x and y can range from -2 to 2 in this particular example. For example, in the 3D case, having a bidirectional change in any direction in the x, y, z directions means 125 different shifts. In this case, when a weighted value is obtained for the shift based on probability, the possible shift vectors are added as a weighted sum. As explained, the concept of shift can be generalized from a single number to a multidimensional vector in a simple and direct way.
[0121] Figure 5A FIG. 2 shows the structure of a comparator in one embodiment. It should be noted here that Figure 5A and 5B The values given in are exemplary (e.g., the size of the convolution kernel). The comparator unit receives a feature map as a first input from one side (e.g., the left side) of the feature map hierarchy. The feature map is Figure 5A denoted by reference numeral 80 and is labeled “reference.” The comparator unit receives as a second input a shifted feature map (denoted by reference numeral 82) from the other side of the feature map hierarchy at the same level (eg, the right side here if the other is the left side).
[0122] exist Figure 5A and 5B In the example above, the operation of the comparator unit is explained for parallax. However, it is obvious that the same can be applied in general by choosing the appropriate dimension. Figure 5A and 5B scheme to obtain the displacement map. In addition, Figure 5A and 5B In , a batch dimension (denoted by N) is introduced. The batch dimension is used when multiple instances of a certain type of input are available that should be processed in parallel in the same way. For example, when more than one stereo camera pair is arranged in a car to generate parallax, then more than one instance of each of the left and right input images is available. These input data pairs can then usually be processed in parallel using the same processing pipeline.
[0123] Therefore, in Figure 5A and 5B In the embodiment of , the comparator unit produces at its output a disparity map, which in this example is a single channel tensor with the same spatial dimensions as the input features (i.e., each feature map in the feature map pair). In general, the displacement map can be generated in the same way. Figure 5A , a disparity map 105 of dimensions N×1×H×W is obtained at the output of the comparator unit, while the reference has dimensions N×C×H×W. In this case, the spatial dimensions are the height (H) and the width (W), which are the same for the reference feature map and the disparity. In the comparator unit, the number of channels (C) is reduced from C to 1 (i.e., to the number of channels of the disparity, which is 1 in this example), see below for corresponding details. According to the approach applied throughout this document, the feature channels correspond to the feature maps, while the coordinate channels correspond to the displacement maps (here the disparity maps). Accordingly, the number of feature channels is reduced from C to the number of coordinate channels (i.e., to the number of channels of the displacement map (here the disparity map)), which is ultimately 1 in this example.
[0124] The comparator unit performs a comparison between the feature map selected as the reference and another feature map at the same level to which several different shifts have been applied. Figure 2 In the example illustrated in , the disparity map does not have multiple coordinate channels, the disparity map has a single number (disparity value) for each of its pixels. In contrast, in the case of 2D optical flow, the displacement map has two values per pixel, i.e. two coordinate channels, while in the case of 3D medical image registration, the displacement map has three values per pixel, i.e. three coordinate channels. Thus, in this case, the tensor of the disparity map is obtained according to the channel reduction performed in the comparator unit (see for the change in dimensionality Figure 5A ) and a pixel matrix with one channel. The disparity map generated by the comparator unit shows for each of its pixels the shift value (disparity, displacement) between the reference and another feature map ( Figure 5B Different possibilities for obtaining parallax in the comparator unit are shown, see below for details).
[0125] exist Figure 5A and 5B In the embodiment illustrated in FIG. 1 , the comparator unit has two different building blocks (subunits): a first subunit ( Figure 5A The comparison map generation unit C1 in , which can also be called the first block) compares each possible shift (i.e., feature map with different shifts) with the reference (i.e., with the feature map in the level that has been selected as the reference), and another subunit considers these comparison results and outputs the disparity ( Figure 5A The first result shift generation unit C2_A in Figure 5B The first, second and third result shift generation units C2_A, C2_B and C2_C in the block; which may also be referred to as the second block).
[0126] In the following, the Figure 5A An exemplary implementation of an embodiment of a comparator unit of Figure 5AIn the example shown in , a reference feature map 80 represented by a tensor with dimensions N×C×H×W is given as a first input to a comparison map generation unit C1 in a comparator unit (this branch on the right hand side in the figure is also referred to as a first calculation branch). A shifted feature map 82 (labeled as shifted feature) represented by a tensor with dimensions N×S×C×H×W is given as a second input to the comparator unit (where S is the number of shifts applied; this branch on the left hand side in the figure is also referred to as a second calculation branch).
[0127] In the first calculation branch, a first comparison convolution unit 84 is applied to the reference feature map 80 (the wording comparison in the name of the unit 84 only indicates that the unit is applied in a comparator, it can be referred to as a convolution unit, i.e., a unit that performs a convolution [operation]); as a result of this operation, the channel dimension (C) of the reference feature map 80 becomes C' (thereby, the dimension of the first intermediate data is N×C'×H×W). C' is equal to or not equal to C. In the illustrated example, the comparison convolution unit 84 (labeled as #1) has a kernel of size 3×3 pixels (referred to as 3×3 size). In addition, in this example, a two-dimensional convolution operator is applied, and accordingly, the kernel scans the so-called width and height dimensions of the image.
[0128] In the second computation branch, a "merge to batch" unit 92 is applied to the shifted feature map of the input, which enables the computation to be efficiently performed in a parallelized manner by a parallel processing architecture for several different shifts (the dimensions are transformed to NS×C×H×W in the fourth intermediate data 93). In other words, since there is a shifted version of the input feature map for each possible shift, these feature maps can be regarded as a batch (i.e., a computation branch that can be calculated simultaneously), and then they are processed simultaneously using the convolution unit (parallel processing, applying the same convolution operation to each element in the batch). The second comparison convolution unit 94 is applied to the "batch-like" input, resulting in the fifth intermediate data 95 having dimensions NS×C'×H×W. In the illustrated example, the second comparison convolution unit 94 (labeled as #2) has the same properties as the first comparison convolution unit 84 (having a 3×3 kernel and being a 2D convolution operation). Thus, the main parameters of units 84 and 94 are the same (both transform the number of channels from C to C'), however, the weights learned during their learning process are generally different. During the learning process, the learnable parameters of these convolutional units 84, 94 converge to values whereby the comparator function can be optimally performed.
[0129] These comparison convolution units 84 and 94 are actually "half" of a single convolution unit that was originally applied to the reference feature map and the channel-concatenated version of each shifted feature map, respectively, because the results are added together by an addition unit 98 in the comparator unit, which preferably performs a broadcast-addition operation. They are separated from each other here to emphasize the fact that half of them (the branch with the comparison convolution unit 94) needs to be performed S times, where S is the number of possible shift offsets; while the other half (the branch with the comparison convolution unit 84) only needs to be performed once on the reference feature map. After the addition unit 98, the result is the same as if a single convolution was performed on the concatenated version of the shifted feature map and the reference map for each shift offset. Therefore, from the perspective of computational cost, this separation of the comparison convolution units 84 and 94 is very advantageous.
[0130] After the corresponding comparison convolution units 84, 94 in the calculation branch, the data becomes a data shape compatible with each other. Accordingly, in the second calculation branch, a data reshaping unit 96 is applied (transforming the data into the N×S×C'×H×W format in the sixth intermediate data 97), with the help of which the data is transformed back from the "batch-type" format (units 92 and 96 perform inverse operations on each other, as is apparent from the dimensionality change in the second calculation branch).
[0131] In order to have a data format compatible with this in the first calculation branch, a dimension extension unit 86 (marked as "Extended Dimension") is first applied in the first intermediate data 85 to obtain second intermediate data 87 of dimension N×1×C×'H×W, i.e. the dimension is extended using the shift dimension (whose value is 1 at this stage). Afterwards, the copy operation is performed by the copy unit 88 in the first calculation branch. With the help of the copy unit 88 (which can also be called a broadcast unit; marked as "Broadcast (Tile / Copy)"), the data in the first calculation branch are prepared for the addition unit 98 (which has N×S×C'×H×W dimensions in the third intermediate data 89), i.e. (virtually) copied / extended according to the number of shifts applied in the second calculation branch.
[0132] In general, this should not be considered an actual copy of the bytes in memory, but only a conceptual (i.e., virtual) copy. In other words, consider a broadcast addition operation, where the left hand side (LHS) and the right hand side (RHS) of the addition have different dimensions (where one side (e.g., the LHS) consists of only 1 plane in a certain dimension, while the other side can have multiple planes depending on the number of shifts). During the addition operation, for example, the LHS is not actually copied to match the shape of the RHS, but rather the LHS is added to each plane of the RHS, and the data associated with the LHS is not actually copied (copied) in computer memory. Thus, the copy unit preferably symbolizes this preparation for the addition, rather than the actual copying of data.
[0133] Therefore, since they have the same data format (N×S×C'×H×W), the data obtained at the end of the two calculation branches can be broadcasted and added by the addition unit 98 (as mentioned above, the broadcast operation is preferably virtual in the copy unit 88, whereby the operation of the copy unit facilitates the broadcast addition operation of the addition unit 98). The addition unit 98 gives the output of the comparison map generation unit C1 in the comparator unit; this output includes the comparison of the characteristic maps for different shifts. Thus, the task of the comparison map generation unit C1 is to compare each shift with the reference.
[0134] The calculation structure applied in the comparison map generation unit C1 in the comparator unit gives the same result as the calculation of comparing the reference with each shift individually (which will be much less computationally cost-effective than the above scheme). In other words, according to the calculation scheme described in detail above, the comparison applied in the comparison map generation unit C1 is mathematically divided into two calculation branches, wherein, in the first branch, the convolution is applied only to the reference, and in the second branch, the convolution is applied to the shifted data. This separation facilitates that the calculation for the reference will be performed only once, rather than individually for each shift. In view of the computational cost effectiveness, this produces a huge advantage. Applying the appropriate data shaping operation as described in detail above, the addition unit 98 produces the same result as the case of the comparison performed individually for each shift. The output of the addition unit 98 is referred to as the result comparison data map (whose height and width are the same as the height and width of both the feature map to which the comparator unit is applied and the disparity map output as the result of the comparator unit). The comparison data map is a simple auxiliary intermediate data set (except for the first and second intermediate comparison data maps in the two branches of the comparator, see below), and can therefore also be referred to as some kind of intermediate data. Figure 5A The comparison data map is not shown in FIG. 1 , but only the seventh intermediate data 101 obtained from the comparison data map by applying the nonlinear unit 99 which is the last unit of the comparison map generating unit C1. The data structure in the seventh intermediate data 101 maintains the same N×S×C′×H×W.
[0135] Thus, in the result shift generation unit C2_A, the "to channel" data shaping operation is completed, that is, the output data of the addition unit 98 is stacked by the stacking unit 102 (marked as "merge to channel"). This operation stacks the channel dimensions of the sub-tensors corresponding to different displacements, that is, the shift dimensions are merged into the channel dimensions. Accordingly, the data format in the eighth intermediate data 103 is N×SC'×H×W. In this embodiment, the appropriate disparity value is obtained from the data tensor by means of the shift study convolution unit 104 (since the disparity value is generated by the convolution unit trained for comparison in this method, it can also be obtained by Figure 5A The comparator unit shown in generates a non-integer number as the disparity value, even though the shift / shift value used by the shifter unit is an integer).
[0136] Thus, at this stage, the results with different shifts are compared (and stacked into channels). Figure 5A As explained, in this preferred example, a convolution unit with a smaller kernel (in this example, a kernel 1×1 in the shift study convolution unit 104) is applied to a larger amount of data (i.e., to stacked data), and a convolution unit with a larger kernel (larger compared to another convolution) is applied to a batch of shifted feature maps (i.e., Figure 5A The fourth intermediate data 93 in block C1; the batched feature map in block C1 constitutes a smaller amount of data than the full stacked data in the result shift generation unit C2_A). From the perspective of computational cost effectiveness, this ratio of the kernel size of the convolution unit is advantageous.
[0137] It should be noted that the comparator units in the disparity unit or in the disparity refinement unit (in general, in the displacement unit or in the displacement refinement unit) are applied to different numbers of feature channels at different levels. In addition, the comparator unit usually outputs an initial or a corrected disparity (in general, a displacement) map, in this example, the number of channels of both is reduced to 1, or in general to the number of coordinate channels. In other words, if Figure 5A As explained in , the comparison for obtaining disparity information is performed for all feature channels. Accordingly, the comparator unit is sensitive to the feature information stored in different feature map levels. Thereby, the disparity map generation process, which takes advantage of machine learning, becomes particularly efficient, using not only low-level features (such as edges and areas with the same color or pattern in the left and right images), but also higher-level features (such as possible internal semantic representations of, for example, road scenes), which means that the same objects on the two images can be matched to each other more accurately.
[0138] The shift study convolution unit 104 performs a convolution operation preferably with a 1×1 kernel. The function of this convolution is to calculate the optimal displacement value from the information obtained by the feature map comparison performed by C1. According to the dimensions applied in this example, a two-dimensional convolution (sweeping the height and width dimensions) is preferably applied. The shift study convolution unit 104 outputs a shift value for each single pixel, that is, a disparity value in the corresponding position. The shift study convolution unit 104 studies one pixel at a time and is taught to output an appropriate shift (e.g., a shift value) as a result shift based on learning based on the data stored in the channel dimension.
[0139] Figure 5A The correspondence between the input feature map and the output disparity is explained, so the resulting shift can be obtained for each pixel position of the disparity map. Figure 5A The output of the result shift generation unit C2_A in the comparator unit is the disparity map of the corresponding level, from which the reference feature map and the feature map on which the shift operation has been performed are selected. In summary, in the result shift generation unit C2_A, the optimal shift offset is determined for each data element (pixel) by means of a preferably 1×1 convolution.
[0140] accomplish Figure 5A The advantage of the above detailed description of the embodiment of the comparator is its high efficiency. It should be noted that the addition unit 98 adds the convolved reference (the reference to which the convolution has been applied) to each possible shifted convolved variant separately. As mentioned above, this operation is equivalent to cascading the reference feature with the output shift one by one, convolving it, and cascading the result. However, the latter approach will require much more computing resources than the approach of the above embodiment, and the reference will be convolved multiple times (this is because the cascade is performed as the first step of the approach), resulting in redundant calculations based on the multiple convolution operations applied to the reference. In contrast, in the approach of the above embodiment, only one convolution operation is applied to the reference.
[0141] Note that the convolution of the comparison graph generation unit C1 (in Figure 5B The same comparison map generation unit C1 is applied in the embodiment of the comparison map generation unit C1) without nonlinearity, that is, there is no nonlinear layer after the comparison convolution units 84, 94 in the comparison map generation unit C1. This is why the convolution units 84, 94 can be split in this way. The nonlinearity is preferably applied in the nonlinear unit 99 after the broadcast addition operation (for example, after the addition unit 98 at the output of the comparison map generation unit C1). The nonlinearity is, for example, a ReLU nonlinearity, that is, a ReLU nonlinearity layer is applied to the output of the addition unit 98.
[0142] As described in detail above, in the result shift generation unit C2_A, a convolution with a 1×1 kernel is performed. The reason for this is that it works on a much higher channel count (because the shift is merged into the channel dimension here), and it will be desirable to avoid using large convolution kernels for these channels. Given the high number of channels, larger convolution kernels (e.g., 3×3) will be much slower. Since a 3×3 kernel is preferably applied in the comparison map generation unit C1, it is unnecessary to use a larger kernel in the result shift generation unit C2_A. A further advantage of a 1×1 kernel is that information from a single data element (pixel) is required, which also makes it unnecessary to use a larger kernel in this convolution. However, in the comparison map generation unit C1, the channel size is much smaller because the shift dimension is merged into the batch dimension. Since the number of operations is linearly proportional to the batch size and quadratically proportional to the channel size, this results in a performance gain when the kernel applied in the shift study convolution unit 104 in the result shift generation unit C2_A is much smaller than the kernels applied in the comparison convolution units 84 and 94 in the comparison image generation unit C1.
[0143] In summary, in Figure 5A and 5B In an embodiment of the method illustrated in , in a displacement (e.g., disparity) generation operation for generating an output displacement map for use as an initial displacement map and / or in a displacement refinement operation for generating an output displacement map for use as a correction displacement map,
[0144] - Matching of a plurality of shifted feature maps 82 and another feature map used as a reference feature map 80 (i.e., another feature map plays the role of a reference feature map, and accordingly, the other feature map is referred to by this name hereinafter) is performed in the following steps, wherein the shift number of these shifted feature maps is the number of different shifts:
[0145] - applying a first comparison convolution unit 84 to the reference feature map 80 to obtain a first intermediate comparison data map,
[0146] - applying a second comparison convolution unit 94 to each of the plurality of shifted feature maps 82 to obtain a plurality of second intermediate comparison data maps,
[0147] - adding the first intermediate comparison data map replicated according to the number of different shifts (i.e., virtually replicated as described in detail above, or physically replicated) to the plurality of second intermediate comparison data maps in an addition operation to obtain a result comparison data map,
[0148] - Generate a corresponding result shift for each data element of the result comparison data map in all dimensions of at least one spatial dimension and / or time dimension thereof, and assign all corresponding result shifts to corresponding data elements in the output displacement map (i.e., data elements in the result comparison data map according to the corresponding result shift).
[0149] In the above embodiments (such as Figure 5A In the embodiment illustrated in (i.e., in the variant including the result shift generation unit C2_A), preferably, the corresponding result shift of each data element of the result comparison data map in all dimensions of at least one spatial dimension and / or time dimension thereof is generated by applying the shift study convolution unit 104 to the result comparison data map. In addition, in this embodiment, preferably, the feature map has one or more feature channels, and the result comparison data map is stacked by the one or more feature channels of the feature map before applying the shift study convolution unit 104.
[0150] In a further variation of the above embodiment (e.g. in an embodiment comprising a result shift generating unit C2_B), the corresponding result shift of each data element of the result comparison data map in all dimensions of at least one spatial dimension and / or time dimension thereof is generated by selecting the best matching shift from the plurality of different shifts (for details and further options see Figure 5B description).
[0151] In yet another variation of the above embodiment (such as in an embodiment including a result shift generating unit C2_C), the corresponding result shift of each data element of the result comparison data map in all dimensions of at least one spatial dimension and / or time dimension thereof is generated by the following operations:
[0152] -Create multiple displacement bins for each different shift value,
[0153] - generating a displacement probability for each displacement bin, the displacement probability being calculated based on a result comparison data graph, and
[0154] - The resulting shift is obtained by weighting the shift values by the corresponding shift probabilities (see for details and further options Figure 5B description).
[0155] Preferably, in any variation of the above embodiment, a non-linear layer is applied to the result comparison data map before generating a corresponding result shift for each element of the result comparison data map in all dimensions of at least one spatial dimension and / or time dimension thereof.
[0156] Figure 5BThree alternative possibilities for generating parallax are shown (all result shift generation units C2_A, C2_B and C2_C output a parallax map of dimension N×1×H×W; these approaches can of course be generalized to displacement generation). Accordingly, the result shift generation units C2_A, C2_B and C2_C are different possibilities that can be applied to the seventh intermediate data 101 to generate parallax. All of these possibilities can be present and one of them can be selected based on the settings. Alternatively, only the result shift generation unit C2_A (such as Figure 5A ), one of C2_B and C2_C.
[0157] Figure 5B The result of the shift generation unit C2_A is Figure 5A Same as explained in (see above for details). Figure 5B Another alternative in is the result shift generation unit C2_B. As indicated in the result shift generation unit C2_B, in this alternative, C' must be 1, i.e., the number of channels must be reduced from C to 1 (C') in the comparison map generation unit C1 with the aid of the comparison convolution units 84 and 94. This can be done by setting the number of output channels of these convolutions to 1. In the result shift generation unit C2_B, the first intermediate unit 106 is applied to the seventh intermediate data 101, wherein the channel dimension (where C' is 1) is squeezed, i.e., eliminated without changing the underlying data itself (resulting in a data format N×S×H×W), after which, in the same first intermediate unit 106, argmax is applied in S ( Figure 5B ) or argmin operation, depending on whether the maximum comparison value or the minimum comparison value as the comparison result in the comparison map generation unit C1 should correspond to the shift with the best correspondence. Either argmax or argmin can be used, but one of them must be decided at the beginning before training the network. The use of argmax or argmin determines whether the comparison map generation unit C1 will output a similarity score (argmax) or a dissimilarity score (argmin), because during the training phase of the comparator unit, the selection of argmax or argmin will cause the comparison convolution unit to learn to output a larger or smaller score for a better shift. By applying argmax or argmin, the shift dimension is reduced from S to 1. Accordingly, the argmax or argmin function selects the following shift: with the help of which the best correspondence is achieved between the reference feature map and another feature map to which a different shift has been applied. In other words, in the result shift generation unit C2_B, the best shift is selected for each data element of the disparity 110 having dimensions N×1×H×W.
[0158] In by Figure 5BIn the alternative embodiment illustrated by the result shift generation unit C2_C in , the disparity is generated based on combining the shift offsets with probability weights, that is, the result shift (shift value or shift vector in a larger dimension) assigned to the corresponding pixel position of the disparity map is estimated (generated, calculated) using a probabilistic approach and discrete bins. In this alternative, the number of channels C' must be 1, similar to the result shift generation unit C2_B. Afterwards, the channel dimension is squeezed in the second intermediate unit 108 (i.e., similar to the result shift generation unit C2_B), and the data format becomes N×S×H×W after squeezing.
[0159] In the case of the result shift generation unit C2_C, the number of output channels corresponds to the number of possible integer values of the disparity at the studied level, i.e., the number of shifts applied in the comparison map generation unit C1; for each of these values, a discrete bin is established. Since the channel variables are squeezed and the number of possible shifts is S, the dimensions of the probability data 114 are N×S×H×W, i.e., for each batch, each shift value and for each pixel (data element) in the height and width dimensions there is a probability. In the shift offset data 112, a tensor with the same variables is established, which has the dimensions 1×S×1×1. In the shift dimension, the tensor has all possible shift values. The shift offset data 112 and the probability data 114 are multiplied with each other in the multiplication unit 116, and the multiplication results are summed in the summation unit 118, and a result can be obtained for each data element of the disparity, which is a weighted sum of the possible shift values and the probabilities corresponding to the bins of the corresponding shift values. Thus, according to the method described in detail above, the shift dimension of the data is reduced to 1, and a disparity 115 of dimensions N×1×H×W is obtained.
[0160] Based on the labels of the second intermediate unit 108, the probabilities of the probability data 114 are obtained by means of a Softmax function, i.e., the so-called Softmax function is applied by shifting the dimension in the present approach. It is expected that the output of the Softmax function will give the probability that the disparity has the value corresponding to the bin. For each data element of each batch element, the Softmax function is applied to the shifted dimension (S) of the seventh intermediate data 101, and the Softmax function is defined as:
[0161]
[0162] Wherein the input to the Softmax function is the score corresponding to the available shift value indicated by zi(i:1..S), which comprises the shift dimension of the seventh intermediate data 101, and pi(i:1..S) is the probability, i.e., the output of the Softmax operation. Here, S is the number of different shift values, i.e., the number of bins. The resulting displacement value (or vector in general) is then obtained by multiplying the probability by the disparity value of the corresponding bin for each bin in unit 116, and then summing the weighted disparity values in unit 118. See the following example.
[0163] In one example, if the possible disparity (shift) of a pixel is between 0 and 3, and the output probability is [0.0, 0.2, 0.7, 0.1], the output of the comparator of the pixel under study is calculated as sum([0.0, 0.2, 0.7, 0.1] * [0, 1, 2, 3]) = 1.9, where * indicates that the elements of the vector are multiplied piece by piece. Accordingly, by this approach, a good estimate of the disparity value of each pixel can be obtained.
[0164] Figure 6 An embodiment of a disparity refinement unit is shown. Figure 6 In FIG. 4 , it should be noted that the illustrated unit is the disparity refinement unit 40 d, but the unit may be any other unit from the disparity refinement units 40 b, 40 c, 40 d, since these units preferably have the same structure. The illustrated disparity refinement unit 40 d receives the left disparity map 46 e and the right disparity map 48 e from the previous hierarchical level and generates as its output new refined left disparity map 46 d and right disparity map 48 d (input and output are in Figure 6 The structure of the parallax refinement unit 40d is similar to Figure 3 The parallax unit 40e.
[0165] First, the input disparity maps 46e, 48e are upscaled (upsampled) to the next scale (in the illustrated exemplary case by scale 2, whereby the units 120, 130 are labeled by "x2"). In the illustrated example, the upscaling is done in the spatial dimensions (i.e., in height and width) according to the other directional scale up-scaling (downscaling) of the feature maps between the two levels. If some levels are skipped from the hierarchy, the upscaling factor can be any power of 2, for example. Powers of 2 are typical choices for upscaling factors, but other numbers (particularly integers) can also be used as upscaling factors (taking into account the scales of the corresponding feature map pairs, the scales of the feature map pairs applied for refinement must correspond to the displacement maps of the corresponding levels; see also below for possible upscaling of the output disparity map). The upscaling (upsampling) of the input disparity map can be done by any upsampling method (e.g., bilinear or nearest neighbor interpolation, deconvolution, or a method including nearest neighbor interpolation followed by depthwise convolution, etc.). Upscaling should consist of multiplying the disparity value (or displacement vector in general) by the upscaling factor, since the feature map ( Figure 6 20d and 30d in ) than the previous feature map used in computing the input disparity ( Figure 6 46e and 48e) in , is some multiple larger, so the displacement between two previous matching data elements should also be multiplied by the same upscaling factor.
[0166] Hereinafter, the operation of the disparity refinement unit 40d is described in the following route: starting from the left upscaling unit 120, passing through the right warping unit 134 and the right shifter unit 132, and ending at the left comparator unit 126 and the left addition unit 128. Since the structure of the disparity refinement unit 40d is symmetrical from left to right and vice versa, the description can also be applied to the route starting from the right upscaling unit 130, passing through the left warping unit 124 and the left shifter unit 122, and ending at the right comparator unit and the right addition unit 138.
[0167] The output of the left upscaling unit 120 is provided to a warping unit 134 for warping the right feature map 30d, which is also forwarded to the addition unit 128. The addition unit generates an output left disparity map 46d of the disparity refinement unit 40d based on the upscaled version of the disparity map 46e and the disparity (refinement) generated by the (left) comparator unit 126. The right disparity map 48d of the next level is generated in a similar manner by applying the addition unit 138 to the output of the upscaling unit 130 and the output of the (right) comparator unit 136. The warping units 124 and 134 perform the warping operations defined above, i.e., roughly speaking, in these operations the disparity map is applied to the feature map.
[0168] The output of the left upscaling unit 120 is routed to the right warping unit 134. In the warping unit 134, its input (i.e., the right feature map 30d) is warped by an upscaled version of the disparity map 46e of the previous level. In other words, the right feature map of the current level is warped with the help of the upscaled disparity of the lower adjacent level, i.e., the disparity of the lower level in the figure is applied to the right feature map of the current level to have a good comparability with the left feature map, because the lower level disparity map is a good approximation of the final disparity (in this way, non-reference features become spatially closer to the reference features). This warped version of the right feature map 30d of the current level is provided to the right shifter unit 132, which generates the appropriate number of shifts and forwards the shifted set of warped right feature maps to the left comparator unit 126 (similar to the disparity unit 40e, cross connections 142 and 144 are applied between the shifter unit and the other side comparator unit, see Figure 6 ).
[0169] Using a warp unit before the corresponding shifter unit allows to apply a relatively small number of symmetric shifts (because the warped version of the right feature map is very close to the left feature map at the same level and vice versa). In other words, the shifter unit (shift block) produces symmetric shifts here, so it can improve the coarse disparity map in any direction. This makes the disparity refinement unit very efficient (in other words, the comparator only needs to process a small number of shifts, which leads to improved performance) for the following reasons.
[0170] The left comparator unit gets as input the left feature map 20d as a reference and a shifted version of the right feature map 30d distorted by the left disparity map 46e of the previous level (lower level in the figure). Because this type of distortion is used, the distortion unit 134 outputs a good approximation of the current left feature map, which is then compared with the true left feature map 20d of the current level. Therefore, the output of the comparator will be a disparity map correction as follows, by which the coarser estimate of the previous level (which is a good estimate at its own level) is refined with the help of finer resolution data (i.e., with the help of feature maps with higher resolution). In the simplified example, it can be seen at the coarser level that the object has a shift of 4 pixels on the smaller feature map pair corresponding to that level. Thus, when refining the next level, the shift will be 16 pixels (using a scaling factor of 4), but this is then refined because the shift can be better observed on the feature map of the next level; accordingly, the prediction of the shift at the next level can be 15.
[0171] exist Figure 2 In the implementation of Figure 6In the case where the embodiment of the present invention is applied to all disparity refinement units 40b, 40c, 40d, the size of the final output disparity maps 46 and 48 is downscaled four times compared to the input images (left image 20a and right image 30a). Of course, it is possible to obtain a disparity map having the same size as the input image. If such a disparity map is to be obtained, the upscaling unit (similar to Figure 6 wherein the upscaling unit applies an upscaling factor of 4, thereby increasing the resolution of the disparity map and simultaneously the magnitude of the displacement vectors (eg, disparity values) by the same upscaling factor.
[0172] Accordingly, in an addition unit 128, the disparity map correction obtained in the comparator unit 126 and an upscaled version of the disparity map 46e of the previous level are added to give a refined disparity map. According to the above approach, the disparity map of the previous level is corrected on a pixel-by-pixel basis in the disparity refinement unit 40d.
[0173] In summary, a hierarchical structure is applied in the method and apparatus according to the present invention, so that only a small number of possible shifts between corresponding feature maps need to be processed at each level in the embodiments. Therefore, the method and apparatus of various embodiments provide a fast way to calculate a disparity map between two images. At each feature scale, the left and right features are compared in a special way that is fast, easy to learn, and robust to small errors in camera correction.
[0174] In the method and apparatus according to the invention, the hierarchical structure exploits the features of the base network at different scales. Summarizing the above considerations, the comparator at the next scale only needs to process the residual correction factor, which leads to a drastic reduction in the number of possible shifts to be applied, resulting in a much faster runtime. When comparing feature maps to each other, one of the left and right feature maps is selected as the reference (F ref ), and compare it with another feature map (F C ). At a given level of the feature hierarchy (typically up to five possible disparity values per level), Fc is shifted by all possible disparity values (i.e., the shifts can also be represented by vectors), and at each shift Fc is compared to ref Make a comparison.
[0175] As mentioned above, some embodiments of the invention relate to an apparatus for generating a displacement map, for example for generating a disparity map of a stereo image pair.The above detailed embodiments of the method according to the invention may also be described as embodiments of the apparatus according to the invention.
[0176] Accordingly, in an embodiment of the device (see Figure 3 ):
[0177] - The displacement unit 40e further comprises a first shifter unit ( Figure 3 a disparity unit 40e and a shifter unit 62, 72 in the feature map basis pair), the first shifter unit being adapted to generate a plurality of shifted feature maps of the first feature map in the feature map basis pair by applying a plurality of different shifts to the first feature map in the feature map basis pair, and
[0178] - First comparator unit ( Figure 3 The comparator units 64, 74 in the embodiment are adapted to obtain an initial displacement map by generating a resulting shift for each data element (pixel) position of the second feature map in the feature map basis pair based on a match between the multiple shifted feature maps of the first feature map in the feature map basis pair and the second feature map in the feature map basis pair.
[0179] In a further embodiment of the device (which can be combined with the previous embodiments; see Figure 6 ):
[0180] - comprising a second shifter unit ( Figure 6 The displacement refinement unit ( Figure 2 a disparity refinement unit 40b, 40c, 40d in the feature map refinement pair, or one or more of these units), adapted to generate a plurality of shifted feature maps of a first feature map in the feature map refinement pair by applying a plurality of different shifts to a warped version of the first feature map in the feature map refinement pair, and
[0181] - Second comparator unit ( Figure 6 The comparator units 126, 136 in the feature map refinement pair are adapted to obtain a correction displacement map by generating a resulting shift for each data element (pixel) position of the second feature map in the feature map refinement pair based on a match between the multiple shifted feature maps of the first feature map in the feature map refinement pair and the second feature map in the feature map refinement pair.
[0182] Similar to the first input data set and the second input data set, the "first" and "second" names mentioned in the first / second comparator unit or the first / second shifter unit only refer to the fact that both the displacement unit and the displacement refinement unit (e.g. the disparity unit and the disparity refinement unit) have their own such sub-units. This proposal does not mean that the internal structure (implementation) itself should be different for the respective first and second units. In fact, the internal structure of the respective first and second shifter units is preferably the same; in the comparator unit, for example, the convolution weights (see Figure 5A and 5B The first and second comparison convolution units 84, 94) in may be different.
[0183] In any of the previous two embodiments (see Figure 5A and 5B ), preferably to match the following:
[0184] - a plurality of shifted feature maps, the number of shifts of the feature maps being the number of different shift values, and
[0185] a second characteristic map used as reference characteristic map 80,
[0186] a first comparator unit for generating an output displacement map for use as an initial displacement map, and / or a second comparator unit for generating an output displacement map for use as a corrected displacement map comprising:
[0187] a first comparison convolution unit 84 adapted to be applied to the reference feature map 80 to obtain a first intermediate comparison data map,
[0188] a second comparison convolution unit 94 adapted to be applied to each of the plurality of shifted feature maps to obtain a plurality of second intermediate comparison data maps,
[0189] an adding unit 98 adapted to add the first intermediate comparison data map replicated (virtually or physically) according to different shift numbers to the plurality of second intermediate comparison data maps to obtain a result comparison data map,
[0190] - Result shift generation unit (see for example Figure 5B The output displacement map further comprises result shift generating units C2_A, C2_B, C2_C in the result comparison data map, which are suitable for generating corresponding result shifts for each data unit of the result comparison data map in all dimensions of at least one spatial dimension and / or time dimension, and for assigning all corresponding result shifts to corresponding data elements in the output displacement map (i.e., data elements in the result comparison data map according to the corresponding result shifts).
[0191] Accordingly, the first comparator unit and / or the second comparator unit comprises units 84, 94, 98 and a result shift generation unit. The purpose of these units is to generate an output displacement (e.g. disparity) map for different purposes, i.e. the output displacement map will be the initial displacement map itself in the first comparator unit of the displacement unit and will be the corrected displacement map in the second comparator unit of the displacement refinement unit.
[0192] Preferably, in the previous embodiment, the result shift generation unit (the result shift generation unit C2_A is an example of this embodiment) includes a shift study convolution unit 104, which is suitable for being applied to the result comparison data map to generate corresponding result shifts for each data element in the result comparison data map in all dimensions of at least one spatial dimension and / or time dimension, and is suitable for assigning all corresponding result shifts to corresponding data elements in the output displacement map (i.e., according to their data elements in the result comparison map).
[0193] Specifically, in the previous embodiment, the feature map has one or more feature channels, and the first comparator unit and / or the second comparator unit includes a stacking unit 102, which is suitable for stacking the result comparison data map through the one or more feature channels of the feature map before applying the shift study convolution unit 104.
[0194] Using an alternative scheme of a result shift generation unit, a corresponding result shift for each data element of the result comparison data graph in all dimensions of at least one spatial dimension and / or time dimension thereof is generated by a result shift generation unit (result shift generation unit C2_B is an example of this embodiment) by selecting the best matching shift from the multiple different shifts.
[0195] In a further alternative scheme using a result shift generation unit, a corresponding result shift for each data element of the result comparison data map in all dimensions of at least one spatial dimension and / or time dimension thereof is generated by the result shift generation unit (the result shift generation unit C2_B is an example of this embodiment) by the following operations:
[0196] - establish multiple displacement bins for each different possible displacement value,
[0197] - generating a displacement probability for each displacement bin, the displacement probability being calculated based on the result comparison data graph, and
[0198] - The resulting shift is obtained by weighting the possible shift values with the corresponding shift probabilities.
[0199] Preferably, in any of the previous five embodiments, the device (i.e., the comparison graph generation unit C1) includes a non-linear layer, which is suitable for being applied to the result comparison data graph before generating a corresponding result shift for each data element of the result comparison data graph in all dimensions of at least one spatial dimension and / or time dimension thereof.
[0200] In an embodiment of the device according to the invention, preferably, the feature map has one or more feature channels, and the displacement unit and / or the displacement refinement unit is adapted to perform the matching by taking the one or more feature channels of the feature map into account and to generate the initial displacement map and the corrected displacement map (in Figure 2 In the case of an embodiment of the present invention, the disparity map has a single coordinate channel, since the disparity map may only include shifts in one direction [the shift is a single number]. Accordingly, without being otherwise constrained by the task, the initial displacement map and the corrected displacement map (which have the same number of channels) have as many coordinate channels as the number of spatial dimensions and / or temporal dimensions.
[0201] In a further embodiment of the device according to the invention (for reference numerals see Figure 2 ), the displacement unit (e.g., the disparity unit 40e) is adapted to generate a pair of first initial displacement maps and a second initial displacement map, and the displacement refinement unit (the disparity refinement units 40b, 40c, 40d) is adapted to generate a pair of first updated displacement maps and a second updated displacement map based on the pair of first initial displacement maps and the second initial displacement map and a pair of first corrected displacement maps and a second corrected displacement map generated by their corresponding second comparator units.
[0202] In an embodiment of the apparatus (as in the illustrated embodiment), the input data set pair is a stereo image pair, the displacement map is a disparity map, the displacement unit is a disparity unit, and the displacement refinement unit is a disparity refinement unit.
[0203] Of course, the present invention is not limited to the preferred embodiments described in detail above, but further variations, modifications and developments are possible within the scope of protection determined by the claims. In addition, all embodiments defined by any combination of dependent claims belong to the present invention.
Claims
1. A method for generating a displacement map of a first input data set and a second input data set in an input data set pair, each input data set having at least one spatial dimension and / or a temporal dimension, the method comprising the following steps: - processing the first input data set and the second input data set by a neural network-based feature extractor (25) to generate a feature map hierarchy (50), the feature map hierarchy (50) comprising a feature map base pair (20e, 30e) and a feature map refinement pair (20b, 30b, 20c, 30c, 20d, 30d), each pair of feature maps (20b, 30b, 20c, 30c, 20d, 30d, 20e, 30e) constituting a level of the feature map hierarchy (50), the feature map refinement pair (20b, 30b, 20c, 30c, 20d, 30d) having a resolution that is less coarse than the resolution of the feature map base pair (20e, 30e) in all dimensions of at least one spatial dimension and / or a temporal dimension; - generating an initial displacement map in a displacement generating operation based on matching a first feature map in the feature map basis pair (20e, 30e) with a second feature map in the feature map basis pair (20e, 30e), the initial displacement map being for the feature map basis pair (20e, 30e) in the feature map hierarchy (50); -During displacement refinement operation: - upscaling the initial displacement map in all dimensions in at least one spatial dimension and / or temporal dimension thereof to the scale of the feature map refinement pair (20b, 30b, 20c, 30c, 20d, 30d) in the feature map level (50) using corresponding upscaling factors, and multiplying the values of the initial displacement map by the corresponding upscaling factors to generate an upscaled initial displacement map, - performing a warping operation by applying the upscaled initial displacement map to a first feature map in the feature map refinement pair (20b, 30b, 20c, 30c, 20d, 30d) in the feature map hierarchy (50), generating a warped version of the first feature map in the feature map refinement pair (20b, 30b, 20c, 30c, 20d, 30d), - matching the warped version of the first feature map in the feature map refinement pair (20b, 30b, 20c, 30c, 20d, 30d) with the second feature map in the feature map refinement pair (20b, 30b, 20c, 30c, 20d, 30d) to obtain a corrected displacement map for the warped version of the first feature map in the feature map refinement pair (20b, 30b, 20c, 30c, 20d, 30d) and the second feature map in the feature map refinement pair (20b, 30b, 20c, 30c, 20d, 30d), the corrected displacement map being added to the upscaled initial displacement map to obtain an updated displacement map.
2. The method according to claim 1, characterized in that: - at least two feature map refinement pairs (20b, 30b, 20c, 30c, 20d, 30d) are included in the feature map hierarchy (50), wherein a first feature map refinement pair (20d, 30d) closest to the feature map base pair (20e, 30e) in the feature map hierarchy (50) has a resolution that is less coarse than the resolution of the feature map base pair (20e, 30e) and is coarser than the first feature map refinement pair (20d, 30d) each successive feature map refinement pair (20b, 30b, 20c, 30c) that is less close to the feature map base pair (20e, 30e) in the feature map level (50) has a resolution that is less coarse than the resolution of an adjacent feature map refinement pair (20d, 30d) that is closer to the feature map base pair (20e, 30e) in the feature map level (50) than the corresponding successive feature map refinement pair (20b, 30b, 20c, 30c), - the displacement refinement operation is performed using the first feature map refinement pair (20d, 30d), and a respective further displacement refinement operation is performed for each successive feature map refinement pair (20b, 30b, 20c, 30c), wherein in each further displacement refinement operation an updated displacement map obtained for the adjacent feature map refinement pair (20d, 30d) closer to the feature map base pair (20e, 30e) in the feature map hierarchy (50) than the respective successive feature map refinement pair (20b, 30b, 20c, 30c) is used as the initial displacement map, the initial displacement map being upscaled to the scale of the respective successive feature map refinement pair (20b, 30b, 20c, 30c) during upscaling in the respective displacement refinement operation using a respective upscaling factor, and the values of the updated displacement map are multiplied by the respective upscaling factor.
3. The method according to claim 1, characterized in that In the displacement generation operation: a plurality of shifted feature maps (82) of the first feature map in the feature map basis pair (20e, 30e) are generated by applying a plurality of different shifts to the first feature map in the feature map basis pair (20e, 30e), and the initial displacement map is obtained by generating a resulting shift for each data element position of the second feature map in the feature map basis pair (20e, 30e) based on studying the matching between the plurality of shifted feature maps (82) of the first feature map in the feature map basis pair (20e, 30e) and the second feature map in the feature map basis pair (20e, 30e), and / or in the displacement refinement operation: - a plurality of shifted feature maps (82) of a first feature map in the feature map refinement pair (20b, 30b, 20c, 30c, 20d, 30d) are generated by applying a plurality of different shifts to the distorted version of the first feature map in the feature map refinement pair (20b, 30b, 20c, 30c, 20d, 30d), - the correction displacement map is obtained by generating a resulting shift for each data element position of the second feature map in the feature map refinement pair (20b, 30b, 20c, 30c, 20d, 30d) based on studying the matching between the plurality of shifted feature maps (82) of the first feature map in the feature map refinement pair (20b, 30b, 20c, 30c, 20d, 30d) and the second feature map in the feature map refinement pair (20b, 30b, 20c, 30c, 20d, 30d).
4. The method according to claim 3, characterized in that In the displacement generating operation for generating an output displacement map for use as the initial displacement map and / or in the displacement refining operation for generating the output displacement map for use as the corrected displacement map, The matching of the plurality of shifted feature maps (82) and the second feature map used as a reference feature map (80) is performed in the following steps, wherein the number of shifts of the plurality of shifted feature maps (82) is the number of different shifts: - applying a first comparison convolution unit (84) to the reference feature map (80) to obtain a first intermediate comparison data map, - applying a second comparison convolution unit (94) to each of the plurality of shifted feature maps (82) to obtain a plurality of second intermediate comparison data maps, - adding the first intermediate comparison data map copied according to the number of different shifts to the plurality of second intermediate comparison data maps in an addition operation to obtain a result comparison data map, - generating, for each data element of the result comparison data map, respective result shifts in all dimensions of at least one spatial dimension and / or a temporal dimension thereof, and assigning all respective result shifts to corresponding data elements in the output displacement map.
5. The method according to claim 4, characterized in that The corresponding result shift for each data element in the result comparison data map in all dimensions of at least one spatial dimension and / or time dimension thereof is generated by applying a shift study convolution unit (104) to the result comparison data map.
6. The method according to claim 5, characterized in that The feature map hierarchy has one or more feature channels, and the result comparison data map is stacked by the one or more feature channels of the feature map hierarchy before applying the shift study convolution unit (104).
7. The method according to claim 4, characterized in that A corresponding result shift for each data element in the result comparison data map in all dimensions of at least one spatial dimension and / or a temporal dimension thereof is generated by selecting a best matching shift from the plurality of different shifts.
8. The method according to claim 4, characterized in that The corresponding result shifts for each data element in the result comparison data graph in all dimensions of at least one spatial dimension and / or time dimension thereof are generated by the following operations: -Create multiple displacement bins for each different shift value, - generating a displacement probability for each displacement bin, the displacement probability being calculated based on the result comparison data map, and - The resulting shift is obtained by weighting the shift values by the corresponding shift probabilities.
9. The method according to claim 4, characterized in that A non-linear layer is applied to the result comparison data map before generating a corresponding result shift for each data element in the result comparison data map in all dimensions of at least one spatial dimension and / or a temporal dimension thereof.
10. The method according to claim 1, characterized in that The feature map level has one or more feature channels, and in the displacement generation operation and / or in the displacement refinement operation, the matching is performed by taking into account the one or more feature channels of the feature map level, and the initial displacement map and the corrected displacement map are generated using the same or fewer number of coordinate channels compared to the dimension of at least one spatial dimension and / or time dimension of the first input data set and the second input data set, respectively.
11. The method according to claim 1, characterized in that A pair of first and second initial displacement maps are generated in the displacement generating operation, and a pair of first and second updated displacement maps are generated in the displacement refining operation based on the pair of first and second initial displacement maps and the pair of first and second corrected displacement maps generated in the displacement refining operation.
12. The method according to claim 1, characterized in that The pair of input data sets is an image pair of stereo images, the displacement map is a disparity map, the displacement generation operation is a disparity generation operation, and the displacement refinement operation is a disparity refinement operation.
13. An apparatus for generating a displacement map of a first input data set and a second input data set in an input data set pair, each input data set having at least one spatial dimension and / or a temporal dimension, the apparatus comprising: a neural network-based feature extractor (25) adapted to process the first input data set and the second input data set to generate a hierarchy of feature maps (50), the hierarchy of feature maps (50) comprising a base pair of feature maps (20e, 30e) and a pair of feature map refinements (20b, 30b, 20c, 30c, 20d, 30d), each pair of feature maps (20b, 30b, 20c, 30c, 20d, 30d, 20e, 30e) constituting a level of the hierarchy of feature maps (50), the resolution of the feature map refinement pair (20b, 30b, 20c, 30c, 20d, 30d) being less coarse than the resolution of the base pair of feature maps (20e, 30e) in all of its at least one spatial dimension and / or temporal dimension; a displacement unit comprising a first comparator unit (64, 74) adapted to match a first feature map of the feature map basis pair (20e, 30e) with a second feature map of the feature map basis pair (20e, 30e) to generate an initial displacement map for the feature map basis pair (20e, 30e) in the feature map hierarchy (50); -Displacement refinement unit, which includes: an upscaling unit (120, 130) adapted to upscale the initial displacement map in all dimensions in at least one spatial dimension and / or temporal dimension thereof to the scale of the feature map refinement pair (20b, 30b, 20c, 30c, 20d, 30d) in the feature map level (50) using a corresponding upscaling factor, and adapted to multiply the values of the initial displacement map by the corresponding upscaling factor to generate an upscaled initial displacement map, a warping unit (124, 134) adapted to perform a warping operation on a first feature map in the feature map refinement pair (20b, 30b, 20c, 30c, 20d, 30d) in the feature map hierarchy (50), generating a warped version of the first feature map in the feature map refinement pair (20b, 30b, 20c, 30c, 20d, 30d), - a second comparator unit (126, 136) adapted to match the distorted version of the first feature map in the feature map refinement pair (20b, 30b, 20c, 30c, 20d, 30d) with the second feature map in the feature map refinement pair (20b, 30b, 20c, 30c, 20d, 30d) to obtain a corrected displacement map for the distorted version of the first feature map in the feature map refinement pair (20b, 30b, 20c, 30c, 20d, 30d) and the second feature map in the feature map refinement pair (20b, 30b, 20c, 30c, 20d, 30d), and - an adding unit (128, 138) adapted to add the corrected displacement map to the upscaled initial displacement map to obtain an updated displacement map.
14. The device according to claim 13, characterized in that: - comprising at least one further displacement refinement unit, and at least two feature map refinement pairs (20b, 30b, 20c, 30c, 20d, 30d) are included in the feature map hierarchy (50), wherein a first feature map refinement pair (20d, 30d) closest to the feature map base pair (20e, 30e) in the feature map hierarchy (50) has a resolution that is less coarse than the resolution of the feature map base pair (20e, 30e) and is identical to the first feature map refinement pair (20d, 30d) 0d, 30d), the resolution of each successive feature map refinement pair (20b, 30b, 20c, 30c) not close to the feature map base pair (20e, 30e) in the feature map level (50) is less coarse than the resolution of an adjacent feature map refinement pair (20d, 30d) that is closer to the feature map base pair (20e, 30e) in the feature map level (50) than the corresponding successive feature map refinement pair (20b, 30b, 20c, 30c), - the displacement refinement unit is applied to the first feature map refinement pair (20d, 30d), and a respective further displacement refinement unit is applied to each successive feature map refinement pair (20b, 20b, 20c, 30c), wherein in each further displacement refinement unit the updated displacement map obtained for the neighbouring feature map refinement pair (20d, 30d) closer to the feature map base pair (20e, 30e) in the feature map hierarchy (50) than the respective successive feature map refinement pair (20b, 20b, 20c, 30c) is used as the initial displacement map, the initial displacement map being upscaled to the scale of the respective successive feature map refinement pair (20b, 20b, 20c, 30c) during upscaling in the respective displacement refinement operation using a respective upscaling factor, and the values of the updated displacement map are multiplied by the respective upscaling factor.
15. The device according to claim 13, characterized in that: - the shifting unit further comprises a first shifter unit (62, 72) adapted to generate a plurality of shifted feature maps (82) of the first feature map in the feature map basis pair (20e, 30e) by applying a plurality of different shifts to the first feature map of the feature map basis pair (20e, 30e), wherein the first comparator unit (64, 74) is adapted to obtain the initial shifted map by generating a resulting shift for each data element position of the second feature map in the feature map basis pair (20e, 30e) based on studying a match between the plurality of shifted feature maps (82) of the first feature map in the feature map basis pair (20e, 30e) and the second feature map in the feature map basis pair (20e, 30e), and / or - the displacement refinement unit further comprises a second shifter unit (122, 132) adapted to generate a plurality of shifted feature maps (82) of the first feature map of the feature map refinement pair (20b, 20b, 20c, 30c, 20d, 30d) by applying a plurality of different shifts to the distorted version of the first feature map of the feature map refinement pair (20b, 20b, 20c, 30c, 20d, 30d), wherein the second comparator unit (126, 136) is ... The correction displacement map is obtained by studying the matches between the plurality of shifted feature maps (82) of the first feature map in the feature map refinement pair (20b, 20b, 20c, 30c, 20d, 30d) and the second feature map in the feature map refinement pair (20b, 20b, 20c, 30c, 20d, 30d) to generate a result shift for each data element position of the second feature map in the feature map refinement pair (20b, 20b, 20c, 30c, 20d, 30d).
16. The device according to claim 15, characterized in that To match the following: - the plurality of shifted feature maps (82), wherein the number of shifts of the plurality of shifted feature maps (82) is the number of different shifts, and - said second characteristic map serving as a reference characteristic map (80), The first comparator unit (64, 74) for generating an output displacement map for use as the initial displacement map and / or the second comparator unit (126, 136) for generating the output displacement map for use as the corrected displacement map comprises: - a first comparison convolution unit (84) adapted to be applied to said reference feature map (80) to obtain a first intermediate comparison data map, - a second comparison convolution unit (94) adapted to be applied to each of the plurality of shifted feature maps (82) to obtain a plurality of second intermediate comparison data maps, - an adding unit (98) adapted to add the first intermediate comparison data map replicated according to the number of different shifts to the plurality of second intermediate comparison data maps to obtain a result comparison data map, - A result shift generation unit (C2_A, C2_B, C2_C), which is suitable for generating a corresponding result shift for each data element of the result comparison data map in all dimensions of at least one spatial dimension and / or time dimension thereof, and is suitable for assigning all corresponding result shifts to corresponding data elements in the output displacement map.
17. The device according to claim 16, characterized in that The result shift generation unit (C2_A) comprises a shift study convolution unit (104), which is suitable for being applied to the result comparison data map to generate corresponding result shifts for each data element of the result comparison data map in all dimensions of at least one spatial dimension and / or time dimension thereof, and is suitable for assigning all corresponding result shifts to the output displacement map according to the data elements of the result comparison data map in which the corresponding result shifts are located.
18. The device according to claim 17, characterized in that The feature map hierarchy has one or more feature channels, and the first comparator unit (64, 74) and / or the second comparator unit (126, 136) includes a stacking unit (102), which is suitable for stacking the result comparison data map through the one or more feature channels of the feature map hierarchy before applying the shift study convolution unit (104).
19. The device according to claim 16, characterized in that The corresponding result shift for each data element in the result comparison data map in all dimensions of at least one spatial dimension and / or time dimension is generated by the result shift generation unit (C2_B) by selecting the best matching shift from the multiple different shifts.
20. The device according to claim 16, characterized in that The corresponding result shifts of each data element in the result comparison data graph in all dimensions of at least one spatial dimension and / or time dimension are generated by the result shift generation unit (C2_C) through the following operations: - create multiple displacement bins for each possible displacement value, - generating a displacement probability for each displacement bin, the displacement probability being calculated based on the result comparison data map, and - obtaining the resulting shift by weighting the possible shift values by the corresponding shift probabilities.
21. The device according to claim 16, characterized in that A non-linear layer is included which is adapted to be applied to the result comparison data map before generating a respective result shift for each data element of the result comparison data map in all dimensions of at least one spatial dimension and / or a temporal dimension thereof.
22. The device according to claim 13, characterized in that The feature map level has one or more feature channels, and the displacement unit and / or the displacement refinement unit are suitable for performing the matching by taking the one or more feature channels of the feature map level into consideration, and are suitable for generating an initial displacement map and a corrected displacement map using the same or fewer number of coordinate channels compared to the dimension of at least one spatial dimension and / or time dimension of the first input data set and the second input data set, respectively.
23. The device according to claim 13, characterized in that The displacement unit is adapted to generate a pair of first and second initial displacement maps, and the displacement refinement unit is adapted to generate a pair of first and second updated displacement maps based on the pair of first and second initial displacement maps and a pair of first and second corrected displacement maps generated by respective second comparator units (126, 136) of the displacement refinement unit.
24. The device according to claim 13, characterized in that The input data set pair is an image pair (20a, 30a) of stereo images, the displacement map is a disparity map, the displacement unit is a disparity unit (40e), and the displacement refinement unit is a disparity refinement unit (40b, 40c, 40d).
Citation Information
Patent Citations
Joint bilateral upsampling
US20080267494A1
System and method of processing stereo images
US20110176722A1
Method of time-efficient stereo matching
US20120008857A1
Disparity Estimation for Misaligned Stereo Image Pairs
US20140147031A1
Process for estimating disparity between the monoscopic images making up a sterescopic image
US5727078A