DEVICE AND METHOD FOR PROCESSING THE DATA OF A NEURAL NETWORK
Patent Information
- Application Number
- DE602017089313
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2017-03-24
- Publication Date
- 2025-05-07
- Estimated Expiration
- 2037-03-24
AI Technical Summary
Existing neural network technologies for data processing, such as convolutional neural networks (CNNs), lack the ability to adaptively use position-dependent kernels, which limits their effectiveness in tasks like image upscaling and deconvolution.
A new type of neural network layer that computes up-scaled data using individual up-scaling weights learned for each spatial position, where these weights are calculated as a function of position-dependent weights or similarity features and position-independent learned weight kernels.
This approach allows for better adaptation of up-scaling weights to input data, improving the quality of upscaling and deconvolution processes by utilizing position-dependent kernels that can vary across different spatial positions.
Description
TECHNICAL FIELD
[0001] Generally, the present invention relates to the field of machine learning or deep learning based on neural networks. More specifically, the present invention relates to a neural network data processing apparatus and method, in particular for processing data in the fields of audio processing, computer vision, image or video processing, classification, detection and / or recognition.BACKGROUND
[0002] Guided up-scaling, which is commonly used in many signal processing applications, including especially image up-scaling methods for image quality improvement, super-resolution and many others [Kaiming He, Jian Sun, Xiaoou Tang, "Guided Image Filtering", ECCV 2010], is a process in which input data is being combined with additional input in form of up-scaling weights that control the influence of each input data value on the result to form the output data.
[0003] In deep-learning, a common approach recently used in many application fields is the utilization of convolutional neural networks (CNNs). Generally, a specific part of such convolutional neural networks is at least one convolution (or convolutional) layer which performs a convolution of input data values with a learned kernel K producing one output data value per convolution kernel for each output position [J. Long, E. Shelhamer, T. Darrell, "Fully Convolutional Networks for Semantic Segmentation", CVPR 2015]. For the two-dimensional case used, for instance, in image processing the convolution using the learned kernel K can be expressed mathematically as follows: out x y = ∑ i = − r r ∑ j = − r r in x − i , y − j ⋅ K i j + B , wherein out(x, y) denotes the array of output data values, in(x - i, y - j) denotes a sub-array of input data values and K(i, j) denotes the kernel comprising an array of kernel weights or kernel values of size (2r+1)×(2r+1). B denotes an optional learned bias term, which can be added for obtaining each output data value. The weights of the kernel K are the same for the whole array of input data values in(x, y) and are generally learned during a learning phase of the neural network which, in case of 1st order methods, consists of iteratively back-propagating the gradients of the neural network output back to the input layers and updating the weights of all the network layers by a partial derivative computed in this way. An extension of CNNs are deconvolutional neural networks (DNNs) with a specific element that extends their functionality relative to CNNs that is called deconvolution. Deconvolution can be interpreted as an "inversed" convolution known from classical CNNs.
[0004] XP032866372 (Riegler Gernot et al: "Conditioned Regression Models for Non-blind Single Image Super-RESOLUTION", 2015 IEEE International Conference on Computer VISION(ICCV), IEEE, 7 December 2015, pages 522-5300) introduces different blur kernels can be used for different images, and does not disclose that different blur kernels can be used for different spatial positions of the same image.
[0005] XP080732861 (Di Kang et al: "Crowd Counting by Adapting Convolutional Neural Networks with Side Information", 2016 ARXIV.ORG, Cornell University Library, NY 14853, 21 November 2016) introduces filter weights are generated by the filter manifold network (FMN) from the auxiliary input z (i.e. side information), and does not disclose that different filter weights can be used for different spatial positions of the input map.
[0006] XP055368215 (Kan Chen et al: "ABC-CNN: An Attention Based Convolutional Neural Network for Visual Question Answering", November 2015) introduces convolutional kernels are applied on the image feature maps to generate question-guided attention maps, and does not disclose that different convolutional kernels can be used for different spatial positions of the image.
[0007] XP033021632 (Jampani Varun et al: " Learning sparse high dimensional filters: image filtering, dense CRFs and Bilateral neural networks", 2016 IEEE International Conference on Computer VISION and Pattern Recognition(CVPR), IEEE, 27 June 2016, pages 4452-4461) introduces bilateral filters depending on the image content, and does not disclose that different bilateral filters can be used for different spatial positions of the image. XP0055711953 (Bert De Brabandere et al: "Dynimic Filter Networks", 31 May 2016) introduces dynamic filtering layer takes images as inputs (page 4, section 3.2 of D1), also discloses a specific filter is associated with each position of the input.
[0008] XP0055711953 (Liang-Chien Chen et al: "DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs", 2016 ARXIV.ORG, Cornell University Library, NY 14853, 2 June 2016) introduces filter weights information), and does not disclose that different filter weights can be used for different spatial positions of the input map) introduces an alternative direction based on coupling the recognition capacity of DCNNs and the fine-grained localization accuracy of fully connected CRFs, which using contrast sensitive potentials in conjunction to local-range CRFs and fully connected CRF model, which employs the energy function.SUMMARY
[0009] It is an object of the invention to provide an improved data processing apparatus and method based on neural networks.
[0010] The foregoing and other objects are achieved by the subject matter of the independent claims. Further implementation forms are apparent from the dependent claims, the description and the figures.
[0011] Generally, embodiments of the invention provide a new approach for deconvolution or upscaling of data for neural networks that is implemented into a neural network as a new type of neural network layer. The neural network layer can compute up-scaled data using individual up-scaling weights that are learned for each individual spatial position. Up-scaling weights can be computed as a function of position dependent weights or similarity features and position independent learned weight kernels, resulting in individual up-scaling weights for each input spatial position. In this way a variety of sophisticated position dependent or position adaptive kernels learned by the neural network can be utilized for better adaptation of the up-scaling weights to the input data.
[0012] More specifically, according to a first aspect the invention relates to a data processing apparatus comprising one or more processors configured to provide a neural network. The data to be processed by the data processing apparatus can be, for instance, two-dimensional image or video data or one-dimensional audio data.
[0013] The neural network provided by the one or more processors of the data processing apparatus comprises a neural network layer being configured to process an array of input data values, such as a two-dimensional array of input data values in(x, y), into an array of are generated by the filter manifold network (FMN) from the auxiliary input z (i.e. side information), and does not disclose that different filter weights can be used for different spatial positions of the input map) introduces an alternative direction based on coupling the recognition capacity of DCNNs and the fine-grained localization accuracy of fully connected CRFs, which using contrast sensitive potentials in conjunction to local-range CRFs and fully connected CRF model, which employs the energy function.SUMMARY
[0014] It is an object of the invention to provide an improved data processing apparatus and method based on neural networks.
[0015] The foregoing and other objects are achieved by the subject matter of the independent claims. In particular, there is provided a data processing apparatus according to claim 1, a computer-implemented data processing method according to claim 13, and a computer program according to claim 14. Further implementation forms are apparent from the dependent claims, the description and the figures.
[0016] Generally, embodiments of the invention provide a new approach for deconvolution or upscaling of data for neural networks that is implemented into a neural network as a new type of neural network layer. The neural network layer can compute up-scaled data using individual up-scaling weights that are learned for each individual spatial position. Up-scaling weights can be computed as a function of position dependent weights or similarity features and position independent learned weight kernels, resulting in individual up-scaling weights for each input spatial position. In this way a variety of sophisticated position dependent or position adaptive kernels learned by the neural network can be utilized for better adaptation of the up-scaling weights to the input data.
[0017] More specifically, according to a first aspect the invention relates to a data processing apparatus comprising one or more processors configured to provide a neural network. The data to be processed by the data processing apparatus can be, for instance, two-dimensional image or video data or one-dimensional audio data.
[0018] The neural network provided by the one or more processors of the data processing apparatus comprises a neural network layer being configured to process an array of input data values, such as a two-dimensional array of input data values in(x, y), into an array of output data values, such as a two-dimensional array of output data values out(x, y). The neural network layer can be a first layer or an intermediate layer of the neural network.
[0019] The array of input data values can be one-dimensional (i.e. a vector, e.g. audio or other e.g. temporal sequence), two-dimensional (i.e. a matrix, e.g. an image or other temporal or spatial sequence), or N-dimensional (e.g. any kind of N-dimensional feature array, e.g. provided by a conventional pre-processing or feature extraction and / or by other layers of the neural network).
[0020] The array of input data values can have one or more channels, e.g. for an RGB image one R-channel, one G-channel and one B-channel, or for a black / white image only one grey-scale or intensity channel. The term "channel" can refer to any "feature", e.g. features obtained from conventional pre-processing or feature extraction or from other neural networks or neural network layers of the same neural network. The array of input data values can comprise, for instance, two-dimensional RGB or grey scale image or video data representing at least a part of an image, or a one-dimensional audio signal. In case the neural network layer is implemented as an intermediate layer of the neural network, the array of input data values can be, for instance, an array of similarity features generated by previous layers of the neural network on the basis of an initial, i.e. original array of input data values, e.g. by means of a feature extraction.
[0021] The neural network layer is configured to generate from the array of input data values the array of output data values on the basis of a plurality of position dependent, i.e. spatially variable kernels and a plurality of different input data values of the array of input data values. Each kernel comprises a plurality of kernel values (also referred to as kernel weights). For a respective position or element of the array of input data values a respective kernel is applied thereto for generating a respective sub-array of the array of output data values. In an implementation form, the plurality of kernel values of a respective position dependent kernel can be respectively multiplied with a respective input data value for generating a respective sub-array of the array of output data values having the same size as the position dependent kernel, i.e. the array of kernel values. Generally, the size of the array of input data values can be smaller than the size of the array of output data values.
[0022] A "position dependent kernel" as used herein means a kernel whose kernel values can depend on the respective position or element of the array of input data values. In other words, for a first kernel used for a first input data value of the array of input data values the kernel values can differ from the kernel values of a second kernel used for a second input data value of the array of input data values. In a two-dimensional array the position could be a spatial position defined, for instance, by two spatial coordinates x,y. In a one-dimensional array the position could be a temporal position defined, for instance, by a time coordinate t.
[0023] Thus, an improved data processing apparatus based on neural networks is provided. The data processing apparatus allows upscaling or deconvolving the input data in a way that can better reflect mutual data similarity. Moreover, the data processing apparatus allows adapting the kernel weights for different spatial positions of the array of input data values. This, in turn, allows, for instance, minimizing the influence of some of the input data values on the result, for instance the input data values that are associated with another part of the scene (as determined by semantic segmentation) or a different object that is being analysed.
[0024] In a further implementation form of the first aspect, the neural network comprises at least one additional network layer configured to generate the plurality of position dependent kernels on the basis of an original array of original input values of the neural network, wherein the original array of original input values of the neural network comprises the array of input values or another array of input values associated to the array of input values. The original array of original input values can be the array of input data values or a different array.
[0025] In a further implementation form of the first aspect, the neural network is configured to generate the plurality of position dependent kernels based on a plurality of learned position independent kernels and a plurality of position dependent weights (also referred to as similarity features). Generally, the position independent kernels can be learned by the neural network and the position dependent weights (i.e. similarity features) can be computed, for instance, by a further preceding layer of the neural network. This implementation form allows minimizing the amount of data being transferred to the neural network layer in order to obtain the kernel values. This is because the kernel values are not transferred directly, but computed from the plurality of position dependent weights (i.e. similarity features) substantially reducing the amount of data for each element of the array of output data values. This can minimize the amount of data being stored and transferred by the neural network between the different network layers, which is especially important during the learning process on the basis of the mini-batch approach as the memory of the data processing apparatus (GPU) is currently the main bottleneck. Moreover, this implementation form allows for a better adaption of the kernel values to the processed data and utilizing more sophisticated similarity features. For instance, information about object shapes or object segmentations can be utilized in order to better preserve better object boundaries or even increase the level of details in the higher-resolution output. In this way, information about some small details from the original array of original input values not present in the possibly low-resolution array of input data values can be combined with the array of input data values in order to create higher-resolution array of output data values.
[0026] In a further implementation form of the first aspect, the neural network is configured to generate a kernel of the plurality of position dependent kernels by adding the learned position independent kernels each weighted by the associated non-learned position dependent weights (i.e. similarity features). This implementation form provides a very efficient representation of the plurality of position dependent kernels using a linear combination of position independent "base kernels".
[0027] In a further implementation form of the first aspect, the plurality of position independent kernels are predetermined or learned, and wherein the neural network comprises at least one additional neural network layer or "conventional" pre-processing layer configured to generate the plurality of position dependent weights (i.e. similarity features) based on an original array of original input values of the neural network, wherein the original array of original input values of the neural network comprises the array of input values or another array of input values associated to the array of input values. The original array of original input values can be the array of input data values or a different array. In an implementation form the at least one additional neural network layer or "conventional" pre-processing layer can generate the plurality of position dependent weights (i.e. similarity features) using, for instance, bilateral filtering, semantic segmentation, per-instance object detection, and data importance indicators like ROI (region of interest).
[0028] In a further implementation form of the first aspect, the array of input data values and the array of output data values are two-dimensional arrays, and the convolutional neural network layer is configured to generate the plurality of position dependent kernels w L (x, y, i, j) on the basis of the following equation: w L x y i j = ∑ f = 1 N f F f x y ⋅ K f i j , wherein F f (x, y) denotes the plurality of N f position dependent weights (i.e. similarity features) and K f (i,j) denotes the plurality of position independent "base" kernels.
[0029] In a further implementation form of the first aspect, the neural network layer is a deconvolutional network layer or an upscaling network layer.
[0030] In a further implementation form of the first aspect, the array of input data values and the array of output data values are two-dimensional arrays, wherein the neural network layer is a deconvolution network layer configured to generate the array of output data values on the basis of the following equations: out x y c o = 1 W L ′ x y c o ∑ c i = 1 C i ∑ x ′ , y ′ : x ′ − i = x , y ′ − j = y in x ′ , y ′ , c i w L x ′ , y ′ , c o , c i , i , j , W L ′ x y c o = ∑ c i = 1 C i ∑ x ′ , y ′ : x ′ − i = x , y ′ − j = y w L x ′ , y ′ , c o , c i , i , j , i ∈ − r , … , r , j ∈ − r , … , r , wherein x, y, x', y', i, j denote array indices, out(x, y, c o ) denotes the multi-channel array of output data values, in(x', y', c i ) denotes the array of input data values, r denotes a size of each kernel of the plurality of position dependent multi-channel kernels w L (x', y', c o , c i , i, j) and W L '(x, y, c 0 ) denotes a normalization factor. In an implementation form the normalization factor W L '(x, y, c 0 ) can be set equal to 1.
[0031] In a further implementation form of the first aspect, the array of input data values and the array of output data values are two-dimensional arrays, wherein the neural network layer is an upscaling network layer configured to generate the array of output data values on the basis of the following equations: out x y = 1 W L ′ x y ∑ x ′ , y ′ : x ′ − i = x , y ′ − j = y in x ′ , y ′ w L x ′ , y ′ , i , j , W L ′ x y = ∑ x ′ , y ′ : x ′ − i = x , y ′ − j = y w L x ′ , y ′ , i , j , i ∈ − r , … , r , j ∈ − r , … , r , wherein x, y, x', y', i, j denote array indices, out(x, y) denotes the array of output data values, in(x', y') denotes the array of input data values, r denotes a size of each kernel of the plurality of position dependent kernels w L (x', y', i, j) and W L '(x, y) denotes a normalization factor. In an implementation form the normalization factor W L '(x, y) can be set equal to 1. As will be appreciated, the sum in the equation above extends over every possible position (x',y') of the array of input data values, where x' and y' meet the conditions: x' - i = x and y' - j = y. In this way, overlapping positions of different position dependent kernels are obtained that are summed to generate the final output data value out(x, y).
[0032] In a further implementation form of the first aspect, the array of input data values and the array of output data values are two-dimensional arrays and the neural network layer is configured to generate the array of output data values on the basis of the following equations: out x y = 1 W L ′ x y ∑ x ′ , y ′ : x ′ − i = x , y ′ − j = y in x ′ , y ′ sel x ′ , y ′ , i , j , W L ′ x y = ∑ x ′ , y ′ : x ′ − i = x , y ′ − j = y sel x ′ , y ′ , i , j , i ∈ − r , … , r , j ∈ − r , … , r , sel x y i j = 1 , w L x y i j is max or min weight of all w L x y k l , k ∈ − r , … , r , l ∈ − r , … , r 0 , otherwise wherein x, y, x', y', i,j, k, l denote array indices, out(x, y) denotes the array of output data values, in(x', y') denotes the array of input data values, r denotes a size of each kernel of the plurality of position dependent kernels w L (x, y, i, j), sel(x, y, i, j) denotes a selection function and W L '(x, y) denotes a normalization factor. In an implementation form the normalization factor W L '(x, y) can be set equal to 1.
[0033] In a further implementation form of the first aspect, the array of input data values and the array of output data values are two-dimensional arrays and the neural network layer is configured to generate the array of output data values on the basis of the following equations: out x y = 1 W L ′ x y ∑ x ′ , y ′ : x ′ − i = x , y ′ − j = y in x ′ , y ′ sel x , y , x ′ , y ′ , i , j , W L ′ x y = ∑ x ′ , y ′ : x ′ − i = x , y ′ − j = y sel x , y , x ′ , y ′ , i , j , i ∈ − r , … , r , j ∈ − r , … , r , sel x , y , x ′ , y ′ , i , j = 1 , w L x ′ , y ′ , i , j is maximum weight of all w L x " , y " , k , l , x " , y " : x " − k = x , y " − l = y , k ∈ − r , … , r , l ∈ − r , … , r 0 , otherwise wherein x, y, x', y', x", y", i, j, k, l denote array indices, out(x, y) denotes the array of output data values, in(x', y') denotes the array of input data values, r denotes a size of each kernel of the plurality of position dependent kernels w L (x' , y' , i, j), sel(x, y, x', y', i, j) denotes a selection function and W L '(x, y) denotes a normalization factor. In an implementation form the normalization factor W L '(x, y) can be set equal to 1.
[0034] According to a second aspect the invention relates to a corresponding data processing method comprising the step of generating by a neural network layer of a neural network from an array of input data values an array of output data values based on a plurality of position dependent kernels and a plurality of different input data values of the array of input data values.
[0035] In a further implementation form of the second aspect, the method comprises the further step of generating the plurality of position dependent kernels by an additional neural network layer of the neural network based on an original array of original input values of the neural network, wherein the original array of original input values of the neural network comprises the array of input values or another array of input values associated to the array of input values.
[0036] In a further implementation form of the second aspect, the step of generating the plurality of position dependent kernels comprises generating the plurality of position dependent kernels based on a plurality of position independent kernels and a plurality of position dependent weights.
[0037] In a further implementation form of the second aspect, the step of generating the plurality of position dependent kernels comprises the step of adding, i.e. summing the position independent kernels weighted by the associated position dependent weights.
[0038] In a further implementation form of the second aspect, the plurality of position independent kernels are predetermined or learned and the step of generating the plurality of position dependent weights comprises the step of generating the plurality of position dependent weights by an additional neural network layer or a processing layer of the neural network based on an original array of original input values of the neural network, wherein the original array of original input values of the neural network comprises the array of input values or another array of input values associated to the array of input values.
[0039] In a further implementation form of the second aspect, the array of input data values and the array of output data values are two-dimensional arrays, and the step of generating a kernel of the plurality of position dependent kernels w L (x, y, i, j) is based on the following equation: w L x y i j = ∑ f = 1 N f F f x y ⋅ K f i j , wherein F f (x, y) denotes the plurality of N f position dependent weights (i.e. similarity features) and K f (i, j) denotes the plurality of position independent kernels.
[0040] In a further implementation form of the second aspect, the neural network layer is a deconvolutional network layer or an upscaling network layer.
[0041] In a further implementation form of the second aspect, the array of input data values and the array of output data values are two-dimensional arrays, wherein the neural network layer is a deconvolution network layer and the step of generating the array of output data values comprises generating the array of output data values on the basis of the following equations: out x y c o = 1 W L ′ x y c o ∑ c i = 1 C i ∑ x ′ , y ′ : x ′ − i = x , y ′ − j = y in x ′ , y ′ , c i w L x ′ , y ′ , c o , c i , i , j , W L ′ x y c o = ∑ c i = 1 C i ∑ x ′ , y ′ : x ′ − i = x , y ′ − j = y w L x ′ , y ′ , c o , c i , i , j , i ∈ − r , … r , j ∈ − r , … r , wherein x, y, x', y', i,j denote array indices, out(x, y, c o ) denotes the multi-channel array of output data values, in(x', y', c i ) denotes the array of input data values, r denotes a size of each kernel of the plurality of position dependent multi-channel kernels w L (x', y', c o , c i , i, j) and W L '(x, y, c o ) denotes a normalization factor. In an implementation form the normalization factor W L '(x, y, c 0 ) can be set equal to 1.
[0042] In a further implementation form of the second aspect, the array of input data values and the array of output data values are two-dimensional arrays, wherein the neural network layer is an upscaling network layer and the step of generating the array of output data values comprises generating the array of output data values on the basis of the following equations: out x y = 1 W L ′ x y ∑ x ′ , y ′ : x ′ − i = x , y ′ − j = y in x ′ , y ′ w L x ′ , y ′ , i , j , W L ′ x y = ∑ i = − r r ∑ j = − r r w L x ′ , y ′ , i , j , x ′ , y ′ : x ′ − i = x , y ′ − j = y , i ∈ − r , … , r , j ∈ − r , … , r wherein x, y, x', y', i, j denote array indices, out(x, y) denotes the array of output data values, in(x', y') denotes the array of input data values, r denotes a size of each kernel of the plurality of position dependent kernels w L (x', y', i, j) and W L '(x, y) denotes a normalization factor. In an implementation form the normalization factor W L '(x, y) can be set equal to 1.
[0043] In a further implementation form of the second aspect, the array of input data values and the array of output data values are two-dimensional arrays and the step of generating the array of output data values comprises generating the array of output data values on the basis of the following equations: out x y = 1 W L ′ x y ∑ x ′ , y ′ : x ′ − i = x , y ′ − j = y in x ′ , y ′ sel x ′ , y ′ , i , j , W L ′ x y = ∑ x ′ , y ′ : x ′ − i = x , y ′ − j = y sel x ′ , y ′ , i , j , i ∈ − r , … , r , j ∈ − r , … , r , sel x y i j = 1 , w L x y i j is max or min weight of all w L x y k l , k ∈ − r , … , r , l ∈ − r , … , r 0 , otherwise wherein x, y, x', y', i,j, k, l denote array indices, out(x, y) denotes the array of output data values, in(x', y') denotes the array of input data values, r denotes a size of each kernel of the plurality of position dependent kernels w L (x, y, i, j), sel(x, y, i, j) denotes a selection function and W L '(x, y) denotes a normalization factor. In an implementation form the normalization factor W L '(x, y) can be set equal to 1.
[0044] In a further implementation form of the second aspect, the array of input data values and the array of output data values are two-dimensional arrays and the step of generating the array of output data values comprises generating the array of output data values on the basis of the following equations: out x y = 1 W L ′ x y ∑ x ′ , y ′ : x ′ − i = x , y ′ − j = y in x ′ , y ′ sel x , y , x ′ , y ′ , i , j , W L ′ x y = ∑ x ′ , y ′ : x ′ − i = x , y ′ − j = y sel x , y , x ′ , y ′ , i , j , i ∈ − r , … , r , j ∈ − r , … , r , sel x , y , x ′ , y ′ , i , j = 1 , w L x ′ , y ′ , i , j is maximum weight of all w L x " , y " , k , l , x " , y " : x " − k = x , y " − l = y , k ∈ − r , … , r , l ∈ − r , … , r 0 , otherwise wherein x, y, x', y', x", y", i, j, k, l denote array indices, out(x, y) denotes the array of output data values, in(x', y') denotes the array of input data values, r denotes a size of each kernel of the plurality of position dependent kernels w L (x' , y' , i, j), sel(x, y, x', y', i, j) denotes a selection function and W L '(x, y) denotes a normalization factor. In an implementation form the normalization factor W L '(x, y) can be set equal to 1.
[0045] According to a third aspect the invention relates to a computer program comprising program code for performing the method according to the second aspect, when executed on a processor or a computer.
[0046] The invention can be implemented in hardware and / or software.BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Further embodiments of the invention will be described with respect to the following figures, wherein: Fig. 1 shows a schematic diagram illustrating a data processing apparatus based on a neural network according to an embodiment; Fig. 2 shows a schematic diagram illustrating a neural network provided by a data processing apparatus according to an example not encompassed by the claims but useful for understanding the invention; Fig. 3 shows a schematic diagram illustrating the concept of up-scaling of data implemented in a data processing apparatus according to an embodiment; Fig. 4 shows a schematic diagram illustrating an up-scaling operation provided by a neural network of a data processing apparatus according to an embodiment; Fig. 5 shows a schematic diagram illustrating different aspects of a neural network provided by a data processing apparatus according to an embodiment; Fig. 6 shows a schematic diagram illustrating different aspects of a neural network provided by a data processing apparatus according to an embodiment; Fig. 7 shows a schematic diagram illustrating different processing steps of a data processing apparatus according to an embodiment; Fig. 8 shows a schematic diagram illustrating a neural network provided by a data processing apparatus according to an embodiment; Fig. 9 shows a schematic diagram illustrating different aspects of a neural network provided by a data processing apparatus according to an embodiment; Fig. 10 shows a schematic diagram illustrating different processing steps of a data processing apparatus according to an embodiment; and Fig. 11 shows a flow diagram illustrating a neural network data processing method according to an embodiment.
[0048] In the various figures, identical reference signs will be used for identical or at least functionally equivalent features.DETAILED DESCRIPTION OF EMBODIMENTS
[0049] In the following description, reference is made to the accompanying drawings, which form part of the disclosure, and in which are shown, by way of illustration, specific aspects in which the present invention may be placed. It is understood that other aspects may be utilized and structural or logical changes may be made. The following detailed description, therefore, is not to be taken in a limiting sense, as the scope of the present invention is defined by the appended claims.
[0050] For instance, it is understood that a disclosure in connection with a described method may also hold true for a corresponding device or system configured to perform the method and vice versa. For example, if a specific method step is described, a corresponding device may include a unit to perform the described method step, even if such unit is not explicitly described or illustrated in the figures. Further, it is understood that the features of the various exemplary aspects described herein may be combined with each other, unless specifically noted otherwise.
[0051] Figure 1 shows a schematic diagram illustrating a data processing apparatus 100 according to an embodiment configured to process data on the basis of a neural network. To this end, the data processing apparatus 100 shown in figure 1 comprises a processor 101. In an embodiment, the data processing apparatus 100 can be implemented as a distributed data processing apparatus 100 comprising more than the one processor 101 shown in figure 1.
[0052] The processor 101 of the data processing apparatus 100 is configured to provide a neural network 110. As will be described in more detail further below, the neural network 110 comprises a neural network layer being configured to generate from an array of input data values an array of output data values based on a plurality of position dependent kernels and a plurality of different input data values of the array of input data values. As shown in figure 1, the data processing apparatus 100 can further comprise a memory 103 for storing and / or retrieving the input data values, the output data values and / or the kernels.
[0053] Each kernel comprises a plurality of kernel values (also referred to as kernel weights). For a respective position or element of the array of input data values a respective kernel is applied thereto for generating a respective sub-array of the array of output data values. Generally, the size of the array of input data values is smaller than the size of the array of output data values. A "position dependent kernel" as used herein means a kernel whose kernel values depend on the respective position or element of the array of input data values. In other words, for a first kernel used for a first input data value of the array of input data values the kernel values can differ from the kernel values of a second kernel used for a second input data value of the array of input data values. In a two-dimensional array the position could be a spatial position defined, for instance, by two spatial coordinates x,y. In a one-dimensional array the position could be a temporal position defined, for instance, by a time coordinate t.
[0054] The array of input data values can be one-dimensional (i.e. a vector, e.g. audio or other e.g. temporal sequence), two-dimensional (i.e. a matrix, e.g. an image or other temporal or spatial sequence), or N-dimensional (e.g. any kind of N-dimensional feature array, e.g. provided by a conventional pre-processing or feature extraction and / or by other layers of the neural network 110). The array of input data values can have one or more channels, e.g. for an RGB image one R-channel, one G-channel and one B-channel, or for a black / white image only one grey-scale or intensity channel. The term "channel" can refer to any "feature", e.g. features obtained from conventional pre-processing or feature extraction or from other neural networks or neural network layers of the neural network 110. The array of input data values can comprise, for instance, two-dimensional RGB or grey scale image or video data representing at least a part of an image, or a one-dimensional audio signal. In case the neural network layer 120 is implemented as an intermediate layer of the neural network 110, the array of input data values can be, for instance, an array of similarity features generated by previous layers of the neural network on the basis of an initial, i.e. original array of input data values, e.g. by means of a feature extraction, as will be described in more detail further below.
[0055] As will be described in more detail below, the neural network layer 120 can be implemented as an up-scaling layer 120 configured to process each channel of the array of input data values separately, e.g. for an input array of R-values one (scalar) R-output value is generated. The position dependent kernels may be channel-specific or common for all channels. Moreover, the neural network layer 120 can be implemented as a deconvolution (or deconvolutional) layer configured to "mix" all channels of the array of input data values. For instance, in case the generated array of output data values is an RGB image, i.e. a multi-channel array, every single channel of a multi-channel input data array is used to generate all three channels of the multi-channel array of output data values. The position dependent kernels may be channel-specific, i.e. multi-channel arrays, or common for all channels.
[0056] Figure 2 shows a schematic diagram illustrating elements of the neural network 110 provided by the data processing apparatus 100 according to an example not encompassed by the claims but useful for understanding the invention. In the example shown in figure 2, the neural network layer 120 is implemented as a up-scaling layer 120. In a further example, the neural network layer 120 can be implemented as a deconvolution layer 120 (also referred to as deconvolutional layer 120), as will be described in more detail further below. As indicated in figure 2, in this example the up-scaling layer 120 is configured to generate a two-dimensional array of output data values out(x, y) 121 on the basis of the two-dimensional array of input data values in(x, y) 117 and the plurality of position dependent kernels 118 comprising a plurality of kernel values or kernel weights.
[0057] In an example, the up-scaling layer 120 of the neural network 110 shown in figure 2 is configured to generate the array of output data values out(x, y) 121 on the basis of the array of input data values in(x, y) 117 and the plurality of position dependent kernels 118 comprising the kernel values w L (x, y, i, j) using the following equations: out x y = 1 w L ′ x y ∑ x ′ , y ′ : x ′ - i = x , y ′ − j = y in x ′ , y ′ w L x ′ , y ′ , i , j , w L ′ x y = ∑ x ′ , y ′ : x ′ - i = x , y ′ − j = y w L x ′ , y ′ , i , j , i ∈ − r , … , r , j ∈ − r , … , r , wherein x, y, x', y', i, j denote array indices, out(x, y) denotes the array of output data values 121, in(x', y') denotes the array of input data values 117, r denotes a size of each kernel of the plurality of position dependent kernels w L (x', y', i, j) 118 (in this example, each kernel has (2r+1)*(2r+1) kernel values) and W L '(x, y) denotes a normalization factor and can be optionally set to 1. As will be appreciated, the sum in the equation above extends over every possible position (x', y') of the array of input data values 117, where x' and y' meet the conditions: x' - i = x and y' - j = y. In this way, overlapping positions of different position dependent kernels 118 are obtained that are summed to generate the final output data value out(x, y).
[0058] In other examples, the normalization factor can be omitted, i.e. set to one. For instance, in case the neural network layer 120 is implemented as a deconvolutional network layer the normalization factor can be omitted. For upscaling the normalization factor allows to keep the DC component. This is usually not necessary in the case of the deconvolutional network layer 120.
[0059] As will be appreciated, the above equations for a two-dimensional input array and a kernel having a quadratic shape can be easily adapted to the case of an array of input values 117 having one dimension or more than two dimensions and / or a kernel having a rectangular shape, i.e. different horizontal and vertical dimensions.
[0060] For an example, where the neural network layer 120 is implemented as a deconvolution layer and the array of input data values in(x, y, c i ) 117 is a two-dimensional array of input data values the deconvolutional layer 120 is configured to generate the array of output data values 121 as a multi-channel array of output data values out(x, y, c o ) 117, an array having more than one channel c o . In this case, also the plurality of position dependent kernels 118 will have the corresponding number of channels, wherein each multi-channel position dependent kernel comprises the kernel values w L (x', y', c o , c i , i, j). For instance, the deconvolutional layer 120 could be configued to deconvolve a monochromatic image into an RGB image with higher resolution using a plurality of position dependent kernels 118 having three channels.
[0061] In an example, the deconvolutional layer 120 is configured to generate the multi-channel array of output data values out(x, y, c o ) 121 on the basis of the array of input data values in(x, y, c i ) 117 having one or more channels and the plurality of multi-channel position dependent kernels 118 comprising the kernel values w L (x', y', c o , c i , i, j) using the following equations: out x y c o = 1 w L ′ x y c o ∑ c i = 1 C i ∑ x ′ , y ′ : x ′ - i = x , y ′ − j = y in x ′ , y ′ , c i w L x ′ , y ′ , c o , c i , i , j , W L ′ x y c o = ∑ c i = 1 C i ∑ x ′ , y ′ : x ′ - i = x , y ′ − j = y w L x ′ , y ′ , c o , c i , i , j , i ∈ − r , … , r , j ∈ − r , … , r , wherein x, y, x', y', i, j denote array indices, r denotes a size of each kernel of the plurality of positon dependent kernels 118 and W L '(x, y, c o ) denotes a normalization factor. In other examples, the normalization factor can be omitted, i.e. set to one.
[0062] In an example, the neural network layer 120 is configured to generate the array of output data values 121 with a larger size than the array of input data values 117. In other words, in an example, the neural network 110 is configured to perform an up-step or upscaling operation of the array of input data values 117 on the basis of the plurality of position dependent kernels 118. Figure 3 illustrates an up-step or upscaling operation provided by a neural network 110 of the data processing apparatus 100 according to an embodiment. Using an up-step or upscaling operation allows increasing the receptive field, enables processing the data with a cascade of smaller filters as compared with a single layer with a kernel covering an equal receptive field, and also enables the neural network 110 to better analyse the data by finding more sophisticated relationships among the data.
[0063] In the up-step or upscaling operation illustrated in figure 3 the neural network layer 120 can up-scale the input data produced by a preceding cascade of down-layers for generating an array of output data values having an increased resolution. This upscaling operation can be performed by deconvolving every channel of each spatial position of the array of input data values with position dependent kernels with a stride S greater than 1, producing a data volume of increased resolution. The stride S specifies the spacing between neighboring input spatial positions for which deconvolutions are computed. If the stride S is equal to 1, the deconvolution is performed for each spatial position. If the stride S is an integer greater than 1, deconvolution is performed for every S spatial position, increasing the output resolution by a factor of S for each spatial dimension.
[0064] In the exemplary embodiment shown in figure 3, the neural network layer 120 up-scales every element of the array of input data values 117 into a respective sub-array of the array of output data values 121 with a size of (2r+1)×(2r+1) (defined by the size of the position dependent kernels 118). In this way, the input data values 117 can be up-scaled to the higher resolution array of output data values 121.
[0065] According to an embodiment, the upscaling operation performed by the neural network layer 120 for the exemplary case of two-dimensional input and output arrays 117, 121 comprises multiplying a respective input data value of the array of input data values 117 with the plurality of kernel weights w L (x, y, i, j) of a respective position dependent kernel 118. In case the respective position dependent kernel 118 has an exemplary size of (2r+1)×(2r+1) this operation will generate a sub-array of the array of output data values 121 (which can also be considered as an interpolation area) having also a size of (2r+1)×(2r+1). As will be appreciated, depending on the selected stride S, the interpolation areas of neighboring input data values may overlap. In order to handle such case, according to an embodiment, the values from all overlapping interpolation areas 122 located at the spatial position (x, y) (i.e. overlapping spatial position) can be aggregated and (optionally) normalized by a normalization factor producing the final output data value out(x, y). This operation is illustrated in figure 4 for the exemplary case of having R sub-arrays or interpolation areas at the spatial position (x, y).
[0066] In the example shown in figure 2, the neural network 110 comprises one or more preceding layers 115 preceding the neural network layer 120 and one or more following layers 125 following the neural network layer 120. In an embodiment, the neural network layer 120 could be the first and / or the last data processing layer of the neural network 110, i.e. in an embodiment there could be no preceding layers 115 and / or no following layers 125.
[0067] In an embodiment, the one or more preceding layers 115 can be further neural network layers, such as a convolutional network layer, and / or "conventional" pre-processing layers, such as a feature extraction layer. Likewise, in an embodiment, the one or more following layers 125 can be further neural network layers and / or "conventional" post-processing layers.
[0068] As shown in the example shown in figure 2, one or more of the preceding layers 115 can be configured to provide, i.e. to generate the plurality of position dependent kernels 118. In an embodiment, the one or more layers of the preceding layers 115 can generate the plurality of position dependent kernels 118 on the basis of an original array of original input data values. As indicated in figure 2, in an example, the original array of original input data values can be an array of input data 111 being the original input of the neural network 110. In another embodiment, the one or more preceding layers 115 could be configured to generate just the plurality of position dependent kernels 118 on the basis of the original input data 111 of the neural network 110 and to provide the original input data 111 of the neural network 110 as the array of input data values 117 to the neural network layer 120.
[0069] As indicated in figure 2, in a further examplethe one or more preceding layers 115 of the neural network 110 are configured to generate the plurality of position dependent kernels 118 on the basis of an array of guiding data 113. A more detailed view of the processing steps of the neural network 110 of the data processing apparatus 100 according to such an embodiment is shown in figure 5 for the exemplary case of two-dimensional input and output arrays. The array of guiding data 113 is used by the one or more preceding layers 115 of the neural network 110 to generate the plurality of position dependent kernels w L (x, y) 118 on the basis of the array of guiding data g(x, y) 113. As already described in the context of figure 2, the neural network layer 120 is configured to generate the two-dimensional array of output data values out(x,y) 121 on the basis of the two-dimensional array of input data values in(x, y) 117 and the plurality of position dependent kernels w L (x, y) 118, which, in turn, are based on the array of guiding data g(x, y) 113.
[0070] In an embodiment, the one or more preceding layers 115 of the neural network 110 are neural network layers configured to learn the plurality of position dependent kernels w L (x, y) 118 on the basis of the array of guiding data g(x, y) 113. In another embodiment, the one or more preceding layers 115 of the neural network 110 are pre-processing layers configured to generate the plurality of position dependent kernels w L (x, y) 118 on the basis of the array of guiding data 113 using one or more pre-processing schemes, such as feature extraction.
[0071] In an embodiment, the one or more preceding layers 115 of the neural network 110 are configured to generate the plurality of position dependent kernels w L (x, y) 118 on the basis of the array of guiding data g(x, y) 113 in a way analogous to up-scaling based on bilateral filters, as illustrated in figure 6. In image processing, a common approach to perform data up-scaling is to use bilateral filter weights [M. Elad, "On the origin of bilateral filter and ways to improve it", IEEE Transactions on Image Processing, vol. 11, no. 10, pp. 1141-1151, October 2002] as a sort of guiding information for interpolating the input data. The usage of bilateral filter weights has the advantage of decreasing the influence of input data values on some spatial positions of the interpolation results, while amplifying its influence for others. As illustrated in figure 6, the weights 618 utilized for up-scaling the array of input data values 617 adapt to input data using the guiding image data g 613 which provides additional information to control the up-scaling process. In the up-scaling process, a single input data value of the array of input data values in(x, y) 617 is multiplied by the kernel w 618 of size (2r+1)×(2r+1) creating an interpolated area of output data out(x ± r, y ± r) 521 of size (2r+1)×(2r+1). As will be appreciated, however, the interpolation areas of neighbouring input positions may overlap. In order to handle such cases, values from different overlapping interpolation areas located at the spatial position x, y can be aggregated and normalized by a normalization factor W'(x, y) producing the final output value out(x, y). If the stride S is greater than 1, the spatial resolution of the output data created by the interpolation areas will be increased. Mathematically, this can be expressed in the following way: out x y = 1 W ′ x y ∑ x ′ , y ′ : x ′ - i = x , y ′ − j = y in x ′ , y ′ w x ′ , y ′ , i , j where: W ′ x y = ∑ x ′ , y ′ : x ′ - i = x , y ′ − j = y w x ′ , y ′ , i , j , i ∈ {-r, ...,r}, j ∈ {-r, ... , r}.
[0072] In an embodiment, the bilateral filter weights 618 are defined by the following equation: w x y i j = e − x − i 2 + y − j 2 2 w r e − d g x − i , y − j , g x y 2 w d , wherein d(.,.) denotes a distance function. Thus, the bilateral filter weights 618 can take into account the distance of the value within the kernel from the center of the kernel and, additionally, the similarity of the data values with data in the center of the kernel.
[0073] Figure 7 shows a schematic diagram highlighting the main processing stage 701 of the data processing apparatus 100 according to an embodiment, for instance, the data processing apparatus 100 providing the neural network 110 shown in figure 2. As already described above, in a first processing step 703 the neural network 110 can generate the plurality of position dependent kernels w L (x, y) 118 on the basis of the array of guiding data g(x, y) 113. In a second processing step 705 the neural network 110 can generate the array of output data values out(x, y) 121 on the basis of the array of input data values in(x, y) 117 and the plurality of position dependent kernels w L (x, y, i, j) 118.
[0074] Figure 8 shows a schematic diagram illustrating the neural network 110 provided by the data processing apparatus 100 according to a further embodiment. As will be described in more detail in the following, the main difference to the example shown in figure 2 is that in the embodiment shown in figure 8 the neural network 110 is configured to generate the plurality of position dependent kernels based on a plurality of position independent kernels 119b (shown in figure 9) and a plurality of position dependent weights F f (x, y) 119a (also referred to as similarity features 119a). In an embodiment, the similarity features 119a could indicate higher-level knowledge about the input data, including e.g. semantic segmentation, per-instance object detection, data importance indicators like ROI (Region of Interest) and many others - all learned by the neural network 110 itself or being an additional input to the neural network 110. In an embodiment, the neural network 110 of figure 8 is configured to generate the plurality of position dependent kernels 118 by adding the position independent kernels 119b weighted by the associated position dependent weights F f (x, y) 119a.
[0075] In an embodiment, the plurality of position independent kernels 119b can be predetermined or learned by the neural network 110. As illustrated in figure 8, also in this embodiment the neural network 110 can comprise one or more preceding layers 115, which precede the neural network layer 120 and which can be implemented as an additional neural network layer or a pre-processing layer. In an embodiment, one or more layers of the preceding layers 115 are configured to generate the plurality of position dependent weights F f (x, y) 119a on the basis of an original array of original input data values. The original array of original input data values of the neural network 110 can comprise the array of input data values 117 to be processed by the neural network layer 120 or another array of input data values 111 associated to the array of input data values 117, for instance, the initial array of input data 111.
[0076] In the exemplary embodiment shown in figure 8, the array of input data values in(x, y) 117 and the array of output data values out(x, y) 121 are two-dimensional arrays and the neural network layer 120 is configured to generate a respective kernel of the plurality of position dependent kernels w L (x, y, i, j) 118 on the basis of the following equation: w L x y i j = ∑ f = 1 N f F f x y ⋅ K f i j , wherein F f (x, y) denotes the set of N f position dependent weights (or similarity features) 119a and K f (i, j) denotes the plurality of position independent kernels 119b, as also illustrated in figure 9.
[0077] Figure 10 shows a schematic diagram highlighting the main processing stage 1001 implemented in the data processing apparatus 100 according to an embodiment, for instance, the data processing apparatus 100 providing the neural network 100 illustrated in figures 8 and 9. As already described above, in a first processing step 1003 the neural network 110 can generate the plurality of position dependent weights or similarity features F f (x, y) 119a on the basis of the array of guiding data g(x, y) 113. In a second processing step 1005 the neural network 110 can generate the plurality of position dependent kernels w L (x, y, i, j) 118 on the basis of the plurality of position dependent weights or similarity features F f (x, y) 119a and the plurality of position independent kernels K f (i, j) 119b. In a further step (not shown in figure 10, but similar to the processing step 705 shown in figure 7) the neural network layer 120 can generate the array of output data values out(x, y) 121 on the basis of the array of input data values in(x, y) 117 and the plurality of position dependent kernels w L (x, y, i, j) 118.
[0078] In a further embodiment, the neural network layer 120 is configured to process the array of input data values 117 on the basis of the plurality of position dependent kernels 118 using an "inverse" maximum or minimum pooling scheme. More specifically, according to such an embodiment, the array of input data values 117 and the array of output data values 121 are two-dimensional arrays and the neural network layer 120 is configured to generate the array of output data values 121 on the basis of the following equations: out x y = 1 W L ′ x y ∑ x ′ , y ′ : x ′ − i = x , y ′ − j = y in x ′ , y ′ sel x ′ , y ′ , i , j , W L ′ x y = ∑ x ′ , y ′ : x ′ − i = x , y ′ − j = y sel x ′ , y ′ , i , j , i ∈ − r , … , r , j ∈ − r , … , r , sel x y i j = 1 , w L x y i j is max or min weight of all w L x y k l , k ∈ − r , … , r , l ∈ − r , … , r 0 , otherwise wherein x, y, x', y', i,j, k, l denote array indices, out(x, y) denotes the array of output data values 121, in(x', y') denotes the array of input data values 117, r denotes a size of each kernel of the plurality of position dependent kernels w L (x, y, i, j) 118, sel(x, y, i, j) denotes a selection function and W L '(x, y) denotes a normalization factor. In an implementation form the normalization factor W L '(x, y) can be set equal to 1.
[0079] In this embodiment, the neural network layer 120 can be considered to adaptively guide data from the array of input data values 117 to a spatial position of a sub-array of the array of output data values 121 (i.e. the interpolated area) based on the individual position dependent kernel values 118. In this way a sort of more intelligent data un-pooling can be performed. In an embodiment, the input data value corresponding to the spatial position (x, y) is copied to the position (x - i max / min , y - j max / min ) of the sub-array of output data values (i.e. the interpolated area) of size (2r+1)×(2r+1), where (i max / min , j max / min ) are the indices of the individual kernel values with the largest (max) or slowest (min) value among all individual kernel values. As can be taken from the equations above, in this embodiment, other values can be set to zero or, in an alternative embodiment, remain unset. Additionally, an aggregation of overlapping sub-arrays, i.e. interpolated areas can be performed, as in the embodiments described above.
[0080] In another embodiment, the array of input data values 117 and the array of output data values 121 are two-dimensional arrays and the neural network layer 120 is configured to generate the array of output data values 121 on the basis of the following equations: out x y = 1 W L ′ x y ∑ x ′ , y ′ : x ′ − i = x , y ′ − j = y in x ′ , y ′ sel x , y , x ′ , y ′ , i , j , W L ′ x y = ∑ x ′ , y ′ : x ′ − i = x , y ′ − j = y sel x , y , x ′ , y ′ , i , j , i ∈ − r , … , r , j ∈ − r , … , r , sel x , y , x ′ , y ′ , i , j = 1 , w L x ′ , y ′ , i , j is maximum weight of all w L x " , y " , k , l , x " , y " : x " − k = x , y " − l = y , k ∈ − r , … , r , l ∈ − r , … , r 0 , otherwise wherein x, y, x', y', x", y", i, j, k, l denote array indices, out(x, y) denotes the array of output data values 121, in(x', y') denotes the array of input data values 117, r denotes a size of each kernel of the plurality of position dependent kernels w L (x' , y' , i, j) 118, sel(x, y, x', y', i, j) denotes a selection function and W L '(x, y) denotes a normalization factor. In an implementation form the normalization factor W L '(x, y) can be set equal to 1.
[0081] In this embodiment, the neural network layer 120 can be considered to adaptively select output data out(x, y) from input data guided into position (x, y) without performing a weighted average, but selecting as the output data value out(x, y) the input data value in(x', y') of the array of input data values 117 which corresponds to the maximum or minimum kernel value w L (x' , y' , i, j). As a result, the output is computed as the input data value which would originally contribute the most (or in the alternative embodiment the least) to the weighted average.
[0082] Figure 11 shows a flow diagram illustrating a data processing method 1100 based on a neural network 110 according to an embodiment. The data processing method 1100 can be performed by the data processing apparatus 100 shown in figure 1 and its different embodiments described above. The data processing method 1100 comprises the step 1101 of generating by the neural network layer 120 of the neural network 110 from the array of input data values 117 the array of output data values 121 based on a plurality of position dependent kernels 118 and a plurality of input data values of the array of input data values 117. As will be appreciated, further embodiments of the data processing method 1100 result directly from the embodiments of the corresponding data processing apparatus 100 described above. Embodiments of the data processing methods may be implemented and / or performed by one or more processors as described above.
[0083] In the following some further details about various aspects and embodiments (aggregation network layer, convolution network layer, correlation network layer and normalization) are provided.Upscaling
[0084] In embodiments the proposed guided aggregation can be applied for feature map up-scaling (spatial resolution increase). Input values which are features of the feature map are up-scaled one-by-one forming overlapping output sub-arrays of values which are than aggregated and optionally normalized to form output data array. Due to additional guiding information in form of position dependent kernels, the up-scaling process for each input value can be performed in a controlled way, enabling addition of higher resolution details, e.g. object or region borders, that was originally not present in the input low-resolution representation. Here, guiding data represents information about object or region borders in higher resolution, and can be obtained by e.g. color-based segmentation, semantic segmentation using preceding neural network layers or an edge map of a texture image corresponding to processed feature map.Deconvolution
[0085] In embodiments the proposed guided deconvolution can be applied for switchable feature extraction or mixing. Input values which are features of the feature map are deconvolved with adaptable filters which are formed from the input guiding data in form of position dependent kernels. This way, each selected area of the input feature map can be processed with filters especially adapted for that area producing and mixing only features desired for these regions. Here, guiding data in form of similarity features represents information about object / region borders, obtained by e.g. color-based segmentation, semantic segmentation using preceding neural network layers, an edge map of a texture image corresponding to processed feature map or a ROI (region of interest) binary map.Normalization
[0086] In general, normalization is advantageous if the output values obtained for different spatial positions are going to be compared to each other per-value, without any intermediate step. As a result, preservation of the mean (DC) component is beneficial. If such comparison is not performed, normalization is not necessary but increases complexity. Additionally, one can omit normalization in order to simplify the computations and compute only an approximate result.
[0087] While a particular feature or aspect of the disclosure may have been disclosed with respect to only one of several implementations or embodiments, such feature or aspect may be combined with one or more other features or aspects of the other implementations or embodiments as may be desired and advantageous for any given or particular application. Furthermore, to the extent that the terms "include", "have", "with", or other variants thereof are used in either the detailed description or the claims, such terms are intended to be inclusive in a manner similar to the term "comprise". Also, the terms "exemplary", "for example" and "e.g." are merely meant as an example, rather than the best or optimal. The terms "coupled" and "connected", along with derivatives may have been used. It should be understood that these terms may have been used to indicate that two elements cooperate or interact with each other regardless whether they are in direct physical or electrical contact, or they are not in direct contact with each other.
[0088] Although specific aspects have been illustrated and described herein, it will be appreciated by those of ordinary skill in the art that a variety of alternate implementations may be substituted for the specific aspects shown and described.
[0089] Although the elements in the following claims are recited in a particular sequence with corresponding labeling, unless the claim recitations otherwise imply a particular sequence for implementing some or all of those elements, those elements are not necessarily intended to be limited to being implemented in that particular sequence.
[0090] Many alternatives, modifications, and variations will be apparent to those skilled in the art in light of the above teachings. Of course, those skilled in the art readily recognize that there are numerous applications of the invention beyond those described herein. While the present invention has been described with reference to one or more particular embodiments, those skilled in the art recognize that many changes may be made thereto.
Claims
1. A data processing apparatus (100) comprising: a processor (101) configured to provide a neural network (110), wherein the neural network (110) comprises a neural network layer (120) being configured to generate from an array of input data values (117) an array of output data values (121) based on a plurality of position dependent kernels (118) and a plurality of input data values of the array of input data values (117), wherein a respective position dependent kernel is a kernel whose kernel values depend on a respective position of the array of input data values; and the array of input data values represents a part or a whole of an image, wherein the neural network (110) comprises one or more preceding neural network layers (115) configured to generate the plurality of position dependent kernels (118) based on an array of guiding data (113) representative of information about object or region borders; characterized in that: the data processing apparatus (100) is for creating a higher-resolution image of an input image; the neural network (110) is configured to generate the plurality of position dependent kernels (118) based on a plurality of position independent kernels (119b) and a plurality of position dependent weights (119a); and the neural network (110) is configured to generate the plurality of position dependent weights (119a) based on the array of guiding data (113).
2. The data processing apparatus (100) of claim 1, wherein the neural network (110) comprises an additional neural network layer (115) configured to generate the plurality of position dependent kernels (118) based on an original array of original input values (111, 117) of the neural network (110), wherein the original array of original input values (111, 117) of the neural network (110) comprises the array of input values (117) or another array of input values (111) associated to the array of input data values (117).
3. The data processing apparatus (100) of claim 1 or 2, wherein the neural network (110) is configured to generate a kernel of the plurality of position dependent kernels (118) by adding the position independent kernels (119b) weighted by the associated position dependent weights (119a).
4. The data processing apparatus (100) of any one of claim 1 to 3, wherein the plurality of position independent kernels (119b) are predetermined or learned and wherein the neural network (110) comprises an additional neural network layer (115) or processing layer (115) configured to generate the plurality of position dependent weights (119a) based on an original array of original input data values (111, 117) of the neural network (110), wherein the original array of original input data values (111, 117) of the neural network (110) comprises the array of input data values (117) or another array of input data values (111) associated to the array of input data values (117).
5. The data processing apparatus (100) of any one of claims 1 to 4, wherein the array of input data values (117) and the array of output data values (121) are two-dimensional arrays and the neural network layer (120) is configured to generate a kernel WL(x, y, i, j) (118) of the plurality of position dependent kernels (119b) on the basis of the following equation: w L x y i j = ∑ f = 1 N f F f x y ⋅ K f i j , wherein Ff(x, y) denotes the plurality of Nf position dependent weights (119a) and Kf (i, j) denotes the plurality of position independent kernels (119b).
6. The data processing apparatus (100) of any one of the preceding claims, wherein the neural network layer (120) is a deconvolutional network layer or an upscaling network layer.
7. The data processing apparatus (100) of any one of the preceding claims, wherein the array of input data values (117) and the array of output data values (121) are two-dimensional arrays and wherein the neural network layer (120) is a deconvolution network layer configured to generate the array of output data values (121) on the basis of the following equations: out x y c o = 1 W L ′ x y c o ∑ c i = 1 C i ∑ x ′ , y ′ : x ′ − i = x , y ′ − j = y in x ′ , y ′ , c i w L x ′ , y ′ , c o , c i , i , j , W L ′ x y c o = ∑ c i = 1 C i ∑ x ′ , y ′ : x ′ − i = x , y ′ − j = y w L x ′ , y ′ , c o , c i , i , j or W L ′ x y c o = 1 , i ∈ − r , … , r , j ∈ − r , … , r , wherein x, y, x' , y' , i, j denote array indices, out(x, y, co) denotes the array of output data values (121) having one or more channels, in(x' ,y' , ci) denotes the array of input data values (117), r denotes a size of each kernel of the plurality of position dependent kernels wL(x' , y' , co, ci, i, j) (118) having one or more channels and WL'(x, y, co) denotes a normalization factor.
8. The data processing apparatus (100) of any one of claims 1 to 6, wherein the array of input data values (117) and the array of output data values (121) are two-dimensional arrays and wherein the neural network layer (120) is an upscaling network layer configured to generate the array of output data values (121) on the basis of the following equations: out x y = 1 W L ′ x y ∑ x ′ , y ′ : x ′ − i = x , y ′ − j = y in x ′ , y ′ w L x ′ , y ′ , i , j , W L ′ x y = ∑ x ′ , y ′ : x ′ − i = x , y ′ − j = y w L x ′ , y ′ , i , j or W L ′ x y = 1 , i ∈ − r , … , r , j ∈ − r , … , r , wherein x, y, x' , y' , i, j denote array indices, out(x, y) denotes the array of output data values (121), in(x' ,y' ) denotes the array of input data values (117), r denotes a size of each kernel of the plurality of position dependent kernels wL(x' , y' , i, j) and WL'(x, y) denotes a normalization factor.
9. The data processing apparatus (100) of any one of claims 1 to 6, wherein the neural network layer (120) is configured to generate the array of output data values (121) on the basis of overlapping interpolation areas (122), wherein each overlapping interpolation area is generated on the basis of the input data value of the array of input data values (117) and the respective kernel of the plurality of position dependent kernels (118) by assigning to the overlapping interpolation area (122) the input data value of the array of input data values (117) at the position corresponding to the position of the maximum or minimum value of the respective kernel of the plurality of position dependent kernels (118) and zero otherwise, wherein the overlapping interpolation areas (122) are located at an overlapping spatial position.
10. The data processing apparatus (100) of any one of claims 1 to 6, wherein the array of input data values (117) and the array of output data values (121) are two-dimensional arrays and the neural network layer (120) is configured to generate the array of output data values (121) on the basis of the following equations: out x y = 1 W L ′ x y ∑ x ′ , y ′ : x ′ − i = x , y ′ − j = y in x ′ , y ′ sel x ′ , y ′ , i , j , W L ′ x y = ∑ x ′ , y ′ : x ′ − i = x , y ′ − j = y sel x ′ , y ′ , i , j or W L ′ x y = 1 , i ∈ − r , … , r , j ∈ − r , … , r , sel x y i j = 1 , w L x y i j is max or min weight of all w L x y k l , k ∈ − r , … , r , l ∈ − r , … , r 0 , otherwise wherein x, y, x', y', i,j, k, l denote array indices, out(x, y) denotes the array of output data values (121), in(x', y') denotes the array of input data values (117), r denotes a size of each kernel of the plurality of position dependent kernels wL(x, y, i, j) (118), sel(x, y, i, j) denotes a selection function and WL'(x, y) denotes a normalization factor.
11. The data processing apparatus (100) of any one of claims 1 to 6, wherein the neural network layer (120) is configured to generate the array of output data values (121), wherein each value of the array of output data values (121) at an overlapping spatial position is generated on the basis of the input data values of the array of input data values (117) for which values of the respective kernels of the plurality of position dependent kernels (118) at the overlapping spatial position are the maximum or minimum value among all the values of the respective kernels of the plurality of position dependent kernels (118) at the overlapping spatial position.
12. The data processing apparatus (100) of any one of claims 1 to 6, wherein the array of input data values (117) and the array of output data values (121) are two-dimensional arrays and the neural network layer (120) is configured to generate the array of output data values (121) on the basis of the following equations: out x y = 1 W L ′ x y ∑ x ′ , y ′ : x ′ − i = x , y ′ − j = y in x ′ , y ′ sel x , y , x ′ , y ′ , i , j , W L ′ x y = ∑ x ′ , y ′ : x ′ − i = x , y ′ − j = y sel x , y , x ′ , y ′ , i , j or W L ′ x y = 1 , i ∈ − r , … , r , j ∈ − r , … , r , sel x , y , x ′ , y ′ , i , j = 1 , w L x ′ , y ′ , i , j is maximum weight of all w L x " , y " , k , l , x " , y " : x " − k = x , y " − l = y , k ∈ − r , … , r , l ∈ − r , … , r 0 , otherwise wherein x, y, x', y', x", y", i, j, k, l denote array indices, out(x, y) denotes the array of output data values 121, in(x', y') denotes the array of input data values 117, r denotes a size of each kernel of the plurality of position dependent kernels wL(x' , y' , i, j) 118, sel(x, y, x', y', i, j) denotes a selection function and WL'(x, y) denotes a normalization factor.
13. A computer-implemented data processing method (1100) comprising: generating (1101) by a neural network layer (120) of a neural network (110) from an array of input data values (117) an array of output data values (121) based on a plurality of position dependent kernels (118) and a plurality of input data values of the array of input data values (117), wherein a respective position dependent kernel is a kernel whose kernel values depend on a respective position of the array of input data values; and the array of input data values represents a part or a whole of an image, wherein the method comprises: generating, by one or more preceding neural network layers (115) of the neural network (110), the plurality of position dependent kernels (118) based on an array of guiding data (113) representative of information about object or region borders; characterized in that: the computer-implemented data processing method (1100) is for creating a higher-resolution image of an input image; and the method further comprises: generating, by the neural network (110), the plurality of position dependent kernels (118) based on a plurality of position independent kernels (119b) and a plurality of position dependent weights (119a); and generating, by the neural network (110), the plurality of position dependent weights (119a) based on the array of guiding data (113).
14. A computer program comprising program code for performing the method (1100) of claim 13, when executed on a computer and / or a processor.