Image reconstruction and processing

FR3159461A1Active Publication Date: 2025-08-22COMMISSARIAT A LENERGIE ATOMIQUE ET AUX ENERGIES ALTERNATIVES
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
FR2024001481
Authority / Receiving Office
FR · FR
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-15
Publication Date
2025-08-22
Estimated Expiration
2044-02-15

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Image reconstruction and processing The present description relates to a method for training a neural network (202) comprising: - generating a modified image, by a modified data generator, on the basis of a first image, the modified image comprising at least one aberrant pixel value with respect to the first image; - providing the modified image to the network; - generating, by the network, a corrected image; - providing the corrected image, by the network, and providing an indication of the position of the at least one modified pixel, by the generator, to a calculation circuit (210) - generating an error value, on the basis of the application of a cost function, by the calculation circuit, taking as input the first image, the indication and the corrected image; - correcting parameters associated with the network by backpropagation of the error in the network. Figure for abstract: Fig. 2
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Image reconstruction and processing Technical field

[0001] The present description relates generally to methods and devices for the reconstruction and processing of images and in particular to methods and devices for the demosaicing of images. Prior art

[0002] Raw images acquired by image sensors, such as for example sensors dedicated to visible and / or infrared imaging, exhibit alterations due to sensor malfunctions. In particular, image sensors implementing canonical imaging systems, using unit pixel matrixing methods, may exhibit, over their lifetime, malfunctions on isolated pixels, or on columns of pixels and / or loss of captured data for isolated pixels or columns of pixels.

[0003] It is desirable to improve the processing and reconstruction of images acquired by sensors comprising dysfunctional pixels. Summary of the invention

[0004] One embodiment provides a method of training a neural network configured to perform image processing operations, the method comprising: - generating a modified image, by a modified data generator, based on a first image, the modified image comprising at least one pixel value that is an outlier compared to the first image; - providing the modified image to the network; - the generation, by the network, of a corrected image; - the supply of the corrected image, by the network, and the supply of an indication of the position of the at least one modified pixel, by the generator, to a calculation circuit - the generation of an error value, based on the application of a cost function, by the calculation circuit, taking as input the first image, the indication and the corrected image; - correction of parameters associated with the network by backpropagation of the error in the network.

[0005] According to one embodiment, the generation of the corrected image by the network comprises: - the generation, by a first sub-module, of an intermediate data value on the basis of the modified image, the intermediate data being at least a partial reconstruction of the first image; - the generation, by a second sub-module of the network, of attention data on the basis of one of the intermediate data and / or the modified image, the attention data comprising an estimate of the location of the at least one modified pixel; - the generation of the corrected image, by the first sub-circuit or by a third sub-circuit, on the basis of the intermediate data and the indication data.

[0006] According to one embodiment, the generation of the attention data comprises the execution of a convolutional neural network configured for image segmentation, on the basis of the intermediate data and / or the corrected image.

[0007] According to one embodiment, the convolutional neural network is a U-net type network or a variant of a U-net type network.

[0008] According to one embodiment, the generation of the attention data further comprises a normalization of the NL1 norm type, based on the data generated by the convolutional neural network.

[0009] According to one embodiment, the generation of the corrected image comprises a multiplexing operation and / or a multiplication operation, for example a point-to-point multiplication, on the basis of the attention data and the intermediate data.

[0010] According to one embodiment, the generation of an error value comprises the application of a focused cost function taking, as input data, the modified image, the first image and the indication of location of the at least one modified pixel.

[0011] According to one embodiment, generating an error value further comprises the application: - a cost function based on an average calculation taking, as input data, the modified image and the first image; and / or - a regularization function taking the modified image as input data.

[0012] According to one embodiment, the first image is included in a database.

[0013] According to one embodiment, the generation of the modified image comprises, for each pixel: - determining whether the pixel is to be modified; and - if the pixel is to be modified, the replacement of the value associated with the pixel by an aberrant value.

[0014] According to one embodiment, the generation of the modified image, by the modified data generator, is further carried out on the basis of a matrixing pattern.

[0015] According to one embodiment, the model of the modified data generator is a neural network model previously trained for the generation of modified images, and configured for the implementation of so-called “style transfer” techniques.

[0016] One embodiment provides an image processing method comprising: - the capture of an image scene, by an imager of an image processing device, the captured image being an image matrixed according to a matrixing pattern - providing the matrixed image to a neural network of the processing device trained according to the above training method; and - the generation of a corrected image, by the network, on the basis of the matrixed image.

[0017] According to one embodiment, the generation of the corrected image by the network comprises: - providing the matrixed image to a first sub-module of the network configured to generate intermediate data by performing a first dematrixing of the matrixed image; - providing the intermediate data to a second sub-module of the network configured to generate attention data based on the intermediate data, the attention data comprising estimates of location of malfunctioning pixels of the imager; and - the generation of the corrected image, by a third sub-module, on the basis of the intermediate data and the attention data.

[0018] According to one embodiment, the matrix pattern corresponds to a matrix pattern used when generating a modified image during training of the network.

[0019] One embodiment provides an image processing device comprising: - an imager configured to capture image scenes, according to a matrix pattern; - a neural network trained according to the above training method, configured to generate a corrected image based on the matrixed image.

[0020] According to one embodiment, the imager comprises one or more dysfunctional pixels. Brief description of the drawings

[0021] These characteristics and advantages, as well as others, will be explained in detail in the following description of particular embodiments given without limitation in relation to the attached figures among which:

[0022] [Fig.l] represents examples of matrixing patterns implemented by image sensors;

[0023] [Fig.2] is a block diagram illustrating a learning architecture of a neural network, according to an embodiment of the present description;

[0024] [Fig.3] is a block diagram illustrating an architecture of the network configured to perform inference operations, according to an embodiment of the present disclosure;

[0025] [Fig.4] is an example of architecture of a processing network, according to an embodiment of the present description;

[0026] [Fig.5] is an example of architecture of an attention module, according to an embodiment of the present description; and

[0027] [Fig.6] is an example of architecture of a processing network, according to an embodiment of the present description. Description of the embodiments

[0028] The same elements have been designated by the same references in the different figures. In particular, the structural and / or functional elements common to the different embodiments may have the same references and may have identical structural, dimensional and material properties.

[0029] For the sake of clarity, only the steps and elements useful for understanding the described embodiments have been shown and are detailed. In particular, the operation and implementation of different types of neural network layers, such as dense layers, convolution layers, etc., are not described in detail and are known to those skilled in the art.

[0030] Unless otherwise specified, when referring to two elements connected to each other, this means directly connected without intermediate elements other than conductors, and when referring to two elements connected (in English "coupled") to each other, this means that these two elements can be connected or be connected by means of one or more other elements.

[0031] In the following description, when reference is made to absolute position qualifiers, such as the terms "front", "back", "top", "bottom", "left", "right", etc., or relative position qualifiers, such as the terms "above", "below", "upper", "lower", etc., or to orientation qualifiers, such as the terms "horizontal", "vertical", etc., reference is made unless otherwise specified to the orientation of the figures.

[0032] Unless otherwise specified, the expressions "about", "approximately", "substantially", and "of the order of" mean to within 10%, preferably to within 5%.

[0033] [Fig. 1] represents examples of matrix patterns implemented by image sensors, or imagers.

[0034] A matrix pattern is a pattern comprising several pixels and acting as a spectral filter replicated, in mosaic, on the surface of an imager. The imager then measures, for each pixel of the mosaic, a value associated with a channel defined by the matrix pattern. The matrix pattern indicates, for each pixel, which unique channel among, for example, color and / or infrared channels is measured. For example, each pixel has a spectral signature determining an interval of wavelengths comprising the spectral response recorded by the pixel.

[0035] In the example of Figure 1, a matrix pattern 100 is a Bayer type pattern. This pattern is 2x2 pixels in size and allows the measurement of the wavelengths associated with the green color (G) for the pixels at the top left and bottom right of the pattern. The Bayer pattern also allows the measurement of the wavelengths associated with the red color (R) of the pixel at the top right as well as the measurement of the wavelengths associated with the blue color (B) at the bottom left of the pattern.

[0036] A 102 matrix pattern is a QuadBayer type pattern. This pattern is 4x4 pixels in size and allows the measurement of the wavelengths associated with the green color (G) for the four pixels at the top left and bottom right of the pattern. The 102 pattern also allows the measurement of the wavelengths associated with the red color (R) for the pixels at the top right as well as the measurement of the wavelengths associated with the blue color (B) for the four pixels at the bottom left of the pattern.

[0037] A 104 pattern is a 3-cell Bayer pattern. This pattern is 6x6 pixels in size and has the same structure as the 100 and 102 patterns except that the red, green, and blue color channels are measured in groups of 3 x 3 pixels.

[0038] A 108 pattern is a 2x2 RGBIR type pattern. This pattern is 2x2 pixels in size and allows the measurement of the wavelengths associated with the green color for the top left pixel of the pattern. The 108 pattern also allows the measurement of the wavelengths associated with the blue color of the top right pixel as well as the measurement of the wavelengths associated with the red color at the bottom left of the pattern. The 108 pattern also allows the measurement of the wavelengths associated with the infrared (IR) of the bottom right pixel.

[0039] A pattern 110 is a 4x4 RGBIR type pattern. This pattern is 4x4 pixels in size and allows the measurement of the wavelengths associated with the green color of the first pixel from the left on the first line, the third pixel from the left on the third line, the first and second pixels from the left on the second line and the third and fourth pixels from the left on the fourth line. The pattern 110 further allows the measurement of the wavelengths associated with the infrared of the second and fourth pixels from the left of the first and third lines.The pattern 110 further allows the measurement of the wavelengths associated with the blue color of the first pixel starting from the left of the third line and of the first and second pixels starting from the left of the fourth line, as well as the measurement of the wavelengths associated with the red color of the third pixel starting from the left of the first line and of the third and fourth pixels starting from the left of the second line.

[0040] These examples of matrix patterns are given by way of example and are of course not limiting. Other matrix patterns, of different size and arrangement, are, of course, conceivable. In other examples, the replicated matrix pattern on the surface of the imager includes a pixel configured to measure a distance between one or more targets in the captured scene and the imager. Such imagers are color image sensors that also capture so-called depth information; these are called monolithic RGBZ imagers.

[0041] A raw image captured by the imager comprises, for example, for each pixel, a single measured value. Demosaicing then consists of estimating, for each pixel, the value of the unmeasured channels using, for example, the spectral correlation between the measured channels as well as the spatial correlation between the value of the same channel for two adjacent pixels.

[0042] However, it happens that some pixels are dysfunctional. For example, one or more isolated pixels become saturated, or unusable. This is called a dead pixel. For example, the so-called Dark signal, or dark current, representing a type of noise in the absence of luminous flux measured for isolated pixels, deviates from the nominal statistics and the measurements associated with these pixels are then aberrant. Other sources of noise can cause aberrant behavior on a pixel or on a group of pixels. For example, some isolated pixels are subject to telegraphic noise (in English "Random Telegraphy Signal" - RTS). For example, following a malfunction of a reading circuit of the imager, columns of pixels are missing, or clamped.In other examples, a digital defect, such as for example a loss of synchronization in the measurements, or a bit flip, results in a loss of data associated with isolated pixels and / or columns and / or rows of pixels and / or one or more pixel frames. In the following, a bad pixel will be called any malfunctioning pixel, whether isolated or in a dysfunctional column, row, or frame. The measurements associated with bad pixels are then aberrant measurements, not approaching reality.

[0043] In addition, bad pixels will, for example, influence the estimation of adjacent pixels during demosaicing. For example, the final image, obtained following demosaicing that does not correct bad pixels, will, for example, have black spots larger than the bad pixel or group of bad pixels.

[0044] However, aberrant measurements associated with bad pixels are relatively infrequent. Indeed, the above-mentioned malfunctions generally have a low occurrence, of the order of 1 / 1000 or less. However, in imagers comprising a very large number of pixels, these malfunctions can be significant with respect to the visual rendering and to the perception of the image. For example, these malfunctions impair the rendering of the image in a much more visible manner than disorders due to homoscedastic noise whose distribution is centered around zero.

[0045] When a deep neural network is trained to demosaice raw images, aberrant measurements, associated with bad pixels, are drowned out by other noises, due to their rare occurrence. Moreover, conventional cost functions, based mainly on averaging calculations and applying canonically regardless of the pixel position, fail to identify and correct this type of aberrant errors.

[0046] [Fig. 2] is a block diagram illustrating a learning architecture of a deep neural network, according to an embodiment of the present description. In particular, the learning of the neural network described in relation to [Fig. 2] allows the consideration of aberrant errors, for example due to bad pixels, in the image reconstruction.

[0047] According to one embodiment, an augmentation module 200 (DATA AUGM) allows for example the execution of a data augmentation model. For example, the module 200 is configured to receive, as input data, an image Y_true corresponding to a ground truth, that is to say to an image not including an error resulting from any processing. For example, each pixel of the image Y_true comprises the ground values ​​of several channels. The module 200 is further configured to generate a modified image Y_bad, on the basis of the image Y_true. The modified image Y_bad then corresponds to the image Y_true, in a matrixed format, to which one or more aberrant errors are added.

[0048] By way of example, the module 200 is configured to model noises of the pixel reading and column reading type, photonic noises, fixed spatial noises, for example structured, telegraphic type noises, and dead or saturated pixels.

[0049] By way of example, the module 200 is configured to add to each pixel, with a probability at most equal to 1 / 1000, an aberrant value. In other words, the module 200 performs, pixel by pixel, a Bernoulli test, of parameter p, where p is at most equal to 1 / 1000. When the Bernoulli test succeeds, the module 200 is configured to modify the value of the associated pixel towards an aberrant value. When the test fails, in the majority of cases, the value of the pixel remains identical to that for the image Y_true. In another example, when the test fails, one or more noises are added to the value of the pixel. By way of example, the aberrant value is a constant value. In another example, the aberrant value comes, for each pixel concerned, from a random draw. The aberrant value is for example added to the value measured for said pixel. In another example, the outlier is multiplied by the measured value for said pixel.In yet another example, the outlier is substituted for the measured value for said pixel. In another example, for each pixel concerned, several outliers are for example determined, randomly or not, and the measured value of the pixel is replaced by a combination of at . less an addition of an outlier with the measured value and / or a multiplication of an outlier with the measured value.

[0050] For example, the generation of the outlier is carried out according to a highly peaked probability distribution, i.e. having a high moment of order 4, such as, for example, a Laplace law, a Cauchy law, etc.

[0051] The methods and circuits for generating such a modified image are known to those skilled in the art and the given example of generation is of course not limiting.

[0052] The module 200 is further configured to generate a Y_aug data item, for example taking the form of a tensor, comprising indications of the positions of the modified pixels. By way of example, the Y_aug data item further comprises an indication of the outlier values ​​that have been assigned.

[0053] By way of example, the terrain images provided to the module 200 are included in a pre-recorded database, for example in a non-volatile memory, such as a server, etc.

[0054] The module 200 is for example configured to provide the modified image Y_bad to a neural network 202 (NEURAL NETWORK) in order to train it. For example, the network 202 is configured to perform regression tasks in order to reconstruct and dematrix a matrixed image provided to it while taking into account possible bad pixels. For example, parameters of the network, such as weights etc., are initially set to initial values.

[0055] According to one embodiment, the matrixing pattern used, by the module 200, to matrix the image Y_true corresponds to the matrixing pattern which will be used by an image sensor of a device comprising the trained network 202.

[0056] In another embodiment, the module 200 comprises a neural network previously trained to generate images, such as for example matrixed images. For example, the generation of the images is carried out using so-called “style transfer” techniques, such as, for example, the techniques described in the publication “Unsupervised Image-to-Image Translation: A review” by Hoyer, H. et al. and published in Sensors in 2022 and / or in the popular article on cycle GANs “A Gentle Introduction to CycleGAN for Image Translation” written by Brownlee, J. on August 17, 2017.

[0057] The network 202 comprises, for example, a processing sub-module 204 (FIRST DEMOSAICING MOD) configured to generate an intermediate data item Y_tmp, based on the modified image Y_bad. The intermediate data item Y_tmp corresponds, for example, to a first demosaicing and reconstruction of the image. For example, the module 204 is a convolutional network integrating several neural layers.

[0058] The network 202 further comprises, for example, an attention sub-module 206 (ATTENTION MOD) configured to receive the intermediate data Y_tmp, or a sub-part thereof, provided by the module 204 and / or the modified image Y_bad provided by the augmentation module 200. The sub-module 206 is further configured to generate attention data Y_pwa having, for example, the form of a tensor. The data Y_pwa is generated by the sub-module 206 so as to carry the information of the modified image pixel by pixel. In particular, the sub-module 206 is designed to implement, firstly, a segmentation of the image, pixel by pixel, in order to determine pixel by pixel a type of demosaicing to be carried out. The sub-module 206 is further configured to generate the data Y_pwa so that it includes information of a type of processing to be carried out pixel by pixel, such as for example a masking operation (in English "gating").Masking is for example performed on the data of the Y_tmp tensor using the Y_pwa signal in order to weight each type of reconstruction present in the Y_tmp tensor according to the Y_pwa tensor.

[0059] The network 202 further comprises, for example, a refinement sub-module 208 (REFINEMENT MOD). For example, the sub-module 208 is configured to generate a corrected image Y_pred on the basis of the intermediate data Y_tmp provided by the sub-module 204 and the attention data Y_pwa provided by the sub-module 206.

[0060] In another example, the sub-module 208 is configured to perform a masking of Y_tmp in the axis of the channels, on the basis of the data Y_pwa.

[0061] During the training of the network 202, the corrected image Y_pred is provided to a calculation circuit 210 (LOSS). The calculation circuit 210 is further configured to receive the terrain image Y_true and the data Y_aug comprising the indications of the positions where the outliers were added by the module 200.

[0062] For example, the data Y_aug is a tensor with a value in {0, 1}, the pixels assigned the value 1 indicate for example the positions of the bad pixels.

[0063] According to one embodiment, the calculation circuit 210 is configured to calculate an error value Err by applying a cost function (in English "loss function") to the modified image Y_pred, the terrain image Y_true and the data Y_aug. By way of example, the cost function applied by the calculation circuit 210 combines several fidelity cost functions, such as a cost function based on average error calculations without a priori and applied between the data Y_pred and Y_true, or a focused cost function (in English "Focal Loss Function" - LFF) applied to the data Y_pred, Y_true and Y_aug for which the focusing is, for example, controlled by the tensor Y_aug. By way of example, the cost function applied by the calculation circuit 210 further comprises a regularization function applied only to the corrected image Y_pred.For example, the regularization function is determined from assumptions about the nature of the signal of the corrected image Y_pred. .

[0064] For example, the error value is such that Err - LF M(Y _pred, Y _true) +^uf LF F {Y _pred, Y_true, Y_aug) +ÀlK x LR(Y _pred) where the LFM function is a cost function based on mean error calculations, for example based on a mean square error calculation, the LFF function is a focused cost function and the LR function is a regularization function, and where Rlfm, and ^lr are weighting coefficients. For example, the coefficients ^lff and ^lr are used to emphasize the importance of one cost function over another.

[0065] As an example, the average error LFM{ Y _pred, Y_true) is such that: [Math 1] LFM( Y_pred, Y_true) = |CNN( Y^pred) - CNN( Y_true) | + Â2^\CNN(Y_pred) - CNN( Y_mui) |2, where CNN(.) is an intermediate output of a neural network model that has been previously trained to perform certain inference tasks. For example, CNN(.) was trained in a supervised manner to feed a sub-model dedicated to image classification tasks. For example, the neural network relies on a latent space from a neural network model trained to perform classification operations on databases, such as the ImageNET database. For example, the CNN(.) function is a vgg(.) function where vgg(-) is a function returning a latent space, for example of type vggl6, the coefficients Xj and X2 are weighting coefficients between distance calculations of type L1 and L2. In particular, the sum is performed on all elements, that is, on all pixels and all channels.

[0066] As an example, the focused error LFF( Y _pred, Y truc, Y _aug^ is such that: [Math 2] LFF( Y_pred, Y_tme, Y_aug) = E ( (dilate^ Y_aug) ) O ( Y_pred- Y_true) y, where dilate^) defines a morphological dilation function, for example having a 5x5 circular kernel and where the operator O corresponds to point-by-point multiplication. In particular, the function dilate^) makes it possible to increase the impact radius of the cost function in the vicinity of outlier pixels, thus making it easier to correct the effects due to large receptive fields of the convolutional neural network used to reconstruct the image.

[0067] As an example, the regularity error LR{ Y_pred} is such that: [Math 3] LR{ Y pred) — TV( Y_pred), where TVÇ) corresponds, for example, to the calculation of the total variation, or to a variant of the total variation. The use of the total variation as a regularization function is given as an example and is of course not limited to mitative.

[0068] According to one embodiment, during the training of the network 202, the error value obtained is back-propagated in the network 202 in order to re-evaluate and update the parameters associated with the network 202. For example, the updating of the parameters is carried out by implementing a gradient back-propagation method based on the error value. Deep learning methods, based on gradient back-propagation, are known to those skilled in the art and are therefore not described in more detail here.

[0069] According to one embodiment, the focused cost function has the role of indicating to the network 202, during the backpropagation of the error, which pixels are highly noisy. For example, the coefficient ^lff is greater than the coefficient ^lfm-

[0070] [Fig. 3] is a block diagram illustrating an architecture of the network 202 configured to perform regression operations, according to one embodiment of the present disclosure.

[0071] For example, once trained, the network 202 is implemented in an image processing device. For example, the processing device further comprises an imager (not shown) configured to capture image scenes based on a matrix pattern such as, for example, one of the patterns 100, 102, 104, 108 or 110 described in relation to [Fig. 1]. For example, the imager comprises one or more bad pixels, or malfunctioning pixels. For example, the device is a camera, a camcorder, a smartphone, etc.

[0072] Once the training of the network 202 is complete, i.e. following the performance of the training described in relation to [Fig.2] on a plurality of ground truth images Y_true, the network 202 is for example configured to perform regression operations on an input image Y_in, replacing the data Y_bad, comprising for example one or more bad pixels. By way of example, the training is completed following the processing of a large number, for example at least 1000, ground truth images. In another example, the training ends when the learning algorithm converges, i.e. when the error value becomes lower than a threshold value.

[0073] The regression operation, performed by the trained network 202, then consists of generating a dematrixed and reconstructed image Y_out on the basis of a matrixed input image Y_in, provided for example by the imager. The trained network 202 then comprises, for example, the sub-modules 204, 206 and 208 whose parameters are fixed following the training. The matrixed image Y_in is then provided directly to the network 202, without passing through the augmentation module 200. Similarly, the dematrixed and reconstructed image Y_out is not provided to the calculation circuit 210. The augmentation module 200 and the calculation circuit 210 are, for example, not part of the processing device in which the trained network 202 is implemented.

[0074] The trained network 202 is then, for example, configured to perform regression, interpolation, denoising, and detection and correction operations of the bad pixels detected in the input image Y_in and to generate the corrected image Y_out on the basis of these operations. In particular, following the training, the processing sub-module 204 of the network 202 is, for example, configured to generate an intermediate data item Y_tmp by making a first estimation of the missing or noisy information, due at least in part to the bad pixels.

[0075] Following the training, the sub-module 206 is for example configured to generate an attention data item Y_pwa, on the basis of the input image Y_in and / or the intermediate data item Y_tmp. The sub-module 206 is then for example configured to detect, individually for each pixel, whether this pixel is a bad pixel or not. The success rate of the detection of bad pixels obviously depends on the training and the level of expressiveness, linked to its internal structure, of the network 202. The sub-module 206 then generates a data item Y_pwa comprising indications on the aberrations detected as well as on their location in the input image Y_in. For example, the detection is a multi-scale detection and the data item Y_pwa provides a context, for example an indication of category among, among others, a type of contour, texture, presence and level of noise and positions of the aberrant pixels, pixel by pixel.

[0076] Following the training, the sub-module 208 is for example configured to generate the dematrixed image Y_out on the basis of the intermediate data Y_tmp and the attention data Y_pwa. For example, the sub-module 208 is configured to correct the areas identified, by the data Y_pwa, as being outside the main statistic, that is to say as the areas comprising one or more bad pixels.

[0077] [Fig.4] is an example of architecture of the processing sub-module 204 according to an embodiment of the present description.

[0078] By way of example, the sub-module 204 comprises a plurality of sub-networks 400. By way of example, the sub-module 204 comprises a number Nd, Nd being an integer greater than 2, preferably greater than or equal to 5 and less than or equal to 12, of sub-networks 400.

[0079] For example, each sub-network 400 is configured to receive the input image Y_in and to generate, each, an intermediate data item Y_tmpi. The Nd intermediate data items Y_tmpi are then concatenated, for example in the channel axis or in an additional axis.

[0080] Each sub-network 400 comprises, for example, a plurality of layers. For example, the input image Y_in is provided to a convolution layer 402 (F- CONV KxK). For example, layer 402 is a two-dimensional convolution layer with F output channels, F being an integer greater than or equal to the number of pixels in the matrix pattern, and comprising kernels of size K x K, where K is an integer strictly greater than the width and / or height of the matrix pattern.

[0081] For example, the output generated by layer 402 is provided to a layer 404 (MOSAIC2CHANNEL). For example, layer 404 is a so-called pixel shuffle layer configured to spatially rearrange the channels closely in the data provided to it. In particular, layer 404 is configured to perform a so-called “space to depth” operation. In addition, the operation performed is parameterized by the size of the matrix pattern considered.

[0082] For example, the output generated by layer 404 is provided to a layer 406 (DW CONV 1x1). For example, layer 406 is a depth-wise convolution layer comprising kernels of size 1x1, thus allowing each channel to be weighted independently.

[0083] For example, the output generated by layer 406 is provided to a layer 408 (MOSAIC2CHANNEL). For example, layer 408 is a pixel mixing layer, for example similar to layer 404.

[0084] For example, the output generated by layer 408 is provided to a layer 410 (DW CONV KxK). For example, layer 410 is a depthwise convolution layer comprising kernels of size KxK.

[0085] For example, the output generated by layer 410 is provided to a layer 412 (MOSAIC2CHANNEL). For example, layer 412 is a pixel mixing layer, for example similar to layers 404 and 408.

[0086] For example, the output generated by layer 412 is provided to a layer 414 (DW CONV 1x1) similar to layer 406.

[0087] As an example, the output generated by layer 414 is provided to a layer 416 (CHANNEL2MOSAIC) called pixel mixing. In particular, layer 416 is configured to perform a so-called “depth to space” operation. In addition, the operation performed is parameterized by the size of the matrix pattern considered.

[0088] For example, the output generated by layer 416 is provided to a layer 418 (1 - CONV KxK). For example, layer 418 is a two-dimensional, single-channel output convolution layer comprising kernels of size KX K.

[0089] For example, a concatenation aggregation layer 420 (CONCAT) is configured to perform a concatenation of the output data of layers 418 and 408. For example, the output data of layer 408 is provided to layer 420 via a so-called “connection hop” branch (in English "skip connection"). The connection is then configured to provide, for example, ¾ of the output channels of layer 408 directly to layer 420.

[0090] The concatenation generated by the layer 420 is for example provided to a convolution layer 422 (3 - CONV KxK). For example, the layer 422 is a two-dimensional convolution layer with three output channels and comprising kernels of size K x K. For example, the three output channels are the three RGB color channels. The layer 422 is then configured to generate the output data Y_tmpi, corresponding to, for example, a first demosaicing and reconstruction of the input image.

[0091] As an example, the data Y_pwa is injected into the model of the network 206 in order to calculate all or part of the intermediate data Y_tmp.

[0092] According to one embodiment, the parameters of the sub-networks 400 are learned during the learning of the network 202 described in relation to [Fig.2].

[0093] [Fig.5] is an example of architecture of the attention module 206, according to an embodiment of the present description.

[0094] Since the sub-module 206 has the task of detecting, pixel by pixel, bad pixels, an architecture used for image segmentation is suitable for the implementation of the sub-module 206.

[0095] For example, the sub-module 206 then comprises a network 500 (U-NET) of the U-net type. In other examples, the network 500 is a type of network other than the U-net configured to perform image-to-image operations. For example, the network 500 is a variant and / or an improvement of a U-net type network. For example, the network 500 is further configured to implement internal attention mechanisms and / or to manipulate 2- or 3-dimensional data. For example, the network 500 further comprises a dense interconnection structure.

[0096] By way of example, the network 500 is a U-net type network having a typical depth equal to 4 and comprising convolution blocks of 5 x 5 receptive fields with a ReLU type activation. In one example, the U-net type network 500 is based on convolution blocks comprising two 3x3 convolutions per group, each followed by a mixing of the channels. This example is specific and makes it possible to limit the complexity of the network 500 and to apply 5 x 5 receptive fields. The person skilled in the art will know how to adapt and modulate the sizes of the receptive fields as well as the number of convolution groups. By way of example, the network 500 is configured to further perform a sub-sampling operation, such as a so-called “stride” operation, for example implemented by a so-called MaxPooling layer. For example, subsampling is subsampling by 2, corresponding to a 2x2 MaxPooling operation.For example, the network 500 further comprises normalization stages making it possible to limit the impact of the absolute amplitude of the . data on inference operations. In another example, when the absolute magnitude of the data is an element taken into account in the detection of bad pixels, the network 500 comprises at least one non-normalized path. The skip connections are implemented by an aggregation operation, for example a concatenation. For example, an upsampling operation is carried out by a convolution having a stride of 2.

[0097] For example, the network 502 is configured to provide the data that it generates to a layer 504 (L1 NORM). The layer 504 is configured to normalize the received value, for example by applying a normalization performed in the axis of the channels and noted NL1. The application of the NL1 normalization allows the sum of the Nd intermediate data to be equal to 1. For example, the application of the NL1 normalization is carried out by means of an activation function of the softmax, or softargmax, type. Indeed, the softargmax(.) function is equal to NL1(exp(.)). The argmax function therefore also returns an output whose sum of the elements is equal to 1. Thus, the different reconstructions of the intermediate data Y_tmp are summed in a normalized manner.For example, the layer 504 is configured to receive data having the form of a tensor T e I and J being, respectively, the number of rows and columns of pixels in the image and C being the number of channels of the tensor. For example, for the module 204, the values ​​C and Nd are equal. By noting the value of the tensor T at the position [i,j] in the image and for a channel c, T[i,j,c], the layer 504 is for example configured to normalize, for each position, the tensor performing the transformation: . [Math 1] In addition, NL1 standardization is applied to positive data, from an activation function of the network 500.

[0098] In another example, in order to avoid a potential division by 0, the L1 norm of the tensor, at position [i,j] is defined such that [Math 2] VL1(T[UC]) =r[î-,y,c] / (EX1|7'[ / . / c]| + where* is a real number, strictly greater than 0.

[0099] In yet another example, the normalization NL1 of the tensor, at position [i,j] is defined such that NL1 ( T [ i, j, c ] ) - 0 or NL1 ( T [ z, j, c ] ) - 1 / C, if -HAS . 2V£l(T[z, j, c] ) = T[i, j, c] / £.|T[U, c] | 500 segmentation network.

[0101] [Fig.6] is an example of architecture of the refinement network 508, according to an embodiment of the present description. ​

[0102] As described in relation to Figures 2 and 3, the sub-module 208 is configured to generate the output Y_pred, during the training of the network 202, or Y_out during the execution of the trained network 202. The input data is then used to infer the output data. The data Y_pwa is, for example, used to select the outputs Y_tmpi of the sub-networks 600. The data Y_pwa is, for example, further used to indicate the type of processing, for example masking, to be carried out on the data Y_tmp resulting from the first demosaicing and reconstruction, carried out by the sub-module 204.

[0103] By way of example, the sub-module 208 comprises a layer 600 (MULT) configured to perform a multiplication operation, for example point to point, of the tensor Y_tmp on the basis of the values ​​of the attention tensor Y_pwa. In another example, the layer 600 is configured to perform a multiplexing operation of the tensor Y_tmp on the basis of the values ​​of the attention tensor Y_pwa. In the example where the layer 600 is configured to perform a multiplexing operation, an activation function of the module 206 is of type argmax. The output of the layer 600 is then, for example, provided to a dense layer 602. The dense layer 602 is then configured to generate an output L, on the basis of the data provided by the layer 600.For example, the output Y] corresponds to a list of D linear combinations of the outputs of layer 600, thus generating a tensor of size WXH x Nc XD, where W is the pixel width of the image, H the pixel height of the image and where Nc refers for example to the three color channels and where D is the dimension in axis 4 of the tensor. The value of D depends for example on the topology, as well as on the hyperparameters of the network 202. .

[0104] The sub-module 208 further comprises a convolution layer 604 (CONV 2D) configured to perform a convolution operation on the tensor Y_pwa. The result of the layer 604 is then provided to a layer 606 (HARDSIGMOID). For example, the layer 606 is configured to apply an activation function, for example a HardSigmoid type function, to the received data.

[0105] Layer 606 is then configured to provide its output to a layer 608 (RESHAPE). For example, layer 608 is configured to resize the data provided by layer 606 into a tensor Y2 of size W x H x Nc x D.

[0106] As an example, the tensor Y 2 is provided to a layer 610 (x<-lx) configured to generate a tensor F3 by inverting, with respect to 1, the values ​​of the tensor Y2. The tensor Y2 and the data Y ] are further provided to a layer 612 (MULT) similar to the layer 600. As an example, the layer 612 is configured to generate a data E4 based on data Yj and Y2.

[0107] The data Y4 is provided to a layer 616 (CNN-CORR). The layer 616 is, for example, a convolution layer, or a cascade of convolution layers, configured to generate a data Y5 of the detected defects using the tensor Y_pwa. Thus, the parts of the input image being detected as being to be corrected are smoothed.

[0108] Layer 616 is further configured to provide data ^5 to a layer 618 (MULT). Layer 618 is further configured to receive data K3, provided by layer 610, and to generate data Y6 by performing an operation similar to that performed by layers 600, 612 and 614.

[0109] The sub-module 208 further comprises a layer 620 (ADD). For example, the layer 620 is an additional aggregation layer. The layer 620 is for example configured to receive the data Y4, the output data from the layer 618, and to add them.

[0110] The sub-module 208 further comprises a layer 622 (CNN-COM) configured to generate the output data Y_pred or Y_out based on the output of the layer 620. For example, the layer 622 is a convolution layer, or a cascade of convolution layers, configured to shape the output data. For example, the shaping comprises one or more rotations and / or colorimetric adjustments and / or smoothing and / or shifting operations, etc.

[0111] According to one embodiment, the use of the activation function, of HardSigmoid type, allows on / off type detection of bad pixels and to carry out an operation similar to multiplexing using multiplicative aggregation layers and an additional aggregation layer.

[0112] Various embodiments and variations have been described. Those skilled in the art will understand that certain features of these various embodiments and variations could be combined, and other variations will occur to those skilled in the art. In particular, variations are possible with respect to the types of cost functions combined with the focused cost function. Similarly, the design, as well as the topology, of the modules 204, 206, and 208 may vary.

[0113] Finally, the practical implementation of the described embodiments and variants is within the reach of the person skilled in the art from the functional indications given above. In particular, the practical implementation of a desired type of imager is within the reach of the person skilled in the art.

Claims

Claims

1. A method for training a neural network (202) configured to perform image processing operations, the method comprising: - generating a modified image (Y_bad), by a modified data generator (200), based on a first image (Y_true), the modified image comprising at least one aberrant pixel value relative to the first image; - providing the modified image to the network; - generating, by the network, a corrected image (Y_pred); - providing the corrected image, by the network, and providing an indication (Y_aug) of the position of the at least one modified pixel, by the generator, to a calculation circuit (210) - generating an error value, based on the application of a cost function, by the calculation circuit, taking as input the first image, the indication and the corrected image; - correction of parameters associated with the network by backpropagation of the error in the network.

2. Method according to claim 1, wherein, the generation of the corrected image (Y_pred) by the network (202) comprises: - the generation, by a first sub-module (204), of an intermediate data value (Y_tmp) on the basis of the modified image (Y_bad), the intermediate data being an at least partial reconstruction of the first image; - the generation, by a second sub-module (206) of the network (202), of an attention data item (Y_pwa) on the basis of one of the intermediate data item and / or the modified image, the attention data item comprising an estimate of the location of the at least one modified pixel; - the generation of the corrected image (Y_pred), by the first sub-circuit or by a third sub-circuit (208), on the basis of the intermediate data item and the indication data item.

3. Method according to claim 2, wherein the generation of the attention data (Y_pwa) comprises the execution of a convolutional neural network configured for image segmentation (500), on the basis of the intermediate data (Y_tmp) and / or the corrected image (Y_bad).

4. The method of claim 3, wherein the convolutional neural network (500) is a U-net or a variant of a U-net.

5. Method according to claim 3 or 4, wherein the generation of the attention data (Y_pwa) further comprises a normalization of type NL1, on the basis of the data generated by the convolutional neural network.

6. Method according to any one of claims 2 to 5, wherein the generation of the corrected image comprises a multiplexing operation and / or a multiplication operation, for example a point-to-point multiplication, on the basis of the attention data (Y_pwa) and the intermediate data (Y_tmp).

7. A method according to any one of claims 1 to 6, wherein generating an error value comprises applying a focused cost function (LFF) taking, as input data, the modified image (Y_pred), the first image (Y_true) and the location indication of the at least one modified pixel (Y_aug).

8. Method according to claim 7, wherein the generation of an error value further comprises the application of: - a cost function based on an average calculation (LFM) taking, as input data, the modified image (Y_pred) and the first image (Y_true); and / or - a regularization function (LR) taking, as input data, the modified image (Y_pred).

9. A method according to any one of claims 1 to 8, wherein the first image is included in a database.

10. Method according to any one of claims 1 to 9, in which the generation of the modified image comprises, for each pixel: - determining whether the pixel is to be modified; and - if the pixel is to be modified, replacing the value associated with the pixel with an aberrant value.

11. A method according to any one of claims 1 to 10, wherein the generation of the modified image (Y_bad), by the modified data generator (200), is further performed on the basis of a matrixing pattern.

12. Method according to any one of claims 1 to 11, in which the model of the modified data generator (200) is a neural network model previously trained for the generation of modified images, and configured for the implementation of so-called "style transfer" techniques.

13. An image processing method comprising: - capturing an image scene, by an imager of an image processing device, the captured image being a matrixed image according to a matrixing pattern - providing the matrixed image to a neural network of the processing device trained according to the training method according to any one of claims 1 to 12; and - generating a corrected image, by the network, on the basis of the matrixed image.

14. Image processing method according to claim 13, wherein the generation of the corrected image by the network comprises: - providing the matrixed image to a first sub-module (204) of the network (202) configured to generate an intermediate data item (Y_tmp) by performing a first dematrixing of the matrixed image; - providing the intermediate data item to a second sub-module of the network (206) configured to generate an attention data item (Y_pwa) on the basis of the intermediate data item, the attention data item comprising estimates of location of malfunctioning pixels of the imager; and - generating the corrected image, by a third sub-module (208), on the basis of the intermediate data item and the attention data item.

15. An image processing method according to claim 13 or 14, wherein the matrix pattern corresponds to a matrix pattern used when generating a modified image during training of the network (202).

16. Image processing device comprising: - an imager configured to capture image scenes, according to a matrix pattern; - a neural network (202) trained according to the training method according to any one of claims 1 to 12, configured to generate a corrected image on the basis of the matrixed image.

17. An image processing device according to claim 16 wherein the imager comprises one or more malfunctioning pixels.