Method and system for image to image translation
The method addresses the challenges of comparing images from different remote sensing sources by using a DDPM for patch-wise image-to-image translation, enhancing resolution and quality while maintaining target domain characteristics.
Patent Information
- Application Number
- PCT/EP2024/083476
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-20
- Filing Date
- 2024-11-25
- Publication Date
- 2025-06-26
AI Technical Summary
Existing image processing techniques for remote sensing struggle with comparing images from different sources due to limitations such as poor quality, low resolution, and differences in color perception.
A computer-implemented method using a Denoising Diffusion Probabilistic Model (DDPM) for patch-wise image-to-image translation, which translates a local image patch from a source domain to a target domain by performing channel-wise mean centering and applying a channel-wise color vector to maintain tonality.
The method effectively translates images from one remote sensing sensor to another, enhancing resolution and quality while maintaining the characteristics of the target domain, thereby facilitating more reliable change detection and comparison.
Smart Images

Figure EP2024083476_26062025_PF_FP_ABST
Abstract
Description
METHOD AND SYSTEM FOR IMAGE TO IMAGE TRANSLATIONField of the Invention
[0001] The present invention generally relates to image-to-image translation of an input image from a source domain to a target domain in the technical domain of remote sensing. The invention concerns an image processing method based on a trained neural network that can adapt images from one remote sensing sensor to another, while enhancing their resolution and quality.Background of the Invention
[0002] At present, deep learning techniques based on an artificial neural network have made great progress in fields such as object classification, text processing, engine recommendation, image search, face recognition, age and speech recognition, humanmachine dialogue, emotional computing and so on. With the deepening of an artificial neural network structure and the improvement of algorithms, the deep learning techniques have made breakthrough progress in the field of human-like data perception, and the deep learning techniques can be used to describe image content, recognize objects in complex environments in an image, perform speech recognition in noisy environments, and the like. At the same time, the deep learning techniques can also be used to solve a problem of image generation and fusion.
[0003] In remote sensing, it is not uncommon to compare images from the same region from different image sources. Change detection is often an important objective, where images from the same region acquired at separate times are compared to identify any changes in the area of interest. Unfortunately, the use of different image sources for comparison often suffers from limitations such as poor quality, low resolution of the images and differences in colour perception.Summary of the Invention
[0004] It is an object of the present invention, amongst others, to solve or alleviate the above identified problems and challenges by providing a computer implementedmethod that performs a patch-wise image-to-image translation of a source image to a target image.
[0005] To this aim, according to a first aspect of the invention, there is provided a computer implemented method for translating a local image patch of an image from a source domain having a first spatial resolution to a target domain having a second spatial resolution, thereby translating said local image patch into a target image patch. The computer implemented method comprising the steps of:- obtaining said local image patch and a global image patch of the image as a stacked image patch, said local image patch comprising image data from a zone of interest, said global image patch comprising image data corresponding to an area exceeding said zone of interest and having an identical pixel count as said local image patch,- performing a channel-wise mean centering of said stacked image patch to obtain a whitened stacked image patch,- performing by a Denoising Diffusion Probabilistic Model, or DDPM, N iterations, using a random noise image having said second spatial resolution as an initial input and said whitened stacked image patch as a conditioning input, to obtain a quasi-whitened output image patch,- applying a channel-wise colour vector to said quasi-whitened output image patch thereby obtaining said target image patch, said channel-wise colour vector characterizing the mean of each channel of an image patch from said target domain or said source domain corresponding to said zone of interest.
[0006] The image from the source domain is a digital image that is acquired by an image acquisition device, such as a remote sensing device. An acquisition device or digital image camera produces images that have a set of often unique characteristics. An image from such an acquisition device is amongst others characterized by a spatial resolution or pixel resolution which relates to the dimensions of an area in the real world represented by a pixel, a pixel count which corresponds to the number of pixels of the sensor and of the produced image, and the number of channels, such as for instance color-channels or depth channels. The acquired image may also becharacterized by a specific tonality which defines the overall appearance of an image, and this regarding the range and distribution of tones and the smoothness of gradation between them. Tonality plays an important role in technical photography, such as for instance in remote sensing photography.
[0007] In remote sensing photography it is not uncommon to find that the spatial resolution, pixel count and tonality of images acquired by two different devices are very different. The method of the invention advantageously allows to translate the image data represented in a source image onto the context of a target image. This process is generally called image-to-image translation or I2l-translation; it is the process of transforming an image from one domain to another, where the goal is to learn the mapping between an input image and an output image. The method of the invention not only spatially maps the source image to the target image, but it also maps the image information while maintaining the characteristics of the target image. In this process, the tonality of the source image is transformed into the tonality of the target image.
[0008] When trying to map a portion or patch of a source image, a local image patch, onto a target image, the corresponding location of the imaged area represented by the local image patch is identified in the target image. The imaged area by the local image patch is called a zone of interest in the context of this invention. This zone of interest is the area in the real world that is represented by the image information in the local image patch. The target image patch is another representation of this zone of interest, but it is the representation in the target domain. The local image patch may therefore have a different pixel count and different tonality compared to the target image patch.
[0009] The method of the invention obtains a local image patch for processing which defines a small portion of the original image of the source domain. Processing the original image of the source domain in patches rather than in its entirety offers the advantage that less resources such as memory and processing power are required to process the entire original image. The method of the invention will combine this local image patch with another image patch that covers a larger area around the local image patch: the global image patch. The global image patch will therefore comprise the same image information as the local image patch, but will also comprise image information covering the area surrounding the area covered by the local image patch. Theadditional surrounding area should at least exceed the local image patch surface with 5% and ideally exceeds it with 100% upto 1000%. The global image patch therefore comprises image information of the surrounding context of the local image patch allowing the algorithm to advantageously make more reliable guesses for the target context. The global image patch is however scaled to match the same pixel count as the local image patch. Both image patches are then stacked into a stacked image patch, before they are fed into the neural network.
[0010] The stacked image patch therefore has an identical pixel count as the local image patch and the global image patch, but comprises all channels of both patches. The stacked image now is processed in a whitening step, which may be performed as a channel-wise mean centering of the stacked image patch to obtain a whitened stacked image patch. Other algorithms neutralizing the overall colour tonality from all channels of the stacked image patch may be considered here as alternatives.
[0011] The mean centering operation is thus performed individually on each individual channel of the stacked image patch. Channel-wise mean centering implies that first the mean value of all pixel values of a channel is calculated, after which each pixel value of this channel is subtracted by this mean value, resulting in a centered version of the channel around zero. This operation is then performed for each channel individually or channel-wise. In other words, after applying the mean centering operation to all channels of the stacked image patch, all values in each channel of the stacked image patch will be centered around zero, such that it can be presented to the neural network as colour neutralized, normalized and therefore standardized data.
[0012] Optionally, an additional normalization step may be applied over all channels simultaneously, meaning that the pixel values in all channels are divided by the maximum pixel value of all channels, such that all pixel values in the stacked image patch are distributed between -1 and +1 .
[0013] The neural network that is used in this method for processing the image-to- image transformation may be a Denoising Diffusion Probabilistic Model or DDPM. DDPM’s are a class of latent variable models inspired by considerations from nonequilibrium thermodynamics. DDPM’s have recently been shown to produce high- quality images.
[0014] Random noise images are used as initial input for the trained DDPM or trained model. The random noise images are required to have the same pixel count as the target image patch, and may be obtained by a random noise generating algorithm. The random noise may be Gaussian noise, white noise, pink noise, Brownian noise or alike.
[0015] The whitened stacked image patch is then fed into the model as a conditioning input. The model performs N iterations before it outputs the quasi-whitened output image patch, which advantageously has the same pixel count as the target image patch.
[0016] In order to obtain the final target image patch, the colour information of either the target image patch or the local image patch is applied by applying a colour vector to the quasi-whitened output image patch. The colour vector corresponds to a vector comprising the mean value calculated for each channel of an image patch from the target domain or the source domain that corresponds to the zone of interest. The choice between selecting the colour information of either the target image patch or the local image patch depends on which colouring the user wants to obtain for the end result in the target image patch. The channel-wise colour vector characterises the mean of each channel of an image patch from the target domain or the source domain corresponding to the zone of interestin other words, the colour vector comprises a channel-wise calculated mean value of the image patch in the target domain or the source depending on the users’ choice, which could be used to perform a whitening operation on the respective image patch. The colour vector therefore is obtained as the inverse whitening operation on the image patch in the target domain, or again the source domain, that corresponds to the same zone of interest as the local image patch. The colour vector represents the tonality information of the zone of interest in the target domai or the source domainn.
[0017] When the colour vector is applied to the quasi-whitened output image patch, the tonality of the output image patch will advantageously match the tonality of the image patch in the target domain or the source domain, depending on which image patch the colour vector is calculated, and that corresponds to the same zone of interest as the local image patch.
[0018] In another embodiment, an optimization is achieved for the generation of the random noise image that is used in the method as initial input for the model. Rather than using a single random noise image as initial input for the model, in this embodiment, the random noise image is selected from a stack of r randomly generated noise images. The stack of r randomly generated noise images comprise a set of randomly generated images that are having the same identical pixel count as the target image patch, and may be obtained by a random noise generating algorithm. The random noise may be Gaussian noise.
[0019] In order to find the best random noise image from the stack, all noise images from the stack are presented to the DDPM as initial input. Again the whitened stacked image patch is used as a conditioning input for the model. After a limited number of iterations, namely N / 7 iterations, the model outputs a stack of r quasi-whitened output image patches of lower quality due to the limited number of iterations performed. The number of iterations that the model executes is deliberately lower than the optimal N iterations that is used to produce the high-quality target image patch itself. The value of / is selected to be a positive natural number larger than one, in order to reduce the number of iterations for the optimization step.
[0020] The method then determines the image patch of the stack of r quasi-whitened output image patches with the highest quality, by determining the image patch of the stack with the highest or maximal quality metric value when comparing the image patch with the original local image patch, or a whitened version of the original local image patch. The method of the invention is advantageous in that a reduced number of iterations still allows a reliable evaluation of the image quality metric, and a reliable selection of the image patch of the stack of r quasi-whitened output image patches.
[0021] In this way, the particular random noise image from the stack of r randomly generated noise images that was involved in the generation of the image patch of the stack of r quasi-whitened output image patches with the highest quality, is then used by the method of claim 1 to generate the high-quality target image patch.
[0022] The invention is advantageous in that the selection of an ideal or optimal random noise image is proposed that will result in an optimal target image patch.
[0023] In an other embodiment of the invention, the determination of the image patch of the stack of r quasi-whitened output image patches with the highest quality is determined by determining the image patch of the stack with the highest or maximal image quality metric when comparing the image patch with the original local image patch, or a whitened version of the original local image patch.
[0024] In yet another embodiment, the invention provides a method that performs a patch-wise image-to-image translation of a source image to a target image. When addressing a translation of a complete source image to a target image, the proposed method will tackle this by splitting up a source image into a series of contiguous patches that will be processed according to the method of claim 1 . An area of a source image is then spatially mapped to a corresponding area in the target image, in other words: an area of an image from a source domain is spatially mapped to a corresponding area of the target domain. The image from the source domain is compartmented or split up into small patches that are preferably square. The method of the invention will generate an equivalent number of target image patches that are recomposed into the target image.
[0025] The invention therefore is advantageous in that the image-to-image translation method allows for a patch-wise processing of a source image into a target image. The compartmentation of the processing lowers the processing and memory requirements for the processing hardware, and eventually speeds up the processing as a whole.
[0026] In another embodiment of the invention, a method for training the DDPM is proposed by obtaining a training dataset for training said DDPM, the training dataset consisting of a plurality of pairs of a noisy version of a whitened target image patch of a zone of interest of an image from said target domain, said whitened target image patch obtainable by performing a channel-wise mean centering of said image patch of said zone of interest of said image from said target domain, and said noisy version obtainable by an multiplicative addition of a randomly generated noise image e to said whitened target image patch, and a stacked image patch of a corresponding zone of interest to the zone of interest of said noisy version of said whitened target image patch, wherein said stacked image patch is obtained by stacking of an image patch from said corresponding zone of interest of said image from said source domain and a down sampled image patch from said source domain corresponding to an areasurrounding said corresponding zone of interest, training said neural network by: presenting to said DDPM as a diffusion input said noisy version of said whitened target image patch of said zone of interest of said image from said target domain, while presenting to said DDPM as a conditioning input said stacked image patch of said corresponding zone of interest from said source domain, to produce a predicted noise image matrix calculating a loss function L using the predicted noise image matrix as a prediction and said randomly generated noise image as a ground truth, calculating a loss value of parameters of the DDPM through said loss function L, modifying said parameters of the DDPM according to said loss value, and in a case where the loss function L satisfies a predetermined condition, obtaining a trained DDPM, and in a case where the loss function L does not satisfy the predetermined condition, continuing to input said noisy version of said whitened image patch of said zone of interest of said image from said target domain and said stacked image patch of said corresponding zone of interest from said source domain so as to repeatedly perform the above training process.
[0027] In yet another embodiment of the invention, the noisy version of the whitened target image patch is obtained by the multiplicative addition of a randomly generated noise image e to the whitened target image patch, and is calculated by applying the formula:wherein ytis a noise level value that can be selected by the user between 0 and 1 , wherein e is said randomly generated noise image and y0is said whitened image patch from said target domain.
[0028] The application of this formula advantageously allows the selection of a user selectable value of the noise level that is added to the whitened image patch.
[0029] In another embodiment, an additional channel-wise mean centering step is applied to said quasi-whitened output image patch before applying said channel-wise colour vector. This additional step provides for an additional colour neutralization step prior to the inverse recolouring step that is applied to obtain the output image patch.
[0030] In another embodiment, an additional channel-wise mean centering step is applied to said quasi-whitened output image patch before comparing it to the channelwise mean centered version of said local image patch to determine the output patch from said stack of r quasi-whitened output image patches for which an image quality metric value is maximal. The additional whitening step leads to better end results produced by the method.
[0031] The embodiments of the systems and methods described herein may be implemented in hardware or software, or a combination of both. However, preferably, these embodiments are implemented in computer programs executing on programmable computers each comprising at least one module component which comprises at least one processor (e.g. a microprocessor), a data storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device. For example and without limitation, the programmable computers (referred to here as computer systems) may be a personal computer, laptop, personal data assistant, and cellular telephone, smart-phone device, tablet computer, and / or wireless device. Program code is applied to input data to perform the functions described herein and generate output information. The output information is applied to one or more output devices, in known fashion.
[0032] Each program is preferably implemented in a high level procedural or object oriented programming and / or scripting language to communicate with a computer system. However, the programs can be implemented in assembly or machine language, if desired. In any case, the language may be a compiled or interpreted language. Each such computer program is preferably stored on a storage media or a device (e.g. ROM or magnetic diskette) readable by a general or special purpose programmable computer, for configuring and operating the computer when the storage media or device is read by the computer to perform the procedures described herein. The subject system may also be considered to be implemented as a computer- readable storage medium, configured with a computer program, where the storage medium so configured causes a computer to operate in a specific and predefined manner to perform the functions described herein.
[0033] Specific examples and preferred embodiments are set out in the dependent claims. Further advantages and embodiments of the present invention will become apparent from the following description and drawings.
[0034] The present invention can be implemented as a computer program product adapted to carry out the steps as set out in the description. The computer executable program code adapted to carry out the steps set out in the description can be stored on a computer readable medium.Briefof the
[0035] Fig. 1 shows a representation of two images which are the subject of an image- to-image translation method, and which comprise the same zone of interest in two different domains.
[0036] Fig. 2 illustrates a flow chart illustrating a single-pass inference method of the invention.
[0037] Fig. 3 illustrates a flow chart illustrating the full-patch inference method of the invention.
[0038] Fig. 4 illustrates a flow chart illustrating the training method of the invention.
[0039] Fig. 5 illustrates a suitable computing system for executing the methods of the invention.Detailed Description of Embodiment(s)
[0040] Fig. 1 shows an example of two images 100, 200 which are the subject of an image-to-image translation method. A first image 100 represents a complete image belonging to a source domain D1 . The source domain has a number of characteristics attributed to it, such as for instance a spatial resolution defining the number of pixels of the image, and a tonality defining the colouring characteristics of the image. The local image patch 110 is a sub-element of the image 100 that comprises a number of image features 111. The local image patch 110 is surrounded by a global image patch 120, wherein the size of the global image patch 120 exceeds the size of the local image patch 110. The additional surrounding area should at least exceed the local imagepatch surface 110 with 5% and ideally exceeds it with 100% upto 1000%. The global image patch 120 therefore provides supplementary image context data around or surrounding the local image patch 110. This supplementary image context data supplements the deep learning model with data of the surroundings of the local image patch to be translated to the context of domain D2.
[0041] A second image 200 represents another image belonging to the target domain D2. The purpose of the method of the invention is to translate a portion or patch 1 10 of the image 100 to the target domain D2, in other words, the method casts the image content of the local image patch 110 in the source domain D1 to the context of the local image patch 210 in the target domain D2. The method therefore produces a target image patch 180, in Fig. 2-4, of which the image content is derived from local image patch 110 of the source domain D1 , but which receives the characteristics of the target domain D2. These characteristics being a.o. the spatial resolution or pixel resolution, pixel count, and tonality.
[0042] In a typical embodiment of the invention, the image 100 in source domain D1 origins from an image source that has a lower resolution compared to the resolution of the image 200 that originates from a target mage source with target domain D2. A typical application or use case of the method is the upscaling of a low-resolution image to match the high-resolution domain aspects of a high-resolution image, such that the upscaled and modified low-resolution image can be directly compared with the high- resolution image. The method can handle images from different image sensors that have different resolutions and tonalities, and makes direct comparison between images from different sources easier.
[0043] The method takes a low-resolution satellite image, for instance an optical satellite image, as input and divides it into several patches. Each patch is then fed to a CNN-based model, which generates a translated version of the patch, with the same content but in the style and resolution of a higher-resolution sensor. The translated patches are then stitched together to form the output image, which has improved resolution and quality.
[0044] Fig. 2 shows a flow chart illustrating a single-pass inference method 400 of the invention according to an aspect of the invention, wherein a local image patch 110 anda global image patch 120 in combination with a random noise image 150 are used as input to provide a target image patch 180 as an output.
[0045] In a first step, the global image patch 120 is resized to the same pixel count as the local image patch 110. The resizing is performed by any known image resizing algorithm known in the art. Both images are then stacked together into a so-called a stacked image patch 130, which essentially means that this will result in an image-like container file having the double amount of channels of the original local image patch 110. For instance, in case where a local image patch 110 was an RGB-colored image patch with a pixel count of 200 x 200 pixels, and the global image patch 120 was an RGB-colored image patch with a pixel count of 500 x 500 pixels, the global image patch 120 would be resized into an RGB-colored image patch with a pixel count of 200 x 200 pixels before it would be combined into a stacked image patch 130 having a pixel count of 200 x 200 pixels and 6 channels.
[0046] The stacked image patch 130 now is processed in a whitening step 310, which may be performed as a channel-wise mean centering of the stacked image patch 130 to obtain a whitened stacked image patch 140. Other algorithms neutralizing the overall colour tonality from all channels of the stacked image patch may be considered here as alternatives. The whitened stacked image patch 140 inherits the same pixel and channel count as the stacked image patch 130.
[0047] Next to the above steps, the method generates a random noise image 150 that has same pixel count as the target image patch 180. The random noise image 150 may be generated by any computer implemented random noise generating algorithm known in the art, or alternatively may be obtained through a random noise image optimization algorithm such as the one depicted in Fig. 3.
[0048] The random noise image 150 is used as initial input for a trained Denoising Diffusion Probabilistic Model or DDPM 300. The whitened stacked image patch 140 is then fed into the trained DDPM 300 as a conditioning input. The model performs N iterations before it outputs the quasi-whitened output image patch 160, which advantageously has the same pixel count as the target image patch 180.
[0049] The quasi-whitened output image patch 160 may again be submitted to a whitening step 310 before it is processed by a colouring block 320 that applies colourvector 170 to obtain the final target image patch 180. As explained above, the colour vector 170 corresponds to a vector comprising the mean value calculated for each channel of an image patch from the target domain that corresponds to the zone of interest. In other words, the colour vector 170 represents the tonality information of the zone of interest in the target domain.
[0050] Fig. 3 illustrates a flow chart illustrating the full-patch inference method of the invention according to an aspect of the invention, wherein an optimized random noise image 150 is obtained by means of a random noise image optimization loop 500.
[0051] The flow chart comprises and shows the same single-pass inference method 400 to calculate the final target image patch 180. But before executing the single-pass inference method 400, the flow chart also illustrates that an optimized random noise image 150 is calculated by means of evaluating a stack of r randomly generated noise images 155 by testing the image quality of the DDPM-produced r lower-quality quasiwhitened output image patches 190 against a whitened version of the original local image patch 110. The random noise image optimization method 500 thus identifies and selects the optimized random noise image 150 that is used by the same singlepass inference method 400 to calculate the final target image patch 180.
[0052] In a first step, the global image patch 120 is again resized to the same pixel count as the local image patch 110. The resizing is performed by any known image resizing algorithm known in the art. Both images are then stacked together into a so- called a stacked image patch 130, which essentially means that this will result in an image-like container file having the double amount of channels of the original local image patch 110. The stacked image patch 130 is again processed in a whitening step 310, which may be performed as a channel-wise mean centering of the stacked image patch 130 to obtain a whitened stacked image patch 140. Other algorithms neutralizing the overall colour tonality from all channels of the stacked image patch may be considered here as alternatives. The whitened stacked image patch 140 inherits the same pixel and channel count as the stacked image patch 130.
[0053] Next, the random noise image optimization loop 500 initially generates a stack of r randomly generated noise images 155, these noise images 155 having the same pixel count as the target image patch 180. The stack of r randomly generated noiseimages is in fact a set or collection of r different noise images having the same pixel count. The optimization loop 500 will process and evaluate all r randomly generated noise images in order to identify the randomly generated noise image from the stack providing the best image quality in the DDPM-produced r lower-quality quasi-whitened output image patches 190.
[0054] The optimization loop 500 therefore uses each of the r randomly generated noise images 155 as initial input for the trained Denoising Diffusion Probabilistic Model or DDPM 300. The whitened stacked image patch 140 is then fed into the trained DDPM 300 as a conditioning input. The DDPM 300 is however executed using a limited number of iterations compared to the number of iterations used in the single-pass inference method 400. The number of iterations is restricted to N / i, wherein i is an integer greater than 1 .
[0055] Limiting the number of iterations reduces the calculation time, but also produces a lower quality quasi-whitened output image patch 190 compared to the image quality that can be achieved when using the full number of iterations N. However the lower quality of the produced lower quality quasi-whitened output image patches 190 is sufficient to evaluate their image quality against the whitened version of the original local image patch 110, and to determine the particular randomly generated noise image from the stack 155 providing this best image quality.
[0056] The identification of the particular randomly generated noise image from the stack of tested r randomly generated noise images 155 providing the best image quality, is performed by determining the lower quality quasi-whitened output image patch of the stack 190 with the highest or maximal pSNR-value when comparing 330 this lower quality quasi-whitened output image patch with the original local image patch, or a whitened version of the original local image patch 110. The comparison is performed by an image quality metric comparison module 330. Other image quality metric values may be considered as alternatives, such as the Structural Similarity Index Measure, SSIM, that compares the local patterns of pixel intensities in an image, and considers the luminance, contrast, and structure of the image. Alternatively, the Mean Squared Error, MSE, may be considered, which calculates the average of the squared differences between the original and the distorted image pixels. Or even theLearned Perceptual Image Patch Similarity, LPIPS, which is based on the features extracted by pretrained image classification networks, such as AlexNet or VGG.
[0057] Next, the randomly generated noise image from the stack providing the best image quality 150 is used as initial input for a trained Denoising Diffusion Probabilistic Model or DDPM 300. The whitened stacked image patch 140 is then fed into the trained DDPM 300 as a conditioning input. The model performs N iterations before it outputs the quasi-whitened output image patch 160, which advantageously has the same pixel count as the target image patch 180. The quasi-whitened output image patch 160 may again be submitted to a whitening step 310 before it is processed by a colouring block 320 that applies colour vector 170 to obtain the final target image patch 180.
[0058] Fig. 4 is a flow chart illustrating the training method according to an aspect of the invention. In a first step, a training dataset for training the DDPM 300 is created. The training dataset comprises a plurality of pairs 230 of a noisy version of a whitened target image patch 240 and a corresponding stacked image patch 130. In a second step, training the DDPM by presenting said training dataset and by calculating a loss value of parameters of the DDPM, and modifying the parameters of the DDPM according to the loss function.
[0059] In the first step, the training dataset for training the DDPM is calculated and obtained. To this end, a plurality of pairs 230 of a noisy version of a whitened target image patch 240 and a corresponding stacked image patch 130 are created or calculated.
[0060] The noisy version of the whitened target image patch 240 of a zone of interest 211 of an image 200 from the target domain D2 is calculated by performing a multiplicative addition 250 of a randomly generated noise image e 270 to a whitened target image patch 220. This whitened target image patch 220 is calculated and obtained by performing a channel-wise mean centering 310 of an image patch 210 of a zone of interest 211 of an image from the target domain D2.
[0061] In a preferred embodiment, the noisy version of the whitened target image patch 240 is obtained by the multiplicative addition 250 of a randomly generated noise image e 270 to the whitened target image patch 220, and is calculated by applying the formula:wherein ytis a noise level value 260 that can be selected by the user between 0 and 1 , wherein e is said randomly generated noise image 270 and y0is said whitened image patch 220 from said target domain.
[0062] The corresponding stacked image patch 130 is obtained by identifying the zone of interest in the source domain that corresponds to the zone of interest of the noisy version of said whitened target image patch 240. The stacked image patch 130 is obtained by stacking the image patch 110 from said corresponding zone of interest of said image from said source domain and a down sampled image patch 120 from said source domain corresponding to an area surrounding said corresponding zone of interest.
[0063] As explained above, the noisy version of the whitened target image patch 240 is combined with the corresponding stacked image patch 130 in a pair 230. The full training dataset comprises a plurality of these pairs 230 and are presented to the DDPM model.
[0064] The DDPM is trained by presenting the paired patches 230 to the DDPM: the paired noisy version of the whitened target image patch 240 and the corresponding stacked image patch 130. Both patches in a pair are corresponding to the same zone of interest. The noisy version of the whitened target image patch 240 of said zone of interest of said image from said target domain is presented to the DDPM as a diffusion input. The stacked image patch 130 of said corresponding zone of interest from said source domain is presented to the DDPM as a conditioning input. The DDPM produces a predicted noise image matrix 280.
[0065] Subsequently a loss function L is calculated using the predicted noise image matrix 280 as a prediction and randomly generated noise image 270 as the ground truth. The parameters of the DDPM are optimized by continuously modifying them according to a calculated loss value through the loss function L, such that in a case where the loss function L satisfies a predetermined condition, a trained DDPM is obtained.
[0066] Whereas, in case where the loss function L does not satisfy the predetermined condition, the above training process is repeatedly performed by continuing to input said pairs of a noisy version of said whitened image patch 240 of said zone of interest of said image from said target domain and said stacked image patch 130 of said corresponding zone of interest from said source domain into the DDPM.
[0067] Fig. 5 shows a suitable computing system 600 enabling to implement embodiments of the method for image-to-image translation of an input image from a source domain to a target domain according to the invention. Computing system 600 may in general be formed as a suitable general-purpose computer and comprise a bus 610, a processor 602, a local memory 604, one or more optional input interfaces 614, one or more optional output interfaces 616, a communication interface 612, a storage element interface 606, and one or more storage elements 608. Bus 610 may comprise one or more conductors that permit communication among the components of the computing system 600. Processor 602 may include any type of conventional processor or microprocessor that interprets and executes programming instructions. Local memory 604 may include a random-access memory (RAM) or another type of dynamic storage device that stores information and instructions for execution by processor 602 and / or a read only memory (ROM) or another type of static storage device that stores static information and instructions for use by processor 602. Input interface 614 may comprise one or more conventional mechanisms that permit an operator or user to input information to the computing device 600, such as a keyboard 620, a mouse 630, a pen, voice recognition and / or biometric mechanisms, a camera, etc. Output interface 616 may comprise one or more conventional mechanisms that output information to the operator or user, such as a display 640, etc. Communication interface 612 may comprise any transceiver-like mechanism such as for example one or more Ethernet interfaces that enables computing system 600 to communicate with other devices and / or systems, for example with other computing devices 681 , 682, 683. The communication interface 612 of computing system 600 may be connected to such another computing system by means of a local area network (LAN) or a wide area network (WAN) such as for example the internet. Storage element interface 606 may comprise a storage interface such as for example a Serial Advanced Technology Attachment (SATA) interface or a Small Computer System Interface (SCSI) for connecting bus 610 to one or more storage elements 608, such as one or more localdisks, for example SATA disk drives, and control the reading and writing of data to and / or from these storage elements 608. Although the storage element(s) 608 above is / are described as a local disk, in general any other suitable computer-readable media such as a removable magnetic disk, optical storage media such as a CD or DVD, - ROM disk, solid state drives, flash memory cards, ... could be used.
[0068] As used in this application, the term "circuitry" may refer to one or more or all of the following: (a) hardware-only circuit implementations such as implementations in only analog and / or digital circuitry and (b) combinations of hardware circuits and software, such as (as applicable): (i) a combination of analog and / or digital hardware circuit(s) with software / firmware and (ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions) and (c) hardware circuit(s) and / or processor(s), such as microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g. firmware) for operation, but the software may not be present when it is not needed for operation.
[0069] This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in a server, a cellular network device, or other computing or network device.
[0070] Although the present invention has been illustrated by reference to specific embodiments, it will be apparent to those skilled in the art that the invention is not limited to the details of the foregoing illustrative embodiments, and that the present invention may be embodied with various changes and modifications without departing from the scope thereof. The present embodiments are therefore to be considered in all respects as illustrative and not restrictive, the scope of the invention being indicated by the appended claims rather than by the foregoing description, and all changes which come within the meaning and range of equivalency of the claims are therefore intendedto be embraced therein. In other words, it is contemplated to cover any and all modifications, variations or equivalents that fall within the scope of the basic underlying principles and whose essential attributes are claimed in this patent application.
[0071] It will furthermore be understood by the reader of this patent application that the words "comprising" or "comprise" do not exclude other elements or steps, that the words "a" or "an" do not exclude a plurality, and that a single element, such as a computer system, a processor, or another integrated unit may fulfil the functions of several means recited in the claims. Any reference signs in the claims shall not be construed as limiting the respective claims concerned. The terms "first", "second", third", "a", "b", "c", and the like, when used in the description or in the claims are introduced to distinguish between similar elements or steps and are not necessarily describing a sequential or chronological order. Similarly, the terms "top", "bottom", "over", "under", and the like are introduced for descriptive purposes and not necessarily to denote relative positions. It is to be understood that the terms so used are interchangeable under appropriate circumstances and embodiments of the invention are capable of operating according to the present invention in other sequences, or in orientations different from the one(s) described or illustrated above.
Claims
CLAIMS1. Computer implemented method for translating a local image patch (110) of an image (100) from a source domain (D1 ) having a first spatial resolution to a target domain (D2) having a second spatial resolution, thereby translating said local image patch (110) into a target image patch (180), the method comprising the steps of:- obtaining said local image patch (110) and a global image patch (120) of the image as a stacked image patch (130), said local image patch (110) comprising image data from a zone of interest (110), said global image patch (120) comprising image data corresponding to an area exceeding said zone of interest (110) and having an identical pixel count as said local image patch (110),- performing a channel-wise mean centering (310) of said stacked image patch (130) to obtain a whitened stacked image patch (140),- performing by a Denoising Diffusion Probabilistic Model (330), DDPM, / V iterations, using a random noise image (150) having said second spatial resolution as an initial input and said whitened stacked image patch (140) as a conditioning input, to obtain a quasi-whitened output image patch (160),- applying a channel-wise colour vector (170) to said quasi-whitened output image patch (160) thereby obtaining said target image patch (180), said channel-wise colour vector (170) characterizing the mean of each channel of an image patch (210) from said target domain or said source domain corresponding to said zone of interest (211 ).
2. The method according to Claim 1 , further comprising determining said random noise image (150) by performing the steps of:- calculating a stack of r randomly generated noise images (155) having an identical spatial resolution as said image patch (210) from said target domain (D2),- performing by a Denoising Diffusion Probabilistic Model (300), DDPM, N / i iterations, wherein / is an integer > 0, thereby using said stack of r randomly generated noise images (155) as an initial input and said whitened stacked image patch (140) as a conditioning input, to obtain a stack of r quasi-whitened output image patches (190),- determining an output patch from said stack of r quasi-whitened output image patches (190) for which an image quality metric value is maximal (330) when comparing said local image patch (110) with said stack of r quasi-whitened output image patches (190),- selecting a random noise image from said stack of r randomly generated noise images (155) that generated said output patch for which said best image quality metric value is calculated as the random noise image (150).
3. The method according to Claim 2, wherein said image quality metric value is a pSNR-value when comparing said local image patch (110) with said stack of r quasi-whitened output image patches (190).
4. Method for adapting an image from a source domain (100) to an image of a target domain (200) for comparison, comprising the steps of:- mapping an area of said image from said source domain to an area of said target domain, and compartmenting said image in small squared patches (110),- adapting said patches of said image from said source domain to said target domain by performing the method of Claim 1 or 2.
5. Method for training said DDPM (300) according to any of the preceding Claims, by:- obtaining a training dataset for training said DDPM, the training dataset consisting of a plurality of pairs (230) ofo a noisy version of a whitened target image patch (240) of a zone of interest (211 ) of an image (200) from said target domain (D2), said whitened target image patch (220) obtainable by performing a channel-wise mean centering (310) of said image patch (210) of said zone of interest of said image from said target domain, and said noisy version (240) obtainable by an multiplicative addition (250) of a randomly generated noise image e (270) to said whitened target image patch (220), and o a stacked image patch (130) of a corresponding zone of interest to the zone of interest of said noisy version of said whitened target image patch (240), wherein said stacked image patch (130) is obtained by stacking of an image patch (110) from said corresponding zone of interest of said image from said source domain and a down sampled image patch (120) from said source domain corresponding to an area surrounding said corresponding zone of interest,- training said DDPM by: o presenting to said DDPM as a diffusion input said noisy version of said whitened target image patch (240) of said zone of interest of said image from said target domain, while o presenting to said DDPM as a conditioning input said stacked image patch (130) of said corresponding zone of interest from said source domain, to produce a predicted noise image matrix (280) o calculating a loss function L using the predicted noise image matrix (280) as a prediction and said randomly generated noise image (270) as a ground truth, o calculating a loss value of parameters of the DDPM through said loss function L, modifying said parameters of the DDPM according to said loss value, and in a case where the loss function L satisfies a predetermined condition, obtaining a trained DDPM, and in a case where the loss function L does not satisfy the predetermined condition, continuing to input said noisy version of said whitenedimage patch (240) of said zone of interest of said image from said target domain and said stacked image patch (130) of said corresponding zone of interest from said source domain so as to repeatedly perform the above training process.
6. The method according to Claim 5, wherein said multiplicative addition of said randomly generated noise image matrix (270) to said whitened image patch (220) is obtained by applying the formula:wherein ytis a noise level value selected between 0 and 1 , wherein e is said randomly generated noise image matrix (270) and y0is said whitened image patch (220) from said target domain.
7. The method according to any of the preceding claims, wherein a channel-wise mean centering (310) step is applied to said quasi-whitened output image patch (160) before applying said channel-wise colour vector (170).
8. The method according to any of the preceding claims 2 to 7, wherein said output patch is determined from said stack of r quasi-whitened output image patches (190) for which an image quality metric value is maximal (320) when comparing a channel-wise mean centered version (310) of said local image patch (110) with a channel-wise mean centered version of said stack of r quasi-whitened output image patches (190).
9. A computer readable storage medium including executable instructions, wherein the instructions, when executed by circuitry, cause the circuitry to perform the method according to any one of the preceding method claims.
10. A data processing apparatus comprising means for carrying out the method of any one of the preceding method claims.11 . A computer program product comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method of any one of the preceding method claims.
12. A trained machine-learning model trained in accordance with the method of Claim 5 or 6.