Image processing method for removing white noise and structured noise
By training the image denoising machine learning model, using the modified copy and guide image of the noisy original image, the decision function is optimized, and the removal of random noise and structured noise in the image is solved, improving image quality and reducing costs.
Patent Information
- Application Number
- CN202380087905.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-20
- Filing Date
- 2023-12-20
- Publication Date
- 2025-07-22
AI Technical Summary
The prior art is difficult to effectively remove random noise and structured noise in images, affecting image quality and interpretation capabilities, and has high cost and resource consumption.
By training the image denoising machine learning model, using the modified copy and guide images of the noisy original image, optimizing the decision function, removing random noise and structured noise, and using technical means such as convolutional neural networks and morphological operations.
Improve image quality, reduce cost and resource consumption, simplify image processing flow, effectively remove different types of noise, and preserve image details and resolution.
Smart Images

Figure CN120359537A_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to the field of data processing in imaging applications, and more particularly to improved techniques for obtaining better imaging results for images used in machine learning. Background Art
[0002] Imaging is a technique for acquiring and interpreting images. Imaging is used in a variety of fields, from microscopy in biological and medical applications to telescope image acquisition in astrophysics. Imaging includes the image acquisition itself, as well as the data interpretation that sets the image into the context of, for example, a learning or research topic. Therefore, image quality is important. Images typically contain noise that can limit the interpretability. Improving the image quality will help improve the interpretability. Data acquisition and interpretation require a certain cost, for example, involving computing resources.
[0003] Therefore, the basic problem of the present invention is to provide techniques for improving image quality and image interpretation while seeking to reduce costs and resources, thereby at least partially overcoming the above disadvantages of the prior art. Summary of the Invention
[0004] In one embodiment of the present invention, a computer-implemented method for training an image denoising machine learning model is provided. The method may include the step of providing at least one modified copy of a noisy original image to the machine learning model. The method may include the step of training the machine learning model at least partially based on at least one modified copy of the noisy original image. The method may include the step of using the machine learning model to predict at least one random noise-corrected image from the noisy original image. The random noise may be at least partially corrected in the random noise-corrected image as compared to the noisy original image. The method may include the step of predicting at least one guidance image using the structured noise corrected from the noisy original image. The structured noise may be at least partially corrected in the guidance image as compared to the noisy original image. Alternatively, the method may include the step of predicting at least one guidance image using the structured noise corrected from a copy of the noisy original image. The structured noise may be at least partially corrected in the guidance image as compared to the noisy original image and / or its uncorrected copy. The method may include the step of optimizing the result of a decision function. The decision function may also be referred to as a comparison function. The decision function may be configured to compare at least a portion of the random noise-corrected image with at least a portion of the guidance image. The method allows for improving the image data quality by improving the noise, preferably the structured noise, as well as the random noise.
[0005] Providing at least one modified copy of the noisy original image may preferably refer to transmitting at least one modified copy of the original image in digital form. Additionally or alternatively, the copy may be stored permanently or even temporarily, e.g., stored on a non-tangible computer-readable medium. Furthermore, the copy may be provided as digital information and / or as some data stream for further processing in a corresponding method and / or imaging system and / or imaging device as described herein.
[0006] Training a machine learning model with at least one modified copy of the noisy image may provide the step of inputting at least one modified copy into a non-tangible computer-readable medium. Wherein, at least one training step of the training method of the machine learning model is performed as described below. The modified copy of the noisy image may be a copy of the noisy image in which the information of at least a single pixel has been modified. The copy may be a digital copy of the noisy original image. The information of several pixels may be modified simultaneously or step by step. The information may also relate to the data of the pixel or a part thereof. The copy may be used as input training data for predicting at least one random noise correction image from the noisy original image using the machine learning model as described below. The random noise correction image may be referred to as a random noise optimization image.
[0007] A random noise optimization term may be derived by training a machine learning model. In doing so, a single pixel value may be predicted based on its surroundings. The pixel value may be used to correct the random noise in the noisy original image or a copy of the noisy original image. The receptive field of a pixel may be defined as a square patch around a certain pixel. The square patch may simulate the area around the pixel in which, for example, an image obtained by an imaging device and / or imaging system can be considered to overlap with the pixels of a detector or sensor, and these pixels may be represented in the form of image pixels. Such a representation may be a direct representation or an indirect representation. In an indirect representation, the number of sensor pixels deviates from the number of image pixels. Therefore, a mapping occurs between the two. In other words, the pixels from the receptive field are considered to be exposed to the recorded image. The recorded image may be mapped to the pixels of the corresponding digital image. The square patch may be referred to as an input patch. A copy of the noisy original image may be decomposed into several patches. Alternatively or additionally, input patches may be copied from the noisy original image to represent at least a part of the noisy original image. "Stitching" the (unmodified) patches together may result in the reconstruction of at least a part of the noisy original image. The reconstructed part may be larger than at least a single patch.
[0008] In an embodiment, a single pixel in a square patch can be modified. Thus, it is possible to mimic a blind spot. In such a blind spot, no information is attached to the single pixel or the current pixel information is erased, resulting in a blind pixel. Alternatively or additionally, the blind pixel can be filled by randomly selecting another pixel from the receptive field of the pixel and / or its surroundings to fill the blind pixel with the corresponding information. Using a corresponding copy of the blind spot as training input can train a machine learning denoising model. Using the modified copy as input training data can prevent the machine learning denoising model from learning the identity operation. A so-called blind spot network can be used to allow prediction of the corresponding random noise optimization term. Alternatively or additionally, a machine learning model can be used to predict the random noise optimization term. In both cases, alternatively or additionally, a masking scheme can replace the value at the center of each input patch with a value randomly selected from the surrounding area (especially the receptive field). In this way, the information of the (central) pixel can be effectively erased and the network can be prevented from learning the identity operation. This is based on the random noise being identically and independently distributed noise. As described above, a copy generated from a noisy original image with erased pixels (even in the form of a single patch) can be a modified copy. When starting from a random noise image, the identity operation is an operation that produces an identical image. The corresponding use of the blind spot network allows the training method to run in a completely unsupervised manner.
[0009] In other words, the input and the target can be derived from a single noisy training image. By directly mapping the value at the center of the input patch to the output, simply extracting the patch as input and using its central pixel as the target to train the network only learns that identity. Thus, as described above, the central pixel information can be replaced by the corresponding pixel data from the surrounding pixels. Accordingly, the surroundings can be described as the receptive field. For example, a clean target from a noise-corrected image is not used for training. To allow effective training of a machine learning model, such as a convolutional neural network (CNN), using the method described above, it is possible to provide simultaneous masking of multiple pixels in a larger training patch and joint calculation of their gradients.
[0010] Alternatively or additionally, a machine learning model can be trained to predict the per-pixel intensity distribution. The posterior distribution of the pixels can be computed. This approach can be combined with a general noise model, such as represented as a histogram. Here, at least one single pixel that is modified to obtain a modified image can be generated by training a machine learning model with the corresponding intensity distribution, histogram, and / or noise model. The corresponding histogram can be referred to as the observation likelihood. Additionally or alternatively, the distribution of possible true pixel intensities can be set as prior. They can be represented by a set of predicted samples. The noise model can be computed from a series of noise-calibrated images. The noise model can characterize the distribution of noise pixels around their corresponding ground truth signal values. The ground truth signal values can be derived from the noise-calibrated images. It simulates the actual signal that can underlie the corresponding image when the random noise contribution is corrected. The noise model can be provided as a set of histograms. Implementing the noise model improves the reconstruction quality of the random noise-corrected images.
[0011] In other words, pixels, especially the central pixel on the receptive field on the input patch, are masked using the mask and / or masking scheme as described above to mimic the blind spot. The masking can be obtained through an all-digital process. This can be the case for all masking processes described here or elsewhere in the entire specification of the embodiments. Different from the above method, the blind pixels can be filled based on a random noise model, a histogram distribution map, and / or the corresponding intensity distribution predicted or hypothesized. In this way, a modified copy is obtained and used as the training input to prevent the machine learning denoising model or any specific network from being used to purely learn identity.
[0012] Alternatively or additionally, the histogram-based noise model can be replaced by a parametric noise model. An example of a parametric noise model that can be used is the Gaussian mixture model (GMM). Additionally, even in the absence of calibration data, a bootstrap step can be used to create a suitable noise model. This can enable the determination of the random noise optimization term to be completely unsupervised. The corresponding method can further include the step of training and applying an unsupervised prediction of the random noise optimization term on the available noise image volume. The images denoised by random noise can be treated as if they were the ground truth for prediction. Thereafter, they can be referred to as pseudo ground truth. The corresponding noise and denoised pixel values can be used to construct a histogram-based noise model or determine a parameter-based noise model.
[0013] Alternatively or additionally, the training can be performed directly on the noisy images. According to an embodiment of the present invention, the invisible pixel values are based on their surroundings, and a machine learning model is trained on all the pixels and surroundings of the training image, rather than just on a single pixel and its corresponding surroundings that have been swapped or modified. According to an embodiment, at least one modified copy can be made, which is shifted in at least one direction of the image in the pixel plane with respect to the pixel plane of the noisy original image. In this way, independent and identically distributed noise, so-called IID noise, can be removed and / or reduced.
[0014] Independent and identically distributed noise, so-called IID noise, generally refers to any type of random noise in the context here. For example, an IID noise source can be caused by, for example, white noise. Non-IID noise, so-called structured noise, can be caused by structured noise originating from the imaging device and / or imaging system. This can be caused by its electronics, imaging optics, and / or other device / setup parameters and / or functional groups and / or components. Non-IID noise can be generated in a reproducible manner by the imaging method and / or device. Non-IID noise can leave noise structures and / or patterns on the image, especially such structures and patterns are the same in each image and are not affected by any imaging situation and / or the object being imaged.
[0015] In other words, such structures / patterns may not be caused by something being imaged, but may be caused at least to some extent by systematic errors in the imaging method and / or by systematic errors in the settings of the imaging device, its optics, and / or its electronics. Correcting structured noise can reduce maintenance costs because no physical hardware adjustment needs to be performed. In some cases, hardware modification may be unmanageable or infeasible. This can be illustrated in the digital correction scheme given here.
[0016] The method can include the step of predicting at least one guidance image using structured noise at least partially corrected from the noisy original image. Alternatively, the method can include the step of predicting at least one guidance image using structured noise corrected from a copy of the noisy original image. Structured noise may not be trainable to be eliminated using any of the methods for correcting random noise described above. In particular, structured noise may not be completely trained to be eliminated. However, the corresponding methods described here can also improve the correction of structured noise. Therefore, the image quality can be further improved. The structured noise optimization term can be derived from the guidance image and compared with the underlying noisy original image. Thus, by this method, it is possible to predict the random noise corrected image and the structured noise corrected image. In this way, it is further possible to allow the generation of two terms of the corresponding decision function.
[0017] This method can further remove / reduce non-IID noise by using a structure-guided processing or other guided image processing methods to remove non-IID noise. In this way, noise-free or less noisy images can be generated. These images can be included as additional constraints for machine learning model training. The images after guided image processing can be considered as pseudo-ground truth (pGT) images to train for noise reduction, especially for IID and non-IID noise. Training the model on the original input images can preserve all details. Therefore, not only can IID noise be removed, but non-IID noise can also be removed based on the guided images created by guided image processing. Guided image processing can further remove structured noise. The shape and resolution can be retained simultaneously, especially without blurring, ringing, overshoot, or undershoot potentially caused by phase shift operations or linear filtering operations.
[0018] For a given data feature and application, the size and / or shape of the surrounding environment can be flexibly defined to effectively remove IID noise. Alternatively, or additionally, the guided image processing method can also be flexibly defined. Combining the complementary advantages of image processing and machine learning can produce the best overall noise removal effect.
[0019] In other words, this method does not require image pairing, which allows for correcting pixel differences between two images, especially between a noisy image and a clean target image, where the clean target image includes a representation of ground truth without noise in particular. In this way, the complexity of preparation is reduced, and the processing of this method can be simplified. In addition, this method and / or any of its steps may not be based on assumptions about the noise level and / or noise distribution. In addition, this method does not necessarily rely on a corresponding noise model. Moreover, it does not require any type of parameter optimization or parameter model, such as Gaussian mixture optimization (GMO) and Gaussian mixture model (GMM). In this way, it also allows for predicting non-independent and non-identically distributed noise (non-IID noise).
[0020] This method can include a step of optimizing the result of a decision function. Among them, the decision function can compare at least a part of the random noise-corrected image with at least a part of the guided image. By comparing at least a part of the random noise-corrected image and at least a part of the guided image, preferably, the basic noise original image and / or a copy of the basic noise original image can be improved with respect to the structured noise in the noise original image. The decision function can include a random noise optimization term and a structured noise optimization term. The decision function can be any function that allows for comparing the two noise optimization terms described in this context.
[0021] The steps of optimizing the result of the decision function can correspond to mathematical operations for determining an optimal value. This optimal value can be the global minimum and / or global maximum of the corresponding function. Additionally, or alternatively, local maxima and / or local minima can be determined.
[0022] In some embodiments, a preset value can be set as the convergence threshold before training. In alternative embodiments, the value can be monitored during training. If the value no longer decreases corresponding to reaching a (local) minimum, or if the value no longer increases corresponding to reaching a (local) maximum, training can be stopped.
[0023] In a practical implementation of the method, the decision function can be a loss function to be minimized. Minimizing the loss function allows for the comparison of noise terms to improve image quality. The loss function can also be referred to as a cost function or an error function. It can be a function that maps the values of one or more variables to a real number, which intuitively represents a certain "cost" associated with an event. Here, the cost can be the correction of the noise in the noisy original image or the generation of its copy to obtain a noise-corrected image. The optimization problem seeks to minimize the loss function. In the case of using the loss function, in some embodiments, a preset value can be set as the convergence threshold before training. In alternative embodiments, the loss value can be monitored during training. If the loss no longer decreases corresponding to reaching a (local) minimum, training can be stopped.
[0024] The decision function can be the loss function as described above, or its opposite form. The opposite form of the loss function can be referred to as a reward function, a profit function, a utility function, or a fitness function, which are several equally descriptive names. In a practical implementation of the method, the reward function can be implemented. In this case, the reward function will be maximized. Both the loss function and the reward function can include terms at several levels from the (algorithm) hierarchy. In the case of using the reward function, in some embodiments, a preset value can be set as the convergence threshold before training. In alternative embodiments, the gain value can be monitored during training. If the gain no longer increases corresponding to reaching a (local) maximum, training can be stopped.
[0025] Additionally, or alternatively, the machine learning model includes at least one convolutional neural network. Convolutional neural networks (CNNs) improve the performance of image processing. CNNs are deep learning algorithms responsible for processing images inspired by the animal visual cortex in the form of a grid pattern. The grid can be a representation of the pixel-based image structure. CNNs are designed to automatically detect and segment specific objects and learn the spatial hierarchy of features from low-level to high-level patterns. They can have three main types of layers: at least one convolutional layer, at least one pooling layer, and at least one fully connected layer (FCL). The convolutional layer can be the first layer, followed by additional convolutional layers or pooling layers. The fully connected layer can be the final layer. For each layer, the complexity of the CNN can increase, and a larger portion of the image can be recognized. The early layers focus on simple features such as color and edges. As the image data progresses between the layers of the CNN, it can start to recognize larger elements or shapes of the object until the expected object is finally recognized. The expected object can be a structure / pattern in the noise map and not necessarily something with a representation in the real world in the image, such as the shape or pattern of an imaging object or an imaging scene. The imaging scene can be multiple objects and / or their arrangement in space. Additionally, the imaging scene can evolve over time.
[0026] In a practical implementation of the method, the loss function can be of the type given in formula (1): 。
[0027] In formula (1), I t is the noisy original image, I’ t is the predicted image, I t_n is the guidance image, and 𝜆 is the weighting factor. The loss function includes a random noise optimization term as the first term and a structured noise optimization term as the second term. The loss function allows comparison of the two terms to achieve noise correction for the noisy original image and / or a copy of the noisy original image.
[0028] According to one aspect of the present invention, the predicted image can be a random noise corrected image, and the random noise corrected image can be fed back into the computer-implemented training method to iteratively optimize the result of the decision function. Alternatively, or additionally, the random noise corrected image can be fed back into the computer-implemented training method to train the machine learning model. This allows for the implementation of an iterative training method to further improve the performance of the method. In the corresponding embodiment, the random noise corrected image can be iteratively fed back into the training method. In this way, the predicted image can be iteratively improved because the random noise correction can also be iteratively improved. This further improves the output, i.e., the (random) noise corrected image.
[0029] Alternatively, or additionally, a guidance image can be fed back into a computer-implemented training method, where the guidance image can be a structured noise-corrected version of the noisy original image to iteratively optimize the result of the decision function. Alternatively, or additionally, the guidance image can be fed back into a computer-implemented training method to train a machine learning model. In corresponding embodiments, the image corrected for structured noise (the corresponding guidance image) can be iteratively fed back, in particular replacing the noisy original image as the input image, to iteratively improve the random noise correction as well. The structured noise can leave no imprint on the feedback image, in particular not to the same extent as on the noisy original image. Thus, any potential interference from structured noise with random noise can be reduced. It is particularly meaningful when the random noise can be at the same level as some structural true features of the imaged object or the imaged scene. In these embodiments, in particular, the situation arises where the input image and the random noise-corrected image with the guidance image are iteratively exchanged.
[0030] According to some embodiments of the present invention, there is provided a computer-implemented method for training a machine learning model for denoising random noise images. The method can include the step of providing at least one noisy original image. Additionally, it can include the step of providing at least one non-identical copy of the noisy original image. Wherein, the at least one non-identical copy can be shifted by at least one pixel relative to the noisy original image. The method can include the step of feeding the at least one non-identical copy into the machine learning model to be trained. The method can include the step of training the machine learning model at least partially based on the at least one non-identical copy. Additionally, there can be provided the step of using the machine learning model to predict at least one random noise-corrected image from the at least one non-identical copy.
[0031] The provision of at least one noisy original image has been described elsewhere in the entire specification. For the sake of brevity, it will not be repeated here. This definition can apply here. This definition can be further applied to the provision of at least one non-identical copy of the noisy original image. The non-identical copy can be a copy of the noisy original image in which the information of at least a single pixel has been modified, as described throughout the specification. The non-identical copy can be generated by shifting the copy by at least one pixel relative to the noisy original image. The complete image can be shifted by a single pixel array row. Additionally, or alternatively, taking into account the boundary conditions, the image can be truncated relative to a single pixel at least in the shifting direction or the reverse shifting direction. Additionally, the original noisy original image can be truncated accordingly. This results in a pair of corresponding images with the same pixel size and pixel dimensions. Additionally, truncation can avoid any artifacts related to the boundaries.
[0032] At least one non-identical copy can be fed into a machine learning model for training. Accordingly, feeding at least one non-identical copy into a machine learning model involves, in this context, transferring the data of the copy to the machine learning model. Training the machine learning model can be based at least in part on at least one non-identical copy. The non-identical copy can include several differences relative to the noisy original image, such as it can be shifted by an entire row of pixels. Accordingly, the training itself can be as described throughout the specification. Shifting an image by a single pixel can result in a shift of at least one complete row of pixels (pixel array), and / or by shifting the complete image by the corresponding row of pixels. Based on this, from the noisy original image or its identical copy to the corresponding shifted copy, a large number of pixel changes can occur. Therefore, more than just one pixel can be masked and / or modified. Thus, since a machine learning denoising method configured to train for correcting random noise can train simultaneously on all pixels and their surroundings, there is no need to feed multiple patches into the training method. In this way, the training and operation speed can be increased. By using a machine learning model, an image optimized for at least one random noise can be predicted from at least one non-identical copy. This image can be available for further processing for use in the training method described herein.
[0033] In a practical implementation of the method, eight non-identical copies of the noisy original image can be made. Each non-identical copy can be shifted by at least one pixel in the direction of the noisy original image and fed into the machine learning model to be trained. Accordingly, the non-identical copies can be different from each other and also different from the noisy original image. Alternatively, or additionally, the machine learning model can include at least one convolutional neural network. Further alternatively, or additionally, structured noise can be corrected at least in part based on a guidance image. This can increase the operation speed as well as the training speed. In addition, there is no need to provide multiple patches, where each patch differs at least in a single pixel. Each of the following shifted non-identical copies can be one of the eight copies provided to the training method: shifted vertically upwards, shifted vertically downwards, shifted horizontally to the left, shifted horizontally to the right, shifted diagonally upwards to the left, shifted diagonally upwards to the right, shifted diagonally downwards to the left, shifted diagonally downwards to the right. Accordingly, since all pixels in the image are shifted relative to the copy of the noisy original image or its unshifted copy, multiple non-identical pixels are provided for training. The machine learning denoising model can train simultaneously on all pixels and their surroundings. This can increase the operation speed of the training and reduce the necessary resources, since there is no need to provide multiple patches with single-pixel modifications.
[0034] In a further embodiment of the present invention, a computer-implemented image denoising method may be provided. The method may include the step of providing at least one exact copy of the noisy original image. It may also include the step of feeding the at least one exact copy into a trained machine learning model. The machine learning model has preferably been trained based on the above training method. The method may include the step of restoring an output image with optimized noise. The method allows for processing image denoising and may allow for the application of a trained machine learning model. In doing so, it would be unnecessary to provide a noise model. This increases the performance of the method as the model does not have to be obtained, stored, and / or provided before and / or during the operation. Additionally, assumptions about the noise distribution and noise level may not have to be provided. This improves the output result. Furthermore, since the image quality can be improved, the interpretability of the denoised image is enhanced.
[0035] In a practical implementation of the method, eight exact copies of the noisy original image may be made. The number of copies may correspond to the respective number of non-identical copies provided in the training method as described above. The number of copies can be referred to as the potency of the model. Based on the guidance image, the structured noise can be at least partially corrected.
[0036] Alternatively, or additionally, the machine learning model may include at least one convolutional neural network. The convolutional neural network has been introduced in detail above. Its respective advantages and characteristics also apply here. In particular, the convolutional neural network architecture is the same as the training method.
[0037] According to certain embodiments of the present invention, the guidance image may be predicted according to at least one of the following steps: structured noise removal, random noise removal, or normalization. Determining the structured noise may further provide for corresponding random noise removal. Any of the above methods may predict the guidance image and may also implement and configure the guidance image. The guidance image can be used as a basis for determining and correcting the structured noise, as detailed above.
[0038] According to certain embodiments of the present invention, normalization may include at least one of the following: increasing the signal intensity while suppressing noise; adjusting the normalization scale such that the signal intensity increases; adjusting the normalization scale such that the signal intensity increases as much as possible without saturation; or adjusting the normalization scale such that the normalization scale is stabilized between the 2nd and 99.98th percentiles. Normalization can avoid saturation of certain pixels and / or avoid so-called truncated pixels, which are set to a certain level in order to implement data truncation at a certain intensity minimum and / or maximum. Therefore, normalization can improve the image quality and suppress artifacts.
[0039] According to certain embodiments of the present invention, structured noise can be removed according to different sizes of structural elements in morphological operations. In mathematical morphological operations, a structural element is a shape used to detect or interact with a given image. From it, conclusions can be drawn about how the shape matches or misaligns with the shapes in the image. Morphological operations can be dilation, erosion, opening, and closing operations, or a hit-or-miss transform. Alternatively, or additionally, structured noise can be removed according to kernels of different sizes in morphological operations. Alternatively, or additionally, structured noise can be removed according to morphological operations with at least one straight line as a structural element and / or as a kernel. The straight line can perform row-by-row operations on the image and execute corresponding morphological operations in the horizontal and / or vertical directions. Alternatively, or additionally, structured noise can be removed according to the opening and / or closing operations applied as morphological operations to identify the horizontal and / or vertical components of the image. Alternatively, or additionally, during the process of removing structured noise, an offset can be added to or subtracted from the image. The offset can be based on the overall dynamic range of the data. The offset can be determined empirically to restore the illuminance and brightness to a level suitable for the application. All these functions described here can achieve structured noise correction, thereby improving the image quality. In addition, at least one vertical component and / or at least one horizontal component of the input image can be identified to correct the respective structured noise.
[0040] For example, performing an opening operation can remove small structural features ("speckles") in the image. In the opening operation, erosion can first be applied to remove small speckles, and then dilation can be applied to restore the size of the original object. The opening operation is idempotent. The closing operation can be used to close the "holes" inside an object or connect components together. In the closing operation, erosion can be performed after dilation. The closing operation is idempotent. These operations can be used to identify structured noise in the image and separate it from random noise. The opening operation can be applied to peaks, maxima, or data peaks, while the closing operation can be applied to depressions, minima, or data valleys. As the size of the structural element increases and / or the size of the kernel increases, the sensitivity may decrease. Therefore, the method can start with smaller structural elements and / or smaller kernel sizes. The method can first attempt to remove bright areas and then attempt to remove dark areas. The number of iterations can be limited to a preset value. This can reduce the risk that as the number of iterations increases, the risk of cutting off relevant structured data also increases.
[0041] According to another aspect of the present invention, a data processing device can be provided. Such a data processing device can be a computer device, or a control unit of an imaging device or imaging system. They can include means for performing any of the above methods.
[0042] Embodiments of the present invention also provide a data processing device, in particular a computer device or a control unit for a microscope and / or a telescope, which includes means for performing any of the methods disclosed herein. The present invention also provides a computer program with program code that, when run on a processor, can be used to perform any of the methods disclosed herein.
[0043] Specific aspects of the present invention also provide a data set that can be used for a machine learning model, in particular a training data set, a validation data set, and / or a training data set. The data set may include a plurality of microscope images or their reference images. These images can be obtained using any of the methods disclosed herein.
[0044] Accordingly, embodiments can be based on or allow the use of a machine learning model or a machine learning algorithm. Machine learning can involve algorithms and statistical models that a computer system can use, which do not require the use of explicit instructions, but rely on models and inferences to perform specific tasks. For example, in machine learning, data transformations inferred from the analysis of historical and / or training data can be used instead of rule-based data transformations. For example, a machine learning model or a machine learning algorithm can be used to analyze the content of an image. In order for a machine learning model to analyze the content of an image, training images can be used as input and training content information can be used as output to train the machine learning model. By training a machine learning model using a large number of training images and / or training sequences (such as words or sentences) and associated training content information (such as labels or annotations), the machine learning model "learns" to recognize the content of the image, and thus the machine learning model can be used to recognize the content of an image not included in the training data. The same principle also applies to other types of sensor data: by training a machine learning model using training sensor data and expected outputs, the machine learning model "learns" the transformation between the sensor data and the outputs, and can be used to provide outputs based on non-training sensor data provided to the machine learning model. The data provided (such as sensor data, metadata, and / or image data) can be preprocessed to obtain feature vectors, which are used as input to the machine learning model.
[0045] A machine learning model can be trained using training input data. The above example uses a training method called "supervised learning". In supervised learning, the machine learning model is trained using multiple training samples, where each sample can include multiple input data values and multiple expected output values, i.e., each training sample is associated with an expected output value. By specifying the training samples and the expected output values, the machine learning model "learns" which output value to provide based on input samples similar to those provided during training. In addition to supervised learning, semi-supervised learning can also be used. In semi-supervised learning, some of the training samples lack the corresponding expected output values. Supervised learning can be based on supervised learning algorithms (e.g., classification algorithms, regression algorithms, or similarity learning algorithms). Classification algorithms can be used when the output is restricted to a finite set of values (categorical variables), i.e., classifying the input into one of the finite set of values. Regression algorithms can be used when the output can have any numerical value (within a certain range). Similarity learning algorithms can be similar to classification and regression algorithms, but are based on learning from examples using a similarity function that measures the similarity or relatedness of two objects. In addition to supervised or semi-supervised learning, unsupervised learning can also be used to train a machine learning model. In unsupervised learning, (only) the input data can be provided, and unsupervised learning algorithms can be used to find the structure in the input data (e.g., by grouping or clustering the input data to find commonalities in the data). Clustering is the assignment of input data including multiple input values to subsets (clusters) such that the input values within the same cluster are similar according to one or more (predefined) similarity criteria, but are not similar to the input values included in other clusters.
[0046] Reinforcement learning is a third group of machine learning algorithms. In other words, reinforcement learning can be used to train a machine learning model. In reinforcement learning, one or more software actors (called "software agents") are trained to take actions in an environment. Based on the actions taken, a reward is calculated. Reinforcement learning is based on training one or more software agents to select actions to increase the cumulative reward, so that the software agents perform better on a given task (an increase in reward is evidence).
[0047] In addition, some techniques can be applied to some machine learning algorithms. For example, feature learning can be used. In other words, a machine learning model can be trained at least in part using feature learning, and / or a machine learning algorithm can include a feature learning component. Feature learning algorithms (which can be called representation learning algorithms) can preserve the information in their input, but can also transform it in a useful way, typically as a preprocessing step before performing classification or prediction. For example, feature learning can be based on principal component analysis or cluster analysis.
[0048] In some examples, anomaly detection (i.e., outlier detection) can be used, which aims to identify input values that are significantly different from most of the input or training data and thus raise suspicion. In other words, a machine learning model can be trained at least in part using anomaly detection, and / or a machine learning algorithm can include an anomaly detection component.
[0049] In some examples, a machine learning algorithm can use a decision tree as a prediction model. In other words, a machine learning model can be based on a decision tree. In a decision tree, an observation about an item (e.g., a set of input values) can be represented by a branch of the decision tree, while the output value corresponding to that item can be represented by a leaf of the decision tree. Decision trees can support both discrete values and continuous values as output values. If discrete values are used, the decision tree can be represented as a classification tree, and if continuous values are used, the decision tree can be represented as a regression tree.
[0050] Association rules are another technique that can be used in machine learning algorithms. In other words, a machine learning model can be based on one or more association rules. Association rules are created by identifying relationships between variables in a large amount of data. A machine learning algorithm can identify and / or utilize one or more relationship rules that represent knowledge derived from the data. These rules can be used to store, manipulate, or apply knowledge.
[0051] Machine learning algorithms are generally based on machine learning models. In other words, the term "machine learning algorithm" can represent a set of instructions that can be used to create, train, or use a machine learning model. The term "machine learning model" can represent a data structure and / or set of rules that represent the learned knowledge (e.g., based on training performed by a machine learning algorithm). In an embodiment, the use of a machine learning algorithm can imply the use of an underlying machine learning model (or underlying machine learning models). The use of a machine learning model can imply that the machine learning model and / or the data structure / set of rules that is the machine learning model is trained by a machine learning algorithm.
[0052] For example, the machine learning model can be an artificial neural network (ANN). An ANN is a system inspired by biological neural networks, such as those in the retina or the brain. An ANN includes multiple interconnected nodes and multiple connections between the nodes, namely the so-called edges. There are generally three types of nodes: input nodes that receive input values, hidden nodes that are (only) connected to other nodes, and output nodes that provide output values. Each node can represent an artificial neuron. Each edge can transmit information from one node to another. The output of a node can be defined as a (non-linear) function of its inputs (e.g., the sum of its inputs). The inputs of a node can be used in the function based on the "weights" of the edges or the nodes providing the inputs. The weights of the nodes and / or edges can be adjusted during the learning process. In other words, the training of an artificial neural network can include adjusting the weights of the nodes and / or edges of the artificial neural network, i.e., achieving the desired output for a given input.
[0053] Alternatively, the machine learning model can be a support vector machine, a random forest model, or a gradient boosting model. A support vector machine (i.e., a support vector network) is a supervised learning model with an associated learning algorithm that can be used to analyze data (e.g., in classification or regression analysis). A support vector machine can be trained by providing an input with multiple training input values belonging to one of two classes. A support vector machine can be trained to assign new input values to one of the two classes. Or, the machine learning model can be a Bayesian network, which is a probabilistic directed acyclic graph model. A Bayesian network can use a directed acyclic graph to represent a set of random variables and their conditional dependencies. Or, the machine learning model can be based on a genetic algorithm, which is a search algorithm and heuristic technique that mimics the process of natural selection.
[0054] The possible use cases of the results generated by the embodiments of the present invention are multi-scenario. For example, the model can be trained to detect specific cell phenotypes, or target specific organelles, and compare them at different signal-to-noise ratios. This helps to create a model for evaluating phototoxicity detection, or a simpler model than this by using simple cell detection and segmentation (measurement metrics such as cell counting and nuclear area, or whether certain markers are shown).
[0055] According to another aspect of the present invention, an imaging device or an imaging system can be provided. Such an imaging device and / or imaging system can in particular be a microscope or a telescope. They can be configured for any of the above methods. Regarding the imaging device or imaging system, the device or system can provide the same technical advantages as the above methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] The present disclosure can be better understood by referring to the following drawings.
[0057] Figure 1: Schematic diagram of the workflow for training an image denoising machine learning model.
[0058] Figure 2 : Schematic diagram of the workflow for image denoising to remove random noise.
[0059] Figure 3 : Schematic diagram of the workflow for image processing-guided noise removal to remove structured noise.
[0060] Figure 4A : As Figure 3 shown, a schematic diagram of performing 1D opening and closing operations on structured noise using the first and second types of structuring elements in the workflow.
[0061] Figure 4B : As Figure 3 shown, a schematic diagram of performing 1D opening and closing operations on structured noise using the third and fourth types of structuring elements in the workflow.
[0062] Figure 4C : In the workflow as Figure 3 shown, a schematic diagram of the structure reconstructed from the noise image based on the 1D opening and closing operations on the structured noise as shown in Figure 4A and Figure 4B shown.
[0063] Figure 5 : Schematic diagram of the workflow for performing the opening and closing operations shown in Figure 4.
[0064] Figure 6 : As Figure 3 shown, a schematic diagram of the iterative opening and closing operation workflow configured for IID noise removal in the workflow.
[0065] Figure 7 : Schematic diagram of a system that can implement the embodiments of the present invention. Detailed Embodiments
[0066] Now, exemplary embodiments of the present invention will be described, which provide a workflow for training an image denoising machine learning model, a workflow for image denoising to remove random noise, and a workflow for image processing-guided noise removal to remove structured noise.
[0067] Figure 1 A schematic diagram showing the workflow of training an image denoising machine learning model 106 is shown. The workflow can be implemented in a computer-implemented method 100 configured to train an image denoising machine learning model 106. The computer-implemented training method 100 can at least provide the following steps.
[0068] In the step, at least one modified copy 104a... 104h of the noisy original image 102 is provided to the machine learning model 106. The machine learning model 106 is trained with at least one modified copy 104a... 104h of the noisy image 102. The machine learning model 106 is used to predict at least one random noise corrected image 108 from the noisy original image 102.
[0069] In addition, at least one guidance image 114 is predicted using the structured noise prediction 112 corrected from the copy 110 of the noisy original image 102. Based on the copy 110 of the noisy original image 102, the guidance image 114, and the random noise optimized image 108, two terms of the decision function 116 are obtained and filled to optimize the result of the decision function 116.
[0070] The optimization according to the embodiment is any suitable mathematical optimization that allows reaching a local optimum value. According to the embodiment, a global and / or local optimum value may be preferred. The optimum values can all be global and / or local maximum or minimum values. According to the embodiment, a predefined threshold of the result of the decision function 116 can be set to evaluate the parameters of the optimization routine. Alternatively, the optimization can be monitored to check whether the values (loss or gain values) are constantly changing and thus not reaching the best (minimum or maximum) value.
[0071] In an embodiment, the first term of the decision function 116 corresponds to at least a part of the random noise optimized image 108, while the second term corresponds to at least a part of the guidance image 114. In this way, at least a part of the random noise corrected image 108 can be compared with at least a part of the guidance image 114 to optimize the decision function 116.
[0072] In an embodiment, the decision function 116 is a loss function to be minimized. In such an embodiment, the optimization can be implemented as a routine to reach a global or local minimum.
[0073] Alternatively, in other embodiments, a gain function can be implemented as the decision function. In such an embodiment, the optimization can be implemented as a routine to reach a global or local maximum.
[0074] Examples of loss functions can be of the following types: , where I t is the noisy original image 102, I’ t is the predicted image 108, I t_n is the guidance image 114, and 𝜆 is a weighting factor. The predicted image 108 can be a random noise corrected image.
[0075] In an embodiment, the machine learning model 106 is a convolutional neural network (CNN). This improves performance as well as learning and training efficiency. The CNN is trained to provide at least one noisy raw image 102 such that at least one non-identical copy 104a... 104h of the noisy raw image 102 is obtained by shifting at least one pixel relative to the noisy raw image 102. At least one non-identical copy 104a... 104h is fed into the machine learning model 106 for training. Thus, the machine learning model 106 is trained on at least one non-identical copy 104a... 104h. Since the copies are non-identical to each other, the CNN does not learn identity. Thus, the CNN learns to predict at least one randomly noisy optimized image 108 from at least one non-identical copy 104a... 104h. The predicted image can be fed back into the computer-implemented training method 100 to iteratively optimize the result of the decision function 116. Alternatively, or additionally, the predicted image 108 can be used to further train the CNN. The training of the CNN described here applies to other machine learning models in a corresponding manner. The features and steps described here apply to other machine learning models accordingly and can thus be transferred.
[0076] In an exemplary embodiment, eight non-identical copies 104a... 104h of the noisy raw image 102 are made. Each copy 104a... 104h is shifted at least one pixel in the direction of the noisy raw image 102. In particular, since all eight copies 104a... 104h are different from each other, the exemplary embodiment includes setting each subsequently shifted non-identical copy 104a... 104h as one of the eight copies 104a... 104h provided to the training workflow 100a. One copy is shifted vertically upwards, one copy is shifted vertically downwards, one copy is shifted horizontally to the left, one copy is shifted horizontally to the right, one copy is shifted diagonally upwards to the left, one copy is shifted diagonally upwards to the right, one copy is shifted diagonally downwards to the left, and one copy is shifted diagonally downwards to the right. This results in a plurality of non-identical pixels being provided for training since all pixels in the image will be shifted relative to the copy of the noisy raw image or its unshifted copy. The shifted images 104a... 104h can be used as a training set for the training workflow 100a. The training workflow 100a itself can be a computer-implemented method or a computer-implemented subroutine. The corresponding training of the training workflow 100a can be performed to train a machine learning model to correct random noise, so-called identically and independently distributed noise, abbreviated as IID noise.
[0077] In other embodiments, a receptive field mask with blind spots can be used to swap only a single pixel to eliminate pixel information, thereby obtaining non-identical copies. Alternatively, or additionally, the central pixel of the receptive field can be replaced with pixel information from the surrounding receptive fields of the pixels to be swapped. In these embodiments, copies of the noisy original image 102 can be based on patches of the noisy original image 102. Thus, multiple non-identical copies 104a... 104h can also be obtained. Using the non-identical copies 104a... 104h obtained by the shifting method described above, multiple higher-order variant pixels can be generated. Among them, all pixels of the image are changed and are thus simultaneously used to train the machine learning model 106. This reduces the complexity of data processing and improves the performance of the method. In addition, the method does not have to set any assumptions in advance and does not have to load and / or save routines regarding the corresponding noise model. This further improves the performance of the method because fewer resources have to be allocated. According to the description provided above, the corresponding machine learning model can be a CNN.
[0078] In Figure 3 a predictive workflow routine 112 for predicting a guidance image 114 is shown and described in detail. Among them, it is illustrated how to at least partially identify a structure 430 (see Figure 5 ) using the workflow shown and described for it to obtain a structured noise-corrected image as the guidance image 114. Figure 4C ).
[0079] Figure 2 The execution of a method 200 for correcting random noise using a trained machine learning model 206 is shown. In Figure 2 , the computer-implemented image denoising method 200 is illustrated by the following steps. In the step, at least one identical copy 204a... 204h of the noisy original image 202 is provided to feed the at least one identical copy 204a... 204h into the trained machine learning model 206. This training is performed as shown and described in the training workflow 100a as shown relative to Figure 1 . As described relative to Figure 1 , the training workflow 100a is capable of feeding eight non-identical copies 104a... 104h of the noisy original image 102 into the machine learning model 106. In the computer-implemented denoising method 200, the number of identical copies 204a... 204h corresponds to the number of non-identical copies 104a... 104h in the training workflow 100a. The two methods are equally effective. In Figure 2In the illustrated embodiment, eight exact copies can be provided for the denoising method 200. In this way, the same architecture and subroutines can be used for training and execution. The computer-implemented method 200 can recover the output image 208 with optimized noise. This output image 208 with optimized noise can be fed back into the training method 100 and / or the training workflow 100a to iteratively improve the training. Additionally or alternatively, in some embodiments, it can be used as a ground truth (GT) or a pseudo-ground truth (pGT). As described above, the machine learning model 106 can be a CNN. The corresponding parts of the recitation are omitted, and more details can be found in the description of the CNN with respect to Figure 1 the description of the CNN.
[0080] Figure 3 A prediction workflow 112 for predicting a guidance image 114 is shown, which can be applied to any of the above methods. In particular, the prediction workflow 112 can form part of the method 100 configured to train the denoising machine learning model 106. In other embodiments, the prediction workflow 112 can be an independent prediction workflow routine 300, executed independently of any random noise optimization workflow 100a. In any embodiment, based on structured noise removal 302, the guidance image 114 can be predicted in the prediction workflow 112 or in the independent prediction workflow 300. The structure 430 (see Figure 4C ) can be superimposed with the random noise 412 (see Figures 4A to 4C ), and the random noise 412 is superimposed as shown in Figures 4A to 4C . A detailed description is given below, with reference to the corresponding sections.
[0081] Additionally, or alternatively, random noise removal 304 can be applied. As described above, the random noise can be corrected based on the trained computer-implemented method. Additionally, the noise can cover the underlying structure 430 in the image, or it can itself exhibit the structure 430. In the latter case, it is called structured noise or non-identically and non-independently distributed noise, abbreviated as non-IID noise. Figures 4A to 4Cillustrates how a structure 430 can be isolated from an image, where the structure 430 is covered by random noise 412. The identified structure 430 can be structured noise, i.e., so-called IID noise. Alternatively, the structure 430 can be an actual structure, where "actual" refers to the fact that the image shows the structure 430 existing in the object. Alternatively, the structure can be an "artificial" (in the sense of not actual but artificially induced) structure based on noise. This can occur if the structure does not have a representation in the actual imaged object or actual imaging scene, but is caused by non-IID noise from, for example, electronic devices, optical devices, and / or any image acquisition method steps and / or data operation method steps. Further details will be described with respect to Figures 4A to 4C Describe further details.
[0082] In Figure 3 , in addition to or in place of the above steps of the prediction workflow 112 or the independent prediction workflow 300, normalization 306 can be applied to the image 102 or a version of the image 102 preprocessed by the prediction workflow 112 or the independent prediction workflow 300. Normalization 306 can include increasing the signal intensity while suppressing noise, adjusting the normalization scale such that the signal intensity increases, adjusting the normalization scale such that the signal intensity increases as much as possible without saturation, or adjusting the normalization scale such that the normalization scale is fixed to the 2nd to 99.98th percentile. More specifically, the normalization scale can be fixed to the 5th to 95th percentile, and even more specifically, the normalization scale can be fixed to the 10th to 90th percentile. This optimizes the input data to improve visualization and / or reduce any artifacts from data prediction, especially in cases where the guidance image 114 is iteratively used in further method steps.
[0083] In a particular embodiment of the prediction workflow 112 or the independent prediction workflow 300, structured noise removal 302 occurs before random noise removal 304. Random noise removal can occur before normalization 306. This cascaded manner of processing the noisy original image 102 allows first removing the random noise 412 that superimposes on the already recognizable structure 430 ( Figures 4A to 4C shown), to reduce the processing resources required for any further steps. Thereafter, the random noise or IID noise can be compensated to further improve the image quality. Normalization 306 allows obtaining a guidance image 114 suitable for the corresponding application and / or optimized for interpretation. In addition, artifacts generated by any of the above previous method steps can be compensated in this way.
[0084] Figure 4A Schematically shows a structure isolation workflow 400, which workflow 400 can be implemented as in Figure 3Part of the structured noise removal 302 in the illustrated prediction workflow 112 or in the independent prediction workflow 300. In the structure isolation workflow 400, a 1D opening operation 410 can be performed using a first type of structural element 414. A corresponding 1D closing operation 420 can follow the 1D opening operation 410. The closing operation 420 can use a second type of structural element 416. The opening operation 410 and the closing operation 420 can be performed on the random noise 412 of the superimposed structure 430 (see Figure 4C ), where the structure 430 is the edge in the image as shown in Figure 4C . The opening operation 410 and the closing operation 420 can follow the corresponding random noise 412 of the superimposed structure 430. The peaks or (local) maxima are processed using the opening operation 410. The valleys or (local) minima are processed using the closing operation 420. Thus, the corresponding concatenation of the closing operation 420 after the opening operation 410 is based on the specifically illustrated example. In addition, the corresponding example represents a "cut" through a single row of pixels in the x-axis direction and the corresponding intensity distribution diagram in the y-axis direction. Correspondingly, the corresponding operations can also be performed in the 2D image plane. Among them, the corresponding 2D structures can be isolated, and they are actual structures or structured noise. Figures 4A to 4C The corresponding examples illustrated in
[0085] Figure 4B are selected for explanatory purposes and do not mean that their implementation is limited to the 1D case. Figure 4A shows the next step of the structure isolation workflow 400 as shown in
[0086] Figure 4C and described above. Another round of 1D opening operation 410 and closing operation 420 can be applied. The opening operation 410 can use a third type of structural element 418, where the pixel dimension in the x-direction increases. Thus, the structural element 418 can cover a larger structure. The closing operation 420 uses a fourth type of structural element 422, where the pixel dimension in the x-direction increases. Thus, the structural element 422 can cover a larger structure. Figure 3 shows a schematic diagram of the structure 430, which can be based on iterative execution and described above (especially performed in the prediction workflow 112 or the independent prediction workflow 300 as shown in Figure 4A and Figure 4B ). The iterative 1D opening operation 410 and 1D closing operation 420 as shown in Figure 3In the reconstruction in ( ). The dilation operation 410 and the erosion operation 420 allow for incrementally identifying a structure 430, which can be the underlying random noise 412 that superimposes the corresponding structure 430 above. The identified structure 430 can be an actual structure in the image. Additionally, it can be a structured noise artifact that mimics a structure but exists independently of any imaging object or imaging scene in the original noise image 102. Thus, it can also be a structured noise feature. In the latter case, the result of the structure isolation workflow 400 can be used for corresponding structured noise removal 302 in the prediction workflow 112 as shown in Figure 3 or in the independent prediction workflow 300 to compensate for the structured noise. The difference in interpretation as to whether it might be structured noise or an actual feature can be arbitrarily set based on the maximum increment size of the structure elements 414, 416, 418, 422. The maximum increment size of the structure elements can be set based on the corresponding trade-off between the imaging object and the acceptable resolution for the actual features that are isolated and potentially subtracted as structured noise.
[0087] Figures 4A to 4C The operations shown can be summarized as follows. In the morphological operation 500 (see Figure 5 and the description below it), based on the different sizes of the structure elements 414, 416, 418, 422, the structured noise can be isolated from the background (random) noise 412. The morphological operation 500 can be any one of the dilation operation 410 and / or the erosion operation 420. Instead of or in addition to the structure elements 414, 416, 418, 422, kernels with different sizes can be used. The structured noise can be removed based on the morphological operation 500 that applies at least one straight line as the structure elements 414, 416, 418, 422 and / or the kernel. Alternatively, or in addition, the structured noise can be removed based on the dilation operation 410 and / or the erosion operation 420 applied as the morphological operation 500 and the potential identification of the structure 430 of the underlying (random) noise 412 as the structured noise. This can identify the horizontal and / or vertical components of the images 102, 110 (see Figure 1 and the description above it) and the corresponding structure 430 therein. Furthermore, during the structured noise removal 302 (further details see Figure 3 and Figure 5 ), an offset 534 can be added to or subtracted from the images 102, 110 (see Figure 5 and the description below it). This can improve the image output to improve the interpretation of the image content.
[0088] Figure 5 Shows a schematic diagram of the workflow for performing the dilation operation 410 and / or the erosion operation 420 as shown in FIG. 4 and above and compensating for the structured noise based on the result of the corresponding morphological operation 500.Figure 5 The workflow shown represents a schematic layout of the morphological operation 500, which is at least partially represented by the opening operation 410 and / or the closing operation 420. The noisy original image 102 is provided 502 to the opening operation 410 and / or the closing operation 420. As with Figure 4A and Figure 4B described, the opening operation 410 and / or the closing operation 420 are applied iteratively. An iterative method is further illustrated and described with respect to Figure 6 below.
[0089] The morphological operation 500 may include two branches 500a, 500b. One branch 500a may provide structured noise compensation for the horizontal component I horz 506. The other branch 500b may provide structured noise compensation for the vertical component I vert 508.
[0090] The noisy original image 102 may also be provided 504 to either of the branches 500a, 500b of the morphological operation 500 to be transmitted to the subtraction operations 510a, 510b of the respective branches 500a, 500b. In the subtraction operation 510a of the horizontal branch 500a, I horz 506 may be subtracted from the noisy original image 102 provided 504 to the subtraction operation 510a of the horizontal branch 500a. This results in an image I horz ’ 512a, in which the horizontal structured noise has been corrected.
[0091] In the subtraction operation 510b of the vertical branch 500b, I vert 508 may be subtracted from the noisy original image 102 provided 504 to the subtraction operation 510b of the vertical branch 500b. This results in an image I vert ’ 512b, in which the vertical structured noise has been corrected.
[0092] The horizontally corrected image 512a and the vertically corrected image 512b may be provided to a minimum operation 520 and a maximum operation 530. In the minimum operation 520, especially on a pixel-by-pixel basis, the minimum I min 522 may be calculated from the comparison of the horizontally corrected image 512a and the vertically corrected image 512b. For example, from each corresponding pixel pair from the horizontally corrected image 512a and the vertically corrected image 512b, the minimum pixel value may be selected, and this minimum value may be combined as the minimum image I min 522.
[0093] In the maximum operation 530, especially on a pixel-by-pixel basis, the maximum value I can be calculated from the comparison of the horizontally corrected image 512a and the vertically corrected image 512b. max 532. For example, from each corresponding pixel pair from the horizontally corrected image 512a and the vertically corrected image 512b, the maximum pixel value can be selected and combined into the maximum value image I max 532.
[0094] Minimum value I min 522 and the maximum value I max 532 can also be provided to the averaging operation 540. In the averaging operation 540, based on the minimum value I min 522 and the maximum value I max 532, especially on a pixel-by-pixel basis, the average value is calculated to provide the average image I avg 542 to be calculated. The average image I avg 542 is also provided to the summation operation 534, which can add the (positive or negative) offset 536 to it to generate the final output image I out 538. This final output image I out 538 can be corrected for structured noise. This structured noise can be the structure 430 that has been identified as the structure in the noise image, as relative to Figures 4A to 4C described.
[0095] Relative to Figure 5 The operations and the ongoing method steps described can also be represented as pseudocode, as shown in the following formulas (2) to (9). Assume that the noise original image 102 can have 100 × 100 pixels, as an example. The pixel size is N. Alternatively or additionally, the pixel size N can be selected as the cut-off size of the structuring element and / or the cut-off size of the kernel size K to avoid eliminating potentially relevant actual structuring elements represented in the noise original image I in 102. The offset 536 can be set to positive as well as negative values. Here, for example only, it is preset to the value 100.
[0096] N = 121, offset = 100 (2) Iterative opening & closing 1 × K, K = 3...N -> I vert (3) I vert ’ = I in – I vert (4) Iterative opening & closing K × 1, K = 3...N -> I horz (5) I horz’ = I in – I horz (6)I min = Min(I vert ’,I horz ’) (7)I max = Max(I vert ’,I horz ’) (8)I avg = (I min + I max ) / 2 (9)I out = I avg + offset.
[0097] Figure 6 illustrates a schematic diagram of an iterative opening and closing operation workflow 600 for an IID noise removal configuration for the workflow as shown in Figure 5 . In particular, corresponding structural elements 414, 416, 418, 422 or a corresponding kernel of size K can be implemented in the iterative opening operation 410 and / or closing operation 420 in the horizontal branch 500a and / or vertical branch 500b of the morphological operation 500 as shown and correspondingly described in Figure 5 .
[0098] The iterative opening and closing operation workflow 600 can correspond to iterative IID noise removal 304. As described above, IID noise removal can be used to isolate corresponding structural features ( Figure 4C structure 430 in) and / or remove IID noise in an image without relying on any structural information. In this process, for example and according to the above embodiments, 1-D closing operation 420 and 1D opening operation can be iteratively applied. This can be performed in both vertical and horizontal directions, increasing the kernel size in each iteration. The corresponding increments can be as follows. The noisy original image 102 can be provided to the workflow 600. The pixel sizes of the closing operation 420 and opening operation 410 are 1 × kernel size K and / or kernel size K × 1. For example, K is set to increase from 3 to 5. Thus, a 1 × 3 closing operation 602 is followed by a 1 × 3 opening operation 604. Next, a 3 × 1 closing operation 606 can be performed, followed by a 3 × 1 opening operation 608. Thereafter, a 1 × 5 closing operation 610 can be performed, followed by a 1 × 5 opening operation 612. A 5 × 1 closing operation 614 can be followed by a 5 × 1 opening operation 616. When the maximum value of the kernel size K is set to 5, the iterative opening and closing operation workflow ends and the output 618 is passed.
[0099] As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items and may be abbreviated as " / ".
[0100] Although certain aspects have been described in the context of an apparatus, it will be apparent that these aspects also represent a description of the corresponding method, where a block or apparatus corresponds to a method step or a feature of a method step. Similarly, the various aspects described in method steps are also descriptions of corresponding modules or items or features of the corresponding apparatus.
[0101] Some embodiments relate to a microscope 710, which consists of a system for performing a method related to one or more of Figures 1 to 6 Alternatively, the microscope 710 may be part of or connected to a system whose structure and configuration are for performing a method related to one or more of Figures 1 to 6 Some embodiments relate to a telescope (not shown), which includes a system for performing a method related to one or more of Figures 1 to 6 Alternatively, the telescope may be part of or connected to a system whose structure and configuration are for performing a method related to one or more of Figures 1 to 6 FIG. shows a schematic diagram of a system 700 configured to perform the methods described herein. The system 700 includes a microscope 710 and a computer system 720. The microscope 710 is configured to capture images and is connected to the computer system 720. The computer system 720 is configured to perform at least a portion of the methods described herein. The computer system 720 may be configured to execute machine learning algorithms. The computer system 720 and the microscope 710 may be separate entities, but may also be integrated within a common housing. The computer system 720 may be part of the central processing system of the microscope 710, and / or the computer system 720 may be part of a sub-component of the microscope 710, such as a sensor, actuator, camera, or illumination unit of the microscope 310, etc. Figure 7
[0102] The computer system 720 can be a local computer device (such as a personal computer, laptop, tablet, or mobile phone) having one or more processors and one or more storage devices, or can be a distributed computer system (such as a cloud computing system having one or more processors and one or more storage devices distributed at different locations (such as distributed at a local client and / or one or more remote server farms and / or multiple data centers)). The computer system 720 can include any circuit or combination of circuits. In one embodiment, the computer system 720 can include one or more processors, which can be of any type. As used herein, the term "processor" can refer to a microscope or microscope component (such as a camera) or any other type of computing circuit, such as a microprocessor, microcontroller, complex instruction set computing (CISC) microprocessor, reduced instruction set computing (RISC) microprocessor, very long instruction word (VLIW) microprocessor, graphics processor, digital signal processor (DSP), multi-core processor, field programmable gate array (FPGA) of a microscope or computer system. Other types of circuits that can be included in the computer system 720 can be custom circuits, application specific integrated circuits (ASICs), etc., such as one or more circuits (such as communication circuits) for wireless devices (such as mobile phones, tablets, laptops, two-way radios, and similar electronic systems). The computer system 720 can include one or more storage devices, and the storage devices can include one or more storage elements suitable for a specific application, such as main memory in the form of random access memory (RAM), one or more hard disk drives, and / or one or more drives for processing removable media (such as compact discs (CDs), flash memory cards, digital video discs (DVDs), etc.). The computer system 720 can also include a display device, one or more speakers, and a keyboard and / or controller, which can include a mouse, trackball, touch screen, voice recognition device, or any other device that allows a system user to input information to and receive information from the computer system 720.
[0103] Some or all of the method steps can be performed by (or using) a hardware device, such as a processor, microprocessor, programmable computer, or electronic circuit. In some embodiments, one or more of some of the main method steps can be performed by such a device.
[0104] According to certain implementation requirements, embodiments of the present invention can be implemented in hardware or software form. It can be implemented using a non-transitory storage medium (such as a digital storage medium, such as a floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM or FLASH memory), in which electronically readable control signals are stored, and these control signals cooperate with (or are capable of cooperating with) a programmable computer system to perform corresponding methods. Therefore, the digital storage medium can be computer-readable.
[0105] Some embodiments according to the present invention include a data carrier having electronically readable control signals, which can cooperate with a programmable computer system to execute one of the methods described herein.
[0106] Generally, embodiments of the present invention can be implemented as a computer program product having program code, and when the computer program product runs on a computer, the program code is used to execute one of the methods. The program code can be stored, for example, on a machine-readable carrier.
[0107] Other embodiments include a computer program for executing one of the methods described herein, which is stored on a machine-readable carrier.
[0108] In other words, therefore, an embodiment of the present invention is a computer program having program code, and the program code is used to execute one of the methods described herein when the computer program runs on a computer.
[0109] Therefore, another embodiment of the present invention is a storage medium (or data carrier or computer-readable medium) on which a computer program for executing one of the methods described herein when executed by a processor is stored. The data carrier, digital storage medium or recording medium is generally tangible and / or non-transitory. Another embodiment of the present invention is the device described herein, which includes a processor and a storage medium.
[0110] Therefore, another embodiment of the present invention is a data stream or signal sequence representing a computer program for executing one of the methods described herein. For example, the data stream or signal sequence can be configured to be transmitted through a data communication connection (such as through the Internet).
[0111] Further embodiments include a processing device, such as a computer or a programmable logic device, which is configured or adapted to execute one of the methods described herein.
[0112] Further embodiments include a computer on which a computer program for executing one of the methods described herein is installed.
[0113] Another embodiment of the present invention includes a device or system configured to transmit (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver can be, for example, a computer, a mobile device, a storage device, etc. The device or system can include, for example, a file server for transmitting the computer program to the receiver.
[0114] In some embodiments, a programmable logic device (e.g., a field programmable gate array) can be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array can cooperate with a microprocessor to perform one of the methods described herein. Generally, these methods are preferably performed by any hardware device.
[0115] List of reference numerals 100 A computer-implemented method configured to train an image denoising machine learning model 102 Noisy original image 104a……104 Non-identical copies 106 Machine learning model / convolutional neural network 108 Random noise optimized image 110 Copy of the noisy original image 112 Predict at least one guidance image 114 Guidance image 116 Decision function / loss function 200 A computer-implemented denoising method configured to denoise an image 202 Image to be denoised 204a……204h Identical copies 206 Trained machine learning model / trained convolutional neural network 208 Noise-corrected output image 300 Independent prediction workflow 302 Structured noise removal / non-IID noise removal 304 Random noise removal / IID noise removal 306 Normalization 400 Structure isolation workflow 410 Opening operation 412 Random noise covering the structure 414, 416, 418, 422 Structure elements 420 Closing operation 430 Structure 500 Morphological operation 500a Horizontal branch 500b Vertical branch 502 Provide the noisy original image for morphological operations 504 Provide the noisy original image for subtraction operations 506 Horizontal component I horz 508 Vertical component I vert 510 Subtraction operation of the horizontal branch 510b Subtraction operation of the vertical branch 512a Image corrected for horizontal structured noise 512b Image corrected for vertical structured noise 520 Minimum operation 522 Minimum I min 530 Maximum operation 532 Maximum I max 534 Addition operation 536 Offset 538 Final output image I out 540 Average operation 542 Average image I avg 600 Iterative closing and opening operation workflow 602 1×3 closing operation 604 1×3 opening operation 606 3×1 closing operation 608 3×1 opening operation 610 1×5 closing operation 612 1×5 opening operation 614 5×1 closing operation 616 5×1 opening operation 618 Output 700 System 710 Microscope 720 Computer device.
Claims
1. A computer-implemented method (100) for training a machine learning model for image denoising, having at least the following steps: Providing at least one modified copy of a noisy original image (102) to the machine learning model (106); Training the machine learning model (106) at least in part based on the at least one modified copy of the noisy original image (102); Using the machine learning model (106) to predict at least one random noise-corrected image (108) from the noisy original image (102); Predicting at least one guidance image (114) using a structured noise prediction (112) corrected from the noisy original image (102) or from a copy (110) of the noisy original image (102); and Optimizing the result of a decision function (116), wherein the decision function (116) is configured to compare at least a portion of the random noise-corrected image (108) with at least a portion of the guidance image (114).
2. The method (100) according to claim 1, wherein, The decision function (116) is a loss function to be minimized.
3. The method (100) according to claim 2, wherein, The type of the loss function is: ; where I t is the original noise image, I' t is the predicted image (108), I t_n is the guidance image (114), and 𝜆 is the weighting factor.
4. The method (100) according to any one of the preceding claims, wherein, The predicted image is the random noise-corrected image (108), and wherein the random noise-corrected image (108) is fed back into the method (100) to iteratively optimize the result of the decision function (116) and / or train the machine learning model (106); and / or wherein the guidance image (114) is fed back into the method (100) to iteratively optimize the result of the decision function (116) and / or train the machine learning model (106).
5. A computer-implemented method (100a) for training a machine learning model (106) for random noise image denoising, having at least the following steps: Providing at least one noisy original image (102); Providing at least one non-identical copy (104a... 104h) of the noisy original image (102), wherein the at least one non-identical copy (104a... 104h) is shifted by at least one pixel relative to the noisy original image (102); Feeding the at least one non-identical copy (104a... 104h) into the machine learning model (106) to be trained; Training the machine learning model (106) at least in part based on the at least one non-identical copy (104a... 104h); Using the machine learning model (106) to predict at least one random noise-optimized image (108) from the at least one non-identical copy (104a... 104h).
6. The method (100a) according to claim 5, wherein, Making eight non-identical copies (104a... 104h) of the noisy original image (102), each copy (104a... 104h) being shifted by at least one pixel in the direction of the noisy original image (102); and / or wherein the machine learning model (106) includes at least one convolutional neural network; and / or wherein structured noise is at least in part corrected based on a guidance image (116).
7. A computer-implemented image denoising method (200), having at least the following steps: Providing at least one exact copy (204a... 204h) of the noisy original image (202); Feeding the at least one exact copy (204a... 204h) into a trained machine learning model (206), the machine learning model (206) preferably being trained based on the method according to any one of claims 1 to 6; and Generating an output image (208) with corrected noise.
8. The method (200) according to claim 7, wherein, Making eight exact copies (204a... 204h) of the noisy original image (202) and feeding them into the trained machine learning model (206); and / or Wherein the structured noise is at least partially corrected based on a guidance image (116); and / or Wherein the machine learning model (106) includes at least one convolutional neural network.
9. The method (200) according to any one of claims 1 to 4 or the method (200) according to claim 8, wherein, Predicting (112, 300) the guidance image (114) based on at least one of the following steps: Structured noise removal (302); Random noise removal (304); or Normalization (306).
10. The method (200) according to claim 9, wherein The normalization (306) includes at least one of the following: An increase in signal intensity while suppressing noise; Adjusting the normalization scale such that the signal intensity increases; Adjusting the normalization scale such that the signal intensity increases as much as possible without saturation; or Adjusting the normalization scale such that the normalization scale is fixed to the 2nd to 99.98th percentile.
11. The method (100) according to any one of claims 1 to 4 or the method (200) according to any one of claims 8 to 10, wherein, Removing the structured noise based on changes in the size of the structuring elements (414, 416, 418, 422) and / or kernels in morphological operations (500); and / or Wherein the structured noise is removed based on morphological operations (500) applying at least one straight line as the structuring element (414, 416, 418, 422) and / or as the kernel; and / or Wherein the structured noise is removed based on opening and / or closing operations of morphological operations (500) applying the horizontal and / or vertical components of the recognition image (102, 110); and / or Wherein an offset (536) is added to or subtracted from the image (102, 110) during the structured noise removal.
12. A data processing device, especially a computer device (720) or a control unit for an imaging device or imaging system (710), comprising means for performing the method according to any one of claims 1 to 11.
13. A computer program having program code for performing the method according to any one of claims 1 to 11 when the computer program runs on a processor.
14. A data set, especially a training data set, a validation data set, and / or a training data set that can be used for a machine learning model, the data set including a plurality of images (102, 110, 104a... 104h, 114) obtained by using the method (100, 200) according to any one of claims 1 to 11 or references thereof.
15. An imaging device or imaging system (710), in particular a microscope or a telescope, which is configured to be used in the method according to any one of claims 1 to 11.