A computer-implemented method for generating high-resolution synthetic mammographic images by an ensemble of diffusion models
An ensemble of neural networks using a modified diffusion-based AI method generates high-resolution mammographic images, addressing resolution limitations and data scarcity, enabling effective training and education while ensuring image quality and context for lesion detection.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- INSTITUTE FOR ARTIFICIAL INTELLIGENCE RESEARCH & DEVELOPMENT OF SERBIA
- Filing Date
- 2025-12-16
- Publication Date
- 2026-06-25
AI Technical Summary
Existing methods for generating mammography images, such as using classical generative adversarial networks (GANs) and variational autoencoders, are limited to a maximum resolution of 1024x1024 pixels, which is insufficient for effective analysis and diagnosis, and there is a need for high-resolution synthetic mammogram images to train deep neural networks for anomaly detection and radiologist education, while addressing data scarcity and privacy concerns.
A computer-implemented method using an ensemble of three neural networks, including a single-channel Unet for low-resolution global context, a three-channel Unet for local context, and another three-channel Unet for full-resolution patch generation, to create high-quality synthetic mammographic images at 4096x3328 pixels, leveraging a modified diffusion-based artificial intelligence approach.
The method generates high-resolution mammographic images that are realistic and of sufficient quality for training deep neural networks and radiologist education, overcoming the limitations of previous methods by achieving a resolution beyond 1024x1024 pixels and providing accurate context for lesion detection.
Smart Images

Figure RS2025000010_25062026_PF_FP_ABST
Abstract
Description
[0001] A computer-implemented method for generating high-resolution synthetic mammographic images by an ensemble of diffusion models
[0002] Technical field to which the invention pertains
[0003] The invention covers the field of generative artificial intelligence in medicine, more specifically, the invention relates to generating mammography images via a new computer-implemented method based on artificial intelligence models.
[0004] Technical Problem
[0005] The invention solves the problem of generating full-resolution synthetic mammogram images (at a resolution of 4096x3328 pixels).
[0006] To date, mammogram image generation has been performed using classical generative adversarial networks (GANs) and variational autoencoders (VAEs), achieving a maximum resolution of 1024x1024 pixels.
[0007] To effectively and reliably conduct the analysis of mammography images as a diagnostic method on a broader population, an approach that is both inexpensive and efficient, it is necessary to enable semiautomatic image processing and demonstrate the effectiveness of such a solution. Generative artificial intelligence (Al) models represent a modern method that can be applied for these purposes. They enhance diagnostics by creating synthetic data, such as "normal" images for anomaly detection or reconstructions of corrupted images, thereby addressing challenges like data scarcity and variability. Additionally, generative models contribute to the training of traditional Al models by creating extensive and diverse datasets to improve their accuracy. These methods also have the potential to support the education of future radiologists by generating realistic datasets, while simultaneously addressing legal and data privacy concerns by enabling synthetic, anonymized images that are compliant with regulatory requirements.
[0008] The global shortage of qualified diagnosticians creates a need to improve and accelerate the process of analyzing mammography images, and this is currently being done using computer-aided diagnosis systems (CADs). CAD development is far from reaching its full potential, and one of the main problems is the inability to adequately train and test the models due to the lack of a high-quality and sufficiently large database of digital mammography images. The invention also attempts to solve this problem by generating synthetic mammography images through a new method that introduces modified diffusion-based artificial intelligence for image processing.
[0009] The invention solves the problem of meeting the need for a large number of high-resolution mammography images, which are available without legal restrictions regarding data privacy, by generating the required number of synthetic mammography images. These images can further be used to train deep neural networks for automatic and semi-automatic anomaly detection in images, as well as for radiologist education. Since the ensemble models themselves (ensemble models are three neural networks) used to generate the images contain information about what a healthy mammogram should look like, they can also be used directly to detect lesions that occur in patients.
[0010] Mammography images are the highest-resolution medical images, as it is essential to detect lesions and changes in healthy tissue in a timely manner, that is, when they occupy a small space, and this detection enables the generation of such full-resolution images.
[0011] The invention introduces three main phases of the method, which are the training phases of the model (neural networks), and a fourth phase which is the image generation phase. A trained ensemble of three models (three neural networks) enables the generation of a higher-quality image at full resolution compared to that obtained through other approaches, such as GAN neural networks and variational autoencoders.
[0012] State of the Art
[0013] The state of the art includes the following proprietary solutions provided further in the text.
[0014] Patent application WO2022251718A1, titled "Generating high-resolution images using selfattention," published on December 1, 2022, which differs from the proposed invention in the way the model architecture is implemented and in the specific semantics of the channels and the training and inference protocol.
[0015] US patent US11669965B2, titled "Al-based label generating system and methods for use therewith," published on June 6, 2023, relates to labels in a training set, whereas our invention relates to generating entirely new mammography images. The system of this patent does not relate to a multimodel architecture and denoising phases that work with high-quality mammography images.
[0016] Patent application EP4202955A1, titled "A system and method for processing medical images," published on June 28, 2023, deals with assessing the risk of processing medical images, not generating the mammographic images themselves. Additionally, it does not deal with processing high-resolution mammographic images.
[0017] Patent application US20240265505A1 titled "Multimodal diffusion models" published on August 8, 2024. also combines multiple neural networks but does not perform iterative image refinement in the sense of increasing resolution (as our invention does), but rather in the sense of guided generation, which our invention does not reference. Our invention is unimodal and deals only with mammography images.
[0018] A patent titled "Al-based multi-label heat map generating system and methods for use therewith," published on October 10, 2023, relates to a system tasked with finding abnormal tissue parts on medical images by generating risk maps. Although it involves generation, it is not about generating medical or mammogram images, certainly not in high resolution.
[0019] Patent US10346982B, titled "Method and system of computer-aided detection using multiple images from different views of a region of interest to improve detection accuracy," published on July 9, 2019, is aimed at finding regions of interest, not generating images, certainly not in full resolution.
[0020] Patent application WO2021017372A1, titled "Medical image segmentation method and system based on generative adversarial network, and electronic equipment," published on February 4, 2021, relates to high-resolution images but does not relate to their generation. The objective of this solution is to achieve pixel-level segmentation, while the objective of our invention is to generate pixel-level mammogram images.
[0021] Here, we must again note that, when it comes to generating mammogram images, either for creating synthetic data to train other models, or as a basis for anomaly detection, the underlying technology remains GANs, which are a machine learning framework consisting of two competing neural networks, with one learning to generate new content and the other learning to distinguish that content from real images.
[0022] The following works also enter the state of the art in addition to the protected solutions.
[0023] The work titled "A generative adversarial network for synthetization of regions of interest based on digital mammograms," by Oyelade et al., is state of the art and deals with generating specific regions of interest for mammography. Although it is believed to be based on so-called patches in the image, it is not generally designed to generate the entire mammogram image.
[0024] The paper titled "Unsupervised anomaly detection with generative adversarial networks in mammography" by Park et al. focuses on generating entire images of healthy breasts and has shown good performance in detecting cancer as an anomaly. In addition to detecting anomalies, this type of approach provides an opportunity to create pixel-level lesion segmentation masks, using only weak labels in the form of image-level annotations. Relying on the StyleGAN2 generation architecture, it achieves a Frechet inception distance (FID) score of 4.383.
[0025] FID provides a metric for evaluating how well a specific generative model performs in terms of generating realistic and diverse images. This can help in comparing different models or comparing a model's performance over the course of training. The paper differs in its description of the methods it proposes.
[0026] Also included in the state of the art are the works "Diffusion models for medical anomaly detection" and "Anomaly detection with denoising diffusion probabilistic models using simplex noise," which differ from the present invention in the type of methods they propose.
[0027] Diffusion models have recently become one of the so-called hot topics in computer vision, as stated in the paper Diffusion models in vision: A survey, authored by Pan et al., shows their impressive generative capabilities, ranging from high levels of detail to the diversity of generated examples, so their application in the field of radiology can be expected. However, the present invention differs in the description of the proposed method from all the methods cited in this survey by Pan et al.
[0028] A recent study by Pan et al., titled "2D medical image synthesis using transformer-based denoising diffusion probabilistic model," has shown that diffusion models can generate X-ray, MR and CT images that are practically indistinguishable from the originals even to the eye of a professional radiologist, and therefore pass a visual Turing test. Quantitatively, the model achieves a relatively modest FID score of 37.27 when considering both real and synthetic data. The study, however, did not use mammography images, and the resulting images are of a relatively low 256x256 pixel resolution at the output. The paper used a classic diffusion model, as opposed to our finding which uses an ensemble of modified diffusion models (i.e., 3 neural networks) with local and global context and is able to generate high-resolution mammogram images (4096x3328 pixels).
[0029] The following papers also relate to the state of the art, e.g., the paper titled Med-cDiff: Conditional Medical Image Generation with Diffusion Models deals with medical image generation, but its focus is on conditional generation, which is different from our method and does not pertain to generating high- resolution mammography images.
[0030] The paper titled "Generation and Evaluation of Medical Images Based on Diffusion Models" and the paper titled "Advanced image generation for cancer using diffusion models" address the generation of medical images, but they do not deal with architectures that support high-resolution generation.
[0031] The paper titled "MAM-E: Mammographic Synthetic Image Generation with Diffusion Models" deals with mammographic images and generates them using diffusion models, but the problem it does not solve, which our invention does, is generating high-quality images at the full resolution of 4096x3328 pixels.
[0032] The paper titled "Panoptic Segmentation of Mammograms with Text-To-lmage Diffusion Model" deals with the use of diffusion models to generate segmentation masks, not the full-resolution mammogram images themselves.
[0033] In the paper titled "High-resolution image synthesis with latent diffusion models," Latent Diffusion Models (LDM) are described, which are common for high-resolution images.
[0034] Latent diffusion models first compress the image, typically using a so-called off-the-shelf or trained variational autoencoder (VAE) or vector quantized variational autoencoder (VQ-VAE), where the image is mapped to a latent representation, which is significantly smaller in terms of the number of values it contains compared to the original image, and this latent representation is then used for training the diffusion model. When the trained model is used to remove noise and thereby generate a synthetic latent representation, the autoencoder is used to translate this representation into a high-resolution image.
[0035] The approach helps to reduce the computational requirements, which would be too high if the diffusion model were trained directly with high-resolution images. Unfortunately, since no pre-trained variational autoencoder exists that would reliably work for mammographic data, this approach is not yet used in the field of mammography.
[0036] Variational autoencoders are generative models used in machine learning to generate new data in the form of variations of the input data on which they were trained. In addition to this, they also perform tasks common to other autoencoders, such as noise removal.
[0037] In the paper "Patched denoising diffusion models for high-resolution image synthesis" by Ding et al., the Patch-DM method for natural images is described. The work belongs to the state of the art, but it is a classic patch-based diffusion model, unlike the invention which uses an ensemble of modified diffusion models with a three-stage image generation. The resolution of the images generated by the invention's method (4096x3328 pixels) is significantly higher than that of Patch-DM (where it is 1024x512 pixels).
[0038] From all of the above, we can say that diffusion models have shown promising potential as a robust approach for generating high-quality samples from complex image and video datasets, but also that they cannot be directly applied to generate high-resolution samples.
[0039] One way to overcome this limitation is an approach based on applying models (neural networks) to smaller parts of the image (so-called patches) and integrating those results to create a high-resolution output image. The paper proposes one such new approach for generating mammography images, which is based on training an ensemble of three diffusion models (i.e., three neural networks) to model images and image parts (patches) at different levels of detail. This innovative approach allows for the preservation of the local and global context of the patches themselves within the image during modeling and generation, and enables the combination of generated patches to produce a full-resolution mammogram image.
[0040] Presentation of the essence of the invention
[0041] Mammography is the leading method for diagnosing breast cancer.
[0042] If the obtained mammography images can be analyzed efficiently and semi-automatically, then mammography represents a cheap and reliable method.
[0043] The invention includes a new method for generating mammographic images using artificial intelligence, with three training phases in which at least three different models are trained (at least three neural networks forming an ensemble): a first model comprising one channel, a second model comprising three channels, and a third model comprising three channels. The channels are considered as levels of image processing, with the first channel representing the image at a low resolution (global context), the second channel encodes information about the local context, and third about the high-resolution patch that is being generated. The images in the channels will be discussed further in detailed description.
[0044] A computer-implemented method represents a modified generative artificial intelligence (diffusion), which is a novelty because the diffusion is performed in the second and third phases of the method and in a third channel (over partially noisy patches of the image) in both phases (classically, it is performed in a single channel over the entire image) and over patches of the image.
[0045] The novelty of the invention is the application of a modified ensemble of diffusion models (three neural networks) that can generate patches (regions) in an image which are then combined to form a high- resolution synthetic mammogram image.
[0046] Diffusion models are used for image generation, including radiology. On the other hand, only a few studies specifically deal with mammography images, and none of them have successfully generated a high-resolution 4096x3328 pixel radiological, mammography image as a retrieval method.
[0047] Mammography has the highest spatial resolution (image sizes up to 4096x5625 pixels) of all imaging modalities, due to the need to detect the delicate features of the smallest lesion during imaging.
[0048] The invention achieves an outstanding state of the art in performance, both in terms of realism and the resolution of the generated image.
[0049] In the first phase of the method, called the low-resolution global image context generation phase, a single processing channel is used and an image from a mammographic image database with a resolution of 4096x3328 pixels is fed to the input, which has been expanded to a square (an image of the same width and height in pixels) with all-black pixels, then the image is resized, the image is downscaled to a resolution of 256x256 pixels and then a noising step is performed on the input image, followed by a denoising step, which is carried out by a neural network that learns to gradually remove noise from the input image. These two steps are the classic steps of a denoising diffusion probabilistic model (DDPM). The process is iteratively repeated using different input images (original mammograms) until the model (neural network) adequately learns to remove noise from the image. Once the model is trained, based on an input that is just randomly generated noise, it can generate an output, a realistic and synthetic image (in this phase) with a resolution of 256x256 pixels. The neural network architecture used in the first stage denoising step is the so-called Unet, and it is a single-channel Unet.
[0050] The UNet is an architecture primarily designed for image segmentation in the medical domain and is a standard architecture used in diffusion models.
[0051] Only in this first training phase, the inference method does not deviate from the standard Denoising Diffusion Probabilistic Model (DDPM), because it works with whole mammograms, represented at a low resolution, and although mammographic images have only a single channel, this is equivalent to the original DDPM approach where noise is added in equal proportion to each channel and then gradually removed during image generation.
[0052] What is innovative compared to DDPM and the state of the art is that the computationally implemented method of our finding includes training two additional diffusion models (neural networks) in the second and third phases, which aim to generate only a specific part of the image (a patch) and in which additional input and output channels are used to provide the model with additional information about the context (environment) of that image part. Therefore, noise is added during training only to the channel corresponding to the patch itself, and not to the context. All information from the local and global context in the other two channels of both models is retained in its unaltered form, representing a modification of the classic approach to training and applying diffusion models. As mentioned above, the channels are treated as images where the first channel is the low-resolution image (global context), the second channel is an image with the local context, i.e., the region of interest in the image shifted to the center of the image, and the third channel is a partially blurred image of the patch with the local context.
[0053] The invention in its method first trains neural networks for all three phases. In the first phase, called the low-resolution global image context generation phase, a single-channel neural network is trained, (a single-channel Unet) to generate whole images at a lower resolution (standard for diffusion models) of 256x256 pixels, based on a noisy image (which can be so noisy that it contains nothing but noise). Then comes a second phase, called the local context generation phase, where a three-channel neural network (a three-channel Unet) is trained and used, whose third channel is used to represent the patch on the original 768x768 pixel image, reduced to a 256x256 pixel resolution, but the other two input channels are given information about the patch's surroundings, which represent its local context. The network of this phase is trained to remove noise only on the channel corresponding to the patch itself. In the third phase of training and deployment, called full-resolution patch generation, image patches are generated representing 256x256 pixel regions of the full-resolution mammogram image. In this phase, when generating synthetic data, the output of the trained models (neural networks) from the first two phases are used to provide the three-channel model of the third phase with information about the global (output of the first-phase model) and local context (output of the second-phase model). The model, similar to the previous phase, is trained to remove noise only on the channel corresponding to the patch itself in the image. During training, the local and global context are obtained from the original input image by cropping the appropriate regions around the patch and down sampling them to a resolution of 256x256 pixels, in the case of the local context, by cropping the appropriate regions around the patch and reducing them to a resolution of 256x256 pixels, and in the case of the global context, by reducing the resolution of the entire image to 256x256 pixels. The computer-implemented method therefore has three training phases and four image generation phases, because in image generation, the first phase is first called iteratively and then the second phase (the local context generation phase) is called iteratively (in order to cover the entire image with patches and generate the local context). Then, the patch integration step, after which the full-resolution patch generation phase takes place, and then a patch integration step is performed and the output full-resolution mammogram image is generated, which is stored in a new database in the memory of the computer on which the method is implemented, on the computer's processor and graphics card.
[0054] Since the local and global context used for training the second and third phase models are derived directly from the original image, the training of the models (neural networks) in the different phases is independent, and the order in which these three phases are performed is not important.
[0055] The division of an image into patches is a way to represent a high-resolution image as a series of smaller images (256x256 pixel resolution), which allows working with diffusion models whose size (the number of parameters in the neural network) allows training on modern hardware with the amount of data available in the mammography domain, as well as a reasonable generation time for synthetic images. In this sense, the patch-based approach is an alternative to the latent diffusion approach.
[0056] A direct approach based solely on representing a high-resolution mammogram image with a set of regions (patches cut from the image) of low resolution and attempting to train a diffusion model on such a dataset, does not provide enough information about the image's context to generate patches that can be used to reconstruct the entire high-resolution image.
[0057] The computer-implemented method relies on the specifics of mammographic images, which are high-resolution but represented as a single-channel image, to provide the second and third-stage three- channel models with enough information needed to generate medium and high-resolution patches, the latter of which are of sufficient quality to be combined into a full high-resolution mammogram.
[0058] Three phases were experimentally confirmed as an adequate number during the research of method implementation in terms of the speed, accuracy, and simplicity of the model (neural networks), taking into account the performance of modern hardware, primarily graphics cards. This number (three) represents the point where mammogram image quality has been optimized in accordance with computer hardware performance. More than three channels would be computationally demanding, while fewer than three would result in insufficient high-resolution image quality.
[0059] With diffusion models, only a single model is typically used, not an ensemble of multiple models, and there is only one training phase and one image generation phase. This phase allows for the generation of only patches within the image, but not its context within the entire image. The second channel in both the second and third phase allows for more precise patch fusion, while the third channel contains information about the partially blurred patch with local context, and in the second and third stages also allows for the generation of patches that follow the general trends in the image, such as e.g. the lighter and darker parts of the image, the appearance of certain types of tissue, etc.
[0060] The novelty of the invention is the application of a new method for generating high-resolution synthetic mammographic images.
[0061] By generating individual patches, without any context information in the second and third phases of the method, the model is unable to recreate a realistic, high-quality synthetic mammogram image. Even if the patch is not generated from pure noise, but from a partially noisy patch, a part of the original image, the reconstruction is poor.
[0062] Brief description of the images
[0063] Figure 1 represents the first stage 100 of the method, the stage of generating the global context of the image at a low resolution.
[0064] Figure 2 represents the second stage 200 of the method, the phase of generating the local context. Figure 2a illustrates in more detail the step of locating a patch in the image and centering it, the second phase step, second channel, stage 204.
[0065] Figure 3 represents the third stage 300 of the detection method, the stage of generating fullresolution patches.
[0066] Figure 4 represents the phase of generating the full-resolution output mammogram image.
[0067] Figure 4a represents subphase 401 of calling the first phase 100 within the phase 400 of generating the mammogram image.
[0068] Figure 4b represents subphase 405 of iteratively calling second phase 200, within phase 400 of generating the mammogram image.
[0069] Figure 4c represents subphase 412 of iteratively calling third phase 300 within phase 400 of generating the mammogram image.
[0070] Figure 5 represents the diffusion process in accordance with the detection method.
[0071] Figure 6 shows samples of mammography images obtained with the detection method.
[0072] Figure 7 shows the similarity between the original and generated patches in the image.
[0073] Table 1 presents a comparison of the performance metrics of the detection method with existing approaches.
[0074] Figure 8 shows: (a) the entire image where one of the patches of interest is marked with a smaller square (originally at 256x256 pixels resolution), and the local context of the same patch with a larger square (originally 768x768), (b) the image at the same scale, but shifted so that the region of interest is in the center, (c) the patch (marked by a square) with its local context, and (d) the patch itself.
[0075] Figure 9 presents the results - two examples of completely synthetic mammograms generated by invoking the entire ensemble of diffusion models.
[0076] Detailed description of the findings
[0077] Since the invention belongs to the field of artificial intelligence, here we list some of the most important concepts and explanations that are significant for a better understanding of the novelty of the invention method.
[0078] Artificial intelligence includes machine learning, where models are trained on large datasets and programs learn from data and adapt without specific programming; it also includes expert systems that are applied in medicine and generative models that create new content based on patterns learned from the data used for training.
[0079] With respect to this definition of terms, the invention belongs to the field of generative artificial intelligence, whose method uses newly trained artificial intelligence tools to generate a database of new mammographic images.
[0080] GANs are a machine learning framework consisting of two neural networks, one that attempts to learn to generate realistic synthetic data to "fool" the other network, and the other that learns to distinguish that synthetic data from real data. As such, they represent the state of the art for the method of invention.
[0081] In addition to them, there are also convolutional neural networks (Convolutional neural network - CNN) which simulate the actual process of recognition and inference used by humans. They are essential for computer-aided diagnosis due to more accurate diagnosis, for image segmentation, e.g., for tumors, as well as for tissue recognition and classification. Computer-aided diagnostics also represent state of the art with respect to the method of detection.
[0082] Diffusion models have emerged as a powerful tool in the field of artificial intelligence, drawing inspiration from the natural phenomenon of diffusion. These models use the concept of particle dispersion to generate new data samples that resemble existing ones by progressively adding noise and then learn to remove that noise. This process allows models trained in this way to learn the patterns present in the data, without the need for human supervision (data labeling).
[0083] In medical image processing, diffusion models have so far been applied to filling in damaged parts of images in CT scans or MR images, as well as to the segmentation of organs and tumors, noise removal, where diffusion models can effectively remove noise and improve image quality, and also in synthetic image generation, where, for example, diffusion models can generate synthetic medical images.
[0084] As the invention pertains to mammography, which is a desirable diagnostic method and a significant step in the prevention of disease, and which is also inexpensive, reliable, and effective for screening a large portion of the population, the following text will discuss improvements in this field through the lens of a new discovery method.
[0085] Namely, mammographic images are high-resolution images, and it takes both time and human resources to analyze them successfully (e.g., due to the need to recognize fine contrast differences caused by variations in breast density). For these reasons, it is understandable that there is a need to improve and accelerate this process. Until now, this has been done by computer-aided diagnostic systems (CAD), and the invention represents a new solution in this field, achieved through the application of artificial intelligence methods and techniques.
[0086] The advancement of deep learning (DL), the domain to which this invention belongs, has led to significant improvements in performance when it comes to medical image analytics, but it is not yet a widely adopted approach. Given that sufficient (labeled) training data is available, deep learning addresses: image generation, risk prediction, cancer prediction, anomaly detection, segmentation, anomaly localization, etc.
[0087] However, high-quality medical data is very difficult to acquire, often due to patient privacy concerns and the legal difficulties involved in the process.
[0088] However, applications of diffusion models in radiology have so far been typically limited to an image resolution of 256 x 256 pixels, making them limited for use in mammography.
[0089] Mammography images are the highest-resolution images used in medicine, ranging from 2300x1800 pixels to 4096x5625 pixels (the invention works with images at a resolution of 4096x3328), where the pixel size corresponds to an area of 100 microns and 54 microns, respectively. The main reason for this is the need to identify a cancer smaller than 1 cm as accurately as possible. This has a significant impact on the outcome of therapy, because the 10-year survival rate after breast cancer treatment is 95% for patients whose cancer is less than 1 cm, and only about 60% when the cancer size is over 3 cm.
[0090] The mammographic images we work with are single-channel images.
[0091] In the following text, when describing the invention method, channels are specified in the method's phases, one channel in the first phase and three channels in the second and third phases. For a real image, this three-channel concept draws an analogy from an RGB image, but in the invention, it refers to a novelty in the sense that the first channel represents the image in low resolution, i.e., the global context of the image, the second channel represents the image with the local context, i.e., the region of interest in the image shifted to the center of the image, and the third channel represents the image with a partially blurred patch in the image with the local context together, which is shown in Figure la.
[0092] The invention proposes a new approach based on an ensemble of diffusion models (three neural networks) that enables processing individual parts (patches) of the image and generating full-resolution synthetic patches, which are of sufficient quality to allow for merging into a high-quality, high-resolution synthetic image. To achieve this goal, a new approach is proposed that provides the final ensemble model with sufficient information about the local and global (full image) context of the generated patch. Both the computationally implemented method itself and its application to mammographic images are novel. Previously, this was done using a GAN neural network and a variational encoder, but the diffusion model generates better image quality compared to these approaches.
[0093] The full resolution includes 4096x3328 pixels, whereas the previous maximum was 1024x1024. Now, with the discovery method, higher-quality synthetic mammogram images are obtained, i.e., a new database of mammogram images located in the computer's memory, on whose processor and graphics card the discovery method is implemented.
[0094] The computer-implemented method is novel in its application of an ensemble of neural networks (the ensemble consists of three neural networks) trained using the diffusion process, with additional channels being added in the second and third training phases, i.e., the image processing method. The method is also novel because the diffusion process is limited to only one channel in each ensemble model (each neural network), and noise is added to the third channel in the second and third phases of the method. The method is also new for its application in mammography.
[0095] Previous methods with image patches were able to learn the neural network's features during the training phase, but these performances were insufficient for mammogram reconstruction tasks. By generating individual patches without any contextual information, the method is unable to create a realistic, high-quality synthetic mammogram image. Even if a patch is not generated from pure noise but from a partially noisy patch, the reconstruction is poor.
[0096] The retrieval method takes the DDPM diffusion probabilistic denoising model, which is based on patches in the image and a triplane UNet architecture, and provides context by using additional input channels, image processing levels, and these then provide the local and global contexts representing the patch's environment and its position in the image. To generate a fully synthetic, high-resolution mammogram image, both the global and local context of the image patches must be generated.
[0097] The invention includes a patch-based approach to the image, because working with the entire high-resolution image would be too demanding on the performance of contemporary graphics cards, and is certainly more efficient than working with the entire image, even when higher-performance hardware becomes available.
[0098] The global context contains information about whether the entire image or a large region of the image is brighter, darker, information about the type of breast tissue, etc.
[0099] Figure 5 below shows the diffusion process according to the method of discovery. A classic diffusion process would only be present if the third image were shown. Diffusion models are a new method of generative Al that produce better results in terms of the quality of generated synthetic images, compared to those that can be obtained using GAN neural networks.
[0100] Diffusion models are deep generative models that work by gradually adding noise (Gaussian noise) to the training data through a series of steps (also known as the forward diffusion process), and then learning to reverse this Markov process (known as denoising or the reverse diffusion process) to recover the data. The model gradually learns to remove the noise. Once trained, the model is able to generate new, high-quality images from random images (images that contain only random noise).
[0101] The noise removal process is parameterized by a neural network that predicts the noise to be removed at each step of reverse diffusion, in order to remove the noise. The forward diffusion process is defined in (1) as: where kt represents the data at time stept, ?t is the variance schedule, which controls the level of noise added to the data at each step, and N denotes the normal probability distribution. This process gradually adds Gaussian noise to the data over T time steps, resulting in pure noise at the boundary value? -> oo xT .
[0102] {0tG (0,l)}Tt=1is called a variance plan or a noise plan.
[0103] The inverse process is modeled in (2) as: (kt, t) and 19 (kt, t) are learned functions that predict the mean value and variance of an individual step of the reverse Markov process. The goal of the diffusion training model is to optimize the 6 parameters of these functions, which allows for the accurate reconstruction of the original data distribution from noisy data.
[0104] The operation of the computer-implemented method for finding is first carried out through the training phases of three diffusion models (three neural networks). The trained models are applied successively in exploitation to generate synthetic data from noise, or partially noisy real data, per channel, first the low-resolution representation of the entire mammogram (global context) which represents the first channel, then patches corresponding to image regions represented at medium resolution (local context) which represents the second channel, and finally, full-resolution patches of the mammogram which represent the third channel. Noise sampling for training the models used in search is done based on a linear variance schedule and with 1,000 steps in the forward diffusion process. It is trained with the ADAM optimizer, with a learning rate of 5 x 10-5and a batch size of 8 samples. The training and inference of the retrieval methods are performed on an NVIDIA GPU. Training the global context generation model (in the first training phase) takes 72 hours on 2 NVIDIA A100 GPUs. Training the image patch generation method (in the third training phase) takes 144 hours on 4 NVIDIA A100 GPUs. Finally, to obtain the second-stage model (which generates local context), we fine-tuned the model from the third training stage and found that for the second stage, the training takes 7 hours on 4 NVIDIA A100 GPUs. Fine-tuning involves the process of adapting pre-trained models to better suit specific tasks, often with a smaller dataset containing relevant examples.
[0105] During the training and working phase, the method creates patches in the image to cover all the details in the image.
[0106] Since two input channels in the image are used for context, the method deviates from the classic diffusion process, in which noise is added proportionally to each channel and then gradually removed during training and inference. The finding method adds noise only to the patched channel, namely the third channel of the second and third phases, while keeping all information from the local and global context in the other two channels in its unaltered form. This method of adding and removing noise represents a modified diffusion process designed to provide the models with the information needed to generate high-quality results.
[0107] Figure 1 shows the first stage of the retrieval method, stage 100 of generating the global context of the image at a low resolution. In this stage, the input step is step 101 where the full-resolution mammogram image is loaded from the database. The training database is located on the computer.
[0108] The first dataset used is the VinDr mammography dataset. This dataset contains a collection of 20,000 images in DICOM format, of which 2,254 images are labeled as containing different types of lesions, while the remaining images are labeled as benign. Since the number of images with lesions is insufficient for training neural networks to reproduce such structures, the training was limited to healthy images. Of the 17,746 healthy images, the majority (13,942) were acquired with a Siemens machine at a resolution of 3518 x 2800 pixels, while the remainder consisted of images taken from devices of other manufacturers at various resolutions. The training was limited to the images captured on the Siemens device. This part of the dataset contains images of the left and right breasts in CC and MLO projections.
[0109] The second dataset used is the RSNA dataset, which contains 54,706 images captured with different machines. Of these, we use 14,898 images captured with the machine labeled as machine 49, as this machine has the most images. The images are at a resolution of 4096x3328 and are diverse in terms of projections, tissue types, lesions, etc.
[0110] Next, after step 101, is step 102, in which preprocessing of the loaded mammographic images is performed to:
[0111] 1. neutralize (invert) the negative where necessary,
[0112] 2. removed all objects that are not the breast and kept only the breast in the image (the rest of the image is replaced with a solid black background),
[0113] 3. the image was fitted to a square by extending it on the side opposite the breast and
[0114] 4. all images where the breast was on the right side were flipped, so that the entire dataset would consistently contain breasts only on the left side.
[0115] Next, after step 102, step 103 follows, in which the image is resized to a resolution of 256x256 pixels, which is followed by the denoising of the input image in step 104, and then in step 105, the training of the neural network for noise removal. Since the mammogram image is in grayscale, this step uses a single-channel U-Net, a neural network trained using a diffusion probabilistic denoising approach on 256x256 pixel images. After step 105, an output image with a resolution of 256x256 pixels is obtained at the output in step 106.
[0116] As mentioned in phase 100, this is a single-channel training that involves working with an image representing the global context. In Figure 5 below, the concept of the channel is shown, where three images are visible in the time interval X0-XT, where at the final moment XT(by analogy also in Xoand some other Xt), the first image represents the global context (the first channel), the second image represents the local context (the second channel), and the third image represents a fully noisy patch at full resolution. Here, the concept of channels is explained: there are three different images, with different information: the global context (the first image), the local context (the second image), and the third image, which is a completely noisy image.
[0117] Figure 2 shows the second phase 200, the local context generation phase.
[0118] Phase 200 processes the image in three channels, i.e., it processes three images (which are also represented in Figure 5 below) where the first channel represents the global context with a low-resolution image, and in this channel, the input image with a resolution of 4096x3328 pixels is first loaded from a computer database in step 201 and then expanded with black pixels to a square shape (same image width and length). After step 201, step 202 is performed, in which the image is resized, and the image is downscaled, after which, in step 203, an image with a resolution of 256x256 pixels is obtained as the output image in the first channel, which is then used as the input image for further processing as the second channel (local context, the second image in the Figure).
[0119] In step 204, locating the patch in the image and moving it to the center is further performed, i.e., the entire image is moved to the center so that the patch is in the center, and the border portion of the image is cropped (in Figure 2a).
[0120] Figure 2a more detailed illustrates step 204, of the second phase 200, of the second channel, where the locating of the patch 212 in image 211 and the centering of the patch 212 by moving the entire image 211 to the center 213 takes place. Here, first the patch 212 is located in image 211, then the step of moving all patches to the center 213 occurs, and the remnants, the peripheral parts 214 of image 211, are removed. In this way, all patches 212 are highlighted which, if they are in a 768x768 pixel resolution, also contain information about their surroundings (local context), and by cropping the peripheral parts 214 of the image in a further embodiment of the phase of the method, only the image patches remain, which are then integrated to form the full-resolution image.
[0121] Then, in step 205, after step 204, the image from step 201 is overlaid onto the image of the third channel of phase 200, and patches are cropped from the image at a resolution of 768x768 pixels to capture the local context of the patch in the image. Then, in step 206, the patch in the image from step 205 is resized to a resolution of 256x256 pixels. Then, in step 207, the patch in the image is denoised, after which, in step 208, the images from steps 203, 204, and 207 are concatenated. After step 207, a neural network is trained in step 209 for gradual noise removal. This is an iterative process, in which steps 204-210 are repeated, and at the output, it is generated, in step 210, a patch in the image at a resolution of 256x256 pixels, but with local context. This means that the patch, although the resolution of the patch in the output image is 256x256 pixels, carries information about its surroundings in the image (see step 205 to see how it's loaded).
[0122] It is important to note here that, in step 205, patches containing only the background and little or almost no breast tissue are excluded from the training. The decision is made based on the calculation of the average pixel intensity value in the patch, and patches for which this average is less than a predefined threshold (& = 0.001) are discarded.
[0123] The local context corresponds to a 768x768 pixel region of the original image, centered on the target patch. To provide a local context for patches on the image edges, the image is extended with a 256- pixel "border" before the local context is cropped. This border is filled with the lowest pixel value in the image (zeros). In terms of the architecture itself, the implementation of the detection method differs from the other methods in the number of channels. In the first phase 100, it involves 1 channel; in the second phase 200, it involves 3 channels; and in the third phase 300, also 3 channels. It has been stated above that the channels are considered as three images: the first representing the global context, the second representing the local context, and the third representing a completely noisy patch in the image (Figure 2a, time instant XT).
[0124] The three channels are empirically determined and represent a novelty in the field of mammogram image generation using diffusion. If three channels were not used, generated patches would be obtained, but they would not be adequate for integration with surrounding patches, because without local context, the model lacks information about the surrounding patches. Here, it is directly concluded that the detection method uses a minimum of three channels.
[0125] Figure 3 presents the third phase of the retrieval method, the training phase for generating patches in the full-resolution image. This phase contains three channels representing the low-resolution image (global context), the local context (i.e., the region of interest in the image centered at the image's center), and a partially noisy patch image— see Figure 2a.
[0126] Phase 300 of the training for generating a patch in the full-resolution image begins with step 301, where an input image with a resolution of 4096x3328 pixels is retrieved from the database, then step 302 is performed where the image is resized and in step 303 an image with a resolution of 256x256 pixels is obtained which is further used as the input image for the second channel of phase 300. The steps 300, 301, and 302 belong to the first channel of the third phase.
[0127] Over the image representing the second channel of the third phase 300, step 304 of cropping the patch in the image to a 768x768 pixel dimension is performed, but the patch is cropped from the image loaded in step 301. Once this patch is cut out in step 304, step 305 is performed, where the patch's size is also changed, and in step 306, an image with a resolution of 256x256 pixels is obtained, which is then used in the third channel of phase 300.
[0128] Then, after processing the image representing the second channel of the third phase 300 (the region of interest in the image shifted to the center of the image — local context), processing of the image representing the third channel is performed (the fully blurred image of the patch with the local context), where first step 307 is performed, in which the image from step 301 is fed in, then step 308 of adding noise to this image is performed, followed by step 309 of concatenation where the images from steps 301, 306, and 308 are concatenated, after which it proceeds to step 310 where a neural network is trained to gradually remove the noise and finally, in step 311, a patch is obtained in the image at a resolution of 256x256 pixels.
[0129] After the training in phases 100, 200, and 300, the method includes an image generation phase, which is illustrated in Figures 4, 4a, 4b, and 4v.
[0130] Figure 4 illustrates phase 400 of the method, which relates to generating mammographic images.
[0131] Figure 4 shows phase 400 of generating a mammogram image with subphases: sub-phase 401 of generating a global image context at a low resolution, sub-phase 405 of generating a patch with a local context, step 412 of integrating patches within a local context, sub-phase 413 of generating a patch in the full-resolution image, a step 420 of integrating the patches, and a step 421 where an output full-resolution mammogram image of 4096x3328 pixels is obtained, which is stored in a database in the computer's memory. This is a new database of mammographic images that are generated by the detection method, and the invention protects the database thus obtained.
[0132] Figure 4a more detailed illustrates subphase 401 of generating a low-resolution global image context. In subphase 401, a noise is first programmatically generated in step 402, then in step 403 the U- Net neural network is trained to gradually remove the noise from step 401, and finally, in step 404, an output mammogram image with a resolution of 256x256 pixels is obtained, which is stored in the database in the computer's memory. Steps 403 and 404 are iteratively repeated.
[0133] Subphase 405 of generating the patch with local context is illustrated in more detail in Figure 4b. In subphase 405, step 406 first takes place, where the input image loaded at 256x256 resolution from the first phase 100 is used as input, then step 407 of locating the patch 212 in the image 211 and centering the patch by moving the entire image 211 to the center 213 (this also involves cropping the parts 214 of the image that are not of interest, the border parts 214, see Figure 2a), after which step 408 of programmatically generating and inputting noise is performed, followed by step 409 where concatenation is performed of this noise from step 408, the image from step 406, and the centered patch with the cropped image parts from step 407, after which the output from the step 409 concatenation is sent to the input of a neural network in step 410 which i s trained to gradually remove noise and in step 411 the generated patches with local context are obtained (meaning that these patches contain information about their surroundings), and are 256x256 pixels in resolution. Steps 409-411 are iteratively repeated. The images, or patches with local context from step 411, are stored in a database in the computer's memory.
[0134] Steps 409 and 410 are explained in more detail with respect to Figure 9 below.
[0135] Figure 4c shows subphase 413 of generating the patch in the full-resolution image. Within sub-phase 413, step 414 of loading the input image at a resolution of 256x256 pixels from the second phase 200, then in step 415 a patch with local context is cropped from the image of step 414, then step 416 takes place where programmatically generated noise is introduced, then in step 417 the image from the first phase is loaded at a resolution of 256x256, after which in step 418 concatenation is performed of the patch from step 415, the programmatically generated noise from step 416, and the loaded image from step 417, after which step 419 takes place where the neural network gradually removes the noise, and finally, step 420 takes place where the resulting patch is obtained at an output resolution of 256x256, which is stored in the computer's memory. Steps 418 and 419 are explained in more detail in the text below Figure 9.
[0136] Namely, Figure 4v shows that instead of generating the local context for each patch, the patches are integrated with the local context into a medium-resolution image (1280x1280 pixels) in step 414, from which it is then possible to crop regions of interest in step 415 (patches with local contexts for patch generation). This procedure reduces the computational load because fewer model calls are required overall, and it also allows for the application of overlap at this processing level, enabling smooth transitions. This integration is essentially the same as the final patch integration, except there are fewer objects to merge.
[0137] The results of the work of the search methods are presented in Figures 6 and 7 and in Table 1.
[0138] Figure 6 illustrates the results of the method. Figure 6a illustrates the original image, Figure 6b illustrates the image after noise removal (the image was noisy, so the noise was removed), Figure 6c represents the reconstructed image from patches using the context-free finding method. The partially noisy patches of the original image are shown, and here it can be seen that the reconstruction is poor, Figure 6g illustrates the result of generating a mammogram image as a whole, and finally, Figure 6d illustrates the difference between Figures 6a and 6b, the difference between the original and denoised image and shows what can be obtained by using the retrieval method to generate a synthetic mammogram image. Image retrieved from pure noise with the implementation of the retrieval method. Since we wanted to show that the synthetic mammogram image is similar to one of the original mammogram images in our dataset, only the local context and patches were generated from pure noise in this example, and the result is quite good.
[0139] To enable a fair comparison with the approach proposed by the authors Park et al. in the state of the art, which focuses on specific regions of interest in a mammographic image, an additional experimental setup was created. Focusing on a specific region of interest (ROI) corresponds to the situation in a retrievalbased approach application where the focus is on a specific patch in terms of its location. Therefore, another dataset was created using layered sampling in terms of the patches' locations. The dataset consists of pairs of patches sampled at the same location, one taken from the original image and one from an image generated from pure noise but using the original local and global context. We show examples of such pairs in Figure 7.
[0140] Figure 7 shows the original patches cropped from the images in the dataset (first row) and their corresponding (in terms of location) patch pairs generated from pure noise using the local and global context of the original image (second row).
[0141] Figure 6e above and Figure 7 illustrate what the search method can achieve in terms of qualitative results. Figure 7 shows patch noise removal samples, as would be used in the detection method to detect anomalies.
[0142] The retrieval method can generate high-quality images that, in many cases, are visually indistinguishable from the original. Currently, in its experimental phase, the differences are imperceptible to laypersons, but it is expected that qualified radiologists will soon be unable to tell the difference. Only in the medium sample, shown in Figure 7, does the detection method appear unable to remove the noise adequately, resulting in a grainy quality of the reconstructed image.
[0143] The assumption is that this is due to the fact that the original patch has a significant texture that likely cannot be discerned at the scale of the local context, which affects the final performance.
[0144] The results show a significant improvement over state-of-the-art approaches in terms of the common FID (Frechet Inception Distance) criterion. FID is a metric used to evaluate the quality of images generated by a generative model.
[0145] The following is an overview of the metrics for the proposed method and other known methods on the FID metric.
[0146] Table 1 shows several state-of-the-art approaches that are most closely related to the proposed method in terms of a specific part of their functionality. The approach by Park et al. is probably the most relevant because it uses classic generative Al (GAN) to create synthetic patches in a mammogram image and can support full resolution. However, it is not able to generate whole images, so it can only be compared at the patch level.
[0147] The approach of Pan et al. is more relevant in terms of its ability to generate entire radiological images, albeit at a low resolution, so we also compare them with their work.
[0148] Patch-DM is the methodology closest to the proposed approach, but it is not designed to work with radiological images. Table 1 presents the state of the art in performance for similar models on common metrics in the natural image domain for high-resolution, patch-based, but data-rich, domains. In Table 1 below, we indicate the best performance for the radiological model (given in bold).
[0149] To evaluate the capabilities of the retrieval methods to generate high-quality data when good global and local context information is available, it is necessary to generate 15,000 patches using local and global contexts extracted from the image of the original dataset. The results are then compared with the same number of patches randomly (in terms of patch location) sampled from the original images. This yielded a reasonable FID of 10.78.
[0150] The inter-patch FID for the original data was also calculated by randomly sampling two patches from each image in our original mammogram image dataset, splitting the images into two datasets during this random sampling.
[0151] Comparing the two thus-generated datasets yields a FID of 2.14, which is quite close to what the retrieval method achieves in the ROI (region of interest) scenario. To compare with the approach of Pan et al., the global context of the retrieval model is used, designed to generate 256 x 256-pixel versions of entire mammography images. For this, the same number of synthetic images as in the original dataset (approximately 14,000) are generated. Again, the retrieval method performs better than the approach of Pan et al., but of course, it cannot achieve the performance attained in the state of the art on natural images.
[0152] The invention is a new generative Al method, applying diffusion for noise removal, which is patchbased, and whereby the method can generate full-resolution digital mammography images. To provide global and local context that enables semantically plausible generation and reduces boundary effects, separate diffusion models are used, and the method is based on an ensemble of diffusion models. This innovative strategy allows us to generate images of 4096x3328 pixels. This is a significant improvement over existing approaches in this field.
[0153] The three-channel models have 78.25 million parameters in the model, while the single-channel global context model has slightly fewer (78.24 million), which then ranks it, so to speak, near the bottom of patch-based diffusion models. For example, Patch-DM has 154 million, and the total number for our retrieval method, i.e. model, is approximately 235 million, and it is still lower in model parameters than those compared by the authors of Patch-DM. The higher the number of model parameters, the more complex the model is, the more data it needs to be trained, and the more demanding it is in terms of hardware performance fortraining and implementation. In this sense, the inference method represents a more efficient approach compared to the state of the art. Figure 8(a) shows the entire image where one of the patches of interest is marked with a smaller square (originally at 256x256 pixels resolution), and the local context of the same patch with a larger square (originally 768x768). Figure 8(b) shows the image at the same scale, but shifted so that the region of interest is in the center. Figure 8(c) shows the patch (marked by a square) with its local context. Figure 8(d) shows the patch itself.
[0154] Figure 9 shows the results - two examples of fully synthetic mammograms generated by invoking the entire ensemble. The model generates square images to allow for the generation of mammograms of different breast proportions, but here only the regions of interest are cropped for space. These images were generated at a resolution of 3072x3072 on the VinDr model (left) and 3840x3840 on the RSNA model (right).
[0155] In Figures 4 and 4v, concatenation and denoising involving patch overlaps are performed.
[0156] If all patches are generated completely independently, the so-called checkerboard effect occurs, meaning that in the reconstructed image, the transitions between patches are not smooth, but it is easy to notice where the boundaries are. To overcome this, a patch overlap procedure is introduced during generation in step 412. Specifically, only the first patch is generated completely independently, from pure noise. Then, for each subsequent one, a 32-pixel overlap is taken from the previously generated, neighboring patch. On the rest of the patch (256x224, 224x256, or 224x224, depending on whether we take the vertical or horizontal neighboring patch, or both), the pixel values are iteratively updated with the values obtained from the neural network. In the overlap region, noise is also programmatically removed iteratively (as the appearance of this region is known), which provides the neural network with sufficient context to generate smooth transitions. This procedure allows for the generation of realistic patches that can be integrated without any noticeable transitions.
[0157] Throughout the application, the question of the resolution of the input images for model training and the retrieval and generation of output mammographic images arises. Here, we provide a brief explanation of the resolution.
[0158] Specifically, the procedure specifies image resolutions corresponding to the RSNA dataset because that dataset contains higher-resolution images, and the model was trained on it to generate higher- resolution images. The model was trained on the VinDr dataset as well, and examples of both can be seen in Figure 9, where the generated resolutions are 3840x3840 and 3072x3072. (RSNA and VinDr respectively) and they do not exactly match the input resolutions of 4096x3328 and 3518x2800. This is not a problem because the model that generates the patches from which the final image is composed was trained on real-world dimensions, and the images are generated at those dimensions. The differences in image dimensions are caused solely by the number of background pixels (which do not affect the quality of the mammogram).
[0159] Mode of industrial or other application of the invention
[0160] The invention finds application in medicine, for generating images in mammography that are further used as auxiliary tools for adequate diagnosis and for further training of Al models.
Claims
Patent Claims:
1. A computer-implemented method for generating synthetic mammography images via a modified diffusion model that includes three training phases of the modified diffusion model: a phase 100 of training for generating a low-resolution global context of the mammogram image, a phase 200 of training for generating a local image context, and a phase 300 of training for generating patches in the full-resolution image; and a phase of generating an output image via said model; characterized by the use of at least three channels in the training phases (100, 200, and 300), namely: in phase (100) of training for generating a global context, the training proceeds on a perchannel basis with step (101) loading a full-resolution input mammogram image from a database, step (102) preprocessing the input image, step (103) resizing the input image, a step (104) of low- resolution denoising the input image, a step (105) of training the neural network to remove the noise, and a step (106) where a low-resolution 256x256 pixel output mammogram image is obtained; in stage (200) training for generating local context, the training proceeds along three channels, where the first channel is the global context, the low-resolution image obtained in step (203), after the step (201) of loading the full-resolution image from the database and the steps (202) resizing the input image; the second channel is the local context image which, after step (203), is processed in step (204) to locate a patch in the image (212) and move it to the center (212) of the image, whereby the peripheral parts (214) of the image (211) are cut off; and a third channel is the fully blurred patch in the image, which is obtained by processing in step (205) where a 768x768 resolution patch is cropped from the image of step (201), then in step (206) the resizing of the image patch, then in step (207) the denoising of the image patch, and then in step (208) the concatenation of the outputs from steps (203, 204, and 207) after which in step (209) the neural network is trained to remove the noise and at the output in step (210) a patch of the image with local context is generated at a resolution of 256x256, where in steps (204-210) are iteratively repeated; in stage 300 of the training for generating full-resolution patches, image processing is performed on three channels, where the first channel is the low-resolution image obtained in step (203), after the steps (301) loading the full-resolution image from the database and step (302) resizing theinput image; the second channel is the image with the local context obtained in step (306), after step (304) where a patch is cropped from the input image of steps (301) and in step (305) the patch is resized; and the third channel represents the image with the fully noise-added patch and local context, which is obtained in step (311) after step (307), where the patch from step (306) is fed in and then processing occurs in step (308) denoising the loaded patch, after which concatenation of the images from steps (303, 306, and 308) occurs in step (309), after which the neural network is trained to remove the noise in step (310), wherein all steps (301-310) are iteratively repeated in phase (300); finally, in phase (400) of generating the output image, in subphase (401) a global context is generated in such a way that in step (402) noise is programmatically generated and then in step (403) a neural network removes the noise and in step (404) the output image is obtained at a resolution of 256x256; then in subphase (405) a patch with local context is generated by first loading in step (406) the input mammogram images from the first phase (100), then in step (407) by locating the patch in the image and centering it, after which in step (408) noise is programmatically generated and then in step (409) concatenates the images from steps (406, 407, and 408), after which the neural network removes noise in step (410) and a patch with local context is obtained in step (411); then step (412) of integrating the patches with local context is performed, then in sub-phase (413) a full-resolution patch is generated through step (414) of loading the input mammogram image from the first phase (100), step (415) loading the local context patch from the second phase (200), step (416) adding programmatically generated noise, step (417) loading the input image at 256x256 resolution from the first phase (100), then step (418) concatenation of the images from steps (415, 416, and 417), step (419) where the neural network denoises, and step (420) obtaining a patch in the image which is iteratively repeated with step (418), and finally, in step (421), the integration of patches in the image takes place, after which the full-resolution output mammogram image is obtained in step (422) and stored in the computer's memory.
2. A computer-implemented method according to claim 1, characterized in that it is applied for generating mammographic images.
3. The computer-implemented method according to claim 1, characterized in that in the first phase (100) in the image noise addition step (103) and the image denoising step (104) a single-channel UNet neural network is implemented while in the second phase (200) over the image representingthe third channel in step (209), as well as in the third phase (300) over the image representing the third channel, a three-channel U-Net neural network is implemented in step (310).
4. A computer-implemented method according to claim 1, characterized in that a modified diffusion probabilistic denoising model is used, the model being modified by adding noise only in the step (207) of denoising in the second phase (200) generating a local context, over an image representing a third channel, and in the step (308) of noise addition in the third phase (300) of generating patches in a full-resolution image, over an image representing a third channel.
5. A computer-implemented method according to claim 1, characterized in that the channels in phase (100), phase (200) and phase (300) image processing levels, wherein the first channel represents the global context which is the low-resolution image, the second channel represents the image with the local context, and the third channel represents the fully noise-corrupted image patch with the local context.
6. A computer-implemented method according to claim 1, characterized in that the full resolution of the mammographic image is a 4096x3328 pixel resolution in the image.
7. A computer-implemented method according to claim 1, characterized in that in the step (102) preprocessing of the input mammogram images, wherein preprocessing of the loaded mammogram images is performed to neutralize the negative where necessary, removing all artifacts and retaining only the breast in the image, with the background being masked to solid black, after which the image is cropped to a square, with the same image width and height, on the opposite side from the breast side, and all images in which the breast is on the right side are rotated so that the entire dataset contains only left-side breasts.
8. A computer-implemented method according to claim 1, characterized in that the global context represents information as to whether the image is lighter, darker, and information about the type of tissue.
9. A computer-implemented method according to claim 1, characterized in that in step (204), of the second phase (200), of the second channel, locating of the patch (212) in the image (211) and centering of the patch (212) by moving the entire image (211) to the center (213), whereby the peripheral parts (214) of the image are removed. (211) are removed, thereby highlighting all patches (212) in 768x768 pixel resolution, which contain information about their surroundings, the local context, and by cropping the peripheral parts (214) of the image (211) in the phase (400) are further integrated to form a full-resolution image.
10. A computer-implemented method according to claim 1, characterized by obtaining a local context in phase (200) and, in phase (300), both a global and a local context, when patches in the image are returned to the original mammogram image, wherein the imag e representing the second channel in phase (200) and the image representing the second chanr el in phase (300) allow the patches to be joined more precisely, while the image representing the third channel represents a partially noisy patch in the image with the local context in both phases (200, 300) and enables the generation of lighter and darker parts of the image with different types of breast tissue.
11. A computer-implemented method according to claim 1,id in that in step (412) the integration of patches is performed in such a way that the patch overlap is performed smoothly so that only the first patch is generated completely independently, fro m the pure noise, then, for each subsequent patch, an overlap of 32 pixels is taken with the previously generated, adjacent patch and the values of the remaining patch are updated dependin i on whether a vertical or horizontal adjacent patch is taken, or both, at a resolution of 256x224 or 224x256, respectively 224x224, the pixel values are iteratively updated with values obtained from the neural network on the overlapping region, and the noise is iteratively removed beca use it is known in advance how that region should look, which allows the neural network to gene ite smooth transitions.
12. A request-based database 1-11, characterized by containing es obtained in the output step (422).