Image generation method, image synthesis method, computing device and electronic device
By adding noise to the original image sequence and then using an image generation model for denoising, an enhanced image sequence is generated. This solves the problem of balancing image content consistency and sharpness in image processing, and achieves image quality improvement in low-light environments or under specific imaging conditions.
Patent Information
- Application Number
- CN202511394036.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-26
- Publication Date
- 2026-01-13
AI Technical Summary
In the field of image processing and analysis, existing technologies struggle to improve image clarity while maintaining the consistency of image content, especially in low-light environments or under specific imaging conditions, where insufficient image contrast and blurred details remain unresolved.
By acquiring the original image sequence of the target object, adding noise, and then using an image generation model to denoise, an enhanced image sequence is generated, ensuring that the contrast between the target area and other areas is greater than a preset value, thus maintaining the consistency and clarity of the image structure.
It significantly improves image clarity and contrast while maintaining image content consistency, solving the technical problem of balancing image processing before and after processing. The image generation model can learn and understand the spatial relationships of the image, ensuring that structural coherence and spatial positioning are not destroyed.
Smart Images

Figure CN121329804A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of computer technology and artificial intelligence, and more specifically, to an image generation method, an image synthesis method, a computing device, and an electronic device. Background Technology
[0002] In the field of image processing and analysis, in applications requiring high contrast and fine structural identification, raw images often suffer from technical problems such as insufficient contrast and blurred details due to insufficient lighting, sensor limitations, or specific imaging conditions. This affects the effective analysis and subsequent applications of the images. For example, images captured at night in security monitoring often have poor contrast due to low-light environments; in industrial inspection, certain parts of an object may not be clearly visible due to the material or surface characteristics of the object; and in medical imaging, non-contrast computed tomography (CT) has limitations in displaying lesions, and related technologies for sharpening the raw images can cause image content misalignment, leading to inconsistencies. Therefore, it is difficult to balance the requirements for image content consistency and sharpness before and after image processing in related technologies.
[0003] There is currently no effective solution to the above problems. Summary of the Invention
[0004] This application provides an image generation method, an image synthesis method, a computing device, and an electronic device to at least solve the technical problem in the related art of balancing the consistency of image content and the requirement for clarity before and after image processing.
[0005] According to one aspect of the embodiments of this application, an image generation method is provided. The method includes: acquiring an original image sequence of a target object, wherein different original images in the original image sequence correspond to different positions of the target object; adding noise to the original images in the original image sequence to obtain a noisy image sequence; and using an image generation model to denoise the noisy image sequence to generate an enhanced image sequence, wherein the structure of the target object in the enhanced image sequence is the same as the structure in the original image sequence, and the contrast between the target region and other regions in the enhanced image sequence is greater than a preset contrast, and the other regions are regions of the target object other than the target region.
[0006] According to another aspect of the embodiments of this application, an image synthesis method is also provided. The method includes: acquiring an original medical image sequence of a biological object, wherein different original medical images in the original medical image sequence correspond to different tissue locations of the biological object; adding noise to the original medical images in the original medical image sequence to obtain a noisy image sequence; and using an image generation model to denoise the noisy image sequence to generate an enhanced medical image sequence, wherein the anatomical structure of the biological object in the enhanced medical image sequence is the same as the anatomical structure in the original medical image sequence, and the contrast between the vascular system region and other regions in the enhanced medical image sequence is greater than a preset contrast, and the other regions are regions of the biological object other than the vascular system region.
[0007] According to another aspect of the embodiments of this application, an image synthesis method is also provided. The method includes: obtaining an original image sequence of a target object by calling a first interface, wherein the first interface includes a first parameter, the parameter value of the first parameter includes the original image sequence, and different original images in the original image sequence correspond to different positions of the target object; adding noise to the original images in the original image sequence to obtain a noisy image sequence; denoising the noisy image sequence using an image generation model to generate an enhanced image sequence, wherein the structure of the target object in the enhanced image sequence is the same as the structure in the original image sequence, and the contrast between the target region and other regions in the enhanced image sequence is greater than a preset contrast, and the other regions are regions of the target object other than the target region; and outputting the enhanced image sequence by calling a second interface, wherein the second interface includes a second parameter, the parameter value of the second parameter includes the enhanced image sequence.
[0008] According to another aspect of the embodiments of this application, an image generation apparatus is also provided. The apparatus includes: an acquisition module for acquiring an original image sequence of a target object, wherein different original images in the original image sequence correspond to different positions of the target object; a noise processing module for adding noise to the original images in the original image sequence to obtain a noisy image sequence; and a denoising processing module for denoising the noisy image sequence using an image generation model to generate an enhanced image sequence, wherein the structure of the target object in the enhanced image sequence is the same as the structure in the original image sequence, and the contrast between the target region and other regions in the enhanced image sequence is greater than a preset contrast, and the other regions are regions of the target object other than the target region.
[0009] According to another aspect of the embodiments of this application, an image synthesis apparatus is also provided. The apparatus includes: an acquisition module for acquiring an original medical image sequence of a biological object, wherein different original medical images in the original medical image sequence correspond to different tissue locations of the biological object; a noise processing module for adding noise to the original medical images in the original medical image sequence to obtain a noisy image sequence; and a denoising processing module for denoising the noisy image sequence using an image generation model to generate an enhanced medical image sequence, wherein the anatomical structure of the biological object in the enhanced medical image sequence is the same as the anatomical structure in the original medical image sequence, and the contrast between the vascular system region and other regions in the enhanced medical image sequence is greater than a preset contrast, and the other regions are regions of the biological object other than the vascular system region.
[0010] According to another aspect of the embodiments of this application, an image synthesis apparatus is also provided. The apparatus includes: a first calling module, configured to obtain an original image sequence of a target object by calling a first interface, wherein the first interface includes a first parameter, the parameter value of the first parameter includes the original image sequence, and different original images in the original image sequence correspond to different positions of the target object; a noise processing module, configured to add noise to the original images in the original image sequence to obtain a noisy image sequence; a noise reduction processing module, configured to perform noise reduction processing on the noisy image sequence using an image generation model to generate an enhanced image sequence, wherein the structure of the target object in the enhanced image sequence is the same as the structure in the original image sequence, and the contrast between the target region and other regions in the enhanced image sequence is greater than a preset contrast, and the other regions are regions of the target object other than the target region; and a second calling module, configured to output the enhanced image sequence by calling a second interface, wherein the second interface includes a second parameter, the parameter value of the second parameter includes the enhanced image sequence.
[0011] According to another aspect of the embodiments of this application, a computing device is also provided, including: a memory storing an executable program; and a processor for running the program, wherein the program executes the methods in various embodiments of this application when it runs.
[0012] According to another aspect of the embodiments of this application, an electronic device is also provided, including: a memory storing an executable program; and a processor connected to the memory via a bus for running the program, wherein the program executes the methods in various embodiments of this application when it runs.
[0013] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, the computer-readable storage medium including a stored executable program, wherein, when the executable program is running, it controls the device where the computer-readable storage medium is located to perform the methods of various embodiments of this application.
[0014] According to another aspect of the embodiments of this application, a computer program product is also provided, including a computer program that, when executed by a processor, implements the methods of various embodiments of this application.
[0015] According to another aspect of the embodiments of this application, a computer program product is also provided, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the methods in various embodiments of this application.
[0016] According to another aspect of the embodiments of this application, a computer program is also provided, which, when executed by a processor, implements the methods of the various embodiments of this application.
[0017] In this embodiment, firstly, an original image sequence of the target object is obtained, where different original images correspond to different positions of the target object. Next, noise is added to the original images in the original image sequence to obtain a noisy image sequence. Finally, an image generation model is used to denoise the noisy image sequence, generating an enhanced image sequence. The structure of the target object in the enhanced image sequence is the same as that in the original image sequence, and the contrast between the target region and other regions in the enhanced image sequence is greater than a preset contrast. The other regions are areas of the target object other than the target region. By obtaining the original image sequence of the target object, the comprehensiveness and hierarchy of the original data are ensured. Adding noise to the original images in the original image sequence facilitates the image generation model in learning the recovery and enhancement rules of image features during the denoising process, thus better simulating the real situation when generating enhanced images. By using the image generation model to denoise the noisy image sequence, the image generation model removes the random noise introduced by the noise addition, intelligently identifies and enhances the contrast of the target region, and maintains the consistency of the target object's structure in the enhanced image sequence with the original image sequence. Image generation models can learn and understand the spatial relationships of image sequences, ensuring that the inherent structural coherence and spatial positioning of the image sequence are not destroyed when enhancing the sharpness of the target region. By introducing controllable noise and using intelligent denoising in the image generation model, both content consistency and sharpness of the image are guaranteed, thus solving the technical problem in related technologies where it is difficult to balance the requirements of image content consistency and sharpness before and after image processing.
[0018] The above general description and the following detailed description are for illustrative and explanatory purposes only and do not constitute a limitation thereof. Attached Figure Description
[0019] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0020] Figure 1 This is a schematic diagram illustrating an application scenario of an image generation method according to an embodiment of this application;
[0021] Figure 2 This is a flowchart of an image generation method according to an embodiment of this application;
[0022] Figure 3 This is a schematic diagram illustrating the training and image generation process of an optional image generation model according to an embodiment of this application;
[0023] Figure 4 This is a schematic diagram of an optional process for generating an image enhancement sequence using a sliding window according to an embodiment of this application;
[0024] Figure 5 This is a flowchart of an image synthesis method according to an embodiment of this application;
[0025] Figure 6 This is a flowchart of another image synthesis method according to an embodiment of this application;
[0026] Figure 7 This is a schematic diagram of an image generation apparatus according to an embodiment of this application;
[0027] Figure 8 This is a schematic diagram of an image synthesis apparatus according to an embodiment of this application;
[0028] Figure 9 This is a schematic diagram of another image synthesis apparatus according to an embodiment of this application;
[0029] Figure 10 This is a structural block diagram of a computing device according to an embodiment of this application;
[0030] Figure 11 This is a structural block diagram of an electronic device according to an embodiment of this application. Detailed Implementation
[0031] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described below are only some, not all, of the embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative effort should fall within the scope of protection of the present application.
[0032] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in other orders. Other orders herein refer to orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed, or inherent to such processes, methods, products, or apparatus.
[0033] First, some nouns or terms that appear in the description of the embodiments of this application shall be interpreted as follows:
[0034] Computed Tomography (CT): This can be a medical imaging technique.
[0035] CT Angiography (CTA) is a non-invasive medical imaging technique that uses computed tomography (CT) combined with the injection of contrast agents to image the vascular system, such as arteries and veins.
[0036] Non-contrast CT (NCCT): also known as plain computed tomography, is a computed tomography scan that does not use contrast agents for enhancement.
[0037] Contrast-enhanced CT (CECT): also known as enhanced computed tomography, is a computed tomography scan enhanced with contrast agents.
[0038] Generative Adversarial Network (GAN): This can be a type of deep learning model.
[0039] Denoise Diffusion Probabilistic Model (DDPM): This can be a type of deep learning model.
[0040] Mean Reverting Diffusion (MR Diff): This can be a type of deep learning model suitable for graph-to-graph tasks.
[0041] Rolling Diffusion: A diffusion model applicable to video generation of any length.
[0042] According to an embodiment of this application, an image generation method is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0043] The image generation method provided in this application embodiment can be applied to, for example, Figure 1 The application scenarios shown are not limited to these. In, for example... Figure 1 In the application scenario shown, server 10 can connect to one or more client devices 20 via a local area network (LAN), wide area network (WAN), internet connection, or other types of data network. These client devices 20 may include, but are not limited to, smartphones, tablets, laptops, PDAs, personal computers, smart home devices, and in-vehicle devices. Client devices 20 can interact with users through a graphical user interface to implement the methods provided in this application embodiment.
[0044] In this embodiment, the system consisting of a client device and a server can perform the following steps: the client device can interact with the server. The server can obtain the original image sequence of the target object; add noise to the original images in the original image sequence to obtain a noisy image sequence; and use an image generation model to denoise the noisy image sequence to generate an enhanced image sequence.
[0045] It should be noted that with the rapid development of high-performance computing units, the methods provided in this application embodiment can also be applied to model-in-machine systems in other application scenarios. In one optional embodiment, the model-in-machine system has multiple built-in models. Users can select one model to adjust as needed to obtain their own model. The high-performance computing unit built into the model-in-machine system can then directly call the adjusted model to execute the methods provided in this application embodiment. In another optional embodiment, the model-in-machine system has a pre-trained model built-in. Therefore, the high-performance computing unit built into the model-in-machine system can directly call this model to execute the methods provided in this application embodiment.
[0046] Furthermore, when users need to train their own models, they can upload their own datasets via the client. These datasets are then sent to the server, allowing the server to adjust the pre-trained model using the dataset to obtain the user's customized model, which can then be deployed to the production environment. To facilitate users' model adjustment needs, the server provides complete adjustment tools, development frameworks, and processes, supporting multiple adjustment strategies. This allows the adjusted model to better adapt to different application domains and achieve a high degree of customization.
[0047] Under the aforementioned operating environment, this application provides the following: Figure 2 The image generation method shown. Figure 2 This is a flowchart of an image generation method according to an embodiment of this application. For example... Figure 2 As shown, the method may include the following steps:
[0048] Step S202: Obtain the original image sequence of the target object.
[0049] In this sequence of original images, different original images correspond to different locations of the target object.
[0050] The target object mentioned above can refer to the object requiring image enhancement, and can be a biological object, including part or all of the human body, such as the heart, lungs, and vascular system. In other application scenarios, the target object can also be an object that needs to reflect more information through image enhancement, such as mechanical parts in industrial inspection, rock strata structures in geological exploration, or animal models in scientific research. It can be determined according to actual needs.
[0051] The aforementioned raw image sequence can refer to an image sequence containing a series of raw images, which can be arranged in order of physical location. Each raw image can reflect cross-sectional information of the target object at different locations. For example, in the medical field, it can be a set of computed tomography (CT) slice images acquired by a non-contrast-induced computed tomography (CT) scanner. In other application scenarios, raw image sequences can also originate from different imaging technologies, such as industrial X-ray inspection, ultrasound scanning, and magnetic resonance imaging, providing information about the target object from different perspectives or depths.
[0052] The aforementioned raw image can refer to a single image in a sequence of raw images, reflecting a cross-sectional view of the target object at a specific location. It can also be an image representation of a specific slice from a non-contrast computed tomography (CT) scan, containing anatomical and tissue information of the slice's location. In other applications, the raw image can be an unprocessed or unenhanced image, such as product images from a factory production line, geological structure images taken in the field, or cell structure images collected in a laboratory.
[0053] In one alternative embodiment, acquiring the original image sequence of the target object can, in medical applications, involve receiving a series of non-contrast-detector computed tomography (CT) images from a computed tomography (CT) scanner. Each original image can represent a view of the target object at different cross-sections. In industrial inspection scenarios, such as the integrity inspection of aircraft engine blades, acquiring the original image sequence of the target object can be a sequence of images taken from multiple angles. By controlling the camera angle and focal length, it can be ensured that each original image covers different areas of the blade, forming a comprehensive blade image sequence. This process can employ automated and standardized image acquisition procedures to guarantee image quality and comprehensive coverage. In geological exploration scenarios, acquiring the original image sequence can involve continuous drilling and sampling of underground rock structures, followed by 3D scanning or X-ray imaging of the samples to generate a series of images reflecting the properties of rocks at different depths. These images can be sorted according to stratum depth, forming a rock image sequence from the surface to the core, serving as the original image sequence and providing basic data for subsequent geological analysis and resource assessment.
[0054] In the above process, by acquiring the original image sequence of the target object at different locations, the original image sequence provides information about the target object on a two-dimensional plane, and through the correlation and positional information between the images, a complete image in three-dimensional space is constructed.
[0055] Step S204: Add noise to the original images in the original image sequence to obtain a noisy image sequence.
[0056] The aforementioned noisy image sequence refers to an image sequence obtained by adding noise to the original image sequence. Adding noise to the original images in the original image sequence helps the image generation model learn the rules for restoring and enhancing image features during the denoising process, thus enabling it to better simulate real-world situations when generating enhanced images. A noisy image sequence can be an original image subjected to uniform, subtle, and randomly distributed granular interference, blurring details. Through processing with deep learning algorithms, it can be cleaned up and transformed into a valuable enhanced image sequence.
[0057] The aforementioned noise addition process can be achieved by adding random noise to the pixel values of the original images in the original image sequence. The type of noise added can be Gaussian noise, etc., and can be determined according to actual needs. Noise addition can simulate the random interference that images may suffer during acquisition, transmission, or processing, and at the same time provide a learning task for subsequent image generation models to recover a clear enhanced image sequence from a noisy image sequence. In an optional embodiment, noise can be added to the original images in the original image sequence to obtain a noisy image sequence. In medical applications, a series of non-contrast-induced computed tomography images can be processed by adding random noise using a specific algorithm to construct a noisy image sequence. In other applications, such as industrial inspection, for example, detecting tiny cracks in bridge structures, the original image sequence can consist of bridge surface images taken from multiple perspectives by a high-precision camera. Noise addition can introduce random pixel fluctuations into the original image sequence, causing blurring and distortion of the images to simulate natural interferences such as light and dust in the environment, thus forming a noisy image sequence. In environmental monitoring scenarios, such as monitoring forest fires, the original image sequence can consist of forest area images taken by satellites at different times. Subsequent image generation models require a prior distribution for processing. A prior distribution can be a probabilistic description of the distribution of a parameter or variable's values based on existing knowledge, experience, hypotheses, or historical data before data is obtained or observed. This application can use a Gaussian distribution for the prior distribution, i.e., for noise addition processing.
[0058] In the above process, adding noise to the original images in the original image sequence helps the image generation model learn the rules of image feature recovery and enhancement during denoising, thus enabling it to better simulate real-world situations when generating enhanced images. Through noise addition, the image generation model needs to learn how to recover original information from images carrying random interference during training. This training method significantly enhances the robustness of noisy image sequences, allowing the image generation model to maintain good performance when facing images under non-ideal real-world conditions. The addition of random noise forces the image generation model to focus more on important features in the image, helping it to prioritize target regions when generating enhanced images, thus improving the accuracy of recognition and analysis.
[0059] Step S206: Use an image generation model to denoise the noisy image sequence and generate an enhanced image sequence.
[0060] In this process, the structure of the target object in the enhanced image sequence is the same as that in the original image sequence, and the contrast between the target region and other regions in the enhanced image sequence is greater than the preset contrast. The other regions are the regions of the target object other than the target region.
[0061] The aforementioned image generation model refers to a deep learning model used to generate images. This model can understand and learn the feature transformation rules between the original and target images. It can be built on a deep learning framework and employ algorithms combining the principles of rolling diffusion and mean regression diffusion models. The model can recover and enhance target information from noisy image sequences, generating clearer, higher-contrast images. It can be applied to various image enhancement and transformation scenarios, such as generating color photos from black and white photographs, generating high-resolution images from low-resolution images, or generating artistic images from natural images.
[0062] The denoising process described above can also be combined with a sliding window mechanism and an image generation model. When processing noisy image sequences, the sliding window mechanism allows the image generation model to focus on local regions of the noisy image sequence, thereby performing more detailed noise analysis and removal. The sliding window can move along the noisy image sequence, gradually covering different image slices. When the image generation model performs denoising within the sliding window, it can utilize the contextual information of adjacent image slices to enhance 3D continuity and reduce interlayer artifacts. This denoising process combined with a sliding window ensures that the generated enhanced image sequence maintains high quality on a single slice and also maintains structural continuity and integrity along the overall axial direction of the sequence.
[0063] The aforementioned enhanced image sequence can refer to the output of an image generation model, and can be a sequence composed of images that have undergone denoising and enhancement processing. For example, the structure of the target object in the enhanced image sequence can be the same as the structure in a non-contrast-doped computed tomography (CT) image sequence, and the contrast between the target region and other regions is increased, making the details of the target region richer and clearer. In other application scenarios, enhanced image sequences can be image sequences used in industrial inspection to highlight product defects, or image sequences used in scientific research to highlight specific biological markers, etc.
[0064] The target region mentioned above can refer to the area in the original image that needs enhancement processing. Image enhancement techniques can improve the visual recognition and information content of the target region. In medical applications, the target region can be a lesion site, the vascular system, or an anatomical structure that requires detailed observation. In other applications, the target region can be a defective area in industrial products, a fissure or mineral distribution area in geological structures, or a specific cell type or structure in a biological sample.
[0065] The aforementioned preset contrast ratio refers to a predefined contrast threshold, which can be used to determine whether the visibility and distinguishability of the target region in an enhanced image sequence meet certain standards. In medical applications, preset contrast ratio can be used to ensure that lesions provide appropriate visual differences for computed tomography angiography scans. In other applications, preset contrast ratio can be used to highlight specific features or defects; for example, in industrial inspection, preset contrast ratio can be used to distinguish minute scratches on the surface of parts.
[0066] In one optional embodiment, an image generation model can be used to denoise noisy image sequences to generate enhanced image sequences. The denoising process can be performed using a deep learning model, employing a combination of rolling diffusion and mean regression diffusion models to address the anisotropy problem in computed tomography (CT) data. In other applications, such as defect detection of industrial parts, image sequences of parts with random noise can be acquired as noisy image sequences. The image generation model, through learning, can identify and remove this noise while enhancing the contrast of defective areas, generating clear enhanced image sequences that facilitate accurate identification of minute defects by automated detection systems. In environmental monitoring, such as monitoring water pollution, by adding simulated optical interference to acquired underwater images, the image generation model can be trained to more clearly display water quality conditions and the location of potential pollutants when generating enhanced images, improving the efficiency and accuracy of environmental monitoring.
[0067] In the above process, the image generation model can recover lost or disturbed structural information from noisy image sequences, while enhancing the visual contrast of the target region, making the important information of the target region stand out more in the image. Through the denoising process, the image generation model can deeply learn the low-level features and structural relationships of the image, which helps the image generation model to more accurately identify and reconstruct target objects when processing complex images.
[0068] In this embodiment, firstly, an original image sequence of the target object is obtained, where different original images correspond to different positions of the target object. Next, noise is added to the original images in the original image sequence to obtain a noisy image sequence. Finally, an image generation model is used to denoise the noisy image sequence, generating an enhanced image sequence. The structure of the target object in the enhanced image sequence is the same as that in the original image sequence, and the contrast between the target region and other regions in the enhanced image sequence is greater than a preset contrast. The other regions are areas of the target object other than the target region. By obtaining the original image sequence of the target object, the comprehensiveness and hierarchy of the original data are ensured. Adding noise to the original images in the original image sequence facilitates the image generation model in learning the recovery and enhancement rules of image features during the denoising process, thus better simulating the real situation when generating enhanced images. By using the image generation model to denoise the noisy image sequence, the image generation model removes the random noise introduced by the noise addition, intelligently identifies and enhances the contrast of the target region, and maintains the consistency of the target object's structure in the enhanced image sequence with the original image sequence. Image generation models can learn and understand the spatial relationships of image sequences, ensuring that the inherent structural coherence and spatial positioning of the image sequence are not destroyed when enhancing the sharpness of the target region. By introducing controllable noise and using intelligent denoising in the image generation model, both content consistency and sharpness of the image are guaranteed, thus solving the technical problem in related technologies where it is difficult to balance the requirements of image content consistency and sharpness before and after image processing.
[0069] In the above embodiments of this application, adding noise to the original images in the original image sequence to obtain a noisy image sequence includes: copying the first and last original images in the original image sequence multiple times based on a preset window length to obtain an extended image sequence; determining the target image sequence located within a sliding window from the extended image sequence, wherein the sliding window slides from the first extended image to the last extended image in the extended image sequence during multiple denoising processes, and the length of the sliding window is a preset window length; and adding noise to the target image sequence to obtain a noisy image sequence.
[0070] The aforementioned preset window length refers to the length of a pre-defined sliding window. The preset window length can be determined based on the characteristics of the data and the expected processing effect. The preset window length can affect the number of slices processed in each iteration of the image generation model, balancing computational resources and information continuity. For example, in medical imaging, the choice of preset window length can affect the consistency of identifying blood vessels or lesions in consecutive slices. The specific preset window length can be determined according to actual needs.
[0071] The aforementioned extended image sequence refers to an image sequence whose length is increased by copying the first and last images of the original sequence. Obtaining an extended image sequence ensures sufficient contextual information at the boundaries of the image sequence during image processing using a sliding window mechanism, avoiding processing inconsistencies or information loss due to sequence boundaries. By extending the image sequence, the sliding window can maintain a complete contextual environment at both ends of the sequence, thus guaranteeing the uniformity of image processing and consistent results.
[0072] The aforementioned sliding window refers to a data processing technique that moves a window of a preset size across an image sequence. It can be used to process image data segment by segment, ensuring that each processing unit has sufficient contextual information. The use of sliding windows can effectively process large-scale or long sequences of image data, while maintaining information continuity and local consistency during image processing. It is also suitable for processing image data with three-dimensional structures or sequence dependencies.
[0073] In one optional embodiment, a reasonable sliding window length can be pre-set based on the characteristics of the original image sequence, such as slice intervals, the size and shape of the target tissue, to ensure that the area covered by the window reflects sufficient local information without being too long and wasting computational resources. Next, sequence expansion can be performed by copying the first and last images of the original image sequence several times to form an expanded image sequence. The number of copies can be determined by the window length to ensure that the window has an appropriate number of buffer slices to provide contextual information when covering the sequence boundaries. Then, the sliding window can be used to move from the first image of the expanded image sequence at a preset step size until the last image of the sequence. During each window slide, the images within the sliding window can be processed as a whole, ensuring information continuity and local consistency during processing. Finally, noise can be added to the target image sequence located within the sliding window to generate a noisy image sequence.
[0074] In the above process, the sliding window mechanism and sequence expansion method ensure the continuity of information during image processing. Even when the sliding window is at the boundary of the image sequence, high-quality processing results can be obtained, improving the overall effect of image generation. By selecting the preset window length, the target region in the image sequence can be focused on, more effectively enhancing and restoring the features of specific structures, such as blood vessel enhancement, crack detection, or texture refinement. The sliding window mechanism allows the image generation model to flexibly handle image sequences of different lengths and structures, making it highly adaptable and widely applicable to various scenarios from medical imaging to industrial inspection. The strategy of expanding the image sequence effectively avoids the boundary effect that occurs when processing sequence edges, ensuring the uniformity and stability of the image processing effect throughout the entire image sequence and maintaining the coherence of the image sequence.
[0075] In the above embodiments of this application, the noise processing of the target image sequence to obtain a noise image sequence includes: adding noise to the last target image in the target image sequence to obtain a first noise image; determining second noise images of other target images, wherein the other target images are target images in the target image sequence other than the last target image, and the second noise image is obtained by adding noise to the other target images, or by denoising the images using an image generation model; and obtaining a noise image sequence based on the first noise image and the second noise image.
[0076] The aforementioned first noisy image can refer to the image obtained by adding noise to the last image in the target image sequence. The noise addition process can be based on a preset noise distribution pattern or algorithm to add random fluctuations or distortions to the pixel values in the image, thereby simulating the noise interference encountered by the image during acquisition, transmission, or storage.
[0077] The aforementioned second noisy image can refer to an image obtained by adding noise to all images in the target image sequence except the last one, or a denoised image obtained after preliminary denoising by an image generation model. During image sequence processing, other target images are denoised or added with varying degrees of noise to construct a dynamically changing noisy image sequence. This better reflects the state changes of the image during the generation process and helps the image generation model learn its ability to recover image features under different noise levels.
[0078] In one optional embodiment, when processing each target image within the target image sequence, random noise of varying intensities can be added to the target images based on their position within the sequence and a preset noise distribution strategy. The last target image is significantly affected by noise and can be generated as a first noisy image. Other target images can undergo different levels of noise enhancement based on their distance from the last target image, generating corresponding second noisy images. The first noisy image can then be combined with multiple second noisy images to form a noisy image sequence. This sequence can include transitions from high to low noise levels to maintain the continuity of the generated image sequence and improve the parallelism of the image generation model, i.e., generation efficiency. Furthermore, the constructed noisy image sequence can be used to train the image generation model to learn the process of recovering a clear image from a noisy image. During training, the image generation model can gradually learn to identify and compensate for image distortion at different noise levels, ultimately predicting and generating a clear, contrast-enhanced image sequence from the noisy image sequence.
[0079] In the above process, by adding different levels of noise to different target images in the target image sequence, the image generation model can learn to more closely resemble the patterns of noise variation in the real world, thus making the generated images more realistic and reliable during the prediction stage. The noise level of each image in the noisy image sequence can gradually change from high to low. The image generation model can maintain the continuity of the image sequence during the denoising process, ensuring that the generated enhanced image sequence is structurally consistent with the original image sequence, which helps to maintain the coherence of the image sequence.
[0080] In the above embodiments of this application, the noise levels of different noise images in the same noise image sequence are different, and the noise levels of noise images at the same position in different noise image sequences are the same.
[0081] In one optional embodiment, the noise levels of different noisy images within the same noisy image sequence can be set to be different, while the noise levels of noisy images at the same location within different noisy image sequences can be the same. This enhances the image generation model's understanding and processing capabilities of image noise, while ensuring consistency in the noise addition and removal process across different sequences, thereby improving the quality and reliability of the generated images. First, the noise level can be defined as the standard deviation of the Gaussian noise added to the image, directly affecting the image's clarity and information integrity. A higher noise level indicates greater random interference, while a lower noise level indicates less random interference. Within the same noisy image sequence, the noise level can exhibit a hierarchical characteristic, with the first image in the sequence having a lower noise level and the last image having a higher noise level, serving as a reference point for simulating worst-case scenarios. Images in the middle of the noisy image sequence can be assigned different noise levels based on their positions and preset rules, forming a noise gradient from low to high and then gradually returning to low. Between different noisy image sequences, images at the same location can be assigned the same noise level. That is, if the Nth image in two sets of original image sequences is at the same location, then after being converted into a noisy image sequence, the noise intensity of the Nth noisy image will also remain consistent. This design ensures that the image generation model can establish stable and reliable noise processing standards during the training and prediction phases.
[0082] In the above setup, through hierarchical noise management, the image generation model can learn how to recover from a high-noise state to a clear state. It also learns processing techniques under different noise levels, expanding the learning depth and breadth of the image generation model and improving its robustness to various noise interferences. Maintaining consistent noise levels at the same location across noisy image sequences allows the image generation model to eliminate the influence of noise variations when conducting comparative analysis and performance evaluation across different sequences, ensuring the fairness of the evaluation results and the validity of the comparison. During training, the image generation model is exposed to diverse noise patterns and distributions, which helps improve its generalization ability. It can accurately identify and process noise in unseen image sequences, thereby generating high-quality enhanced image sequences.
[0083] In the above embodiments of this application, the noise level of the previous noise image in the same noise image sequence is lower than the noise level of the next noise image.
[0084] In one optional embodiment, the noise level of the preceding noisy image in the same noisy image sequence can be set to be lower than the noise level of the following noisy image. This strategy, by constructing an increasing gradient of noise levels in the image sequence, provides a more complex and hierarchical training environment for the deep learning model, which can improve the denoising capability and fidelity of the generated images. First, the construction strategy of the noisy image sequence is determined, that is, the noise level in the noisy image sequence will gradually increase from the first image until the end of the noisy image sequence. The images are relatively clear at the beginning of the noisy image sequence, while the degree of noise affecting the images gradually increases as the noisy image sequence progresses. Next, based on a preset noise distribution rule, a specific noise level can be assigned to each image in the noisy image sequence. Then, random noise of corresponding intensity can be added to each original image in the target image sequence according to the position of the original image in the sequence, generating a noisy image sequence. At the beginning of the noisy image sequence, the noise intensity is low, and the image details are relatively well preserved; at the end of the noisy image sequence, the noise intensity is high, and the image information is masked.
[0085] In the process described above, by constructing a gradient of noise levels within the same noisy image sequence, the image generation model can learn how to identify and progressively remove noise of varying intensities, thereby gaining a deeper understanding of the image's underlying features and structural information. This hierarchical learning process enhances the noise adaptability and processing capabilities of the image generation model. During the generation and processing of gradient-noise image sequences, the image generation model can learn to remove noise while being guided to maintain the image's basic structure. Because the noise level increases from low to high, when processing high-noise images at the end of the sequence, the image generation model can refer to the clear images at the beginning of the sequence, maintaining the structural consistency of the generated images and avoiding image distortion caused by excessive denoising.
[0086] In the above embodiments of this application, the image generation model is used to denoise the noisy image sequence to generate an enhanced image sequence, including: using the image generation model to denoise the noisy image sequence to generate an enhanced image and a denoised image, wherein the enhanced image corresponds to the first target image in the target image sequence, and the denoised image corresponds to the noisy images in the target image sequence other than the first target image; and an enhanced image sequence is obtained based on the enhanced image obtained from multiple denoising processes.
[0087] The enhanced image mentioned above refers to an image that, after processing by an image generation model, has significantly improved contrast, making the distinction between the target region and other regions more obvious. The enhanced image corresponds to the first original image in the target image sequence. Starting from the first original image, the image generation model generates an image with enhanced contrast and clearer details through denoising and feature enhancement operations.
[0088] The aforementioned denoised image refers to an image that has been processed by an image generation model to remove noise interference and restore the original clarity and detail. The denoised image corresponds to the noisy images in the target image sequence excluding the first original image. That is, the image generation model performs denoising processing on the remaining noisy images in the target image sequence one by one to restore or enhance the original image features and contrast.
[0089] In one optional embodiment, the image generation model needs to load preset parameters and noise distribution strategies before processing, preparing to denoise the noisy image sequence. The image generation model can process the first noisy image in the noisy image sequence. Through deep learning technology, the image generation model can enhance the contrast of the target region and generate the first enhanced image, which may include operations such as feature extraction, contrast enhancement, and detail restoration. Next, the image generation model can continue to process other noisy images in the noisy image sequence, gradually removing random noise from each image, restoring image clarity and original details, and generating a denoised image. During the image generation process, the image generation model can iterate and improve multiple times to ensure good structural and visual coherence between the generated enhanced and denoised images. This continuity ensures the overall quality of the image sequence and the smoothness of the analysis process. Finally, based on the enhanced and denoised images obtained from multiple denoising processes, they can be resynthesized in the order of the original image sequence to obtain a complete enhanced image sequence. The enhanced image sequence has a significant enhancement in contrast, and the coherence and integrity of image information are also guaranteed.
[0090] In the above process, an image generation model is used to obtain an enhanced image sequence based on the enhanced image obtained from multiple denoising processes. The contrast between the target area and other areas is significantly enhanced. In medical applications, small tumors or thrombi that are difficult to identify in the original plain CT scan image can be clearly presented in the enhanced image. In the process of generating enhanced and denoised images, the image generation model follows the anatomical structure and inter-slice continuity of the original image sequence, avoiding morphological distortion or inter-slice break artifacts during image processing, and ensuring the coherence of the image sequence.
[0091] In the above embodiments of this application, the image generation model is used to denoise a noisy image sequence to generate an enhanced image and a denoised image. This includes: using a first attention module in the image generation model to perform attention processing on the noisy images in the noisy image sequence to obtain a first processing result, wherein the first processing result is used to characterize the relationship between different pixels in the noisy image; using a second attention module in the image generation model to perform attention processing on two adjacent noisy images in the noisy image sequence to obtain a second processing result, wherein the second processing result is used to characterize the relationship between different noisy images; and using a generation module in the image generation model to summarize the first processing result and the second processing result to generate an enhanced image and a denoised image.
[0092] The aforementioned first attention module can refer to a module in an image generation model used to identify and focus on local features and details within an image. The first attention module can assign higher weights to highly correlated regions by calculating the similarity or correlation between each pixel and other pixels in the image, thereby prioritizing and preserving information from these highly correlated regions during processing.
[0093] The first processing result mentioned above can refer to the result representing the relationship between different pixels within the noisy image after processing by the first attention module. The first processing result can reflect which parts of the image are more important and which parts need more information recovery or feature enhancement.
[0094] The aforementioned second attention module can refer to a module used to capture and utilize the informational correlations between adjacent images in an image sequence. When processing 3D image sequences, the second attention module can perform attention processing on the same positions of two adjacent images, identifying temporal or spatial consistency and difference features between images, which helps maintain the continuity and structural consistency of the image sequence.
[0095] The second processing result mentioned above can refer to the result representing the relationship between different noisy images after processing by the second attention module. The second processing result emphasizes the coherence of the image sequence at the overall level and the interdependence between each slice, providing important contextual information for the generation process.
[0096] The aforementioned generation module can refer to the module in the image generation model that is responsible for integrating the first processing result and the second processing result. It can utilize deep learning algorithms to generate enhanced images and denoised images.
[0097] In one optional embodiment, when processing each noisy image, the first attention module can perform local attention processing, identifying regions containing important information or details, such as the edges of blood vessels or the outlines of lesions in a medical setting, and assigning these regions higher weights. Simultaneously, the second attention module can focus on the similarities and differences between adjacent images, ensuring that the generated images maintain structural coherence within the sequence. Next, the first and second processing results can be fed into the generation module. The generation module can utilize the information from the first and second processing results, employing a deep learning generation process to remove noise and enhance the contrast of the target region, generating denoised and enhanced images. This process can be iterated multiple times, with each iteration adjusting and improving based on the previous generation result to achieve the desired image quality and contrast level. Finally, the generation module can summarize the first and second processing results to generate enhanced and denoised images.
[0098] In the above process, the use of the first attention module enables the image generation model to accurately identify and restore key structures and details in the image, and effectively extract and enhance features of the target region even in high-noise environments. The second attention module ensures the continuity and structural consistency of the image sequence during processing, avoiding artifacts or structural misalignments between layers in the generated image, thus maintaining the coherence and diagnostic value of the image sequence. By integrating local and sequence information, the generation module can more intelligently generate enhanced and denoised images, improving image clarity and contrast while maintaining structural authenticity. This provides higher accuracy and detection efficiency for applications such as medical imaging and industrial inspection.
[0099] The technical solution proposed in this application is described below with reference to an optional embodiment. This application proposes a method for synthesizing virtual angiography in computed tomography (CT) scans. CT scans possess characteristics such as high spatial resolution and clear display of anatomical structures. Among various CT scan techniques, non-contrast-enhanced CT scans are simple to operate, require no contrast agents, have short examination times, and are non-invasive. During CT imaging, different tissues attenuate X-rays to varying degrees. This difference is quantified into Henlein units and presented in grayscale on the image, thereby achieving visualization of tissue structures. However, non-contrast-enhanced CT scans have certain limitations in clinical applications. The density of certain lesions, such as lung tumors, mediastinal lesions, or pulmonary thrombosis, is similar to that of surrounding normal tissues, resulting in low contrast and blurred boundaries in non-contrast-enhanced CT scans. However, by intravenously injecting iodine-containing contrast agents, contrast-enhanced CT scans can significantly enhance the enhancement effect of blood vessels and lesion tissues, thereby more clearly displaying the extent of the lesion and blood supply. However, the clinical application of contrast-enhanced CT scans is somewhat limited. Iodine-containing contrast agents can trigger allergic reactions and are nephrotoxic, potentially inducing acute kidney injury or even renal failure in high-risk individuals such as those with renal insufficiency, diabetes, or the elderly. Furthermore, the use of contrast agents increases examination costs and the burden on patients, and some patients cannot tolerate contrast-enhanced scans due to contraindications. Therefore, how to obtain diagnostic information similar to contrast-enhanced CT scans from non-contrast-enhanced CT images without relying on additional imaging scans and contrast agent injections has become a pressing technical problem in the field of medical imaging. The technical solution proposed in this application, based on artificial intelligence-based image conversion technology, can directly generate high-quality virtual contrast-enhanced CT images from non-contrast-enhanced CT data. This not only avoids the risks associated with contrast agents and reduces medical costs but also provides an important diagnostic pathway for patients who cannot tolerate contrast-enhanced scans.
[0100] The technical framework of this application, through ingenious mechanism design, achieves high-fidelity and high-continuity conversion from non-contrast-enhanced computed tomography (CT) scans to contrast-enhanced CT scans. To ensure the authenticity of the anatomical structures in the synthesized images, this application employs a mean-regression diffusion framework. In this framework, the original non-contrast-enhanced CT image can be considered as a mean or anchor point. During training, the image generation model learns to progressively degrade the mapping from the target contrast-enhanced CT image to the source non-contrast-enhanced CT image. During image synthesis, the image generation model can start from a noise state close to that of the non-contrast-enhanced CT scan and is continuously guided by the original non-contrast-enhanced CT scan in each denoising step, thereby ensuring that basic anatomical structures such as bone and organ contours are accurately preserved. The entire framework's process includes the training of the image generation model, where the model learns to evolve from clear contrast-enhanced computed tomography (CT) scans to noise distribution centered on non-contrast-enhanced CT scans; and the image generation process, where the model, guided by non-contrast-enhanced CT scans, can gradually restore noisy images to high-quality contrast-enhanced CT scans.
[0101] Figure 3 This is a schematic diagram illustrating the training and image generation process of an optional image generation model according to an embodiment of this application, such as... Figure 3 As shown, from right to left, the training process of the image generation model is shown. The image generation model can use forward stochastic differential equations to evolve from contrast-enhanced computed tomography (CT) images to non-contrast-enhanced CT images. From left to right, the image generation process of the image generation model is shown. The image generation model can use backward stochastic differential equations to generate contrast-enhanced CT images based on non-contrast-enhanced CT images.
[0102] The diffusion and generation processes of the diffusion model can be described by stochastic differential equations (SDEs). The stochastic differential equation corresponding to the classical diffusion probability model (Denoising Diffusion Probabilistic Model, abbreviated as DDPM) determines that the initial state of the sampling process is a normal distribution. The mean regression diffusion model adopted in this application, by designing a new stochastic differential equation, makes the initial state of the sampling process non-contrast-enhanced computed tomography (CT) with Gaussian noise, naturally injecting the information of non-contrast-enhanced CT into the sampling process. The diffusion process of the image generation model describes the transformation of the target distribution, i.e., contrast-enhanced CT, into the prior distribution, i.e., non-contrast-enhanced CT with Gaussian noise. In each training step, a data example x and a time point t can be randomly sampled, and the neural network fits the transformation from x(t) to x(0). The generation process of the image generation model, also known as the sampling process, describes the transformation process from the prior distribution to the target distribution. This process also corresponds to a stochastic differential equation, i.e., numerically solving the stochastic differential equation. This process requires the participation of a neural network.
[0103] To address the anisotropy and interslice fracture artifacts in 3D computed tomography (CT) data along the Z-axis, this application designs a Z-axis continuity assurance and artifact suppression module. A sliding window mechanism moves along the Z-axis, i.e., the slice sequence direction, combined with an asynchronous denoising strategy based on a rolling diffusion model. Slices within the sliding window can have a stepped noise level distribution, with newly entered slices having higher noise intensity and slices nearing completion being clearer. During the asynchronous denoising process, as the sliding window rolls along the Z-axis, a clean, denoised contrast-enhanced CT slice is pushed out of the window, while a slice from the original non-contrast-enhanced CT scan is pulled in and given higher noise, resulting in a stepped noise level distribution among the slices within the sliding window. This design ensures that each slice is generated within a sufficient context, effectively maintaining Z-axis continuity.
[0104] The stepped noise distribution is associated with the sliding window. Since the sampling process of the diffusion model can be simply understood as removing noise from the image, denoising multiple two-dimensional images together allows the neural network to capture the correlation information between the images. Meanwhile, three-dimensional images are large in size, making overall denoising difficult. This application employs asynchronous denoising to balance spatial continuity and temporal efficiency.
[0105] Considering the boundary effect problem encountered by this rolling generation mechanism when processing the beginning and end of the sequence, this application designs a sequence filling sampling strategy to address this technical challenge. Before generation begins, i.e., in the initial stage, the first slice of the original sequence can be copied several times and added to the beginning of the sequence to warm up the sliding window, allowing it to enter a stable stepped noise working mode when processing the first real slice. Similarly, at the end of the sequence, i.e., in the final stage, the last slice can be copied for filling, ensuring that the last few original slices can also be processed within the complete context window. The generated filling portion can eventually be discarded. This strategy cleverly avoids designing special training modes for boundary cases, simplifies the implementation of the image generation model and the sampling process, while maintaining the structural continuity of the boundaries. The above functions can be performed by an anisotropic three-dimensional network architecture. A three-dimensional U-shaped network can be used as the main body, and the three-dimensional U-shaped network has been deeply customized. Considering the resolution difference of computed tomography data in the XY plane and Z axis, the downsampling and upsampling operations in the network can be performed only in the XY plane, thus completely preserving the original correspondence of the Z axis. During the training phase, the neural network takes x and t as inputs and learns the output x(0). Here, x can be a series of adjacent slices, and t can be a time series. Since it is asynchronous denoising, each slice can correspond to a different time t. Because the encoding of time t is injected into each layer of the neural network, the Z-axis dimension of x can remain unchanged and correspond to the time series t.
[0106] To support asynchronous denoising, the time-step embedding module has been extended to provide independent time codes corresponding to the noise level of each slice in the sliding window. Time t can be converted into a vector through sinusoidal position encoding before being input into the neural network, which helps the neural network learn concepts such as relative position and periodicity, and also makes it easier to integrate with information from x.
[0107] Furthermore, this application deeply integrates two attention mechanisms into the neural network: spatial attention, which captures details within each slice, and axial attention, which captures inter-layer dependencies along the Z-axis. Attention mechanisms excel at capturing the relationships between elements within a sequence. Spatial attention can flatten the X and Y axes into a single dimension, capturing the relationships within the XY plane; axial attention treats each slice as an element, capturing the relationships between slices along the Z-axis.
[0108] Figure 4 This is a schematic diagram of an optional process for generating an image enhancement sequence using a sliding window according to an embodiment of this application, such as... Figure 4As shown, this is the process of generating an enhanced image sequence from an original image sequence using a sliding window. The process begins by repeatedly copying the first original image in the original image sequence, resulting in the image sequence (-4, -3, -2, -1, 0, 1, 2, 3). Next, random noise is added to each original image within the sliding window, followed by denoising. The padded portion is discarded from the resulting image sequence, yielding the enhanced image sequence (0, 1, 2, 3). As the sliding window moves, random noise is added to the original image sequence (k-3, k-2, k-1, k, k+1, k+2, k+3) within the sliding window, followed by denoising, resulting in the enhanced image sequence (k-3, k-2, k-1, k, k+1, k+2, k+3). As the sliding window moves to the end of the original image sequence, the last original image in the original image sequence can be copied multiple times to obtain the image sequence (z, z+1, z+2, z+3, z+4); random noise can be added to the original images within the sliding window, and the resulting image sequence can be denoised to obtain the enhanced image sequence (z), at which point the process ends.
[0109] The above series of designs enables neural networks to perform the aforementioned structure preservation and continuity generation tasks efficiently and accurately.
[0110] The proposed technical solution achieves high-fidelity structural alignment using a mean-regression diffusion framework. The original non-contrast-induced computed tomography (CT) scan is considered the anchor point for generation, and each denoising step incorporates the non-contrast-induced CT scan as a mandatory reference, ensuring a one-to-one correspondence between the generated CT scan and the original structure and preventing structural misalignment. For 3D continuity and Z-axis artifact suppression, a sliding window asynchronous denoising mechanism is designed. The noise intensity of different slices within the sliding window increases progressively, allowing each CT scan to rely on its preceding and following context, resulting in a natural transition of the Z-axis structure and effectively eliminating inter-slice artifacts. Boundary handling is improved by using a sequence-filling sampling strategy, filling the beginning and end of the sequence with repeated slices, ensuring the sliding window has a complete context when handling boundaries. Finally, the filled portion is discarded, achieving end-to-end structural continuity.
[0111] This application proposes a generative model designed to address the anisotropy problem of computed tomography (CT) data. An asynchronous sliding window and sequence filling strategy is proposed to ensure the continuity of the Z-axis (slice continuity) and avoid artifacts when processing 3D CT scans. A rolling diffusion model is employed, and a sequence filling strategy is proposed. By filling the beginning and end of the data sequence with repeated slices, the boundary problem of the sliding window method at the beginning and end stages is cleverly solved, thus eliminating the need for complex special training and sampling algorithms and simplifying the process. For the conditional diffusion model in angiography CT synthesis, to address the technical problem that standard diffusion models, generated from pure noise, easily lose key anatomical structures in non-contrast-enhanced CT scans, a method of applying average regression diffusion to angiography CT synthesis is proposed. This innovatively uses the input non-contrast-enhanced CT scan as the regression target of the diffusion process, enabling the image generation model to generate vascular enhancement effects while being strongly constrained by the original anatomical structures, thereby preserving structural fidelity to the maximum extent.
[0112] According to an embodiment of this application, an image synthesis method is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in an order not used here.
[0113] Figure 5 This is a flowchart of an image synthesis method according to an embodiment of this application, such as... Figure 5 As shown, the specific steps may include the following:
[0114] Step S502: Obtain the original medical image sequence of the biological object.
[0115] In this context, different original medical images in the original medical image sequence correspond to different tissue locations of the biological object.
[0116] The aforementioned biological objects can refer to the organisms scanned in medical imaging technology, including but not limited to human bodies, animals, or other biological samples. These biological objects contain lesions that need to be identified, such as tumors, thrombi, and complex vascular systems, which are key areas of focus for analysis in medical imaging.
[0117] The aforementioned raw medical image sequence can refer to a series of images obtained after scanning a biological object using non-invasive medical imaging equipment. The images in the raw medical image sequence can be arranged in temporal or spatial order, and can comprehensively reflect multiple cross-sectional views of the internal structure of the biological object. For example, it can be a non-contrast-enhanced computed tomography scan result, which can contain raw images of different tissue locations of the biological object, but lacks details of blood vessels and lesions under contrast enhancement, and has lower contrast.
[0118] Step S504: Add noise to the original medical images in the original medical image sequence to obtain a noisy image sequence.
[0119] Step S506: Denoise the noisy image sequence using an image generation model to generate an enhanced medical image sequence.
[0120] In this sequence, the anatomical structure of the biological object in the enhanced medical image sequence is the same as that in the original medical image sequence, and the contrast between the vascular system region and other regions in the enhanced medical image sequence is greater than the preset contrast. The other regions are the areas of the biological object other than the vascular system region.
[0121] The aforementioned enhanced medical image sequence refers to the result obtained by processing the original medical image sequence using an image generation model. The images in the enhanced medical image sequence maintain the same anatomical structure as the original sequence, but the contrast between the vascular system region and other regions is significantly improved. The image generation model removes random perturbations introduced by noise processing and enhances the features of the target region. The effect of the enhanced medical image sequence can be similar to that of a computed tomography scan enhanced with contrast agents, but no contrast agents are actually used in the actual generation process; instead, the contrast effect is simulated through algorithms.
[0122] The aforementioned vascular system region refers to the vascular network region within a biological object, which may include arteries, veins, and structures related to blood flow. In medical imaging, improving the clarity of the vascular system region can reveal blood flow status, vascular morphology, and lesions that can affect blood flow, such as tumors and thrombi. In enhanced medical imaging sequences, the contrast between the vascular system region and other regions is significantly enhanced, making blood vessels more clearly discernible.
[0123] In one alternative embodiment, a raw medical image sequence containing different tissue locations within a biological subject can be acquired using non-invasive medical imaging techniques, such as non-contrast-free computed tomography (CT). The raw medical image sequence can display basic anatomical structures, but the vascular system is poorly recognizable due to the lack of contrast enhancement. Next, to train an image generation model to better understand and remove inherent or artificially introduced noise from the images, noise can be added to each image in the raw medical image sequence. This noise addition process can be achieved by adding random noise to the images, constructing a noisy image sequence. Then, an image generation model can be used to denoise the noisy image sequence. The image generation model can remove random noise from the images and significantly enhance the contrast between the vascular system region and other regions while maintaining the realism of the anatomical structures. The image generation model can perform feature extraction, noise estimation and removal, and contrast enhancement on the input image. By combining a mean regression diffusion model and a rolling diffusion model, computed tomography images with large slice intervals can be flexibly processed while ensuring axial continuity.
[0124] In the aforementioned process, enhancing the contrast of the vascular system ensures the authenticity and integrity of the anatomical structures, avoiding structural distortions in image processing and facilitating medical image analysis. Through precise denoising and contrast enhancement, the contrast between the vascular system region and other regions is significantly improved, making vascular details more clearly visible. This avoids the use of contrast agents, reducing the risk of allergic reactions and kidney damage, while also lowering examination costs, making it highly valuable for high-risk populations and resource-limited medical institutions. Combining the characteristics of mean regression diffusion models and rolling diffusion models, an efficient conversion from plain non-contrast-enhanced computed tomography (CT) scans to contrast-enhanced CT scans is achieved.
[0125] According to an embodiment of this application, an image synthesis method is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in an order not used here.
[0126] Figure 6 This is a flowchart of an image synthesis method according to an embodiment of this application, such as... Figure 6 As shown, the specific steps may include the following:
[0127] Step S602: Obtain the original image sequence of the target object by calling the first interface.
[0128] The first interface includes a first parameter, the value of which includes an original image sequence, and different original images in the original image sequence correspond to different positions of the target object.
[0129] Step S604: Add noise to the original images in the original image sequence to obtain a noisy image sequence.
[0130] Step S606: Use an image generation model to denoise the noisy image sequence and generate an enhanced image sequence.
[0131] In this process, the structure of the target object in the enhanced image sequence is the same as that in the original image sequence, and the contrast between the target region and other regions in the enhanced image sequence is greater than the preset contrast. The other regions are the regions of the target object other than the target region.
[0132] Step S608: Output the enhanced image sequence by calling the second interface.
[0133] The second interface includes a second parameter, the value of which includes the enhanced image sequence.
[0134] According to another aspect of the present invention, an image generation apparatus is also provided, which can execute the image generation method of the above embodiments. The specific implementation method and preferred application scenarios are the same as those of the above embodiments, and will not be described in detail here.
[0135] Figure 7 This is a schematic diagram of an image generation apparatus according to an embodiment of this application, such as... Figure 7 As shown, the device includes the following: an acquisition module 702, a noise processing module 704, and a noise reduction processing module 706.
[0136] The acquisition module 702 is used to acquire the original image sequence of the target object, wherein different original images in the original image sequence correspond to different positions of the target object; the noise processing module 704 is used to add noise to the original images in the original image sequence to obtain a noisy image sequence; the denoising processing module 706 is used to denoise the noisy image sequence using an image generation model to generate an enhanced image sequence, wherein the structure of the target object in the enhanced image sequence is the same as the structure in the original image sequence, and the contrast between the target region and other regions in the enhanced image sequence is greater than a preset contrast, and the other regions are the regions of the target object other than the target region.
[0137] The noise-adding module is further configured to: copy the first and last original images in the original image sequence multiple times based on a preset window length to obtain an extended image sequence; determine the target image sequence located within a sliding window from the extended image sequence, wherein the sliding window slides from the first extended image to the last extended image in the extended image sequence during multiple noise reduction processes, and the length of the sliding window is a preset window length; and add noise to the target image sequence to obtain a noisy image sequence.
[0138] The noise processing module is further used to add noise to the last target image in the target image sequence to obtain a first noise image; determine the second noise images of other target images, wherein the other target images are target images in the target image sequence other than the last target image, and the second noise image is obtained by adding noise to the other target images or by denoising the images using an image generation model; and obtain a noise image sequence based on the first noise image and the second noise image.
[0139] Among them, different noise images in the same noisy image sequence have different noise levels, while noise images at the same location in different noise image sequences have the same noise level.
[0140] In the same noisy image sequence, the noise level of the previous noisy image is lower than the noise level of the next noisy image.
[0141] The denoising module is also used to denoise the noisy image sequence using an image generation model to generate an enhanced image and a denoised image. The enhanced image corresponds to the first target image in the target image sequence, and the denoised image corresponds to the noisy images in the target image sequence other than the first target image. An enhanced image sequence is obtained based on the enhanced image obtained from multiple denoising processes.
[0142] The denoising module is further configured to use the first attention module in the image generation model to perform attention processing on the noisy images in the noisy image sequence to obtain a first processing result, wherein the first processing result is used to characterize the relationship between different pixels in the noisy image; use the second attention module in the image generation model to perform attention processing on two adjacent noisy images in the noisy image sequence to obtain a second processing result, wherein the second processing result is used to characterize the relationship between different noisy images; and use the generation module in the image generation model to summarize the first processing result and the second processing result to generate an enhanced image and a denoised image.
[0143] According to another aspect of the present invention, an image synthesis apparatus is also provided, which can perform the image synthesis method of the above embodiments. The specific implementation method and preferred application scenarios are the same as those of the above embodiments, and will not be described again here.
[0144] Figure 8 This is a schematic diagram of an image synthesis apparatus according to an embodiment of this application, such as... Figure 8 As shown, the device includes the following: an acquisition module 802, a noise processing module 804, and a noise reduction processing module 806.
[0145] The acquisition module 802 is used to acquire the original medical image sequence of the biological object, wherein different original medical images in the original medical image sequence correspond to different tissue locations of the biological object; the noise processing module 804 is used to add noise to the original medical images in the original medical image sequence to obtain a noisy image sequence; the denoising processing module 806 is used to denoise the noisy image sequence using an image generation model to generate an enhanced medical image sequence, wherein the anatomical structure of the biological object in the enhanced medical image sequence is the same as the anatomical structure in the original medical image sequence, and the contrast between the vascular system region and other regions in the enhanced medical image sequence is greater than a preset contrast, and the other regions are the regions of the biological object other than the vascular system region.
[0146] According to another aspect of the present invention, an image synthesis apparatus is also provided, which can perform the image synthesis method of the above embodiments. The specific implementation method and preferred application scenarios are the same as those of the above embodiments, and will not be described again here.
[0147] Figure 9 This is a schematic diagram of an image synthesis apparatus according to an embodiment of this application, such as... Figure 9 As shown, the device includes the following: a first calling module 902, a noise processing module 904, a noise reduction processing module 906, and a second calling module 908.
[0148] The system comprises the following modules: a first calling module 902, used to obtain the original image sequence of the target object by calling a first interface, wherein the first interface includes a first parameter, the value of which includes the original image sequence, and different original images in the original image sequence correspond to different positions of the target object; a noise processing module 904, used to add noise to the original images in the original image sequence to obtain a noisy image sequence; a denoising processing module 906, used to denoise the noisy image sequence using an image generation model to generate an enhanced image sequence, wherein the structure of the target object in the enhanced image sequence is the same as the structure in the original image sequence, and the contrast between the target region and other regions in the enhanced image sequence is greater than a preset contrast, and the other regions are regions of the target object other than the target region; and a second calling module 908, used to output the enhanced image sequence by calling a second interface, wherein the second interface includes a second parameter, the value of which includes the enhanced image sequence.
[0149] Optionally, Figure 10 This is a structural block diagram of a computing device according to an embodiment of this application, such as... Figure 10 As shown, the computing device A may include one or more (only one is shown in the figure) processors 102, memory 104 and peripheral interfaces 106, wherein the processors 102, memory 104 and peripheral interfaces 106 are interconnected via a bus 108.
[0150] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the methods and apparatus in the embodiments of this application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby implementing the methods in the above embodiments. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to terminal A via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0151] The processor can access information and applications stored in the memory via a transmission device to execute the steps in each embodiment.
[0152] Embodiments of this application may provide an electronic device. Figure 11 This is a structural block diagram of an electronic device according to an embodiment of this application, such as... Figure 11 As shown, the electronic device may include: an input / output device 1102; a memory 1104; and a processor 1106, wherein the processor 1106 is connected to the input / output device 1102 and the memory 1104 via a bus 1108.
[0153] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the methods and apparatus in the embodiments of this application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby implementing the methods in the above embodiments. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to terminal A via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0154] The processor can invoke an executable program stored in memory via a transmission device to perform the following methods: acquiring an original image sequence of a target object, wherein different original images in the original image sequence correspond to different positions of the target object; adding noise to the original images in the original image sequence to obtain a noisy image sequence; using an image generation model to denoise the noisy image sequence to generate an enhanced image sequence, wherein the structure of the target object in the enhanced image sequence is the same as the structure in the original image sequence, and the contrast between the target region and other regions in the enhanced image sequence is greater than a preset contrast, and the other regions are regions of the target object other than the target region; and executing the methods in the various embodiments of this application.
[0155] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0156] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0157] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, they can also be implemented by hardware. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause the processing unit to execute the methods in the various embodiments of this application.
[0158] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0159] Embodiments of this application also provide a computer-readable storage medium. Optionally, in this embodiment, the computer-readable storage medium can be used to store program code executed by the method provided in the above embodiments.
[0160] Optionally, in this embodiment, the storage medium may be located in a computing device.
[0161] Optionally, in this embodiment, the computer-readable storage medium is configured to store an executable program, which, when the executable program is running, controls the device where the computer-readable storage medium is located to execute the method described in any of the above embodiments.
[0162] Embodiments of this application also provide a computer program product. Optionally, in this embodiment, the computer program product may include a computer program that, when executed by a processor, implements the methods provided in the embodiments described above.
[0163] The aforementioned computer program products can refer to software programs that have been written, tested, and released, and can run on computers or other devices. Computer program products can include application programs, operating systems, utility software, etc., used to achieve specific functions or solve specific problems.
[0164] Embodiments of this application also provide a computer program product. Optionally, the computer program product may include a non-volatile computer-readable storage medium, which can be used to store a computer program that, when executed by a processor, implements the method provided in the above embodiments.
[0165] The aforementioned non-volatile computer-readable storage medium can refer to a medium for storing data. Non-volatile computer-readable storage media can retain data without loss when power is off and can be used to store long-term data, such as operating systems, applications, and user files. Non-volatile storage media can include hard disk drives, solid-state drives, optical disks, and flash memory storage devices, etc.
[0166] Embodiments of this application also provide a computer program. Optionally, in this embodiment, when the computer program is executed by a processor, it implements the method provided in the above embodiments.
[0167] The aforementioned computer program can refer to a set of instructions used to tell the computer to perform specific tasks or operations. Computer programs can be written by programmers using specific programming languages and can include algorithms, data structures, logic, and control flow. Computer programs can be used for a variety of purposes, including application software, operating systems, etc.
[0168] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0169] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.
[0170] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0171] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0172] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.
[0173] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. An image generation method, characterized in that, include: Obtain the original image sequence of the target object, wherein different original images in the original image sequence correspond to different positions of the target object; The original images in the original image sequence are subjected to noise processing to obtain a noisy image sequence; The noisy image sequence is denoised using an image generation model to generate an enhanced image sequence. The structure of the target object in the enhanced image sequence is the same as that in the original image sequence, and the contrast between the target region and other regions in the enhanced image sequence is greater than a preset contrast. The other regions are the regions of the target object other than the target region.
2. The method according to claim 1, characterized in that, The step of adding noise to the original images in the original image sequence to obtain a noisy image sequence includes: Based on a preset window length, the first and last original images in the original image sequence are copied multiple times to obtain an extended image sequence. The target image sequence located within the sliding window is determined from the extended image sequence, wherein the sliding window slides from the first extended image to the last extended image in the extended image sequence during multiple denoising processes, and the length of the sliding window is the preset window length; The target image sequence is subjected to noise processing to obtain the noisy image sequence.
3. The method according to claim 2, characterized in that, The step of adding noise to the target image sequence to obtain the noisy image sequence includes: The last target image in the target image sequence is subjected to noise processing to obtain a first noisy image; A second noise image is determined for other target images, wherein the other target images are target images in the target image sequence other than the last target image, and the second noise image is obtained by adding noise to the other target images, or by performing denoising processing using the image generation model; The noise image sequence is obtained based on the first noise image and the second noise image.
4. The method according to claim 2 or 3, characterized in that, The noise levels of different noise images in the same noisy image sequence are different, while the noise levels of noise images at the same location in different noisey image sequences are the same.
5. The method according to claim 4, characterized in that, In the same noisy image sequence, the noise level of the previous noisy image is lower than the noise level of the next noisy image.
6. The method according to claim 1, characterized in that, The step of using an image generation model to denoise the noisy image sequence and generate an enhanced image sequence includes: The image generation model is used to denoise the noisy image sequence to generate an enhanced image and a denoised image, wherein the enhanced image corresponds to the first target image in the target image sequence, and the denoised image corresponds to the noisy images in the target image sequence other than the first target image; The enhanced image sequence is obtained based on the enhanced image obtained from multiple denoising processes.
7. The method according to claim 6, characterized in that, The step of using the image generation model to denoise the noisy image sequence to generate enhanced and denoised images includes: The first attention module in the image generation model is used to perform attention processing on the noisy images in the noisy image sequence to obtain a first processing result, wherein the first processing result is used to characterize the relationship between different pixels in the noisy image; The second attention module in the image generation model is used to perform attention processing on two adjacent noisy images in the noisy image sequence to obtain a second processing result, wherein the second processing result is used to characterize the relationship between different noisy images; The first processing result and the second processing result are summarized using the generation module in the image generation model to generate the enhanced image and the denoised image.
8. An image synthesis method, characterized in that, include: Obtain raw medical image sequences of biological objects, wherein different raw medical images in the raw medical image sequences correspond to different tissue locations of the biological object; The original medical images in the original medical image sequence are subjected to noise processing to obtain a noisy image sequence; The noisy image sequence is denoised using an image generation model to generate an enhanced medical image sequence. The anatomical structure of the biological object in the enhanced medical image sequence is the same as that in the original medical image sequence. The contrast between the vascular system region and other regions in the enhanced medical image sequence is greater than a preset contrast. The other regions are the regions of the biological object other than the vascular system region.
9. An image synthesis method, characterized in that, include: The original image sequence of the target object is obtained by calling a first interface, wherein the first interface includes a first parameter, the parameter value of the first parameter includes the original image sequence, and different original images in the original image sequence correspond to different positions of the target object; The original images in the original image sequence are subjected to noise processing to obtain a noisy image sequence; The noisy image sequence is denoised using an image generation model to generate an enhanced image sequence. The structure of the target object in the enhanced image sequence is the same as that in the original image sequence, and the contrast between the target region and other regions in the enhanced image sequence is greater than a preset contrast. The other regions are the regions of the target object other than the target region. The enhanced image sequence is output by calling a second interface, wherein the second interface includes a second parameter, and the parameter value of the second parameter includes the enhanced image sequence.
10. A computing device, characterized in that, include: Memory, which stores executable programs; A processor for running the program, wherein the program, when running, performs the method according to any one of claims 1 to 9.
11. An electronic device, characterized in that, include: Memory, which stores executable programs; A processor, connected to the memory via a bus, is used to run the program, wherein the program executes the method according to any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored executable program, wherein, when the executable program is executed, it controls the device on which the computer-readable storage medium is located to perform the method according to any one of claims 1 to 9.
13. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the method of any one of claims 1 to 9.
Citation Information
Patent Citations
Image enhancement method and image enhancement device
CN115760615A
Neutron photography image quality evaluation method based on visual saliency
CN117635541A
AI image super-division detail enhancement system based on film and television production
CN118822854A
Photoacoustic image enhancement method based on time-driven Transform diffusion model
CN119444630A
Multi-domain perceptual contrast enhanced computed tomography image synthesis method and system and electronic equipment
CN119648552A