Medical image enhancement method and system based on diffusion model noise control

By introducing noise manipulation technology into the diffusion model, the attention layer of global and local processes is designed, and the diffusion model lacks important information and high computational consumption in the medical image data enhancement is solved, and more realistic image samples are generated, which improves the performance of downstream tasks and reduces deployment costs.

CN120278906APending Publication Date: 2025-07-08SHANDONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510243499.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing data augmentation method for synthesizing medical images based on diffusion models has the problems of lack of important medical information in synthetic samples, ignoring fine-grained areas, high computing consumption and large storage requirements.

Method used

The diffusion model is customized and improved by using noise control technology. The attention layer is designed through global and local processes, focusing on key areas, and noise manipulation is used to use the Unet network with special attention layer to generate image samples that are more in line with medical reality.

Benefits of technology

The generated medical image samples are more in line with medical practice, reducing computing and storage needs, improving the performance of downstream tasks, and achieving lightweight deployment and high-quality data enhancement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278906A_ABST
    Figure CN120278906A_ABST
Patent Text Reader

Abstract

The invention provides a medical image enhancement method and system based on diffusion model noise control, and relates to the technical field of artificial intelligence and medical images, and the method comprises the steps: designing a suitable local region instruction and a local text instruction according to a downstream task needing to be carried out; inputting the condition global text instruction, the local area instruction and the local text instruction into a framework, starting reasoning by the framework, and preparing initial noise by the framework; performing global process interactive calculation on initial noise and a global text instruction to obtain a global noise graph, and performing local process interactive calculation on initial noise, a local region instruction and a local text instruction to obtain a local noise graph; splicing the global noise map and the local noise map of the noise map, and repeating the process until enhanced data is obtained; according to the method, the intermediate process of the diffusion model is subjected to customized modification, the generation process of the focus area is innovatively realized, and the generated medical image sample is more suitable for medical practice.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence and medical image technology, and particularly to a medical image enhancement method and system based on diffusion model noise manipulation Background Art

[0002] The statements in this part merely provide background technical information related to the present disclosure and do not necessarily constitute prior art

[0003] Medical image analysis is an important field in modern medical tasks. In the past few years, the rapid development of deep learning has achieved important results in medical image analysis tasks. Data-driven neural networks learn the feature representations from a large amount of training data and apply them to various downstream tasks, such as medical image segmentation and medical image classification. However, the scarcity of high-quality labeled data in the medical field has become an obstacle to the development of deep learning technology in medical image analysis. Generally speaking, medical images for training need to be manually labeled by professional physicians, which is expensive and inefficient. In addition, real medical data is often collected from the databases of various hospitals, which may involve patient privacy issues. In summary, the scale of the medical image dataset is limited, posing a huge challenge to the development of the combination of medical image analysis and deep learning technology

[0004] To solve this problem, a common solution is to use data augmentation techniques. Data augmentation is a strategy aimed at transforming and processing existing samples without adding new real samples to increase the diversity and quantity of data. Commonly used data augmentation methods in deep learning include horizontal flipping, image cropping, adding noise, affine transformation, etc. The essence of data augmentation is to amplify the training samples through a large amount of data similar to but different from the original dataset to fit the complex real environment and reduce overfitting in training. However, due to the particularity of the medical scenario, some general-domain data augmentation methods such as color transformation or image cropping are not applicable to medical image analysis. Because the colors or specific organ shapes in medical images are important diagnostic bases, color transformation or image cropping may cause the loss of important information and affect feature learning. Therefore, directly migrating general-domain data augmentation methods is not appropriate

[0005] Another common data augmentation method is to enhance data by mixing two images. Derived from the MixUp technique, it adds sample features and their labels proportionally to create new sample-label pairs. This method improves the generalization ability of the model by telling the model that linear interpolation of sample features corresponds to linear interpolation of labels. Subsequently, a series of such data augmentations were introduced. The CutMix method mixes directly at the pixel level rather than the feature level, and the SaliencyMix method selects the mixing area through bottom-up saliency testing. After training on a dataset with mixed samples added, the model effectively reduces overfitting and performs well in tasks in multiple fields.

[0006] A promising data augmentation method involves synthesizing samples through generative models. These models are good at capturing the latent structure and distribution of data, enabling them to generate samples very similar to actual data instances. Therefore, incorporating these synthetic samples into the dataset is an intuitive and effective strategy. In this augmentation paradigm, the most popular type of generative model is generative adversarial networks (GANs). However, recent research has shown that diffusion models have an advantage over GANs in terms of diversity, so using diffusion to expand the dataset is also a promising direction. Diffusion models are a type of generative model based on Markov chains that generate data by gradually adding Gaussian noise to the data until it reaches a state of pure noise and then reversing the process to restore it to the target state. The DIFFUSEMIX technique customizes a batch of conditional prompts for the generated images, which are mixed with the original images and fractal images to form augmented samples. The ALIA method enhances the training data through text-prompt-guided image editing. In the medical field, some studies have added samples generated by diffusion models to downstream task datasets to improve classification performance.

[0007] However, there are still the following deficiencies in synthesizing medical images based on diffusion models to augment the dataset:

[0008] 1) Although the synthesized samples are similar to real medical images, they lack important medical information.

[0009] 2) When diagnosing based on medical images, physicians pay more attention to fine-grained parts, such as blood vessels and tissues in local areas; while in direct synthesis, such parts are often ignored by the model or synthesized unclearly.

[0010] 3) It needs to be adjusted multiple times to adapt to multiple downstream tasks, and has a high demand for training computing consumption.

[0011] 4) It contains multiple weight parameters, and the deployment in the medical system requires a large amount of storage consumption. Summary of the Invention

[0012] To solve the above problems, the present disclosure proposes a medical image enhancement method and system based on noise manipulation of the diffusion model. By using noise manipulation, the diffusion model's attention to the regions of interest in medical tasks is enhanced, and the intermediate process of the diffusion model is customized and improved. A generation process focusing on key regions is proposed, and a special attention layer in the diffusion model is designed to make the generated medical image samples more in line with medical practice.

[0013] According to some embodiments, the present disclosure adopts the following technical solutions:

[0014] A medical image enhancement method based on noise manipulation of the diffusion model, comprising:

[0015] Obtain the original medical image and preprocess the original medical image;

[0016] Design a global text instruction, a local region instruction, and a local text instruction according to the medical image prediction task, and input the global text instruction, the local region instruction, and the local text instruction into the diffusion model framework;

[0017] Input the preprocessed original medical image and the initial noise into the diffusion model framework for medical image enhancement;

[0018] Among them, the process of enhancing the medical image in the diffusion model framework includes: controlling the global process of medical image synthesis according to the global text instruction, using the attention layer based on the standard Unet network to interact the initial noise and the global text instruction to obtain the global noise map; controlling the local process of generating the key attention region according to the local region instruction and the local text instruction, using the Unet network with a special attention layer to perform interactive calculations on the initial noise, the local region instruction, and the local text instruction to obtain the local noise map; splicing the global noise map and the local noise map to obtain the predicted noise map; continuously predicting the predicted noise added at each step, and gradually subtracting the predicted noise until the input medical image is restored.

[0019] According to some embodiments, the present disclosure adopts the following technical solutions:

[0020] A data acquisition module, configured to obtain the original medical image and preprocess the original medical image;

[0021] A condition acquisition module, configured to design a global text instruction, a local region instruction, and a local text instruction according to the medical image prediction task, and input the global text instruction, the local region instruction, and the local text instruction into the diffusion model framework;

[0022] An image enhancement module, configured to input the preprocessed original medical image and the initial noise into the diffusion model framework for medical image enhancement;

[0023] Among them, the process of enhancing the medical image in the diffusion model framework includes: controlling the global process of medical image synthesis according to the global text instruction, using the attention layer based on the standard Unet network to interact the initial noise and the global text instruction to obtain the global noise map; controlling the local process of generating the key attention area according to the local area instruction and the local text instruction, using the Unet network with a special attention layer to perform interactive calculations on the initial noise, the local area instruction, and the local text instruction to obtain the local noise map; splicing the global noise map and the local noise map to obtain the predicted noise map; continuously predicting the predicted noise added in each step and gradually subtracting the predicted noise until the input medical image is restored.

[0024] According to some embodiments, the present disclosure adopts the following technical solutions:

[0025] A computer program product includes a computer program, and when the computer program is executed by a processor, it implements the medical image enhancement method based on diffusion model noise manipulation described above.

[0026] According to some embodiments, the present disclosure adopts the following technical solutions:

[0027] A non-transitory computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by a processor, the medical image enhancement method based on diffusion model noise manipulation described above is implemented.

[0028] According to some embodiments, the present disclosure adopts the following technical solutions:

[0029] An electronic device includes: a processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory so that the electronic device executes and implements the medical image enhancement method based on diffusion model noise manipulation described above.

[0030] Compared with the prior art, the beneficial effects of the present disclosure are:

[0031] A medical image enhancement method based on noise manipulation of diffusion models adopts noise manipulation technology to customize and modify the intermediate process of diffusion models, innovatively realizing a generation process that focuses on key regions. According to local region instructions and local text instructions, it controls the local process of generating key regions of interest. It uses a Unet network with a special attention layer to interactively calculate the initial noise, local region instructions, and local text instructions to obtain a local noise map. A specially designed attention layer will only mask the regions outside the area of interest, only retaining the interaction between the key regions and local text instructions to obtain a noise map that only contains information about the key regions. The generated medical image samples are more in line with medical practice, that is, they focus on key pathological regions and specific organs.

[0032] A medical image enhancement method based on noise manipulation of diffusion models adopts diffusion models. Compared with other generative models, diffusion models can generate higher-quality domain data. In actual medical decision-making, higher-quality samples help with decision-making. The training process of diffusion models adopts two stages, and only the first stage needs to be trained. In the second stage, it can adapt to different medical scenarios only by changing the input of instructions. Therefore, it can have less storage consumption during deployment and achieve lightweight deployment.

[0033] The data augmentation diffusion model framework proposed by a medical image enhancement method based on noise manipulation of diffusion models does not require a specific neural network structure, that is, any deep neural network can use this data augmentation method to improve the performance of medical tasks, and it has good scalability and transferability. Brief Description of the Drawings

[0034] The specification drawings forming a part of this disclosure are used to provide a further understanding of this disclosure. The schematic embodiments of this disclosure and their descriptions are used to explain this disclosure and do not constitute an improper limitation of this disclosure.

[0035] Figure 1 It is a flowchart of the application of the medical image enhancement method based on noise manipulation of diffusion models in an embodiment of this disclosure;

[0036] Figure 2 It is an implementation architecture diagram of the medical image enhancement method based on noise manipulation of diffusion models in an embodiment of this disclosure;

[0037] Figure 3 It is an implementation architecture diagram of the local process with a special attention layer in an embodiment of this disclosure. Detailed Description of the Embodiments

[0038] The following further describes this disclosure in conjunction with the drawings and embodiments.

[0039] It should be noted that the following detailed description is illustrative and aims to provide further explanation of the present disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present disclosure belongs.

[0040] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present disclosure. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0041] Embodiment 1

[0042] In one embodiment of the present disclosure, a medical image enhancement method based on diffusion model noise manipulation is provided. The objective of the present disclosure is to enhance the attention of the diffusion model to fine-grained regions in medical image analysis, synthesize high-quality medical samples, and expand the medical image analysis task dataset, so as to achieve the enhancement of common deep neural network calculations. The enhancement method steps include:

[0043] Step 1: Obtain the original medical image and preprocess the original medical image;

[0044] Step 2: Design a global text instruction, a local region instruction, and a local text instruction according to the medical image prediction task, and input the global text instruction, the local region instruction, and the local text instruction into the diffusion model framework;

[0045] Step 3: Input the preprocessed original medical image and the initial noise into the diffusion model framework for medical image enhancement;

[0046] Among them, the process of enhancing the medical image in the diffusion model framework includes: controlling the global process of medical image synthesis according to the global text instruction, using the attention layer based on the standard Unet network to interact the initial noise and the global text instruction to obtain the global noise map; controlling the local process of generating the key attention region according to the local region instruction and the local text instruction, using the Unet network with a special attention layer to perform interactive calculations on the initial noise, the local region instruction, and the local text instruction to obtain the local noise map; splicing the global noise map and the local noise map to obtain the predicted noise map; continuously predicting the predicted noise added in each step, and gradually subtracting the predicted noise until the input medical image is restored.

[0047] As an embodiment, the framework introduction of the medical image enhancement method based on diffusion model noise manipulation in the present disclosure includes four parts: (1) The overall process, based on which the framework completes the synthesis of an enhanced sample. (2) The design of a special attention layer in the diffusion model. During the synthesis process of the diffusion model, this special attention layer enables the diffusion model to focus on specific pathological regions. (3) The noise manipulation technique, which realizes the fusion processing of specific pathological regions and general synthesis regions. (4) Downstream task data augmentation, through which the final goal is achieved.

[0048] Among them, the overall process includes two stages. In the first stage, the standard diffusion model is fine-tuned and trained. The diffusion model is trained with a certain number of real medical images to achieve the purpose of migrating the model to the medical field. After this stage, the diffusion model can generate images similar to real medical images through instructions, but most of them are not beneficial for medical image analysis tasks. In the second stage, a specially designed inference generation is performed on the fine-tuned diffusion model. Specifically, the generation process of the standard diffusion model is divided into two parts: the global process and the local process. The global process is responsible for guiding the generation of the global information of the medical image, and the local process focuses on the synthesis of key regions. The specific training implementation process is as follows:

[0049] Step 1: Obtain the original medical image and preprocess it;

[0050] Collect the original medical image and preprocess it, including cleaning. For medical images, image data that is severely damaged, unclear, or has abnormal angles needs to be removed to ensure data quality.

[0051] Divide the preprocessed original medical images into datasets. Randomly divide the datasets into a training set and a validation set according to a certain ratio, where the validation set is used for hyperparameter tuning.

[0052] Step 2: Input the medical image into the diffusion model framework. Using the standard diffusion model training process, the medical image is gradually added noise to become a pure noise image, and then the model starts to predict the noise added at each step and gradually subtract the noise until the input image is restored.

[0053] Specifically, the Unet network is an important component of the diffusion model. It predicts the noise at each time step, designs global text instructions, local region instructions, and local text instructions according to the medical image prediction task, and inputs the global text instructions, local region instructions, and local text instructions into the diffusion model framework;

[0054] Among them, the global text instruction p g: This instruction controls the synthesis process of the global process. Its form is similar to the standard diffusion model instruction. In the framework, to ensure the unity of the generation process, the format is fixed, such as "Generate a chest X-ray image with pneumonia".

[0055] Local area instruction b: This instruction specifies the area to be concerned about in the local process, that is, the area of interest in medical image analysis. For example, in the diagnosis of "pleural effusion", the lower part of the chest and the areas on both sides of the lungs need to be concerned, then the local area instruction will delimit this area. In the diffusion model framework, this instruction is represented by image coordinates.

[0056] Local text instruction p l : This instruction controls the synthesis guidance of the local process. The local text instruction and the local area instruction together complete the synthesis of the local key area. For example, after delimiting the local area instruction for "pleural effusion", the local text instruction is set to "chest with the risk of effusion", emphasizing the pathological risk and the area where it is located.

[0057] Furthermore, the global process and the local process will be completed in parallel in the denoising time steps of the diffusion model. First, the diffusion model prepares an initial noise z t , and gradually restores the required image from time step T to 1. In each time step, the diffusion model predicts the image noise and subtracts this noise.

[0058] For the global process, a processing network containing an attention layer interacts with the initial noise and the global text instruction through the attention mechanism to control the synthesis of an image and obtain a global noise map

[0059] For the local process, in the denoising time step, a specially designed attention layer will only mask the areas outside the concerned area and only retain the interaction between the key area and the local text instruction to obtain a local noise map containing only the information of the key area Using noise manipulation technology on the two noise maps to fuse the global information and the local information, that is, fuse the global noise map and the local noise map to form an intermediate noise map ∈ t This intermediate noise map ∈ t is the predicted noise. Then, using the method of removing the predicted noise from the original image in the standard diffusion process, this time step is completed and preparation is made to start the next time step. Finally, when the time step is 1, a noise-free enhanced medical image is restored, and this enhanced image and the labels contained in the instruction are added to the downstream task dataset for training to improve the performance of the deep neural network.

[0060] Further, in the global process, an attention layer based on the standard Unet network is used to interact the initial noise and the global text instruction to obtain a global noise map. Specifically, the attention layer in Unet receives p g as a condition, and p g is processed by common natural language processing techniques (such as a text encoder) into a feature representation suitable for deep learning. At the same time, the initial noise z t is mapped into a feature representation of the same length as p g . Using p g as the query feature, z t as the key feature and value feature for calculation, the calculation result is the global noise map

[0061] Further, in the local process, according to the local area instruction and the local text instruction control, a local process for generating a region of interest is carried out. A Unet network with a special attention layer is used to interactively calculate the initial noise, the local area instruction, and the local text instruction to obtain a local noise map, including:

[0062] The Unet network is an important component of the diffusion model, which predicts the noise at each time step. The attention layer is a guiding layer in the Unet network for conditional image generation. An attention layer is defined, which supports the interaction between the local area instruction b and the local text instruction p l . Specifically, at time step t, the masked attention processor layer receives the initial noise as input, and the local area instruction b is converted into an attention mask variable m. The conversion process is as follows:

[0063] m = trans(b) / f

[0064] where trans represents a function that converts from coordinate form to a matrix, and f represents the vae factor of the autoencoder.

[0065] Then, based on the attention mask variable m and the initial noise z t , the local process noise map can be obtained through the following equation

[0066]

[0067] In this process, the attention processor layer can only obtain the masked information. Therefore, the local process is focusing on the key region in the cross-attention mechanism. Subsequently, the attention score can be calculated through the masked attention processor in the Unet network as follows

[0068]

[0069] Among them, Q, K, and V are obtained through the following formulas:

[0070]

[0071] Among them, W Q , W K and W V are learnable matrices, represents 's intermediate representation. τ θ represents the conditional text encoder that projects p l into the embedding. In the pipeline, the CLIP text encoder is used. Other Unet network components are consistent with the standard diffusion model. Subsequently, Unet outputs the local noise map

[0072] according to the above formula. Further, noise manipulation is performed. The global noise map and the local noise map are concatenated to obtain the predicted noise map; the predicted noise added at each step is continuously predicted, and the predicted noise is gradually subtracted until the input medical image is restored.

[0073] Specifically, the present invention applies noise processing to data augmentation. During the reverse denoising process, the global noise and the local noise are obtained by passing the global process and the local process at each time step t. Subsequently, a concatenation method is adopted for these two types of noise, and this method is given by the following formula:

[0074]

[0075] Among them, λ1 and λ2 are weight parameters used to control the intensity of each part during the concatenation process. ε t is the predicted noise at this step. Then, the model subtracts the predicted noise ε t from the initial noise z t until the input image is restored, completing the training of the model.

[0076] Step 3: Deploy the trained model;

[0077] Specifically, 1. Model weight saving: Save the model weight file trained in the first stage and clarify the trained parameters.

[0078] 2. Model loading and deployment: Load the model weight file into the model and deploy the model.

[0079] 3. Framework inference generation: Design instruction conditions according to different downstream tasks and generate corresponding enhanced images.

[0080] 4. Downstream task testing and evaluation: Add the enhanced images to the downstream task dataset for training and test with real data to evaluate the effectiveness of the framework.

[0081] 5. Framework go live: Officially deploy the framework to go live and provide auxiliary diagnosis and treatment decision support for doctors. At the same time, continuously monitor the online results of the framework to ensure its stability and rationality.

[0082] As an embodiment, the application process of the medical image enhancement method based on diffusion model noise manipulation of the present disclosure is as follows:

[0083] S1: Design a suitable initial noise z t , full-text instruction p g , local area instruction b and local text instruction p l .

[0084] S2: Normalize the conditions p g , b and p l , and input them into the diffusion model framework.

[0085] S3: In the global process of the diffusion model framework, based on the attention layer interaction of the standard Unet network, calculate the initial noise z t and the full-text instruction p g to obtain the global noise map

[0086] In the local process, based on the Unet network with a special attention layer, perform interactive calculations on the local conditions, that is, interactively calculate the initial noise z t , local area instruction b and local text instruction p l to obtain the local noise map

[0087] S4: Concatenate the two noise maps and to obtain the predicted noise map. Subtract the predicted noise map from the original image to complete this denoising step.

[0088] S5: Repeat steps S3 - S4 until the enhanced image is restored.

[0089] S6: Add the enhanced images to the task dataset to achieve data augmentation and improve the performance of the deep neural network.

[0090] Specifically, after obtaining the CXR enhanced samples, the downstream task dataset is extended in the following form:

[0091] D e = D v ∪ D s

[0092] Among them, D v represents the original data set, and D s represents the sample set synthesized in the previous process. D e is the finally used hybrid augmented data set. Then, the model is retrained using the hybrid augmented data set to improve the performance of the CXR classification task.

[0093] Example 2

[0094] In one embodiment of the present disclosure, a medical image enhancement system based on diffusion model noise manipulation is provided, including:

[0095] A data acquisition module for acquiring the original medical image and preprocessing the original medical image;

[0096] A condition acquisition module for designing a global text instruction, a local region instruction, and a local text instruction according to the medical image prediction task, and inputting the global text instruction, the local region instruction, and the local text instruction into the diffusion model framework;

[0097] An image enhancement module for inputting the preprocessed original medical image and the initial noise into the diffusion model framework for medical image enhancement;

[0098] Among them, the process of enhancing the medical image in the diffusion model framework includes: controlling the global process of medical image synthesis according to the global text instruction, using the attention layer based on the standard Unet network to interact the initial noise and the global text instruction to obtain the global noise map; controlling the local process of generating the key attention area according to the local region instruction and the local text instruction, using the Unet network with a special attention layer to perform interactive calculations on the initial noise, the local region instruction, and the local text instruction to obtain the local noise map; splicing the global noise map and the local noise map to obtain the predicted noise map; continuously predicting the predicted noise added in each step, and gradually subtracting the predicted noise until the input medical image is restored.

[0099] Example 3

[0100] In one embodiment of the present disclosure, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, it implements the medical image enhancement method based on diffusion model noise manipulation.

[0101] Example 4

[0102] In one embodiment of the present disclosure, a non-transitory computer-readable storage medium is provided, and the non-transitory computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by a processor, the medical image enhancement method based on diffusion model noise manipulation is implemented.

[0103] Example 5

[0104] In an embodiment of the present disclosure, an electronic device is provided, including: a processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory so that the electronic device executes the medical image enhancement method based on diffusion model noise manipulation described above.

[0105] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the specified functions in one or more flows and / or one or more blocks. Figure 1 one or more flows and / or Figure 1 one or more blocks.

[0106] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in one or more flows and / or one or more blocks. Figure 1 one or more flows and / or Figure 1 one or more blocks.

[0107] Although the specific implementation manners of the present disclosure have been described above in conjunction with the accompanying drawings, it is not a limitation to the protection scope of the present disclosure. Those skilled in the art should understand that, based on the technical solutions of the present disclosure, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present disclosure.

Claims

1. A medical image enhancement method based on noise manipulation of diffusion models, characterized in that, Including: Obtain the original medical image and preprocess the original medical image; Design a global text instruction, a local region instruction, and a local text instruction according to the medical image prediction task, and input the global text instruction, the local region instruction, and the local text instruction into the diffusion model framework; Input the preprocessed original medical image and the initial noise into the diffusion model framework for medical image enhancement; Among them, the process of enhancing the medical image in the diffusion model framework includes: controlling the global process of medical image synthesis according to the global text instruction, using the attention layer based on the standard Unet network to interact the initial noise and the global text instruction to obtain a global noise map; controlling the local process of generating the key attention region according to the local region instruction and the local text instruction, using the Unet network with a special attention layer to perform interactive calculations on the initial noise, the local region instruction, and the local text instruction to obtain a local noise map; splicing the global noise map and the local noise map to obtain a predicted noise map; continuously predicting the predicted noise added at each step, and gradually subtracting the predicted noise until the input medical image is restored.

2. The medical image enhancement method based on diffusion model noise manipulation according to claim 1, wherein The global process and the local process will be completed in parallel in the denoising time steps of the diffusion model. First, set the initial noise, and gradually restore the required image from time step T to 1. In each time step, the diffusion model will predict the image noise and subtract the predicted noise obtained at each step, gradually restoring and enhancing the image. The basic network structure of the diffusion model is the Unet neural network, and the attention layer based on the standard Unet network is the guiding layer for global condition to image generation in the Unet network.

3. The medical image enhancement method based on diffusion model noise manipulation according to claim 1, wherein A local process for generating a region of interest is controlled according to a local region instruction and a local text instruction. The Unet network with a special attention layer is used to perform interactive calculations on the initial noise, the local region instruction, and the local text instruction, including: at time step t, the masked special attention layer receives the initial noise as input, and the local region instruction b is converted into an attention mask variable m, and the conversion process is as follows: m = trans(b) / f Among them, trans represents the function of converting from coordinate form to matrix, and f represents the vae factor of the autoencoder.

4. The medical image enhancement method based on diffusion model noise manipulation according to claim 3, wherein Based on the attention masking variable m and the initial noise z t , the local process noise map is obtained in the following manner:

5. The medical image enhancement method based on diffusion model noise manipulation according to claim 1, wherein, During the reverse denoising process, the global noise map is obtained by passing the global process and the local process at each time step t and the local noise map They are noise maps generated based on global conditions and local region conditions respectively. Subsequently, a splicing method is adopted for these two types of noise, and this method is given by the following formula: where λ1 and λ2 are weight parameters used to control the intensity of each part in the cascading process, and ε t is the prediction noise of this step.

6. The medical image enhancement method based on diffusion model noise manipulation according to claim 1, wherein In this process, the special attention layer can only obtain the masked information, and the local process focuses on the key region in the cross-attention mechanism, and calculates the attention score through the masked attention in the Unet network.

7. A medical image enhancement system based on noise manipulation of diffusion models, characterized in that, Including: A data acquisition module for obtaining the original medical image and preprocessing the original medical image; A condition acquisition module for designing a global text instruction, a local region instruction, and a local text instruction according to the medical image prediction task, and inputting the global text instruction, the local region instruction, and the local text instruction into the diffusion model framework; An image enhancement module for inputting the preprocessed original medical image and the initial noise into the diffusion model framework for medical image enhancement; Among them, the process of enhancing the medical image in the diffusion model framework includes: controlling the global process of medical image synthesis according to the global text instructions, using the attention layer based on the standard Unet network to interact the initial noise and the global text instructions to obtain the global noise map; controlling the local process of generating the key attention area according to the local area instructions and the local text instructions, using the Unet network with a special attention layer to perform interactive calculations on the initial noise, the local area instructions and the local text instructions to obtain the local noise map; splicing the global noise map and the local noise map to obtain the predicted noise map; continuously predicting the predicted noise added in each step, and gradually subtracting the predicted noise until the input medical image is restored.

8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the medical image enhancement method based on diffusion model noise manipulation according to any one of claims 1-6.

9. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by the processor, it implements the medical image enhancement method based on diffusion model noise manipulation according to any one of claims 1-6.

10. An electronic device, characterized in that, Including: A processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory so that the electronic device executes the medical image enhancement method based on diffusion model noise manipulation according to any one of claims 1-6.