Medical image and pixel-level labeling generation method and device based on diffusion model
Patent Information
- Application Number
- CN202310118369.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-10
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-02-10
AI Technical Summary
[0002]基于机器学习的医学影像自动分割方法通常需要较大规模的医学影像及像素级标注作为训练数据,而实际中大规模像素级标注的获取需要非常可观的人力成本
[0036]通过深度学习的方式自动生成大规模医学影像及像素级标注,有助于在真实数据有限的情况下通过增加训练数据来提高自动分割方法的准确性和鲁棒性,也可以通过仅在生成数据上训练分割模型来避免真实数据泄露造成的隐私问题。
Smart Images

Figure CN116109824B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of medical image processing, and in particular to a method and apparatus for generating medical images and pixel-level annotations based on a diffusion model. Background Technology
[0002] Machine learning-based automatic medical image segmentation methods typically require large-scale medical images and pixel-level annotations as training data. However, obtaining large-scale pixel-level annotations in practice requires considerable human resources. Furthermore, deep segmentation models trained on real-world data risk leaking the training data.
[0003] Diffusion models are a relatively new type of generative model based on deep learning, characterized by high-quality generation, strong diversity, and stable training. In recent years, diffusion models have achieved remarkable results in image generation, video generation, and other fields. Among them, DDPM (Denoising Diffusion Probabilistic Model) is a widely used image generation method. Summary of the Invention
[0004] To address the aforementioned issues, a method and apparatus for generating medical images and pixel-level annotations based on a diffusion model are proposed.
[0005] The first aspect of this application proposes a method for generating medical images and pixel-level annotations based on a diffusion model, including:
[0006] Acquire medical image samples and perform annotation processing on the medical image samples to determine the pixel-level segmentation annotation samples corresponding to the medical image samples;
[0007] The medical image samples are normalized and then stitched together with the pixel-level segmentation and annotation samples to obtain stitched data.
[0008] The spliced data is preprocessed to generate training data;
[0009] The training data is used to obtain a diffusion model, wherein the diffusion model uses U-Net as the network structure;
[0010] Randomly sampled Gaussian noise is input into the diffusion model and, through multiple iterations, medical images and corresponding pixel-level segmentation annotations are generated.
[0011] Optionally, the step of normalizing the medical image samples and stitching them together with the pixel-level segmentation and annotation samples to obtain stitched data includes:
[0012] The pixel-level segmentation and annotation sample is represented as a high-dimensional vector with the same spatial resolution as the corresponding medical image sample, wherein the element at each position represents the category of the pixel at the corresponding position in the medical image sample;
[0013] The categories are assigned values to normalize the medical image samples, wherein the values of the categories are evenly distributed in the range of -1 to 1.
[0014] The pixel-level segmentation and annotation samples are stitched together with the medical image samples along the channel dimension to generate the stitched data.
[0015] Optionally, the preprocessing of the spliced data to generate training data includes:
[0016] The spliced data is scaled to a fixed size and used as training data for the diffusion model.
[0017] Optionally, the network structure of the diffusion model further includes:
[0018] Two-dimensional images are processed using 2D U-Net;
[0019] 3D images are processed using 3D U-Net.
[0020] Optionally, the step of inputting randomly sampled Gaussian noise into the diffusion model and generating medical images and corresponding pixel-level segmentation annotations through multiple iterations includes:
[0021] Randomly sampled Gaussian noise is input into the diffusion model to generate network output data;
[0022] Calculate the next network input data based on the network output data;
[0023] Repeat the above steps until the preset number of iterations is reached, and the generated data is obtained.
[0024] The generated data is processed according to the channel division to obtain the medical image and the encoded pixel-level segmentation annotation.
[0025] Optionally, the method further includes:
[0026] The pixel-level segmentation labels are post-processed, and the category corresponding to the category with the closest numerical value is used as the category label at each pixel position.
[0027] The second aspect of this application proposes a medical image and pixel-level annotation generation device based on a diffusion model, comprising:
[0028] The acquisition module is used to acquire medical image samples, and to perform annotation processing on the medical image samples to determine the pixel-level segmentation annotation samples corresponding to the medical image samples.
[0029] The stitching module is used to normalize the medical image samples and stitch them together with the pixel-level segmentation and annotation samples to obtain stitched data.
[0030] The preprocessing module is used to preprocess the spliced data to generate training data;
[0031] A training module is used to train the training data to obtain a diffusion model, wherein the diffusion model uses U-Net as the network structure;
[0032] The output module is used to input randomly sampled Gaussian noise into the diffusion model and generate medical images and corresponding pixel-level segmentation annotations through multiple iterations.
[0033] A third aspect of this application provides a computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, it implements any of the methods described in the first aspect above.
[0034] The fourth aspect of this application provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in any of the first aspects above.
[0035] The technical solutions provided by the embodiments of this application bring at least the following beneficial effects:
[0036] Automatically generating large-scale medical images and pixel-level annotations through deep learning can help improve the accuracy and robustness of automatic segmentation methods by increasing training data when real data is limited. It can also avoid privacy issues caused by the leakage of real data by training the segmentation model only on the generated data.
[0037] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0038] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:
[0039] Figure 1 This is a flowchart illustrating a method for generating medical images and pixel-level annotations based on a diffusion model, according to an embodiment of this application.
[0040] Figure 2This is a flowchart illustrating a method for generating medical images and pixel-level annotations based on a diffusion model, according to an embodiment of this application.
[0041] Figure 3 This is a flowchart illustrating a method for generating medical images and pixel-level annotations based on a diffusion model, according to an embodiment of this application.
[0042] Figure 4 This is a block diagram illustrating a medical image and pixel-level annotation generation device based on a diffusion model, according to an embodiment of this application.
[0043] Figure 5 It is a block diagram of an electronic device. Detailed Implementation
[0044] The embodiments of this application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.
[0045] Figure 1 This is a flowchart illustrating a method for generating medical images and pixel-level annotations based on a diffusion model, according to an embodiment of this application, including:
[0046] Step 101: Obtain medical image samples and perform annotation processing on the medical image samples to determine the pixel-level segmentation annotation samples corresponding to the medical image samples.
[0047] In this application embodiment, the medical image sample is an image processed by X-ray projection, CT, ultrasound, magnetic resonance imaging, radionuclide, etc. In this application, the medical image sample is obtained through a publicly available medical image database, such as the TCIA database, MedPix database, LONI database, etc.
[0048] In this embodiment, the purpose of pixel-level segmentation annotation is to indicate the location and category of the category of interest (such as lesion or physiological structure) on the medical image. In the computer, it is represented as a high-dimensional vector with the same spatial resolution as the medical image sample. Each pixel / voxel of the medical image sample corresponds one-to-one with a category vector of the pixel-level segmentation annotation sample.
[0049] The category vector indicates the category label of the pixel. For an application scenario with a total of C categories, the category vector is a one-dimensional vector of length C. The category to which the corresponding pixel / voxel belongs is assigned a value of 1, and otherwise it is assigned a value of 0.
[0050] Step 102: Normalize the medical image samples and stitch them together with the pixel-level segmentation and annotation samples to obtain stitched data.
[0051] In this embodiment of the application, the medical image samples and pixel-level segmentation and annotation samples are represented in a form suitable for the diffusion model. Specifically, step 102 further includes:
[0052] Step 201: Represent the pixel-level segmentation annotation sample as a high-dimensional vector with the same spatial resolution as the corresponding medical image sample, where each element at each position represents the category of the pixel at the corresponding position in the medical image sample.
[0053] Step 202: Assign values to categories to normalize medical image samples, where the values of the categories are evenly distributed in the range of -1 to 1.
[0054] In this embodiment of the application, each category is represented by a fixed numerical value, and different categories correspond to different numerical values. The numerical values of all categories are evenly distributed in the range of -1 to 1, thereby normalizing the medical image samples to the range of -1 to 1.
[0055] Step 203: In the channel dimension, the pixel-level segmentation and annotation samples are stitched together with the medical image samples to generate stitched data.
[0056] In this embodiment, medical image samples and corresponding pixel-level segmentation and annotation samples are concatenated into a high-dimensional vector, which serves as the form of data generated by the diffusion model.
[0057] Step 103: Preprocess the spliced data to generate training data.
[0058] In this embodiment of the application, the spliced data is scaled to a fixed size and used as training data for the diffusion model.
[0059] Step 104: Train the training data to obtain the diffusion model, wherein the diffusion model uses U-Net as the network structure.
[0060] In this embodiment, DDPM is used as the diffusion model and U-Net is used as the network structure of the diffusion model. In each denoising process, the data output from the previous step is input into the U-Net network, and denoising is performed based on the noise prediction of the network output.
[0061] Two-dimensional images are processed using 2D U-Net, while three-dimensional images are processed using 3D U-Net.
[0062] In one possible embodiment, the two-dimensional image is an X-ray image, and the three-dimensional image is a CT image.
[0063] Step 105: Input random sampled Gaussian noise into the diffusion model and generate medical images and corresponding pixel-level segmentation annotations through multiple iterations.
[0064] This application considers the inverse process of gradually adding random Gaussian noise to real data, and gradually denoising from random noise to generate real data. Specifically, step 105 also includes:
[0065] Step 301: Input randomly sampled Gaussian noise into the diffusion model and generate network output data;
[0066] Step 302: Calculate the next network input data based on the network output data;
[0067] Step 303: Repeat the above steps until the number of iterations meets the preset number, and obtain the generated data;
[0068] Step 304: Generate data based on channel division, and obtain pixel-level segmentation annotations of medical images and codes.
[0069] In this embodiment, the preset number of times is set to 1000, and the generated data has the same format and size as the training data.
[0070] In addition, the pixel-level segmentation labels are post-processed, and the category corresponding to the category with the closest numerical value is used as the category label at each pixel location.
[0071] The embodiments of this application automatically generate large-scale medical images and pixel-level annotations through deep learning. This helps to improve the accuracy and stability of automatic segmentation methods by increasing training data when real data is limited. It also avoids privacy issues caused by the leakage of real data by training the segmentation model only on the generated data.
[0072] Figure 4 This is a block diagram of a medical image and pixel-level annotation generation device based on a diffusion model, according to an embodiment of this application, including an acquisition module 410, a stitching module 420, a preprocessing module 430, a training module 440, and an output module 450.
[0073] The acquisition module 410 is used to acquire medical image samples, perform annotation processing on the medical image samples, and determine the pixel-level segmentation annotation samples corresponding to the medical image samples.
[0074] The stitching module 420 is used to normalize medical image samples and stitch them together with pixel-level segmented and labeled samples to obtain stitched data.
[0075] The preprocessing module 430 is used to preprocess the spliced data to generate training data;
[0076] Training module 440 is used to train training data to obtain a diffusion model, wherein the diffusion model uses U-Net as the network structure;
[0077] The output module 450 is used to input randomly sampled Gaussian noise into the diffusion model and generate medical images and corresponding pixel-level segmentation annotations through multiple iterations.
[0078] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0079] Figure 5 A schematic block diagram of an example electronic device 500 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0080] like Figure 5 As shown, device 500 includes a computing unit 501, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 502 or a computer program loaded from storage unit 503 into random access memory (RAM) 503. RAM 503 may also store various programs and data required for the operation of device 500. The computing unit 501, ROM 502, and RAM 503 are interconnected via bus 504. Input / output (I / O) interface 505 is also connected to bus 504.
[0081] Multiple components in device 500 are connected to I / O interface 505, including: input unit 506, such as keyboard, mouse, etc.; output unit 507, such as various types of monitors, speakers, etc.; storage unit 508, such as disk, optical disk, etc.; and communication unit 509, such as network card, modem, wireless transceiver, etc. Communication unit 509 allows device 500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0082] The computing unit 501 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as the voice command response method. For example, in some embodiments, the voice command response method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed on device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by the computing unit 501, one or more steps of the voice command response method described above may be performed. Alternatively, in other embodiments, the computing unit 501 may be configured to perform the voice command response method by any other suitable means (e.g., by means of firmware).
[0083] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0084] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0085] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0086] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0087] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), the Internet, and blockchain networks.
[0088] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service ecosystem, addressing the shortcomings of traditional physical hosts and VPS (Virtual Private Server, or simply "VPS") services, such as high management difficulty and weak business scalability. Servers can also be servers for distributed systems or servers incorporating blockchain technology.
[0089] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0090] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for generating medical images and pixel-level annotations based on a diffusion model, characterized in that, include: Acquire medical image samples and perform annotation processing on the medical image samples to determine the pixel-level segmentation annotation samples corresponding to the medical image samples; The purpose of pixel-level segmentation annotation is to indicate the location and category of the category of interest in medical images. In computers, it is represented as a high-dimensional vector with the same spatial resolution as the medical image sample. Each pixel / voxel of the medical image sample corresponds one-to-one with a category vector of the pixel-level segmentation annotation sample. The category vector indicates the category label of the pixel. For an application scenario with a total of C categories, the category vector is a one-dimensional vector of length C. The category that the corresponding pixel / voxel belongs to is assigned a value of 1, otherwise it is assigned a value of 0. The medical image samples are normalized, and the pixel-level segmentation annotation samples are represented as high-dimensional vectors with the same spatial resolution as the corresponding medical image samples. The pixel-level segmentation annotation samples and the medical image samples are then concatenated in the channel dimension to generate concatenated data. The spliced data is preprocessed to generate training data; wherein the spliced data is scaled to a fixed size and used as training data for the diffusion model. The training data is used to obtain a diffusion model, wherein the diffusion model uses DDPM as the diffusion model and U-Net as the network structure. In each denoising step, the data output from the previous step is input into the U-Net network, and denoising is performed based on the noise prediction of the network output. 2D U-Net is used to process two-dimensional images, and 3D U-Net is used to process three-dimensional images. Randomly sampled Gaussian noise is input into the diffusion model to obtain generated data, wherein the generated data has the same format as the training data; the generated data is processed according to channel division to obtain medical images and encoded pixel-level segmentation annotations.
2. The method according to claim 1, characterized in that, The process of normalizing the medical image samples and obtaining stitched data includes: The pixel-level segmentation and annotation sample is represented as a high-dimensional vector with the same spatial resolution as the corresponding medical image sample, wherein the element at each position represents the category of the pixel at the corresponding position in the medical image sample; The categories are assigned values to normalize the medical image samples, wherein the values of the categories are evenly distributed in the range of -1 to 1. The pixel-level segmentation and annotation samples are stitched together with the medical image samples along the channel dimension to generate the stitched data.
3. The method according to claim 2, characterized in that, The method further includes: The pixel-level segmentation labels are post-processed, and the category corresponding to the category with the closest numerical value is used as the category label at each pixel position.
4. A medical image and pixel-level annotation generation device based on a diffusion model, characterized in that, include: The acquisition module is used to acquire medical image samples, and to perform annotation processing on the medical image samples to determine the pixel-level segmentation annotation samples corresponding to the medical image samples. The purpose of pixel-level segmentation annotation is to indicate the location and category of the category of interest in medical images. In computers, it is represented as a high-dimensional vector with the same spatial resolution as the medical image sample. Each pixel / voxel of the medical image sample corresponds one-to-one with a category vector of the pixel-level segmentation annotation sample. The category vector indicates the category label of the pixel. For an application scenario with a total of C categories, the category vector is a one-dimensional vector of length C. The category that the corresponding pixel / voxel belongs to is assigned a value of 1, otherwise it is assigned a value of 0. The stitching module is used to normalize the medical image sample and represent the pixel-level segmentation annotation sample as a high-dimensional vector with the same spatial resolution as the corresponding medical image sample. The pixel-level segmentation annotation sample is stitched with the medical image sample in the channel dimension to generate stitched data. A preprocessing module is used to preprocess the spliced data to generate training data; wherein the spliced data is scaled to a fixed size and used as training data for the diffusion model. The training module is used to train the training data to obtain a diffusion model. The diffusion model uses DDPM as the diffusion model and U-Net as the network structure. In each denoising step, the data output from the previous step is input into the U-Net network, and denoising is performed based on the noise prediction of the network output. Specifically, 2D U-Net is used to process two-dimensional images, and 3D U-Net is used to process three-dimensional images. The output module is used to input randomly sampled Gaussian noise into the diffusion model to obtain generated data, wherein the generated data has the same format as the training data; the generated data is processed according to channel division to obtain medical images and encoded pixel-level segmentation annotations.
5. A computer device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, it implements the method as described in any one of claims 1-3.
6. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-3.
Citation Information
Patent Citations
Method for segmenting radiotherapy image by combining deep neural network and probability graph model
CN110619639A