Medical image synthesis method based on autoregression model

Through the medical image synthesis method based on autoregression model, the residual quantization mechanism and autoregression training paradigm are used to solve the shortcomings of the existing models in terms of stability and detail retention, efficient and accurate medical image synthesis is achieved, and the accuracy of computer-assisted diagnosis is improved.

CN120198299APending Publication Date: 2025-06-24UNIV OF SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510301859.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Existing medical image synthesis models such as GAN and Diffusion have shortcomings in stability and detail retention, resulting in difficult convergence of models and inefficient image synthesis.

Method used

Using a medical image synthesis method based on autoregression model, by constructing a medical image prior encoder and an autoregression medical image synthesis module, the residual quantization mechanism and autoregression training paradigm are used to gradually learn the potential features of medical images and fit the generation process.

Benefits of technology

It improves the accuracy, stability and efficiency of medical images in potential space representations, and enhances the accuracy of computer-assisted medical diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198299A_ABST
    Figure CN120198299A_ABST
Patent Text Reader

Abstract

The invention discloses a medical image synthesis method based on autoregression. The method comprises the following steps: 1, preprocessing a medical image; 2, constructing a medical image prior encoder, and carrying out discretization encoding on the structure and intensity information of the medical image by adopting a residual quantization mechanism; 3, constructing a medical image synthesis module, and fitting the synthesis process of the medical image by using a Transform decoder structure; and 4, synthesizing the medical image by using the provided autoregressive medical image synthesis method. According to the method, the high-quality medical image can be synthesized efficiently, so that the accuracy of computer-aided diagnosis can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image processing, and specifically relates to a medical image synthesis method based on an autoregressive model. Background Art

[0002] With the development of artificial intelligence technology, computers can be used to recognize medical images. In the field of computer-aided diagnosis, the severe shortage of labeled data affects the accuracy and reliability of medical diagnosis models. Data synthesis is a new technical means that can solve the problem of insufficient data to enhance the diagnostic performance of deep learning models.

[0003] Currently, an image synthesis model for synthesizing medical images can be obtained by training GAN (Generative Adversarial Network) and Diffusion (diffusion model). However, the adversarial training process of GAN is unstable, making it difficult for the model to converge. Diffusion is mainly applied to natural images and easily omits image details, especially the skin around the lesions. Moreover, its generation process is to gradually denoise, resulting in low image synthesis efficiency. Summary of the Invention

[0004] The present invention aims to solve the above-mentioned deficiencies in the prior art and proposes a medical image synthesis method based on an autoregressive model, aiming to efficiently synthesize high-quality medical images, thereby improving the accuracy of computer-aided medical diagnosis.

[0005] To achieve the above object, the present invention adopts the following technical solutions:

[0006] The medical image synthesis method based on an autoregressive model of the present invention is characterized in that it is carried out according to the following steps:

[0007] Step 1: Obtain medical images with disease category labels and perform preprocessing to obtain a medical image dataset and its corresponding disease category labels , where represents the th preprocessed medical image, represents the th medical image corresponding disease category label, and ∈[1,N]; represents the number of disease categories; n represents the total number of medical images;

[0008] Step 2: Construct a medical image prior encoder, including: an encoder E, a residual quantization layer R, and a decoder D, and for the Zhang medical image Perform prior encoding to obtain the Zhang reconstructed medical image , thereby constructing the loss function of the variational autoencoder , and use the stochastic gradient descent method to train the variational autoencoder and minimize the loss function , until the loss function converges, thereby obtaining the trained medical image prior encoder;

[0009] Step 3. Construct an autoregressive medical image synthesis module, including: a category embedding layer, a decoder, and process and the optimal discrete feature map output by the trained medical image prior encoding to obtain the i-th prediction vector , thereby constructing the cross-entropy loss function of the medical image synthesis module ; and use the stochastic gradient descent method to train the autoregressive medical image synthesis module and minimize the cross-entropy loss function , until the cross-entropy loss function converges, thereby obtaining the trained autoregressive medical image synthesis module;

[0010] Step 4. Use the trained medical image prior encoder and the trained autoregressive medical image synthesis module to generate medical images of a specified disease category .

[0011] The feature of a medical image synthesis method based on an autoregressive model according to the present invention also lies in that the medical image prior encoder in step 2 obtains an image according to the following steps and constructs a loss function :

[0012] Step 2.1. The encoder E includes: an input convolutional layer, SC downsampling convolutional modules; wherein, each downsampling convolutional module sequentially includes a group normalization layer and a convolutional layer;

[0013] Input into the encoder E, and after being processed by the input convolutional layer, obtain the th convolutional feature map ;

[0014] When , Input into the sc-th downsampling convolutional module, and after being processed by the group normalization layer of the th downsampling convolutional module, obtain the th normalized feature map ;

[0015] After being processed by the convolutional layer of the th downsampling convolutional module, the th downsampled feature map is obtained ;

[0016] When , the th downsampled feature is input into the th downsampling convolutional module for processing, and the th downsampled feature map is obtained , so that the th downsampling convolutional module outputs the th downsampled feature map , and it is used as the th latent feature map ;

[0017] Step 2.2: The residual quantization layer R is composed of SK quantization modules and a reconstruction module. Among them, the quantization module is composed of a downsampling layer, a quantization layer, and an upsampling layer; the quantizer contains a codebook of learnable parameters and an initialized queue , represents the number of feature vectors contained in the codebook , and is the length of the feature vector;

[0018] Step 2.2.1: The SK quantization modules process to obtain the th intermediate feature map;

[0019] When , is input into the th quantization module, and after being processed by the downsampling layer of the th quantization module, the th downsampling result is obtained;

[0020] Then it is input into the quantization layer of the th quantization module, and the feature vector at coordinate in is used to obtain the index of the most similar feature vector in the codebook , so as to obtain the index set , and according to the index set from the codebook Extract the corresponding eigenvectors to form the i-th discrete feature map :

[0021] (1)

[0022] In formula (1), refers to extracting the -th eigenvector from the codebook ;

[0023] After being processed in the upsampling layer of the -th quantization module, the -th upsampling result is obtained, and thus the -th feature map :

[0024] (2)

[0025] When , the -th feature map is input into the -th quantization module for processing, and thus the -th quantization module outputs the -th feature map :

[0026] Step 2.2.2. The reconstruction module obtains the i-th intermediate feature map using formula (3) :

[0027] (3)

[0028] In formula (3), represents a convolution operation; represents the i-th discrete feature map;

[0029] Step 2.3. The decoder consists of an input convolutional layer, upsampling convolutional modules, and an output convolutional layer, and processes to obtain the i-th reconstructed image ;

[0030] Step 2.4. Use formula (4) to construct the loss function of the variational autoencoder :

[0031] (4).

[0032] Furthermore, step 3 is carried out as follows:

[0033] Step 3.1: Input the disease category label into the category embedding layer for processing to obtain the i-th category embedding ; ;

[0034] Step 3.2: Input the input into the trained medical image prior encoder for processing, and output the i-th optimal discrete feature map by the quantization layer, where represents the -th optimal discrete feature;

[0035] Concatenate the category embedding with the previous SK - 1 optimal discrete features to obtain the i-th input feature vector ;

[0036] Step 3.3: Input the input into the decoder for processing to obtain the i-th prediction vector , where represents 's predicted value;

[0037] Step 3.4: Use Equation (5) to construct the cross-entropy loss function of the medical image synthesis module :

[0038] (5).

[0039] Furthermore, Step 4 is carried out as follows:

[0040] Step 4.1: Input the disease category label into the category embedding layer of the trained autoregressive medical image synthesis module for processing to obtain the i-th optimal category embedding , initialize = 1;

[0041] Initialize the -th input queue = { };

[0042] Step 4.2: Input the input into the decoder of the trained autoregressive medical image synthesis module for processing to obtain the -th prediction queue , where represents the k-th optimal predicted value;

[0043] Input The input is processed by the quantization layer in the rd quantization module of the trained medical image prior encoder to obtain the th optimal discrete feature map , and it is added to , so as to obtain the +1 input queues ;

[0044] Step 4.3, After being assigned to , return to step 4.2 and execute sequentially until =K, thus obtaining ={ };

[0045] Step 4.4, use Equation (6) to obtain the final latent layer feature :

[0046] (6)

[0047] In Equation (6), represents the convolution operation;

[0048] Step 4.5, input the final latent layer feature into the decoder D in the trained medical image prior encoder for processing, and obtain the optimally reconstructed medical image .

[0049] An electronic device according to the present invention includes a memory and a processor, characterized in that the memory is used to store a program for supporting the processor to execute the medical image synthesis method, and the processor is configured to execute the program stored in the memory.

[0050] A computer-readable storage medium according to the present invention, characterized in that a computer program is stored on the computer-readable storage medium, and the computer program executes the steps of the medical image synthesis method when being run by a processor.

[0051] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0052] 1. The present invention introduces a variational autoencoder with a residual quantization mechanism to perform discrete coding on medical images. The residual quantization mechanism can gradually learn the latent features of medical images, thereby improving the accuracy of the representation of medical images in the latent space.

[0053] 2. The present invention introduces a new autoregressive training paradigm to fit the generation process of medical images. Compared with traditional medical image synthesis models, the training process is more stable, and the synthesis efficiency of medical images is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 This is the prior encoder diagram of the medical image of the present invention;

[0055] Figure 2 This is the process diagram of the medical image synthesis of the present invention. Detailed implementation manners

[0056] In this embodiment, a medical image synthesis method based on an autoregressive model is carried out according to the following steps:

[0057] Step 1: First, preprocess the medical image. The preprocessing steps include cropping the medical image to a specified size, such as 256*256, and then normalizing the image. Preprocessing the medical image is to standardize the medical image for convenient input into the model. After obtaining a medical image dataset with disease category labels and preprocessing it, and its corresponding disease category label , where represents the th preprocessed medical image, represents the th medical image corresponding disease category label, and ∈[1,N]; represents the number of disease categories; n represents the total number of medical images;

[0058] Step 2: To efficiently synthesize medical images, it is necessary to model the medical images in the latent space. Therefore, a medical image prior encoder is constructed, including: an encoder E, a residual quantization layer R, and a decoder D. Perform prior encoding on the th medical image to obtain the th reconstructed medical image . The medical image encoder is as Figure 1 shown.

[0059] Step 2.1: The encoder E includes: an input convolutional layer, and SC downsampling convolutional modules; where each downsampling convolutional module sequentially includes a group normalization layer and a convolutional layer;

[0060] Input into the encoder E and, after being processed by the input convolutional layer, obtain the th convolutional feature map ;

[0061] When , input into the sc-th downsampling convolutional module and, after being processed by the After the processing of the group normalization layer of the th downsampling convolutional module, the th normalized feature map is obtained;

[0062] After passing through the convolutional layer of the th downsampling convolutional module, the th downsampled feature map is obtained; ;

[0063] When , the th downsampled feature is input into the th downsampling convolutional module for processing, and the th downsampled feature map is obtained; , so that the th downsampling convolutional module outputs the th downsampled feature map; and serves as the th latent feature map; Downsampling the medical image is to compress information and efficiently preserve the information of the medical image.

[0064] Step 2.2. The residual quantization layer R consists of SK quantization modules and a reconstruction module, where the quantization module consists of a downsampling layer, a quantization layer, and an upsampling layer; the quantizer contains a codebook of learnable parameters and an initialized queue , represents the number of feature vectors contained in the codebook , is the length of the feature vector; The residual quantization mechanism can effectively model the medical image at different scales and more accurately represent the medical image in the latent space. Using a codebook to represent image features facilitates modeling the image generation process in an autoregressive manner.

[0065] Step 2.2.1. The SK quantization modules process to obtain the th intermediate feature map;

[0066] When , is input into the th quantization module and, after passing through the downsampling layer of the th quantization module, the th downsampling result is obtained; ;

[0067] Then it is input into the In the quantization layer of a quantization module, and obtain using Equation (1) The coordinates in The eigenvector of And the codebook The index of the eigenvector most similar to that in , thus obtaining an index set , and according to the index set Take the corresponding eigenvectors from the codebook To form the i-th discrete feature map :

[0068] (1)

[0069] In Equation (1), Refers to taking the th eigenvector from the codebook ;

[0070] After passing through the upsampling layer of the th quantization module and being processed, the th upsampling result is obtained, and thus the th feature map is obtained using Equation (2):

[0071] (2)

[0072] When , the feature maps are input into the th quantization module for processing, and thus the th quantization module outputs the th feature map :

[0073] Step 2.2.2. The reconstruction module uses Equation (3) to obtain the i-th intermediate feature map :

[0074] (3)

[0075] In Equation (3), Represents a convolution operation; Represents the i-th discrete feature map;

[0076] Step 2.3. The decoder Consists of an input convolution layer, upsampling convolution modules, and an output convolution layer, and processes to obtain the i-th reconstructed image ;

[0077] Step 2.5: Construct the loss function of the variational autoencoder using Equation (4) , which is used to measure the difference between the original medical image and the medical image recompiled using the variational autoencoder;

[0078] (4)

[0079] Step 2.6: Train the variational autoencoder using the stochastic gradient descent method and minimize the loss function , until the loss function converges, thereby obtaining the trained prior encoder of the medical image. The codebook in this prior encoder of the medical image stores the structural and intensity information of the medical image in the latent space, and the decoder can decode the medical image in the latent space into a medical image.

[0080] Step 3: Construct an autoregressive medical image synthesis module, including: a class embedding layer, a decoder. In order to model the relationship between the discrete feature maps in the prior encoder of the medical image , use structure to fit the relationship between them, and use structure can perform efficient parallel computing during training, thereby improving the computing efficiency.

[0081] Step 3.1: Input the disease category label into the class embedding layer for processing to obtain the i-th class embedding ;

[0082] Step 3.2: Input the into the trained prior encoder of the medical image for processing, and the quantization layer outputs the i-th optimal discrete feature map , where represents the th optimal discrete feature;

[0083] After concatenating the class embedding and the first SK - 1 optimal discrete features, the i-th input feature vector is obtained.

[0084] Step 3.3: Input the into the decoder for processing to obtain the i-th predicted vector , where represents the predicted value of;

[0085] Step 3.4: Construct the cross-entropy loss function of the medical image synthesis module using Equation (5). :

[0086] (5)

[0087] Step 3.5: Train the autoregressive medical image synthesis module using the stochastic gradient descent method and minimize the cross-entropy loss function until the cross-entropy loss function converges, thereby obtaining the trained autoregressive medical image synthesis module. , until the cross-entropy loss function converges, thus obtaining the trained autoregressive medical image synthesis module.

[0088] Step 4: Use the trained medical image prior encoder and the trained autoregressive medical image synthesis module to generate medical images of a specified disease category. The disease categories available for selection are the same as those in the data used for training. The medical image synthesis process is as shown in Figure 2 ;

[0089] Step 4.1: Input the disease category label into the category embedding layer of the trained autoregressive medical image synthesis module for processing to obtain the i-th optimal category embedding , and initialize = 1;

[0090] Initialize the -th input queue ={ };

[0091] Step 4.2: Input into the decoder of the trained autoregressive medical image synthesis module for processing to obtain the -th prediction queue , where represents the k-th optimal prediction value;

[0092] Input into the quantization layer of the -th quantization module of the trained medical image prior encoder for processing to obtain the -th optimal discrete feature map , and add it to to obtain the + 1-th input queue .

[0093] Step 4.3: Assign to , then return to Step 4.2 and execute sequentially until = K, thus obtaining ={ };

[0094] Step 4.4. Obtain the final latent feature by using Equation (6) :

[0095] (6)

[0096] In Equation (6), represents the convolution operation.

[0097] Step 4.5. Input the final latent feature into the decoder D in the trained medical image prior encoder to process it, restore the medical image in the latent space, and obtain the optimally reconstructed medical image , which is the synthesized medical image of the specified category.

[0098] In this embodiment, an electronic device includes a memory and a processor. The memory is used to store a program that supports the processor to execute the above method, and the processor is configured to execute the program stored in the memory.

[0099] In this embodiment, a computer-readable storage medium stores a computer program, and when the computer program is run by a processor, it executes the steps of the above method.

[0100] Example:

[0101] To verify the effectiveness of the method of the present invention, in this embodiment, skin disease images are synthesized. First, skin disease image data of seven categories are collected, and this data is used to train the medical image prior encoder and the autoregressive medical image synthesis module. Then, the method in Step 4 of the present invention is used for image synthesis. The synthesized medical images and real images are combined to form a new dataset, and the computer-aided diagnosis classification model is trained, significantly improving the accuracy of the classification model.

Claims

1. A medical image synthesis method based on an autoregressive model, characterized in that: The steps are as follows: Step 1: Get After preprocessing, we can get a medical image dataset with disease category labels. and its corresponding disease category label ,in, After preprocessing, Medical images, Representative Medical images The corresponding disease category label, and ∈[1,N]; represents the number of disease categories; n represents the total number of medical images; Step 2: Construct a medical image prior encoder, including: encoder E, residual quantization layer R, decoder D, and Medical images Perform a priori encoding and obtain Reconstructed medical image , thereby constructing the loss function of the variational autoencoder , and train the variational autoencoder using stochastic gradient descent to minimize the loss function , until the loss function Until convergence, the trained medical image prior encoder is obtained; Step 3: Construct an autoregressive medical image synthesis module, including: category embedding layer, decoder, and The optimal discrete feature map output by the trained medical image prior encoding is processed to obtain the i-th prediction vector , thereby constructing the cross entropy loss function of the medical image synthesis module ; and using the stochastic gradient descent method to train the autoregressive medical image synthesis module and minimize the cross entropy loss function , until the cross entropy loss function Until convergence, thus obtaining the trained autoregressive medical image synthesis module; Step 4: Generate medical images of specified disease categories using the trained medical image prior encoder and the trained autoregressive medical image synthesis module .

2. The method for synthesizing medical images based on an autoregressive model according to claim 1, characterized in that: The medical image priori encoder in step 2 is performed according to the following steps to obtain the image And construct the loss function : Step 2.1, the encoder E comprises: an input convolution layer, SC downsampling convolution modules; wherein each downsampling convolution module comprises a group normalization layer and a convolution layer in sequence; Input into the encoder E, and after being processed by the input convolution layer, the first Convolutional feature maps ; when hour, Input into the scth downsampling convolution module and pass through the After processing the group normalization layer of the downsampling convolution module, we get Normalized feature map ; After After the convolutional layer of the downsampling convolution module is processed, the Downsampled feature maps ; when When Downsampled features Enter the The downsampling convolution module is used to process the Downsampled feature maps , thus The downsampling convolution module outputs Downsampled feature maps , and as the first Latent feature map ; Step 2.2, the residual quantization layer R is composed of SK quantization modules and a reconstruction module, wherein the quantization module is composed of a downsampling layer, a quantization layer, and an upsampling layer; the quantizer contains a codebook of parameters to be learned and the initialized queue , Representation Codebook The number of eigenvectors contained in is the length of the eigenvector; Step 2.2.1, SK quantization module pair Processing is performed to obtain the i-th intermediate feature map ; when hour, Enter the In the quantization module, and after the After processing the downsampling layer of the quantization module, we get Downsampling results ; Then enter In the quantization layer of the quantization module, and using formula (1) to get The median coordinate is The eigenvector of With codebook The index of the most similar eigenvector in , thus obtaining the index set , and follow the index collection From the codebook Take out the corresponding feature vector to form the i-th discrete feature map : (1) In formula (1), Refer to the codebook Take out feature vectors; After the After processing in the upsampling layer of the quantization module, the Upsampling results , and then use formula (2) to get the Feature Map : (2) when When Feature Map Enter the The quantization module is used to process the The quantization module outputs Feature Map : Step 2.2.2: The reconstruction module uses formula (3) to obtain the i-th intermediate feature map : (3) In formula (3), Represents the convolution operation; represents the i-th discrete feature map; Step 2.3: The decoder By input convolution layer, The upsampling convolution module and the output convolution layer are composed of Processing is performed to obtain the i-th reconstructed image ; Step 2.4: Use formula (4) to construct the loss function of the variational autoencoder : (4)。 3. The method for medical image synthesis based on an autoregressive model according to claim 2, characterized in that: Step 3 is performed as follows: Step 3.1: Disease category label The input category embedding layer is processed to obtain the i-th category embedding ; Step 3.2: The trained medical image is input into the prior encoding for processing, and the quantization layer outputs the i-th optimal discrete feature map ,in, express No. optimal discrete features; Embed categories After concatenating with the first SK-1 optimal discrete features, we get the i-th input feature vector ; Step 3.3: enter The decoder is processed and the i-th prediction vector is obtained accordingly. ,in, represent The predicted value of Step 3.4: Use formula (5) to construct the cross entropy loss function of the medical image synthesis module : (5)。 4. The method for synthesizing medical images based on an autoregressive model according to claim 3, characterized in that: Step 4 is performed as follows: Step 4.1: Label the disease category Input the trained autoregressive medical image synthesis module into the category embedding layer for processing to obtain the i-th optimal category embedding ,initialization =1; Initialize Input Queue ={ }; Step 4.2: Input trained autoregressive medical image synthesis module The decoder is processed to obtain Prediction Queue ,in, represents the kth optimal prediction value; Will Input the trained medical image prior encoder The quantization layer in the quantization module is processed to obtain the Optimal discrete feature map , and add , thus obtaining the +1 input queue ; Step 4.3 Assign to Then return to step 4.2 and execute sequentially until = K, thus we get ={ }; Step 4.4: Use formula (6) to get the final latent features : (6) In formula (6), Represents the convolution operation; Step 4.5: The final latent features Input the decoder D in the trained medical image prior encoder for processing to obtain the optimally reconstructed medical image .

5. An electronic device, comprising a memory and a processor, characterized in that: The memory is used to store a program that supports the processor to execute the medical image synthesis method according to any one of claims 1 to 4, and the processor is configured to execute the program stored in the memory.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the medical image synthesis method according to any one of claims 1 to 4 are executed.