PET image enhancement method and system based on generative model and medium

Through the improved generative model, the problems of insufficient generalization ability and difficulty in evaluating uncertainty in cross-domain PET image processing are solved, and more powerful adaptability and denoising effects are achieved.

CN120070212APending Publication Date: 2025-05-30NANFANG HOSPITAL OF SOUTHERN MEDICAL UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411922563.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When processing cross-domain PET images, the existing convolutional neural network models have limited generalization capabilities, which are difficult to adapt to data diversity and distribution uncertainty, and lack quantitative information for predictive uncertainty.

Method used

Using a PET image enhancement method based on a generative model, a vector quantization variational autoencoder is constructed, and improved based on hollow convolution and attention mechanisms are improved, the model's capture receptive field of cross-domain data is expanded, and the parameters are dynamically updated to adapt to the uncertainty of the input data.

Benefits of technology

The adaptability of the generative model to cross-domain PET images is improved, the generalization performance and denoising effect of the model are enhanced, and the reliability and stability of the prediction can be more accurately evaluated.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070212A_ABST
    Figure CN120070212A_ABST
Patent Text Reader

Abstract

The invention discloses a PET image enhancement method and system based on a generative model, and a medium. The PET image enhancement method comprises the following steps: collecting PET image data sets from different scanners, different tracers and different centers; preprocessing the PET image data set to obtain a PET image training set; constructing a vector quantization variational auto-encoder, and improving the vector quantization variational auto-encoder based on cavity convolution and an attention mechanism to obtain a generative model; training the generative model according to the PET image training set to obtain a generative image enhancement model; and inputting a PET image to be enhanced into the generative image enhancement model to obtain a PET enhanced image. The method can be more suitable for sample diversity and distribution uncertainty of cross-domain PET images, and can be widely applied to the technical field of image processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technologies, and in particular, to a PET image enhancement method, system, and medium based on a generative model. Background Art

[0002] Positron Emission Tomography (PET) is a molecular imaging technology that reflects the metabolic activities in an organism by injecting a tracer labeled with a radionuclide. PET has the characteristics of high sensitivity and high specificity and has been widely used in fields such as oncology, neurology, and cardiology. The whole-body low-count PET image enhancement model is usually implemented using a convolutional neural network. The convolutional neural network can extract the image features of the training data and is widely used for denoising low-count PET images. By using a large number of low-count PET images as training data, a convolutional neural network is used to generate high-quality PET images. However, the effectiveness of the Convolutional Neural Network (CNN) model is largely limited by the deterministic distribution of the training data, and it belongs to a data-driven deterministic estimation model. In practical applications, technical problems such as data uncertainty and model uncertainty are usually faced:

[0003] 1) When PET data comes from different medical institutions, uses different tracers or scanners, and the lesion positions and sizes are different, this kind of cross-domain data increases data diversity, affects the data distribution, and introduces significant uncertainty into the experimental data. If the CNN model fails to fully capture the diversity of this cross-domain data during the training process, it may lead to limited model generalization ability and affect the denoising effect;

[0004] 2) As a deterministic model, the prediction result of the CNN model is single and does not contain quantitative information about prediction uncertainty. This uncertainty may stem from multiple aspects such as the model structure, parameter initialization, and training process. When dealing with cross-domain data, due to the differences between different data domains, the CNN model may not be able to accurately evaluate the reliability and stability of its predictions. Summary of the Invention

[0005] To solve the above technical problems, the purpose of the present invention is to provide a PET image enhancement method, system, and medium based on a generative model, which can adapt to the sample diversity and distribution uncertainty of cross-domain data.

[0006] To achieve the above purpose, one aspect of the embodiments of the present application proposes a PET image enhancement method based on a generative model, including the following steps:

[0007] Collect PET image datasets from different scanners, different tracers, and different centers;

[0008] Preprocess the PET image datasets to obtain a PET image training set;

[0009] Construct a vector quantization variational autoencoder, and improve the vector quantization variational autoencoder based on dilated convolution and attention mechanism to obtain a generative model;

[0010] Train the generative model according to the PET image training set to obtain a generative image enhancement model;

[0011] Input the PET image to be enhanced into the generative image enhancement model to obtain a PET enhanced image.

[0012] In some embodiments, the preprocessing of the PET image datasets to obtain a PET image training set specifically includes:

[0013] Convert each PET image in the PET image datasets to obtain a first PET image set;

[0014] Resample the first PET image set to obtain a second PET image set;

[0015] Perform image enhancement on the second PET image set to obtain the PET image training set.

[0016] In some embodiments, the construction of the vector quantization variational autoencoder specifically includes:

[0017] Construct an encoder, which is used to encode the input data to obtain a continuous feature vector;

[0018] Construct a codebook, which is used to map the continuous feature vector to obtain a discrete codebook vector, and replace the discrete codebook vector with an index;

[0019] Construct a decoder, which is used to decode the index to obtain a reconstructed image.

[0020] In some embodiments, the improvement of the vector quantization variational autoencoder based on dilated convolution and attention mechanism to obtain a generative model specifically includes:

[0021] Improve the encoder based on dilated convolution, improve the codebook through a discrete coding structure, and improve the decoder based on the attention mechanism to obtain the generative model.

[0022] In some embodiments, the improvement of the encoder based on dilated convolution specifically includes:

[0023] The encoder is improved through a hierarchical coding structure. The improved encoder includes an upper-layer encoder and a lower-layer encoder;

[0024] Replace the convolutional layer in the upper-layer encoder with dilated convolution;

[0025] Among them, the upper-layer encoder includes a multi-layer convolutional network, and the lower-layer encoder includes a shallow convolutional structure.

[0026] In some embodiments, the decoder is improved based on the attention mechanism, specifically including:

[0027] Introduce the attention mechanism module into the decoder.

[0028] In some embodiments, training the generative model according to the PET image training set to obtain a generative image enhancement model specifically includes:

[0029] Input the PET image into the improved encoder to obtain multi-scale cross-domain image features;

[0030] Input the multi-scale cross-domain image features into the improved codebook to obtain discrete feature vectors;

[0031] Input the discrete feature vectors into the improved decoder to obtain a cross-domain PET enhanced image;

[0032] Optimize the parameters of the generative model according to the cross-domain PET enhanced image and a preset loss function to obtain the generative image enhancement model.

[0033] To achieve the above object, another aspect of the embodiments of the present application proposes a PET image enhancement system based on a generative model, including:

[0034] A data acquisition module, configured to acquire PET image data sets from different scanners, different tracers, and different centers;

[0035] An image preprocessing module, configured to preprocess the PET image data set to obtain a PET image training set;

[0036] A model improvement module, configured to construct a vector quantization variational autoencoder, and improve the vector quantization variational autoencoder based on dilated convolution and the attention mechanism to obtain a generative model;

[0037] A model training module, configured to train the generative model according to the PET image training set to obtain a generative image enhancement model;

[0038] An image enhancement module, configured to input a PET image to be enhanced into the generative image enhancement model to obtain a PET enhanced image.

[0039] To achieve the above object, on the other hand, an embodiment of the present application provides an electronic device, which includes a memory, a processor, a program stored on the memory and executable on the processor, and a data bus for realizing connection communication between the processor and the memory. When the program is executed by the processor, it implements the PET image enhancement method based on the generative model as described above.

[0040] To achieve the above object, on the other hand, an embodiment of the present application provides a storage medium, which is a computer-readable storage medium for computer-readable storage. The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the PET image enhancement method based on the generative model as described above.

[0041] The beneficial effects of the present invention are as follows: The PET image enhancement method, system and medium based on the generative model of the present invention first collect PET image data sets from different scanners, different tracers and different centers, and preprocess the PET image data sets to obtain a PET image training set. Secondly, a vector quantization variational autoencoder is constructed, and the vector quantization variational autoencoder is improved based on dilated convolution and attention mechanism to obtain a generative model. Then, the generative model is trained according to the PET image training set to obtain a generative image enhancement model. Finally, the PET image to be enhanced is input into the generative image enhancement model to obtain a PET enhanced image. The present invention improves the vector quantization variational autoencoder based on dilated convolution and attention mechanism, which can expand the receptive field of the generative model for capturing cross-domain data, and fuse multi-scale features with the help of the attention mechanism, so that the generative model dynamically updates parameters according to the uncertainty of the input data, so that the trained generative image enhancement model can better adapt to the sample diversity and distribution uncertainty of cross-domain PET images. Description of the Drawings

[0042] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following introduces the drawings required to be used in the embodiments of the present invention. It should be understood that the drawings introduced below are only for conveniently and clearly expressing some embodiments of the technical solutions in the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.

[0043] Figure 1 It is a flowchart of the steps of a PET image enhancement method based on a generative model provided by an embodiment of the present invention;

[0044] Figure 2 Structural schematic diagram of the generative model provided by the embodiment of the present invention;

[0045] Figure 3 Processing flow chart of the generative model provided by the embodiment of the present invention;

[0046] Figure 4 Structural schematic diagram of a PET image enhancement system based on the generative model provided by the embodiment of the present invention;

[0047] Figure 5 Hardware structural schematic diagram of an electronic device provided by the embodiment of the present invention. Detailed implementation manners

[0048] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the embodiments of the present application. They are only examples of devices and methods consistent with some aspects of the embodiments of the present application described in detail in the appended claims.

[0049] It can be understood that the terms "first", "second", etc. used in the present application can be used in this document to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, the first information can also be called the second information, and similarly, the second information can also be called the first information. Depending on the context, as used herein, the words "if", "when" can be interpreted as "when...", "when...", or "in response to determining".

[0050] The terms "at least one", "a plurality", "each", "any one", etc. used in the present application, at least one includes one, two or more than two, a plurality includes two or more than two, each refers to each of the corresponding plurality, and any one refers to any one of the plurality.

[0051] Positron Emission Tomography (PET) is a molecular imaging technique that reflects the metabolic activities in an organism by injecting a tracer labeled with a radionuclide. PET features high sensitivity and high specificity and has been widely applied in fields such as oncology, neurology, and cardiology. The whole-body low-count PET image enhancement model is usually implemented using a convolutional neural network. The convolutional neural network can extract the image features of the training data and is widely used for denoising low-count PET images. By providing a large amount of low-count PET images as training data, high-quality PET images can be generated using a convolutional neural network. However, the effectiveness of the Convolutional Neural Network (CNN) model is largely limited by the deterministic distribution of the training data, and it belongs to a data-driven deterministic estimation model. In practical applications, technical problems such as data uncertainty and model uncertainty are usually faced:

[0052] 1) When PET data comes from different medical institutions, uses different tracers or scanners, and the lesion locations and sizes vary, this kind of cross-domain data increases data diversity, affects the data distribution, and introduces significant uncertainty into the experimental data. During the training process of the CNN model, if the diversity of this cross-domain data fails to be fully captured, it may lead to limited model generalization ability and affect the denoising effect;

[0053] 2) As a deterministic model, the prediction result of the CNN model is single and does not contain quantitative information about prediction uncertainty. This uncertainty may stem from multiple aspects such as model structure, parameter initialization, and training process. When dealing with cross-domain data, due to the differences between different data domains, the CNN model may not be able to accurately evaluate the reliability and stability of its predictions.

[0054] To this end, the embodiments of the present invention propose a PET image enhancement method based on a generative model. First, a PET image dataset from different scanners, different tracers, and different centers is collected, and the PET image dataset is preprocessed to obtain a PET image training set. Secondly, a vector quantization variational autoencoder is constructed, and the vector quantization variational autoencoder is improved based on dilated convolution and attention mechanism to obtain a generative model. Then, the generative model is trained according to the PET image training set to obtain a generative image enhancement model. Finally, the PET image to be enhanced is input into the generative image enhancement model to obtain a PET enhanced image. The present invention improves the vector quantization variational autoencoder based on dilated convolution and attention mechanism, which can expand the receptive field of the generative model for capturing cross-domain data, and fuse multi-scale features with the help of the attention mechanism, so that the generative model dynamically updates parameters according to the uncertainty of the input data, so that the trained generative image enhancement model can better adapt to the sample diversity and distribution uncertainty of cross-domain PET images.

[0055] Referring to Figure 1 , Figure 1 FIG. is a flowchart of the steps of a PET image enhancement method based on a generative model provided by an embodiment of the present invention. The embodiments of the present invention propose a PET image enhancement method based on a generative model, and the method includes steps S101 to S105:

[0056] S101. Collect a PET image dataset from different scanners, different tracers, and different centers;

[0057] Specifically, collect a whole-body low-count PET image dataset from different scanners, different tracers, and different centers.

[0058] S102. Preprocess the PET image dataset to obtain a PET image training set;

[0059] Further as an optional implementation manner, the step of preprocessing the PET image dataset to obtain a PET image training set can be specifically divided into the following steps S1021 to S1023:

[0060] S1021. Convert each PET image in the PET image dataset to obtain a first PET image set;

[0061] S1022. Resample the first PET image set to obtain a second PET image set;

[0062] S1023. Enhance the second PET image set to obtain a PET image training set.

[0063] Specifically, first, convert the PET images in different datasets into SUV PET images to obtain the first PET image set. Converting the PET images into SUV PET images can reduce the dynamic range of the image intensity, making the image intensities between different datasets more comparable, thus facilitating subsequent image analysis and network training. Then, resample the PET images in the first PET image set to unify the image resolution and obtain the second PET image set. Finally, perform image enhancement processing on the PET images in the second PET image set to obtain the PET image training set.

[0064] Among them, the image enhancement processing can include operations such as image rotation, image inversion, image translation, and image offset to improve the quality and diversity of the images, thereby enhancing the model's ability to extract PET image features and generalization performance.

[0065] S103. Construct a vector quantization variational autoencoder, and improve the vector quantization variational autoencoder based on dilated convolution and attention mechanism to obtain a generative model;

[0066] Further, as an optional implementation manner, the step of constructing a vector quantization variational autoencoder can be specifically divided into the following steps S1031 to S1033:

[0067] S1031. Construct an encoder, which is used to encode the input data to obtain a continuous feature vector;

[0068] S1032. Construct a codebook, which is used to map the continuous feature vector to obtain a discrete codebook vector, and replace the discrete codebook vector with an index;

[0069] S1033. Construct a decoder, which is used to decode the index to obtain a reconstructed image.

[0070] In some optional embodiments, the vector quantization variational autoencoder (VQ-VAE model) includes three parts: an encoder, a codebook, and a decoder. The encoder first encodes the input data x into a continuous latent representation vector z e (x). Then, this continuous latent representation vector is mapped to a discrete codebook vector through a quantization layer. The codebook is composed of a set of learned discrete vectors, and each vector corresponds to an index in the codebook. The role of the quantization layer is to find the codebook vector closest to the continuous latent representation vector and replace the original continuous vector with its index. This index is then used as the input to the decoder. The task of the decoder is to reconstruct the input data according to this discrete index, that is, to generate a reconstructed image. In this process, the decoder actually maps the index back to the corresponding vector in the codebook and then uses this vector to reconstruct the data.

[0071] Specifically, the model defines a learnable latent embedding space \(e\in\mathbb{R}\) K×D , also known as the codebook, where \(K\) is the size of the discrete latent space, that is, the number of representation vectors \(e\) i . There are \(K\) embedding vectors \(e\) i \(\in\mathbb{R}\) D , \(i\in1,2,\cdots,K\), and \(D\) is the dimension of each latent embedding vector \(e\) i . The model receives the input image \(x\) and generates the continuous feature vector \(z\) e \((x)\) through the encoder. Assume that the feature map size of the continuous feature vector \(z\) e \((x)\) is \(H\times W\times D\), that is, \(H\times W\) \(D -\)dimensional vectors. Each vector finds the index \(k\) of the vector closest to it in the codebook \(e\) through nearest neighbor search, and obtains the closest vector according to the index, that is, the quantized discrete feature vector \(z\) q \((x)\). The specific process is shown in the following formula:

[0072] \(z\) q \((x)=e\) k , where \(k = \arg\min\) j \(\left\lVert z\right.\) e \((x)-e\) j \(\right\rVert\) 2

[0073] Among them, \(z\) e \((x)\) is the continuous feature vector output by the encoder, and \(z\) q \((x)\) is the quantized discrete feature vector. During the forward propagation process, the index of the codebook space needs to be taken.

[0074] Furthermore, as an optional implementation, the vector quantization variational auto - encoder is improved based on dilated convolution and attention mechanism to obtain the generative model. This step can be specifically divided into the following steps S1034:

[0075] S1034. Improve the encoder based on dilated convolution, improve the codebook through the discrete coding structure, and improve the decoder based on the attention mechanism to obtain the generative model.

[0076] Specifically, the embodiment of the present invention improves the encoder based on dilated convolution. By setting the dilation rate to fill zeros in the convolution, the receptive field of the input image is further increased, providing more context information for the model. Secondly, a discrete coding structure is adopted for the latent variable space (i.e., the codebook), and the continuous variables in the latent space are quantified by finding the nearest neighbor codebook vectors. On the one hand, this discretization process reduces the complexity of the model and makes the training more stable. On the other hand, the discrete structured representation is easier to capture the global and local features of the image. Furthermore, the decoder is improved based on the attention mechanism. The enhanced attention mechanism can process all elements of the sequence in parallel, dynamically adjust the focus of attention, and the model can more flexibly adapt to different data features and task requirements.

[0077] It should be noted that, compared with the traditional convolutional layer, the self-attention mechanism can reduce the parameters of the model without sacrificing performance, represent features more precisely, consider the relationship between each element and other elements, not just the local neighborhood, effectively capture the global dependence of the input data, and improve the generalization ability.

[0078] Further as an optional implementation manner, the step of improving the encoder based on dilated convolution can be specifically divided into the following steps S10341 and S10342:

[0079] S10341. Improve the encoder through a hierarchical coding structure. The improved encoder includes an upper encoder and a lower encoder;

[0080] Among them, the upper encoder includes multiple convolutional networks, and the lower encoder includes a shallow convolutional structure.

[0081] Specifically, as Figure 2 shown in the structural schematic diagram of the generative model, the embodiment of the present invention improves the vector quantization variational autoencoder (VQ-VAE model), and proposes a new hierarchical vector quantization variational autoencoder (LAH-VAE model) to generate PET enhanced images. The upper and lower encoders are used to achieve multi-scale feature representation of the image. The upper encoder is composed of multiple convolutional networks, and its deep structure design expands the receptive field, enabling it to effectively capture the global features of low-count PET images; the lower encoder models the image details through a shallow convolutional structure to capture the local features of low-count PET images.

[0082] Exemplarily, as Figure 3The figure shows the processing flow chart of the generative model. If the input of the model is a low-count PET image of 256×256, the lower encoder downsamples the image by a factor of 4, and the representation is the underlying feature map of 64×64. Then, through the upper encoder, the size of the feature map is further reduced by half, generating the top-level feature map of 32×32. The top-level feature map passes through a quantizer composed of a codebook, and then through the decoder for upsampling to generate a feature map of 64×64. The decoder is also a feed-forward network, which takes the quantized latent feature vector as input and consists of several residual blocks and transposed convolution layers, upsampling the upper-level feature map to the size of the original image. Finally, the feature map of the lower encoder and the feature map of the upper layer that has passed through the encoder, quantizer, and decoder are merged in the channel to form the output of the entire model encoder.

[0083] S10342. Replace the convolutional layer in the upper encoder with a dilated convolution;

[0084] Specifically, as Figure 2 shown, to further increase the multi-scale dimensional feature extraction of the hierarchical encoder, the embodiment of the present invention also replaces the convolutional layer of the upper encoder with a dilated convolution, and further increases the receptive field of the input image by setting the dilation rate to fill zeros in the convolution. The formula for the dilated convolution is as follows:

[0085] y[i,j,k]=∑∑∑w[p,q,r,k]*x[i+a*p,j+b*q,(r-1)d+k]

[0086] where y[i,j,k] represents the output of the convolutional layer, w[p,q,r,k] represents the weight of the convolutional kernel, x[i+a*p,j+b*q,(r-1)d+k] represents the pixel value of the input data, a and b respectively represent the strides in the horizontal and vertical directions, and d represents the dilation size of the dilated convolution.

[0087] It should be noted that compared with the traditional convolutional kernel, the dilated convolution can arbitrarily expand the receptive field without additional parameters. When k is the size of the original convolutional kernel and r is the dilation rate of the dilated convolution parameter, the actual convolutional kernel size corresponding to the dilated convolution is k = k+(k - 1)×(r - 1). For the same 3×3 convolutional kernel, the dilated convolution can achieve an approximate effect of 5×5 or 7×7.

[0088] Further as an optional implementation manner, the step of improving the decoder based on the attention mechanism can be specifically divided into the following steps S10343:

[0089] S10343. Introduce an attention mechanism module in the decoder.

[0090] Specifically, incorporate the output features of the entire hierarchical encoder into the quantizer to obtain the underlying discrete feature vector z bq(x). To effectively combine discrete feature vectors of multiple scales, the discrete feature vector z output by the top-level quantizer is q upsampled on (x) to obtain a global discrete feature vector z bq with the same dimension as (x) tq (x). In the embodiment of the present invention, an enhanced attention mechanism module is used to replace the channel merging operation to combine the global and local discrete feature vectors. The enhanced attention mechanism is defined as passing the global and local discrete feature vectors through the self-attention mechanism module respectively. For each element xi in the discrete feature vector, the three vectors of query Q, key K, and value V are calculated by the following formula:

[0091] Q = xW Q , K = xW K , V = xW v

[0092] where W Q , W K and W v are all learned weight matrices used to map the input vector to different vector spaces. After mapping, Q, K, and V have the same dimension. The dot product Score = QK T is calculated for similarity measurement. To avoid the dot product result from being too large, scale scaling is performed to stabilize the gradient The attention scores of the query vector Q are normalized through the softmax function, and finally the V vector is weighted and summed to obtain the final output vector, as shown in the following formula:

[0093]

[0094] z bq (x) and z tq (x) are respectively processed by the self-attention mechanism for a single vector and then merged through the following formula:

[0095] z Q (x) = Concat(SelfAttention(z bq (x)), SelfAttention(z tq (x)))

[0096] As Figure 2 and Figure 3 shown, the concatenated result is further processed by a convolutional layer to obtain a multi-scale discrete feature vector combined based on the enhanced attention mechanism, and finally a reconstructed PET enhanced image is obtained through a decoder.

[0097] S104. Train the generative model according to the PET image training set to obtain a generative image enhancement model;

[0098] As a further optional implementation, the step of training the generative model according to the PET image training set to obtain the generative image enhancement model can be specifically divided into the following steps S1041 to S1044:

[0099] S1041: Input the PET image into the improved encoder to obtain multi-scale cross-domain image features;

[0100] S1042: Input the multi-scale cross-domain image features into the improved codebook to obtain discrete feature vectors;

[0101] S1043: Input the discrete feature vectors into the improved decoder to obtain cross-domain PET enhanced images;

[0102] Specifically, in the first stage of model training, input the PET image into the improved encoder. The encoder combines dilated convolution and hierarchical encoding structure to generate multi-scale cross-domain image features. In the second stage of model training, input the multi-scale cross-domain image features into the improved codebook, and hierarchically quantize the multi-scale features based on the discrete coding structure codebook to generate discrete feature vectors. In the third stage of model training, input the discrete feature vectors into the improved decoder, and use the attention mechanism module to fuse the multi-scale image features to generate high-quality cross-domain PET enhanced images.

[0103] S1044: Optimize the parameters of the generative model according to the cross-domain PET enhanced image and the preset loss function to obtain the generative image enhancement model.

[0104] Specifically, the model adopts the echelon stopping technique of the direct gradient estimator. Since the dimensionality of the discrete feature vector z q (x) and the continuous feature vector z e (x) is the same, both being H×W×D, during backpropagation, directly copy the gradient of the discrete feature vector z q (x) to the continuous feature vector z e (x), and the copied gradient directly guides the parameter update of the encoder.(x), and the copied gradient directly guides the parameter update of the encoder.

[0105] In the fourth stage of model training, the loss function of the model is designed, and the model parameters are iteratively optimized to make it applicable to cross-domain PET image enhancement. The loss function of the model includes reconstruction loss, embedding loss, and commitment loss. Due to the use of the gradient stop technique, the reconstruction loss can only optimize the parameters of the encoder and decoder to make the reconstructed image closer to the real image, and the codebook cannot be updated. To align the vector space of the codebook with the encoder output, the model introduces an embedding loss and a commitment loss. The embedding loss is used to update the codebook vectors to make the codebook vectors as close as possible to the latent representation output by the encoder. The commitment loss encourages the model to be more firm in the selection of codebook vectors during training. Once the output of an encoder is quantized to the corresponding codebook vector, the mapping relationship is maintained as much as possible to help the model learn a more stable and consistent latent representation. The loss function is shown as follows:

[0106]

[0107] Among them, the first term is the reconstruction loss, the second term is the embedding loss, and the last term is the commitment loss. sg represents the gradient stop technique, e is the codebook vector, and D(e) represents the output of the decoder, that is, the reconstructed image. The decoder is optimized and updated through the reconstruction loss, and the reconstruction loss and the commitment loss will optimize the encoder. The codebook is optimized and updated through the embedding loss. In actual operation, since the exponential moving average (EMA) converges faster than the L2 loss, the EMA technique is usually used to update the embedding space. At each iteration t, the latent vector e k is updated by the following formula:

[0108]

[0109]

[0110]

[0111] where is the eigenvector assigned to e k in k is the number of eigenvectors, γ is the decay parameter, usually set to 0.99.

[0112] S105. Input the PET image to be enhanced into the generative image enhancement model to obtain the enhanced PET image.

[0113] In summary, the processing flow of the PET image enhancement method based on the generative model in the embodiments of the present invention is as follows:

[0114] Step 1: Collect PET image datasets from different scanners, different tracers, and different centers;

[0115] Step 2: Preprocess the collected PET image datasets across scanners, tracers, and centers to obtain a PET image training set;

[0116] Step 3: Build a vector quantization variational autoencoder, which includes an encoder, a codebook, and a decoder;

[0117] Step 4: Improve the encoder in the basic framework of the built vector quantization variational autoencoder. The encoder adopts a hierarchical coding structure, including upper and lower layer coding structures. The lower layer coding structure models the image details through a shallow convolutional structure to capture the local features of low-count PET images, and the upper layer encoder is composed of multiple convolutional networks. Its deep structure design expands the receptive field, enabling it to effectively capture the global features of low-count PET images;

[0118] Step 5: Replace the convolutional layer of the upper layer encoder with a dilated convolution. By setting the dilation rate to fill zeros in the convolution, further increase the receptive field of the input image, providing more context information for the image enhancement model;

[0119] Step 6: Adopt a discrete coding structure for the latent variable space (i.e., the codebook), and quantify the continuous variables in the latent space by finding the nearest neighbor codebook vectors. This discretization process reduces the model complexity on the one hand, making the training more stable, and on the other hand, the discrete structured representation is easier to capture the global and local features of the image;

[0120] Step 7: Introduce an attention mechanism in the decoder. The attention mechanism can process all elements of the sequence in parallel, dynamically adjust the focus of attention, enabling the model to more flexibly adapt to different data features and task requirements. Compared with traditional convolutional layers, the attention mechanism can reduce the model parameters without sacrificing performance, represent features more precisely, consider the relationship between each element and other elements, rather than just the local neighborhood, effectively capture the global dependence of the input data, and improve the generalization ability;

[0121] Step 8: Train based on the PET image training set obtained by preprocessing and the improved generative model to obtain a generative image enhancement model applicable to different domains;

[0122] Step 9: Apply the generative image enhancement model to cross-domain PET image enhancement to obtain PET enhanced images.

[0123] The above description is about the PET image enhancement method based on the generative model in the embodiments of the present invention. It can be recognized that compared with the PET image enhancement methods in the prior art, the embodiments of the present invention have the following advantages:

[0124] First, it provides the ability to dynamically learn the importance of features. By means of dilated convolution and hierarchical coding structure, it expands the receptive field of the model for capturing cross-domain data. With the help of the attention mechanism, it fuses multi-scale features and dynamically learns the importance of features, enabling the model to dynamically update parameters according to the uncertainty of the input data. Learning the importance of features makes the model more adaptable to the sample diversity and distribution uncertainty of cross-domain data.

[0125] Second, the encoder is improved by combining the hierarchical coding structure with dilated convolution. The hierarchical coding structure can effectively capture information at different scales and levels. This structure allows the model to encode images at different levels, thereby improving the reconstruction quality and detail performance, reducing information overlap, and enhancing the efficiency of the latent representation. Dilated convolution extracts richer context information while maintaining the image resolution by increasing the receptive field without downsampling. The hierarchical structure enables the model to perform feature extraction at different levels, while dilated convolution ensures that important spatial information is not lost during the denoising process. Combining the two in the encoder can provide more accurate enhanced PET images.

[0126] Third, the codebook is improved by adopting a discrete coding structure. The discrete representation of latent variables makes the hidden representation of the image more structured, which helps to capture local features and global structures in the image, thus improving the quality and efficiency of image restoration. Discrete coding can more effectively retain information in the image because each discrete codebook vector represents a specific set of features, avoiding information loss that may occur in continuous coding. Since the representation generated by discrete coding consists of a finite number of discretized code vectors, it is easier to interpret and understand, which helps to analyze the model's understanding and restoration process of the image.

[0127] Fourth, the encoder is improved by adopting the attention mechanism. It can process all elements of the sequence in parallel, improving the computational efficiency. It can dynamically adjust the focus of attention, enabling the model to be more flexible in adapting to different data features and task requirements. Compared with traditional convolutional layers, the self-attention mechanism can reduce the parameters of the model without sacrificing performance, represent features more precisely, consider the relationship between each element and other elements, rather than just the local neighborhood, effectively capture the global dependence of the input data, and improve the generalization ability.

[0128] Referring to Figure 4 , the embodiments of the present invention also provide a PET image enhancement system based on the generative model, including:

[0129] A data acquisition module for acquiring PET image datasets from different scanners, different tracers, and different centers;

[0130] An image preprocessing module for preprocessing the PET image datasets to obtain a PET image training set;

[0131] A model improvement module for constructing a vector quantization variational autoencoder and improving the vector quantization variational autoencoder based on dilated convolution and attention mechanism to obtain a generative model;

[0132] A model training module for training the generative model according to the PET image training set to obtain a generative image enhancement model;

[0133] An image enhancement module for inputting the PET image to be enhanced into the generative image enhancement model to obtain a PET enhanced image.

[0134] The content in the above embodiments of the PET image enhancement method based on the generative model is applicable to the embodiments of the PET image enhancement system based on the generative model. The functions specifically implemented by the embodiments of the PET image enhancement system based on the generative model are the same as those in the above embodiments of the PET image enhancement method based on the generative model, and the beneficial effects achieved are also the same as those in the above embodiments of the PET image enhancement method based on the generative model.

[0135] An embodiment of the present invention also provides an electronic device, which includes: a memory, a processor, a program stored on the memory and executable on the processor, and a data bus for realizing the connection and communication between the processor and the memory. When the program is executed by the processor, it implements the above PET image enhancement method based on the generative model. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.

[0136] As Figure 5 shown is a schematic diagram of the hardware structure of the electronic device provided by an embodiment of the present invention. Referring to Figure 5 , an embodiment of the present invention provides an electronic device, including:

[0137] A processor 1001, which can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present invention;

[0138] The memory 1002 can be implemented in the form of a Read Only Memory (ROM), a static storage device, a dynamic storage device, or a Random Access Memory (RAM), etc. The memory 1002 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1002 and are called by the processor 1001 to execute the PET image enhancement method based on the generative model in the embodiments of the present invention;

[0139] The input / output interface 1003 is used to implement information input and output;

[0140] The communication interface 1004 is used to implement communication interaction between this device and other devices. It can achieve communication through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);

[0141] The bus 1005 transmits information between various components of the device (such as the processor 1001, the memory 1002, the input / output interface 1003, and the communication interface 1004);

[0142] Among them, the processor 1001, the memory 1002, the input / output interface 1003, and the communication interface 1004 are communicatively connected to each other inside the device through the bus 1005.

[0143] The embodiments of the present invention also provide a storage medium. The storage medium is a computer-readable storage medium for computer-readable storage. The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the above-mentioned PET image enhancement method based on the generative model.

[0144] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0145] Embodiments of the present invention also disclose a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device may read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes Figure 1 the method shown.

[0146] In some alternative embodiments, the functions / operations mentioned in the block diagrams may not occur in the order mentioned in the operation diagrams. For example, depending on the functions / operations involved, two consecutive blocks shown may actually be executed substantially simultaneously or the above-mentioned blocks can sometimes be executed in the reverse order. In addition, the embodiments presented and described in the flowcharts of the present invention are provided by way of example for the purpose of providing a more comprehensive understanding of the technology. The disclosed method is not limited to the operations and logical flows presented herein. Alternative embodiments are contemplated, in which the order of various operations is changed and the sub-operations described as part of a larger operation are executed independently.

[0147] In addition, although the present invention has been described in the context of functional modules, it should be understood that one or more of the above functions and / or features may be integrated in a single physical device and / or software module unless otherwise stated to the contrary, or one or more functions and / or features may be implemented in separate physical devices or software modules. It can also be understood that a detailed discussion of the actual implementation of each module is not necessary for understanding the present invention. More precisely, considering the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skills of an engineer. Therefore, those skilled in the art can implement the present invention as set forth in the claims without undue experimentation. It can also be understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.

[0148] If the above functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the above method in various embodiments of the present invention. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, etc., which can store program codes of various kinds.

[0149] The logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.

[0150] More specific examples (non-exhaustive list) of computer-readable media include the following: electrical connection parts (electronic devices) having one or more wirings, portable computer disk cartridges (magnetic devices), random access memories (RAMs), read-only memories (ROMs), erasable programmable read-only memories (EPROMs or flash memories), optical fiber devices, and portable compact disc read-only memories (CDROMs). Additionally, the computer-readable medium can even be paper or other suitable media on which the above program can be printed, because the above program can be obtained electronically, for example, by optically scanning the paper or other media, then editing, interpreting, or processing it in other suitable ways as necessary, and then storing it in a computer memory.

[0151] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.

[0152] In the foregoing description of the present specification, the description with reference to the terms "one embodiment / example", "another embodiment / example", or "certain embodiments / examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0153] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the claims and their equivalents.

[0154] The above has specifically described the preferred embodiments of the present invention, but the present invention is not limited to the embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.

Claims

1. A PET image enhancement method based on a generative model, characterized in that: The following steps are involved: Acquisition of PET image datasets from different scanners, different tracers, and different centers; Preprocessing the PET image data set to obtain a PET image training set; A vector quantized variational autoencoder is constructed, and the vector quantized variational autoencoder is improved based on dilated convolution and attention mechanism to obtain a generative model; Training the generative model according to the PET image training set to obtain a generative image enhancement model; The PET image to be enhanced is input into the generative image enhancement model to obtain a PET enhanced image.

2. A PET image enhancement method based on a generative model according to claim 1, characterized in that: The preprocessing of the PET image data set to obtain a PET image training set specifically includes: Converting each of the PET images in the PET image data set to obtain a first PET image set; resampling the first PET image set to obtain a second PET image set; Perform image enhancement on the second PET image set to obtain the PET image training set.

3. A PET image enhancement method based on a generative model according to claim 1, characterized in that: The constructing of a vector quantization variational autoencoder specifically includes: Constructing an encoder, wherein the encoder is used to encode the input data to obtain a continuous feature vector; Constructing a codebook, the codebook being used to map the continuous feature vector to obtain a discrete codebook vector, and replacing the discrete codebook vector with an index; A decoder is constructed, and the decoder is used to decode the index to obtain a reconstructed image.

4. A PET image enhancement method based on a generative model according to claim 3, characterized in that: The vector quantization variational autoencoder is improved based on the dilated convolution and attention mechanism to obtain a generative model, which specifically includes: The encoder is improved based on dilated convolution, the codebook is improved through a discrete coding structure, and the decoder is improved based on an attention mechanism to obtain the generative model.

5. A PET image enhancement method based on a generative model according to claim 4, characterized in that: The improvement of the encoder based on the dilated convolution specifically includes: The encoder is improved by a layered coding structure, wherein the improved encoder comprises an upper layer encoder and a lower layer encoder; Replacing the convolutional layer in the upper encoder with a dilated convolution; Among them, the upper encoder includes a multi-layer convolutional network, and the lower encoder includes a shallow convolutional structure.

6. A PET image enhancement method based on a generative model according to claim 4, characterized in that: The improvement of the decoder based on the attention mechanism specifically includes: The attention mechanism module is introduced into the decoder.

7. A PET image enhancement method based on a generative model according to claim 4, characterized in that: The training of the generative model according to the PET image training set to obtain a generative image enhancement model specifically includes: Inputting the PET image into the improved encoder to obtain multi-scale cross-domain image features; Inputting the multi-scale cross-domain image features into the improved codebook to obtain a discrete feature vector; Inputting the discrete feature vector into the improved decoder to obtain a cross-domain PET enhanced image; According to the cross-domain PET enhanced image and a preset loss function, the generative model is parameter optimized to obtain the generative image enhancement model.

8. A PET image enhancement system based on a generative model, characterized in that: include: A data acquisition module for acquiring PET image datasets from different scanners, different tracers, and different centers; An image preprocessing module, used for preprocessing the PET image data set to obtain a PET image training set; A model improvement module, used for constructing a vector quantized variational autoencoder, improving the vector quantized variational autoencoder based on a dilated convolution and an attention mechanism to obtain a generative model; A model training module, used for training the generative model according to the PET image training set to obtain a generative image enhancement model; The image enhancement module is used to input the PET image to be enhanced into the generative image enhancement model to obtain a PET enhanced image.

9. An electronic device, characterized in that: The electronic device includes a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for realizing connection and communication between the processor and the memory. When the program is executed by the processor, the steps of the generative model-based PET image enhancement method as described in any one of claims 1 to 7 are realized.

10. A storage medium, the storage medium being a computer-readable storage medium, used for computer-readable storage, characterized in that: The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the generative model-based PET image enhancement method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • No-reference low-illumination image enhancement method and system based on generative adversarial network

    CN111798400A