A medical image enhancement processing method and system based on deep learning
By adopting deep learning-based medical image enhancement processing methods in medical image analysis systems, using generative adversarial networks and autoencoders to increase sample diversity, and model training is carried out through federated learning architecture, the sample singularity problem in the existing system is solved, and the generalization ability and data privacy protection of the model are improved.
Patent Information
- Application Number
- CN202510175817.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-02-18
AI Technical Summary
The image samples of existing medical image analysis systems based on artificial intelligence are usually relatively single, lacking diversity, affecting the generalization ability of the model.
Deep learning-based medical image enhancement processing method is adopted to enhance image by generating adversarial networks and autoencoders, increasing sample diversity, and using federated learning architecture to train models on multiple scattered devices or servers to avoid exchanging original medical image data.
It effectively increases the diversity of samples, improves the generalization ability of the model, makes it more robust and accurate in practical applications, and protects data privacy.
Smart Images

Figure CN119671884B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image processing, and in particular to a medical image enhancement processing method and system based on deep learning. Background Art
[0002] Traditional medical image processing mainly relies on mathematical and physical principles to identify different structures of human tissues, and commonly used technologies include X-rays, CT, MRI, and ultrasound. These methods mainly rely on manual operations and rule-based algorithms in image preprocessing, image segmentation, and feature extraction, which is time-consuming and cumbersome.
[0003] The invention patent with the existing publication number CN119006382A proposes a medical image analysis system based on artificial intelligence, including a data preprocessing module, a feature extraction module, a diagnosis support module, and a learning and optimization module. The data preprocessing module preprocesses the medical image data to segment the lesion area, the feature extraction module extracts features from the lesion area to obtain the state information corresponding to the lesion area, the diagnosis support module evaluates the state information based on the trained artificial intelligence model to form an evaluation result, and warns the medical staff of the evaluation result. The learning and optimization module triggers the adjustment of the artificial intelligence model according to the medical staff's score on the evaluation result. Through the cooperation of the feature extraction module and the diagnosis support module, it is ensured that the entire system has the advantages of strong medical image adaptability, high intelligence, reliable recognition efficiency, low labor intensity of medical staff, strong human-computer interaction ability and strong early lesion recognition ability.
[0004] However, the above-mentioned AI-based medical image analysis system mainly focuses on feature extraction and status evaluation through existing medical image data, without explicitly mentioning how to increase sample diversity through technical means. Image samples are usually relatively single and lack diversity. In practical applications, if the input medical image data itself lacks diversity, then even if the system has high performance, the generalization ability of the model may be affected due to the limitations of the training data, and it may not be able to effectively deal with various pathological conditions and image variations. Summary of the invention
[0005] The purpose of the present invention is to provide a medical image enhancement processing method and system based on deep learning, which solves the problem that the image samples of the existing medical image analysis system based on artificial intelligence are usually relatively single and lack diversity, which affects the generalization ability of the model.
[0006] To achieve the above object, the present invention provides a medical image enhancement processing method based on deep learning, comprising the following steps:
[0007] Select the image enhancement model based on whether the number of samples needs to be expanded;
[0008] If the number of samples needs to be expanded, a generative adversarial network is used for image enhancement, otherwise an autoencoder is used for image enhancement;
[0009] The model is trained using a federated learning architecture, and is trained on multiple distributed devices or servers, where the devices or servers hold local medical image data samples without the need to exchange these medical image data.
[0010] If the number of samples needs to be expanded, a generative adversarial network is used for image enhancement, otherwise an autoencoder is used for image enhancement, and the steps further include:
[0011] The generative adversarial network is trained based on a generative adversarial network based on local feature enhancement. On the basis of the traditional generative adversarial network, by adopting multi-layer nonlinear transformation and local feature enhancement mechanism, the generator not only relies on the noise vector, but also integrates the historical generation state and local features of the image.
[0012] Wherein, the training is performed based on a generative adversarial network with local feature enhancement, and the steps further include:
[0013] Initialize the generator and discriminator
[0014] Multi-layer nonlinear transformation and local feature enhancement mechanism are added to the generator network. The nonlinear transformation operation of the generator depends on the input noise and the historical generation state of the image to achieve more complex image generation, which can be expressed as:
[0015]
[0016] In the formula, For the generator The weight matrix of the layer; is the bias term of the generator; For the generator Local features of the generated image of the layer; is the first nonlinear function, specifically Sigmoid, is the second nonlinear function, is the number of layers of the generator
[0017] In each round of training, the discriminator is first trained using real medical images as input, and then a random noise vector is input to the generator to generate images, and the generated images are sent to the discriminator for evaluation. The goal of the discriminator is to distinguish between real images and generated images, while the goal of the generator is to maximize the judgment error of the discriminator. The discriminator and the generator are trained alternately. The loss function of the discriminator is calculated based on the adversarial loss, and an adaptive adjustment mechanism based on the gradient domain is used to enhance the sensitivity of the discriminator to detail features, which can be expressed as:
[0018]
[0019] In the formula, For real medical images, is the weight of the discriminator, is the discriminator function, is the distribution of real medical image data, is the distribution of the noise vector, is the loss function of the discriminator, Express expectations, It means that it obeys a specific distribution. is the distribution of generated medical image data, is the L2 norm, is the gradient of the discriminator to the generated image;
[0020] The diversity of images generated by the generator is increased through the diffusion process and nonlinear approximation method. The diffusion process expands the noise vector of the generator from the low-dimensional space to the high-dimensional space to increase the diversity of sample generation, and enhances the details of the image by gradually optimizing the output image of the generator. The diffusion process is expressed as:
[0021]
[0022] In the formula, For the The image after diffusion of iterative times, For the The diffusion process of times iteration;
[0023] During the training process, the gradient domain adaptive adjustment strategy is adopted to adjust the gradient propagation mode in the network according to the gradient information of the image, and automatically adjust the gradient size to stabilize the convergence. The gradient update rules of the generator and the discriminator are expressed as follows:
[0024]
[0025]
[0026] In the formula, is the gradient adjustment term of the generator, is the gradient adjustment term of the discriminator, The loss function of the generator is the inverse of the loss function of the discriminator. is the loss function of the discriminator, is the adaptive gradient adjustment function of the generator, is the adaptive gradient adjustment function of the discriminator, represents the Hadamard product, is the gradient of the generated image
[0027] In the process of generating the generator, the local feature enhancement loss function is used to dynamically enhance the feature expression of the key information area in the image. The calculation method of the local feature enhancement loss function is expressed as:
[0028]
[0029] In the formula, Enhance the loss function for local features, To generate the image in The representation on a local area, For the real image The representation on a local area, For the The weight of a region measures the importance of the region to the overall image quality. The weight of the region is calculated based on the local gradient and high-frequency features of the image. It is adaptively adjusted according to the size of the gradient and the high-frequency part of the regional features to ensure that the weight of the important region is large. It is expressed as:
[0030]
[0031] In the formula, Indicates area Generating images The gradient on To control the sensitivity of local region importance, Indicates A local area range;
[0032] By continuously iteratively optimizing the parameters of the generator and the discriminator until the generator can produce high-quality images and the discriminator's ability to distinguish between true and false images is optimal, the update method is expressed as:
[0033]
[0034]
[0035] In the formula, For the The learning rate of the generator for the iteration, For the The learning rate of the discriminator at the iteration, is the gradient of the generator’s loss function with respect to the weight parameters, is the gradient of the discriminator’s loss function with respect to the weight parameters, is the parameter update operation, Enhance the loss function for local features, The gradient of the loss function for local feature enhancement with respect to the generator weight parameters, The gradient of the loss function with respect to the discriminator weight parameters is enhanced for local features, and an adaptive dynamic learning rate optimization strategy is used to enhance the accuracy and efficiency of gradient updates during training. The calculation method is expressed as:
[0036]
[0037]
[0038] In the formula, is the initial learning rate of the generator, is the initial learning rate of the discriminator, is the dynamic learning rate adjustment factor, which controls the decay rate of the learning rate. is the learning rate regularization constant, which prevents the learning rate from being too large in the case of extremely small gradients. It is the power exponent of the adjustment factor, which is used to control the nonlinear characteristics of the learning rate decay.
[0039] Wherein, the training is performed based on a generative adversarial network with local feature enhancement, and the steps further include:
[0040] Both the generator and the discriminator use a convolutional neural network architecture to identify the difference between generated images and real images. The generator generates medical images by inputting noise and model weights, which is expressed as:
[0041]
[0042] In the formula, is the random input noise of the generator, is the weight of the generator model, Medical images generated by the generator, is a generator function.
[0043] Wherein, the training is performed based on a generative adversarial network with local feature enhancement, and the steps further include:
[0044] The second nonlinear function is a multi-level local feature mapping performed by combining the weighted ReLU activation function with the convolution kernel. The calculation method is expressed as:
[0045]
[0046] In the formula, is the ReLU activation function, which is used to activate local features and enhance the accuracy of image generation in the local context. is the adjustment factor of the generator.
[0047] If the number of samples needs to be expanded, a generative adversarial network is used for image enhancement, otherwise an autoencoder is used for image enhancement, and the steps further include:
[0048] The autoencoder uses a feature sparse mapping-based autoencoder for medical image data enhancement based on sparse coding and adaptive LeakyReLU activation function.
[0049] Wherein, the medical image data is enhanced by an autoencoder based on feature sparse mapping, and the steps further include:
[0050] Initialize the weights and biases in the autoencoder network. The initialization method is expressed as:
[0051]
[0052] In the formula, represents the initial weight of the autoencoder, is the initialization coefficient of the autoencoder, used to scale the weights, is the dimension of the autoencoder input feature, represents a normal distribution with a mean of 0 and a variance of 1, where the bias term of the autoencoder is initialized to zero, expressed as:
[0053]
[0054] In the formula, Represents the initial bias of the autoencoder. The initialization coefficient of the autoencoder is calculated by the dimension of the input medical image data and the distribution characteristics of the medical image data, which is expressed as:
[0055]
[0056] In the formula, is the initial scaling factor, is the number of features of the input medical image data, is the number of features of the output medical image data of the encoder in the autoencoder;
[0057] The input high-dimensional features are normalized and expressed as:
[0058]
[0059] In the formula, represents the normalized input medical image data, is the original input medical image data, is the mean of the input medical image data, is the standard deviation of the input medical image data;
[0060] The encoder part maps the input medical image data into a lower-dimensional sparse representation space.
[0061] The sparsity loss function is calculated during the encoding process and is expressed as:
[0062]
[0063] In the formula, is the sparsity loss function, is the weight of the encoder’s sparse regularization, represents the L1 norm, controlling the sparsity of features, is the L2 norm, is the encoder weight Column eigenvalues, is the number of encoder layers;
[0064] Perform sample space projection optimization by calculating the projection error of the sample and adjusting the mapping weight of the encoder to optimize the feature representation of the space after dimensionality reduction, which is expressed as:
[0065]
[0066] In the formula, represents the weight matrix of the optimized autoencoder, is the learning rate for spatial projection optimization, is the gradient with respect to the weight matrix, is the projection loss function
[0067] After the encoder reduces the dimensionality of the medical image data, the decoder maps the low-dimensional features back to the high-dimensional space to generate a reconstructed version of the input. The way the decoder reconstructs the features is expressed as
[0068]
[0069] In the formula, is the reconstructed medical image data output by the decoder, that is, the medical image data enhanced by the autoencoder, is the activation function in the decoder, specifically the Sigmoid activation function, is the decoder weight, which is the decoder part of the weight matrix of the optimized autoencoder, is the bias of the decoder;
[0070] Calculate the reconstruction error and update the autoencoder weights based on the reconstruction error. The autoencoder weight update method is expressed as:
[0071]
[0072] In the formula, is the autoencoder weight updated by the back-propagation algorithm, is the learning rate for updating the autoencoder weights via the back-propagation algorithm, is the total loss function of the autoencoder, is the gradient of the total loss function of the autoencoder with respect to the weights, where the calculation method of the total loss function of the encoder is expressed as:
[0073]
[0074] In the formula, is the total loss function of the autoencoder, is the local weighted regularization loss function;
[0075] Repeat the above steps until the preset stop iteration condition is met.
[0076] Wherein, the medical image data is enhanced by an autoencoder based on feature sparse mapping, and the steps further include:
[0077] The input medical image data is mapped to a lower-dimensional sparse representation space through the encoder part. The encoding process of the encoder is expressed as
[0078]
[0079] In the formula, represents the low-dimensional representation of the encoder output, is the activation function in the encoder, specifically the self-adaptively adjusted LeakyReLU activation function, is the weight of the encoder, which is the encoding part of the weight of the autoencoder, is the encoder offset.
[0080] Wherein, the medical image data is enhanced by an autoencoder based on feature sparse mapping, and the steps further include:
[0081] The projection loss function is calculated using an adaptive projection moment based on feature distribution, which is expressed as:
[0082]
[0083] In the formula, It is the sum of squares of each feature calculated by the optimized autoencoder weight matrix, indicating the strength of the feature. is a small constant used to avoid division by zero errors.
[0084] A medical image enhancement processing system based on deep learning, comprising an image enhancement module and a federated learning module, wherein the federated learning module is connected to the image enhancement module;
[0085] The image enhancement module is used to enhance the medical image;
[0086] The federated learning module is used to update the model parameters between multiple clients without exchanging original medical image data.
[0087] The invention discloses a medical image enhancement processing method and system based on deep learning, which adopts a generative adversarial network based on local feature enhancement. Compared with the traditional generative adversarial network model, by adopting multi-layer nonlinear transformation and local feature enhancement mechanism, the generator not only relies on the noise vector, but also can fuse the historical generation state and local features of the image, significantly improves the details and precision of the generated image, and improves the quality of the generated image, especially in the detail part, avoiding the blurring and detail loss problems that are easy to occur when the traditional generative adversarial network generates images. When generating images, the noise vector is expanded from the low-dimensional space to the high-dimensional space through the diffusion process, thereby increasing the diversity of the generated images; at the same time, the local feature approximation of the image is performed by using a nonlinear activation function, so that the generated images are richer and more diverse, effectively improving the stability and quality of image generation, and reducing the gradient explosion or disappearance problems that may occur during the generation process. The federated learning architecture is adopted, so that each client only exchanges model parameters instead of original data, thereby realizing data privacy protection. During the training process, each client uses local data for training, and the updated model parameters are transmitted to the central model for aggregation. Without sharing medical data, the decentralized data can be fully utilized for model training, ensuring data privacy while improving the performance of the model. The autoencoder adopts a strategy based on feature sparse mapping. Through sparse coding and adaptive LeakyReLU activation function, it can effectively remove redundant information and retain key features in medical images, thereby enhancing image quality. Sparse coding retains the high-dimensional information of the image and reduces redundancy. At the same time, the adaptively adjusted activation function enhances the network's ability to express complex data patterns. Using a generative adversarial network, a variety of medical image samples can be generated to improve the diversity of samples. Through sparse coding and feature mapping, the autoencoder can extract key features and generate new samples to increase the diversity of data. Through the federated learning architecture, model training can be performed on multiple clients without exchanging original medical image data. This method can utilize decentralized data sources to enhance the diversity of model training samples while protecting data privacy. Through the above methods, the diversity of samples is effectively increased, thereby improving the generalization ability of the model, making it more robust and accurate in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0088] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art are briefly introduced below.
[0089] Figure 1 It is a flowchart of the medical image enhancement processing method based on deep learning according to the first embodiment of the present invention.
[0090] Figure 2 Schematic diagram of the federated learning architecture of the first embodiment of the present invention.
[0091] Figure 3 This is a step diagram of the medical image enhancement processing method based on deep learning according to the first embodiment of the present invention.
[0092] Figure 4 It is a principle block diagram of a medical image enhancement processing system based on deep learning according to the second embodiment of the present invention.
[0093] In the figure: 201-image enhancement module, 202-federated learning module DETAILED DESCRIPTION
[0094] Embodiments of the present invention are described in detail below. Examples of the embodiments are shown in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, but should not be construed as limiting the present invention.
[0095] The first embodiment of the present application is:
[0096] See also Figures 1 to 3 ,in, Figure 1 It is a flowchart of the medical image enhancement processing method based on deep learning according to the first embodiment of the present invention. Figure 2 Schematic diagram of the federated learning architecture of the first embodiment of the present invention. Figure 3 This is a step diagram of the medical image enhancement processing method based on deep learning according to the first embodiment of the present invention.
[0097] The present invention provides a medical image enhancement processing method based on deep learning, comprising the following steps:
[0098] S101: Select an image enhancement model based on whether the number of samples needs to be expanded;
[0099] S102: If the number of samples needs to be expanded, a generative adversarial network is used for image enhancement, otherwise an autoencoder is used for image enhancement;
[0100] Specifically, medical images are enhanced through a deep learning model; medical image enhancement methods include enhancing by using a generative adversarial network to generate samples and enhancing by using an autoencoder to reconstruct samples; in the medical image enhancement task, the image enhancement model is selected based on whether the number of samples needs to be expanded. Specifically, if the number of samples needs to be expanded, a generative adversarial network is used for image enhancement, otherwise, an autoencoder is selected for image enhancement.
[0101] Furthermore, each client’s model uses generative adversarial networks or autoencoders for medical image enhancement;
[0102] For the generative adversarial network algorithm, the present invention adopts a generative adversarial network based on local feature enhancement for training. On the basis of the traditional generative adversarial network, by adopting multi-layer nonlinear transformation and local feature enhancement mechanism, the generator not only relies on the noise vector, but also integrates the historical generation state and local features of the image, thereby improving the details and accuracy of the generated image, making the generated image higher in quality and richer in details, making the training process more stable, and effectively enhancing the effect of adversarial training.
[0103] Specifically, the training process of the generative adversarial network algorithm based on local feature enhancement is as follows:
[0104] 1. Initialize the generator and discriminator of the generative adversarial network. The generator adopts a convolutional neural network-based architecture, and the discriminator also adopts a convolutional neural network architecture to identify the difference between the generated image and the real image. For the input noise of the generator, after passing through the generator network, the generator outputs an image, which is expressed as:
[0105]
[0106] In the formula, is the random input noise of the generator, is the weight of the generator model, Medical images generated for the generator; is a generator function.
[0107] In order to enhance the details and diversity of the generated images, multi-layer nonlinear transformation and local feature enhancement mechanism are added to the generator network. It is assumed that the nonlinear transformation operation of the generator depends not only on the input noise, but also on the historical generation state of the image to achieve more complex image generation, which can be expressed as:
[0108]
[0109] In the formula, For the generator The weight matrix of the layer; is the bias term of the generator; For the generator Local features of the generated image of the layer; is the first nonlinear function, specifically Sigmoid; is the second nonlinear function; is the number of layers of the generator.
[0110] Furthermore, the second nonlinear function is a multi-level local feature mapping performed by combining a weighted ReLU activation function with a convolution kernel, and the calculation method is expressed as:
[0111]
[0112] In the formula, ReLU is the activation function used to activate local features and enhance the accuracy of image generation in the local context; is the adjustment factor of the generator. Preferably, Set to 0.2.
[0113] 2. In each round of training, the discriminator is first trained using real medical images. The goal of the discriminator is to effectively distinguish between real images and generated images.
[0114] Next, the random noise vector is fed into the generator, which generates an image using the noise, and the generated image is fed into the discriminator for evaluation;
[0115] The goal of the generator is to maximize the judgment error of the discriminator, that is, to make the discriminator unable to distinguish between real and fake images;
[0116] The discriminator and generator are trained alternately to continuously improve their respective capabilities in adversarial learning;
[0117] During the training process, the goal of the discriminator is to maximize its classification accuracy for real images and generated images. The loss function of the discriminator is calculated based on adversarial loss. In order to improve the discriminator's ability to judge generated images and real images, the present invention adopts an adaptive adjustment mechanism based on the gradient domain. The adaptive adjustment mechanism based on the gradient domain enhances the sensitivity of the discriminator to detail features by adopting a gradient domain regularization term, thereby improving the quality of the generated image, which is expressed as:
[0118]
[0119] In the formula, It is a real medical image; is the weight of the discriminator; is the discriminator function; The distribution of real medical image data; is the distribution of noise vector; is the loss function of the discriminator; express expectations; Indicates compliance with a specific distribution; The distribution of medical image data generated; is the L2 norm; is the gradient of the discriminator with respect to the generated image.
[0120] 3. In the generation process of the traditional generative adversarial network, the generator generates images through a single noise vector. The present invention makes the images generated by the generator more diverse through the diffusion process and nonlinear approximation method;
[0121] The diffusion process expands the noise vector of the generator from a low-dimensional space to a high-dimensional space to increase the diversity of sample generation. Specifically, the generator uses the diffusion process to iterate the generated images in multiple steps during the training process to improve the diversity and details of the generated images. The diffusion process is expressed as:
[0122]
[0123] In the formula, For the The image after diffusion of iterative times, For the Iteration of diffusion process.
[0124] Furthermore, the diffusion process enhances the details of the image by gradually optimizing the output image of the generator. Each step of the diffusion process is an iterative operation that depends on the local gradient information of the image. Through the nonlinear approximation method, the local features of the image are approximated based on the nonlinear activation function, which is expressed as:
[0125]
[0126] In the formula, is the weight of the diffusion layer, is the bias term of the diffusion layer; is the third nonlinear function, specifically the Sigmoid activation function; For the The image after diffusion of iterations; It is the fourth nonlinear function, specifically the Swish activation function.
[0127] 4. During the training process, in order to avoid the instability or gradient disappearance of the generator during gradient descent, a gradient domain adaptive adjustment strategy is adopted. According to the gradient information of the image, the gradient propagation method in the network is adjusted, so that the generator can be adaptively adjusted according to the high-frequency features of the image when updating parameters, ensuring that the generator maintains a high resolution when generating details and avoiding blur or distortion of the generated image. Specifically, the gradient adjustment item is used and dynamically adjusted by adjusting the gradients of the generator and the discriminator. The gradient update rules of the generator and the discriminator are expressed as follows:
[0128]
[0129]
[0130] In the formula, is the gradient adjustment term of the generator; is the gradient adjustment term of the discriminator; The loss function of the generator is specifically the inverse of the loss function of the discriminator; is the loss function of the discriminator; Adaptive gradient adjustment function for the generator; is the adaptive gradient adjustment function of the discriminator; represents the Hadamard product; is the gradient of the generated image.
[0131] Furthermore, the adaptive gradient adjustment function of the generator and discriminator is used to automatically adjust the gradient size according to the local details of the image, so that the generator and discriminator of the generative adversarial network can converge stably during the training process and effectively avoid the problem of gradient explosion or gradient disappearance. The calculation method is expressed as:
[0132]
[0133]
[0134] in, To control the intensity parameter of the generator gradient adjustment, is a strength parameter for controlling the discriminator gradient adjustment. Preferably, Set to 2, Set to 3.
[0135] 5. In the process of generating the generator, the local feature enhancement loss function is used to dynamically enhance the feature expression of the key information area in the image, thereby improving the detail retention and structural authenticity of the generated image. Specifically, by performing local area detection on the generated image, the image is divided into multiple local areas, and the importance of each area is measured by an adaptive weight to represent the contribution of the area to the overall image quality. The calculation method of the local feature enhancement loss function is expressed as:
[0136]
[0137] In the formula, Enhance the loss function for local features; To generate the image in Representation on a local area; For the real image Representation on a local area; For the The weight of a region measures the importance of the region to the overall image quality.
[0138] Furthermore, the weight of the region is calculated based on the local gradient and high-frequency features of the image, and is adaptively adjusted according to the size of the gradient and the high-frequency part of the regional features to ensure that the weight of the important area is large, which is expressed as:
[0139]
[0140] In the formula, Indicates area Generating images The gradient on Parameters to control the sensitivity of local region importance; Indicates Preferably, Set to 2.
[0141] 6. During the training process, the parameters of the generator and the discriminator are optimized through continuous iteration until the generator can produce high-quality images and the discriminator's ability to distinguish between true and false images is optimal. In each round of training, the generator and the discriminator are optimized alternately, and the weights are updated through iteration. The update method is expressed as:
[0142]
[0143]
[0144] In the formula, For the The learning rate of the generator for this iteration; For the The learning rate of the discriminator at the iteration, is the gradient of the generator’s loss function with respect to the weight parameters; is the gradient of the discriminator's loss function with respect to the weight parameters; It is the parameter update operation; Enhance the loss function for local features; Enhance the gradient of the loss function with respect to the generator weight parameters for local features; The gradient of the loss function with respect to the discriminator weight parameters is enhanced for local features.
[0145] Furthermore, an adaptive dynamic learning rate optimization strategy is used to enhance the accuracy and efficiency of gradient updates during training, thereby improving the quality of generated images and the discriminant ability of the discriminator. By dynamically adjusting the learning rate during training, the parameters of the generator and discriminator are finely adjusted according to the gradient size at different stages and the complexity of the medical image data features, avoiding the problem of overly fast convergence or local optimality that may be caused by conventional learning rate strategies. The calculation method is expressed as:
[0146]
[0147]
[0148] In the formula, is the initial learning rate of the generator; is the initial learning rate of the discriminator; It is a dynamic learning rate adjustment factor, which controls the decay rate of the learning rate; is the learning rate regularization constant, which prevents the learning rate from being too large in the case of extremely small gradients; is the power exponent of the adjustment factor, which is used to control the nonlinear characteristics of the learning rate decay. Preferably, Set to 2, Set to 0.1, Set to 3.
[0149] After the generative adversarial network training is completed, the generator can generate a variety of high-quality medical image samples by inputting noise vectors, providing the model with richer and higher-quality training medical image data.
[0150] Each client’s model can also use autoencoders for medical image enhancement;
[0151] The present invention adopts an autoencoder based on feature sparse mapping to perform medical image data enhancement. Based on sparse coding and adaptive LeakyReLU activation function, the autoencoder can more effectively capture the key features in medical images while reducing redundant information, thereby improving the expressive power of the model.
[0152] Specifically, the training process of the autoencoder based on feature sparse mapping is as follows:
[0153] 1. Initialize the weights and biases in the autoencoder network. The initialization method is expressed as:
[0154]
[0155] In the formula, represents the initial weight of the autoencoder; is the initialization coefficient of the autoencoder, used to scale the weights; is the dimension of the autoencoder input feature; represents a normal distribution with mean 0 and variance 1.
[0156] Furthermore, the bias term of the autoencoder is initialized to zero, expressed as:
[0157]
[0158] In the formula, Represents the initial bias of the autoencoder.
[0159] Furthermore, the initialization coefficient of the autoencoder is calculated by the dimension of the input medical image data and the distribution characteristics of the medical image data, expressed as:
[0160]
[0161] In the formula, is the initialization scaling factor; is the number of features of the input medical image data; is the number of features of the output medical image data of the encoder in the autoencoder.
[0162] 2. During the training process of the autoencoder, the input high-dimensional features are passed into the network. The high-dimensional features are pixel vectors of medical images. Since the input feature space may contain noise or redundancy, the input medical image data is first normalized, which is expressed as:
[0163]
[0164] In the formula, represents the normalized input medical image data; is the original input medical image data; is the mean of the input medical image data; is the standard deviation of the input medical image data.
[0165] 3. The input high-dimensional features first pass through the encoder part. The encoder part maps the input medical image data to a lower-dimensional sparse representation space through sparse coding. The sparse coding method is to impose regularization constraints so that each input has fewer non-zero elements in the low-dimensional space, thereby retaining the important features of the medical image data and compressing redundant information. The encoding process of the encoder is expressed as:
[0166]
[0167] In the formula, Represents the low-dimensional representation of the encoder's output; It is the activation function in the encoder, specifically the self-adaptively adjusted LeakyReLU activation function; is the weight of the encoder, which is the encoding part of the weight of the autoencoder; is the encoder offset.
[0168] Furthermore, the sparsity loss function is calculated during the encoding process. The purpose is to ensure that the output low-dimensional features have higher expressive power by sparsifying the weight matrix, which is expressed as:
[0169]
[0170] In the formula, is the sparsity loss function; is the weight of the encoder’s sparse regularization; Represents the L1 norm, which controls the sparsity of features; is the L2 norm; is the encoder weight Column eigenvalues, is the number of layers of the encoder. Preferably, Set to 0.1.
[0171] Furthermore, in order to better adapt to the characteristics of medical images, especially in the case of highly nonlinear and noisy medical images, the adaptively adjusted LeakyReLU activation function uses a term that controls the degree of nonlinearity to enhance the network's ability to express complex medical image data features. The calculation method is expressed as:
[0172]
[0173] In the formula, Represents the activation function output in the encoder; Represents the independent variable of the function, corresponding to the encoding process ; It is a parameter that controls the shape of the activation function; is the maximum value function; is a parameter that controls the degree of nonlinearity of the activation function. Preferably, Set to 2, Set to 0.5.
[0174] 4. Perform sample space projection optimization by calculating the sample projection error and adjusting the encoder mapping weight to optimize the feature representation of the space after dimensionality reduction. Specifically, after each sample is transformed into a new low-dimensional space through projection, the projection error will be fed back to the network to drive the adjustment of the projection weight to ensure that the space after dimensionality reduction can better maintain the structure and relevance of medical image data, which can be expressed as:
[0175]
[0176] In the formula, Represents the weight matrix of the optimized autoencoder; is the learning rate of spatial projection optimization; is the gradient with respect to the weight matrix; is the projection loss function. Preferably, Set to 0.01.
[0177] Furthermore, the calculation method of the projection loss function is expressed as:
[0178]
[0179] In the formula, is the projection matrix, which is used to project the input samples into a lower-dimensional space.
[0180] It should be noted that the traditional projection matrix is usually a simple linear matrix. The present invention uses an adaptive projection moment based on feature distribution to calculate the projection loss function. By adaptively adjusting the projection weight of each feature, the reduced-dimensional feature can better reflect the essential structure of the input medical image data, which is expressed as:
[0181]
[0182] In the formula, It is the sum of squares of each feature calculated by the optimized autoencoder weight matrix, indicating the strength of the feature; is a small constant used to avoid division by zero errors. Preferably, Set to 0.01.
[0183] 5. After the encoder reduces the dimensionality of the medical image data, the decoder is responsible for mapping the low-dimensional features back to the high-dimensional space to generate a reconstructed version of the input. The purpose of the decoder is to reconstruct the input medical image data as accurately as possible and minimize information loss. The way the decoder reconstructs features is expressed as:
[0184]
[0185] In the formula, is the reconstructed medical image data output by the decoder, that is, the medical image data enhanced by the autoencoder; It is the activation function in the decoder, specifically the Sigmoid activation function; is the decoder weight, which is the decoder part of the weight matrix of the optimized autoencoder; is the bias of the decoder.
[0186] Furthermore, the reconstruction error is calculated based on the reconstructed medical image data output by the decoder, which is expressed as:
[0187]
[0188] In the formula, It is the number of samples input into the self-compiled selling account in the current batch.
[0189] 6. Update the weights of the autoencoder through the back-propagation algorithm. Update the weights and biases in the network through the gradient descent method combined with the loss function of the autoencoder. The weight update method of the autoencoder is expressed as:
[0190]
[0191] In the formula, is the autoencoder weight updated by the back-propagation algorithm; is the learning rate for updating the autoencoder weights via the back-propagation algorithm; is the total loss function of the autoencoder; is the gradient of the total loss function of the autoencoder with respect to the weights;
[0192] Furthermore, the calculation method of the total loss function of the autoencoder is expressed as:
[0193]
[0194] In the formula, is the total loss function of the autoencoder; is the local weighted regularization loss function.
[0195] Furthermore, the goal of the local weighted regularization loss function is to measure the similarity between each pixel or feature in the image, and to obtain local similarity information through the image neighborhood similarity calculation method. The similarity calculation method between features is expressed as:
[0196]
[0197] In the formula, Represents medical image data and The similarity between the features; is the first reconstructed medical image data output by the decoder eigenvalues; is the first reconstructed medical image data output by the decoder eigenvalues; is the first reconstructed medical image data output by the decoder eigenvalues; is a bandwidth parameter that controls the similarity metric. Preferably, Set to 2.
[0198] Furthermore, a weighted regularization term is constructed through local similarity measurement, which aims to adjust the importance of each feature in training according to the similarity of the local area, forcing similar features to have similar weights, thereby improving the robustness and consistency of local area features. The calculation method of the local weighted regularization loss function is expressed as:
[0199]
[0200] In the formula, represents the weight matrix of the optimized autoencoder elements; represents the weight matrix of the optimized autoencoder elements.
[0201] 7. Repeat the above steps until the preset stop iteration condition is met, indicating that the model training is completed. In one embodiment, the preset stop iteration condition is reaching a preset maximum number of iterations. Preferably, the preset maximum number of iterations is set to 1000 times.
[0202] S103: A federated learning architecture is used to train the model, where the model is trained on multiple distributed devices or servers, where the devices or servers hold local medical image data samples without exchanging the medical image data.
[0203] Specifically, the model training architecture proposed in the present invention adopts a federated learning architecture, and the model is trained on multiple decentralized devices or servers, and the devices or servers hold local medical image data samples without exchanging these medical image data; Specifically, the federated learning architecture proposed in the present invention has multiple clients, labeled as client 0, client 1, client 2 and client 3, etc., these clients represent independent nodes in the federated learning network, and each client has its own local medical image data;
[0204] Moreover, each client has a local model, which is trained on the medical image data of each client. The local model interacts with the central model, that is, the update flow of model parameters. represents the model parameters updated after local training, represents the parameters received from the central model;
[0205] Furthermore, the central model aggregates updates from local models, and the update of the global model is a way for the federated learning process to aggregate individual updates to create a new improved global model. After the global model update is completed, the final trained model is obtained.
[0206] In one embodiment, the model parameters arrive exchanging between the local model and the central model in a manner that allows the central model to aggregate updates and for the local model to receive new, aggregated parameters;
[0207] Further, through multiple iterations, the local model is trained, the updates are sent to the central model, a global update is performed, and then the updated parameters are sent back to the local model;
[0208] Based on this federated learning architecture, data privacy protection can be achieved. During the training process, the original medical image data is not shared between clients or with the central server, but only the model parameters are exchanged for updating. That is, the parameters are synchronized to each client through the central model. Assign a value to each client.
[0209] Furthermore, the medical image data of each client is derived from a medical image database, which contains different types of medical image data, including X-ray, CT image, magnetic resonance imaging and other image data, and the medical image data set used covers cases of different ages, genders, disease types and stages of disease, including cardiovascular and cerebrovascular diseases, tumor diseases, etc.;
[0210] The pixels of medical images are 1024×1024, and the depth of each pixel is usually 16 bits, that is, the grayscale value of each pixel ranges from 0 to 65535;
[0211] Medical images are stored using the standard medical image format DICOM (Digital Imaging and Communications in Medicine).
[0212] The generative adversarial network based on local feature enhancement is adopted. Compared with the traditional generative adversarial network model, by adopting multi-layer nonlinear transformation and local feature enhancement mechanism, the generator not only relies on the noise vector, but also can integrate the historical generation state and local features of the image, significantly improving the details and precision of the generated image, and improving the quality of the generated image, especially in the details, avoiding the blurring and detail loss problems that are easy to occur when the traditional generative adversarial network generates images. When generating images, the noise vector is expanded from low-dimensional space to high-dimensional space through the diffusion process, thereby increasing the diversity of generated images; at the same time, the local feature approximation of the image is adopted by nonlinear activation function, making the generated images richer and more diverse, effectively improving the stability and quality of image generation, and reducing the gradient explosion or disappearance problems that may occur during the generation process. The federated learning architecture is adopted so that each client only exchanges model parameters instead of original data, thereby realizing data privacy protection. During the training process, each client uses local data for training, and the updated model parameters are transmitted to the central model for aggregation. Without sharing medical data, the decentralized data can be fully utilized for model training, ensuring data privacy while improving the performance of the model. The autoencoder adopts a strategy based on feature sparse mapping. Through sparse coding and adaptive LeakyReLU activation function, it can effectively remove redundant information and retain key features in medical images, thereby enhancing image quality. Sparse coding retains the high-dimensional information of the image and reduces redundancy. At the same time, the adaptively adjusted activation function enhances the network's ability to express complex data patterns. Using a generative adversarial network, a variety of medical image samples can be generated to improve the diversity of samples. Through sparse coding and feature mapping, the autoencoder can extract key features and generate new samples to increase the diversity of data. Through the federated learning architecture, model training can be performed on multiple clients without exchanging original medical image data. This method can utilize decentralized data sources to enhance the diversity of model training samples while protecting data privacy. Through the above methods, the diversity of samples is effectively increased, thereby improving the generalization ability of the model, making it more robust and accurate in practical applications.
[0213] The second embodiment of the present application is:
[0214] Based on the first embodiment, please refer to Figure 4 ,in, Figure 4 It is a principle block diagram of a medical image enhancement processing system based on deep learning according to the second embodiment of the present invention.
[0215] The deep learning-based medical image enhancement processing system of this embodiment includes an image enhancement module 201 and a federated learning module 202.
[0216] For this specific implementation, the federated learning module 202 is connected to the image enhancement module 201;
[0217] The image enhancement module 201 is used to enhance the medical image;
[0218] The federated learning module 202 is used to update the model parameters between multiple clients without exchanging original medical image data.
[0219] The image enhancement module 201 sends the enhanced medical image data to the federated learning module 202 as input data for training the federated learning module 202; the federated learning module 202 performs model training locally based on the enhanced medical image data received from the image enhancement module 201, and generates model parameter updates; the federated learning module 202 sends the model parameter updates to a central server or other clients to aggregate the updates and generate global model parameters; the aggregated global model parameters are sent back by the federated learning module 202 to the image enhancement module 201 of each client to update and optimize the model parameters of the image enhancement module 201, thereby improving the accuracy and efficiency of image enhancement.
[0220] A deep learning-based medical image enhancement processing system is used in this embodiment, and the image enhancement module 201 is connected to the federated learning module 202, so that the enhanced medical images generated by the image enhancement module 201 on each client are used to train and update the model parameters in the federated learning module 202. At the same time, the federated learning module 202 feeds back the updated model parameters to the image enhancement module 201 to improve the image enhancement effect.
[0221] What is disclosed above is only one or more preferred embodiments of the present application, and cannot be used to limit the scope of rights of the present application. Ordinary technicians in this field can understand that all or part of the processes of implementing the above embodiments and equivalent changes made according to the claims of the present application are still within the scope covered by the present application.
Claims
1. A medical image enhancement processing method based on deep learning, characterized in that: The following steps are involved: Select the image enhancement model based on whether the number of samples needs to be expanded; If the number of samples needs to be expanded, a generative adversarial network is used for image enhancement, otherwise an autoencoder is used for image enhancement; The generative adversarial network is trained by using a generative adversarial network based on local feature enhancement. On the basis of the traditional generative adversarial network, by adopting multi-layer nonlinear transformation and local feature enhancement mechanism, the generator not only relies on the noise vector, but also integrates the historical generation state and local features of the image. A federated learning architecture is used to train the model, where the model is trained on multiple distributed devices or servers, where the devices or servers hold local medical image data samples without exchanging these medical image data; The generative adversarial network based on local feature enhancement is trained, and the step further includes: Initialize the generator and discriminator; Multi-layer nonlinear transformation and local feature enhancement mechanism are added to the generator network. The nonlinear transformation operation of the generator depends on the input noise and the historical generation state of the image to achieve more complex image generation, which can be expressed as: In the formula, For the generator The weight matrix of the layer; is the bias term of the generator; For the generator Local features of the generated image of the layer; is the first nonlinear function, specifically Sigmoid, is the second nonlinear function, is the number of layers of the generator; In each round of training, the discriminator is first trained using real medical images as input, and then a random noise vector is input to the generator to generate images, and the generated images are sent to the discriminator for evaluation. The goal of the discriminator is to distinguish between real images and generated images, while the goal of the generator is to maximize the judgment error of the discriminator. The discriminator and the generator are trained alternately. The loss function of the discriminator is calculated based on the adversarial loss, and an adaptive adjustment mechanism based on the gradient domain is used to enhance the sensitivity of the discriminator to detail features, which can be expressed as: In the formula, For real medical images, is the weight of the discriminator, is the discriminator function, is the distribution of real medical image data, is the distribution of the noise vector, is the loss function of the discriminator, Express expectations, It means that it obeys a specific distribution. is the distribution of generated medical image data, is the L2 norm, is the gradient of the discriminator to the generated image; Medical images generated for the generator; The diversity of images generated by the generator is increased through the diffusion process and nonlinear approximation method. The diffusion process expands the noise vector of the generator from the low-dimensional space to the high-dimensional space to increase the diversity of sample generation, and enhances the details of the image by gradually optimizing the output image of the generator. The diffusion process is expressed as: In the formula, For the The image after diffusion of iterative times, For the The diffusion process of times iteration; During the training process, the gradient domain adaptive adjustment strategy is adopted to adjust the gradient propagation mode in the network according to the gradient information of the image, and automatically adjust the gradient size to stabilize the convergence. The gradient update rules of the generator and the discriminator are expressed as follows: , In the formula, is the gradient adjustment term of the generator, is the gradient adjustment term of the discriminator, is the loss function of the generator, is the loss function of the discriminator, is the adaptive gradient adjustment function of the generator, is the adaptive gradient adjustment function of the discriminator, represents the Hadamard product, is the gradient of the generated image; In the process of generating the generator, the local feature enhancement loss function is used to dynamically enhance the feature expression of the key information area in the image. The calculation method of the local feature enhancement loss function is expressed as: In the formula, Enhance the loss function for local features, To generate the image in The representation on a local area, For the real image The representation on a local area, For the The weight of a region measures the importance of the region to the overall quality of the image. The weight of the region is calculated based on the local gradient and high-frequency features of the image, and is adaptively adjusted according to the size of the gradient and the high-frequency part of the regional features, expressed as: In the formula, Indicates area Generating images The gradient on To control the sensitivity of local region importance, Indicates A local area range; By continuously iteratively optimizing the parameters of the generator and the discriminator until the generator can produce high-quality images and the discriminator's ability to distinguish between true and false images is optimal, the update method is expressed as: In the formula, For the The learning rate of the generator for the iteration, For the The learning rate of the discriminator at the iteration, is the gradient of the generator’s loss function with respect to the weight parameters, is the gradient of the discriminator’s loss function with respect to the weight parameters, is the parameter update operation, Enhance the loss function for local features, The gradient of the loss function for local feature enhancement with respect to the generator weight parameters, The gradient of the loss function with respect to the discriminator weight parameters is enhanced for local features, and an adaptive dynamic learning rate optimization strategy is used to enhance the accuracy and efficiency of gradient updates during training. The calculation method is expressed as: In the formula, is the initial learning rate of the generator, is the initial learning rate of the discriminator, is the dynamic learning rate adjustment factor, which controls the decay rate of the learning rate. is the learning rate regularization constant, which prevents the learning rate from being too large in the case of extremely small gradients. It is the power exponent of the adjustment factor, which is used to control the nonlinear characteristics of the learning rate decay.
2. The medical image enhancement processing method based on deep learning according to claim 1, characterized in that: Training is performed on a generative adversarial network based on local feature enhancement, the steps further comprising: Both the generator and the discriminator use a convolutional neural network architecture to identify the difference between generated images and real images. The generator generates medical images by inputting noise and model weights, which is expressed as: In the formula, is the random input noise of the generator, is the weight of the generator model, Medical images generated by the generator, is a generator function.
3. The medical image enhancement processing method based on deep learning according to claim 2, characterized in that: Training is performed on a generative adversarial network based on local feature enhancement, the steps further comprising: The second nonlinear function is a multi-level local feature mapping performed by combining the weighted ReLU activation function with the convolution kernel. The calculation method is expressed as: In the formula, is the ReLU activation function, which is used to activate local features and enhance the accuracy of image generation in the local context. is the adjustment factor of the generator.
4. The medical image enhancement processing method based on deep learning according to claim 1, characterized in that: If the number of samples needs to be expanded, a generative adversarial network is used for image enhancement, otherwise an autoencoder is used for image enhancement, and the steps further include: The autoencoder uses a feature sparse mapping-based autoencoder for medical image data enhancement based on sparse coding and adaptive LeakyReLU activation function.
5. The medical image enhancement processing method based on deep learning according to claim 4, characterized in that: Performing medical image data enhancement based on an autoencoder for feature sparse mapping, the steps further comprising: Initialize the weights and biases in the autoencoder network. The initialization method is expressed as: In the formula, represents the initial weight of the autoencoder, is the initialization coefficient of the autoencoder, used to scale the weights, is the dimension of the autoencoder input feature, represents a normal distribution with a mean of 0 and a variance of 1, where the bias term of the autoencoder is initialized to zero, expressed as: In the formula, Represents the initial bias of the autoencoder. The initialization coefficient of the autoencoder is calculated by the dimension of the input medical image data and the distribution characteristics of the medical image data, which is expressed as: In the formula, is the initial scaling factor, is the number of features of the input medical image data, is the feature number of the output medical image data of the encoder in the autoencoder; the input high-dimensional features are normalized and expressed as: In the formula, represents the normalized input medical image data, is the original input medical image data, is the mean of the input medical image data, is the standard deviation of the input medical image data; The input medical image data is mapped into a low-dimensional sparse representation space through the encoder part; The sparsity loss function is calculated during the encoding process and is expressed as: In the formula, is the sparsity loss function, is the weight of the encoder’s sparse regularization, represents the L1 norm, controlling the sparsity of features, is the L2 norm, is the encoder weight Column eigenvalues, is the number of encoder layers; Perform sample space projection optimization by calculating the projection error of the sample and adjusting the mapping weight of the encoder to optimize the feature representation of the space after dimensionality reduction, which is expressed as: In the formula, represents the weight matrix of the optimized autoencoder, is the learning rate for spatial projection optimization, is the gradient of the weight matrix, is the projection loss function; After the encoder reduces the dimensionality of the medical image data, the decoder maps the low-dimensional features back to the high-dimensional space to generate a reconstructed version of the input. The way the decoder reconstructs the features is expressed as In the formula, is the reconstructed medical image data output by the decoder, that is, the medical image data enhanced by the autoencoder, is the activation function in the decoder, specifically the Sigmoid activation function, is the decoder weight, which is the decoder part of the weight matrix of the optimized autoencoder, is the bias of the decoder; Calculate the reconstruction error and update the autoencoder weights based on the reconstruction error. The autoencoder weight update method is expressed as: In the formula, is the autoencoder weight updated by the back-propagation algorithm, is the learning rate for updating the autoencoder weights via the back-propagation algorithm, is the total loss function of the autoencoder, is the gradient of the total loss function of the autoencoder with respect to the weights, where the calculation method of the total loss function of the encoder is expressed as: In the formula, is the total loss function of the autoencoder, is the local weighted regularization loss function; Repeat the above steps until the preset stop iteration condition is met.
6. The medical image enhancement processing method based on deep learning according to claim 4, characterized in that: Performing medical image data enhancement based on an autoencoder for feature sparse mapping, the steps further comprising: The input medical image data is mapped into a low-dimensional sparse representation space through the encoder part. The encoding process of the encoder is expressed as In the formula, represents the low-dimensional representation of the encoder output, is the activation function in the encoder, specifically the self-adaptively adjusted LeakyReLU activation function, is the weight of the encoder, which is the encoding part of the weight of the autoencoder, is the encoder offset.
7. A medical image enhancement processing system based on deep learning, applicable to the medical image enhancement processing method based on deep learning according to any one of claims 1 to 6, characterized in that: It includes an image enhancement module and a federated learning module, wherein the federated learning module is connected to the image enhancement module; The image enhancement module is used to enhance the medical image; The federated learning module is used to update the model parameters between multiple clients without exchanging original medical image data.
Citation Information
Patent Citations
Medical image processing method, computer equipment and storage medium
CN119006382A
Federal learning-based data enhancement method and system
CN118821869A
MRI image enhancement system and method based on double-domain MTGAN
CN118864279A