Multi-Stage Progressive Super-Resolution Method for Medical Images Based on Optical Amplification Principle
Through multi-stage progressive multi-attention network models, the shortcomings of the super-resolution reconstruction method of the medical image super-resolution reconstruction method in the prior art in feature extraction and multi-scale utilization are solved, and higher quality medical image generation is achieved.
Patent Information
- Application Number
- CN202411918507.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-12-25
AI Technical Summary
The existing optical amplification principle based on neural networks. The super-resolution reconstruction method of medical images has weak local and global features combined with weak spatial and channel redundancy during feature extraction. The single-stage idea cannot fully utilize multi-scale detailed information, resulting in the generated medical images not being clear enough.
A multi-stage progressive super-resolution method is designed. By constructing a multi-stage progressive multi-attention network model, a shallow feature extraction module, an initial feature extraction module and a supervisory attention module are used to gradually extract and fuse image features to generate higher resolution medical images.
Through the combination of multi-stage design and multiple attention mechanisms, deep features and multi-scale information of the image can be more effectively extracted, and a clearer and more in line with clinical needs can be generated.
Smart Images

Figure CN119359546B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image processing, and in particular, to a multi-stage progressive super-resolution method for medical images based on the principle of optical magnification. Background Art
[0002] In the current field of medical images, medical images based on the principle of optical magnification have become an indispensable tool in clinical diagnosis and treatment. For example, dermoscopes, otoscopes, and ophthalmoscopes, etc. These images are widely used in multiple fields such as dermatology, retinal diseases, otolaryngology, etc., helping doctors to observe lesions more accurately and diagnose diseases. They not only provide an intuitive observation of disease lesions but also reveal many potential pathological features. Doctors can analyze the texture, structure, and color of the lesion area in the image to further obtain rich pathological information and make more accurate judgments and diagnoses. Specifically, actinic keratosis, basal cell carcinoma, benign keratosis-like lesions (seborrheic keratosis and lichen planus-like keratosis), skin fibroma, melanoma, melanocytic nevus, and vascular lesions (hemangioma, angiokeratoma), cataract, glaucoma, diabetic retinopathy, acute otitis media, chronic suppurative otitis media, ear ventilation tube, excessive earwax, foreign body in the ear, otitis externa are all common diseases in daily life. These skin, ear, and retinal diseases are numerous, including both benign and malignant lesions, seriously affecting the appearance and quality of life of patients. Therefore, accurate diagnosis plays a crucial role in improving symptoms and preventing damage.
[0003] Medical imaging devices based on the principle of optical magnification such as dermoscopes, ophthalmoscopes, and otoscopes all play an indispensable role in the diagnosis and treatment of diseases. Generally speaking, doctors' decisions often rely on the imaging results of these medical devices. High-quality medical images can better help doctors identify the type of disease and locate the position of the disease. However, although these medical devices based on the principle of optical magnification have been widely used clinically, they still face many challenges, including device limitations, changes in lighting conditions, and patient individual differences, etc. These factors often result in lower-resolution medical images with some details blurred, thus causing certain difficulties in the diagnosis and treatment of patients' conditions. Therefore, clear and accurate medical images based on the principle of optical magnification play a crucial role in the treatment of lesions in related positions.
[0004] Super-resolution is a method for generating high-resolution images, aiming to generate high-resolution images from low-resolution images, which can effectively overcome the problem of insufficient clarity of medical image resolution. With the successful application of deep learning in super-resolution tasks, the model can be trained on a large number of datasets and finally generate high-resolution images from low-resolution inputs. Currently, a series of medical image super-resolution tasks based on the principle of optical magnification mainly use neural networks as the main technical method, and these methods can directly learn the end-to-end non-linear mapping regression between low-resolution and high-resolution images, with good representation ability.
[0005] In previous medical image super-resolution reconstruction methods based on neural networks and the principle of optical magnification, there are two obvious drawbacks. First, the ability of these methods to combine local and global features is weak, and there are phenomena of spatial and channel redundancy in the feature extraction process, resulting in an unsatisfactory final performance of the model. Second, most methods are designed with a single-stage idea and cannot make full use of detailed information at different scales, thus restricting the generation of higher-resolution medical images and unable to provide doctors with clearer medical images. Therefore, how to design a new method to overcome the drawbacks of previous methods, balance the order of feature extraction strategies, and further utilize multi-scale information to achieve better medical image reconstruction based on the principle of optical magnification remains a challenge.
[0006] Therefore, the present invention provides a multi-stage progressive super-resolution method for medical images based on the principle of optical magnification to solve the above problems. Summary of the Invention
[0007] In view of the deficiencies of the prior art, the present invention has developed a multi-stage progressive super-resolution method for medical images based on the principle of optical magnification, with the main purpose of repairing details and refining medical images, and then generating clearer medical images based on the principle of optical magnification.
[0008] The technical solution for the present invention to solve the technical problem is a multi-stage progressive super-resolution method for medical images based on the principle of optical magnification, which is specifically as follows:
[0009] S1. Collect medical images for the principle of optical magnification to form a dataset. The medical images include low-resolution images and original high-resolution images, preprocess the low-resolution medical images in the dataset to obtain a preprocessed dataset, and then divide the dataset into a training set, a validation set, and a test set;
[0010] S2. Construct a multi-stage progressive multi-attention network model, which includes three stages. The first stage and the second stage both contain a shallow feature extraction module, an initial feature extraction module, and a supervised attention module. The third stage contains a shallow feature extraction module, an initial feature extraction module, and an image reconstruction module. Output the low-resolution image to this model for feature extraction, and finally generate a high-resolution medical image based on the optical magnification principle;
[0011] S3. Input the data in the training set, validation set, and test set into the multi-stage progressive multi-attention network model in sequence. Input the data in the training set into the model and train the model in combination with the loss function. Then input the data in the validation set into the trained model and adjust the model parameters in combination with the evaluation index. Finally, input the data in the test set into the adjusted model to test the model and output the final result.
[0012] S1 is specifically as follows:
[0013] S1.1. Construct a data set:
[0014] Select medical images according to the scaling factor, and construct a data set based on the medical images , the medical images specifically include low-resolution medical images for the optical magnification principle and original high-resolution images. The data set is denoted as , where represents the th medical image, represents the th low-resolution medical image for the optical magnification principle, with a size of 64×64, represents the th original high-resolution picture for the optical magnification principle, with a size of 64S×64S, where S represents the scaling factor, represents the number of medical images in the data set, represents the index of the medical image, ;
[0015] S1.2. Data set preprocessing and partitioning:
[0016] Perform horizontal flipping and rotation operations on the low-resolution medical images in the data set for data augmentation, and then divide the preprocessed data set into a training set, a validation set, and a test set according to a ratio.
[0017] S2 is specifically as follows:
[0018] The low-resolution image goes through three stages in sequence. The low-resolution image is input into the first stage. The output of the first stage and the low-resolution image are used as the input of the second stage. Then, the output of the second stage and the low-resolution image are used as the output of the third stage. The third stage finally outputs a high-resolution medical image based on the optical magnification principle.
[0019] The process in the first stage of the multi-stage progressive multi-attention network model is as follows:
[0020] S2.1. First, the input low-resolution image is divided according to small scales, and then the division results are respectively input into the shallow feature extraction module of the first stage to obtain small-scale shallow features 、 、 and . Then, the results of the shallow feature extraction module are respectively input into the initial feature extraction module of the first stage for preliminary feature capture. The captured small-scale initial features are then stitched together in terms of image size to obtain initial features. Finally, the initial features and the low-resolution image are input into the supervised attention module of the first stage to obtain the final output of the first stage;
[0021] Among them, , represents the height, represents the width, represents the number of channels of the input low-resolution image, 、 、 、 , represents the number of channels of the features output by the shallow feature extraction module of the first stage,
[0022] (1) The small-scale division operation is as follows:
[0023] ,
[0024] Among them, represents the small-scale division operation, represents the result of the small-scale division;
[0025] (2) The shallow feature extraction module of the first stage consists of a convolution with a kernel size of 3 and a stride of 1. The calculation in the shallow feature extraction module of the first stage is as follows:
[0026] ,
[0027] Among them, represents the convolution operation;
[0028] (3) The initial feature extraction module in the first stage is essentially a variety of attention groups, including the global channel attention module, the multiple convolutional feature fusion module, and the dual attention module. The multiple convolutional feature fusion module includes a channel reconstruction unit, a spatial reconstruction unit, and a standard convolutional group. First, the small-scale shallow features are respectively input into the global channel attention module to capture pixel information. Then, the output of the global channel attention module is input into the multiple convolutional feature fusion module for local and global information fusion. Finally, the output of the multiple convolutional feature fusion module is input into the dual attention module to capture context information through different windows, obtaining the small-size initial features. ;
[0029] (4) Concatenate the small-size initial features , and then input the concatenation result and the low-resolution image together into the supervised attention module of the first stage to obtain the final output of the first stage. The specific calculation is as follows:
[0030] ,
[0031] ,
[0032] ,
[0033] ,
[0034] ,
[0035] ,
[0036] ,
[0037] where, represents the tensor concatenation method of the scale, and represent two different concatenation results, represents the set of concatenation results, represents the supervised attention module, represents the intermediate output of the supervised attention module, and represent two different outputs of the supervised attention module in the first stage.
[0038] The specific initial feature extraction module in the first stage is as follows:
[0039] (3-1) The specific calculation in the global channel attention module is as follows:
[0040] ,
[0041] ,
[0042] ,
[0043] ,
[0044] Among them, represents the intermediate output of the overall channel attention module, represents the th small-scale shallow feature, , represents the convolution operation, represents the activation function operation, represents the output of the overall channel attention module, represents the operation of represents the convolution operation, represents the operation of the channel attention module, represents the activation function operation, represents the activation function operation, represents the average pooling operation;
[0045] The calculation in the (3 - 2) multi-convolution feature fusion module is as follows:
[0046] 1) Spatial reconstruction unit:
[0047]
[0048] ,
[0049] ,
[0050] ,
[0051] ,
[0052] Among them, represents the normalization operation, represents the result of processing, and represent the trainable scaling parameter and translation parameter respectively, and represent the mean and standard deviation of represents the constant term to maintain the stability of division, represents the normalization-related weight, Represent two different weights obtained after threshold processing, Represent threshold gating, with the threshold set to 0.5, Represent And The result obtained after the tensor cutting operation after multiplication, Represent And The result obtained after the tensor cutting operation after multiplication, Represent the tensor cutting operation, Represent the tensor splicing operation, Represent the operation of the spatial reconstruction unit, Represent the output of the spatial reconstruction unit;
[0053] 2) Channel reconstruction unit:
[0054] ,
[0055] ,
[0056] ,
[0057] ,
[0058] ,
[0059] ,
[0060] Among them, Represent the result obtained after the tensor cutting operation on , Represent grouped convolution, Represent the result obtained after the grouped convolution operation on , Represent the result obtained after the tensor splicing operation on , Represent the tensor splicing operation, Represent And The result obtained after the tensor splicing operation, Represent the softmax operation, Represent the weight of the softmax operation, Represent the intermediate feature of the fusion strategy process, Represent the operation of the channel reconstruction unit, Represent the output of the channel reconstruction unit;
[0061] 3) Standard convolution group:
[0062] ,
[0063] ,
[0064] ,
[0065] ,
[0066] ,
[0067] ,
[0068] ,
[0069] ,
[0070] Among them, represents the space-channel operation, represents the channel-space operation, represents the result of the space-channel operation, represents the result of the channel-space operation, represents the result of the space-channel standard convolution, represents the result of the channel-space standard convolution, represents the input of the module, , represents the result of the n-th standard convolution, represents the total number of standard convolutions, represents the index of the number of standard convolutions, represents the sum branch result, represents the difference branch result, represents the final output feature of the multi-convolution feature fusion module;
[0071] The calculation in the (3-3) dual attention module is as follows:
[0072] ,
[0073] ,
[0074] ,
[0075] ,
[0076] ,
[0077] The internal operation of the DAB module is as follows:
[0078] ,
[0079] ,
[0080] Among them, represents the layer normalization result, represents the deformable size window self-attention mechanism, represents the double routing attention mechanism, represents the intermediate output of the double attention module, represents the operation of the multi-layer perceptron, represents the output result of the double attention module, represents the overall channel attention, represents the multiple convolution feature fusion module, represents the double attention module, represents the th multiple attention group operation performed in the first stage, represents the number of multiple attention group operations, , represents the final feature output of the continuous multiple attention group modules in the first stage.
[0081] The process in the second stage of the multi-stage progressive multiple attention network model is as follows:
[0082] S2.2. First, the input low-resolution image is divided according to the medium scale, and then the division results are respectively input into the shallow feature extraction module of the second stage to obtain the medium-scale shallow features and . Then, the results of the shallow feature extraction module and the output of the supervised attention module in the first stage are respectively input into the initial feature extraction module of the second stage for preliminary feature capture. Then, the captured medium-scale initial features are stitched together in terms of image size to obtain the initial features. Finally, the initial features and the low-resolution image are input into the supervised attention module of the second stage to obtain the final output of the second stage;
[0083] Among them, 、 , represents the number of channels of the features output by the shallow feature extraction module in the second stage;
[0084] (1) The medium-scale division operation is as follows:
[0085] ,
[0086] Among them, represents the medium-scale division operation, represents the result of the medium-scale division;
[0087] (2)The shallow feature extraction module in the second stage consists of a convolution with a kernel size of 3 and a stride of 1. The calculations in the shallow feature extraction module of the second stage are as follows:
[0088] ,
[0089] Among them, represents the convolution operation;
[0090] (3)The structure of the initial feature extraction module in the second stage is the same as that in the first stage. The medium-scale shallow features and the output of the supervised attention module in the first stage are respectively input into the initial feature extraction module in the second stage to obtain the output of the initial feature extraction module in the second stage. The specific calculations are as follows:
[0091] ,
[0092] ,
[0093] ,
[0094] ,
[0095] Among them, represents the method of tensor concatenation of channels, and represent two different results of tensor concatenation of channels, and represent two different medium-scale initial features output by the initial feature extraction module in the second stage;
[0096] (4)The medium-scale initial features and are concatenated, and then the concatenated result and the low-resolution image are jointly input into the supervised attention module in the second stage to obtain the final output in the second stage. The specific calculations are as follows:
[0097] ,
[0098] ,
[0099] Among them, represents the method of tensor concatenation of scales, represents the operation of the supervised attention module, represents the final output in the second stage.
[0100] The process in the third stage of the multi-stage progressive multi-attention network model is as follows:
[0101] S2.3. First, divide the input low-resolution image according to large scales, and then input the division results into the shallow feature extraction module of the third stage respectively to obtain large-scale shallow features . Then, input the results of the shallow feature extraction module and the output of the supervised attention module in the second stage into the initial feature extraction module of the third stage for preliminary feature capture. Next, splice the captured large-scale initial features in terms of image size to obtain initial features. Finally, input the initial features and the low-resolution image into the supervised attention module of the third stage to obtain the final output of the third stage, that is, the high-resolution medical image based on the optical magnification principle;
[0102] Among them, , represents the number of channels of the features output by the shallow feature extraction module in the third stage;
[0103] (1) The large-scale division operation is as follows:
[0104] ,
[0105] Among them, represents the large-scale division operation, represents the result of the large-scale division, that is, the input low-resolution image ;
[0106] (2) The shallow feature extraction module in the third stage consists of a convolution with a kernel size of 3 and a stride of 1. The calculation in the shallow feature extraction module in the third stage is as follows:
[0107] ,
[0108] Among them, represents the convolution operation;
[0109] (3) The structure of the initial feature extraction module in the third stage is the same as that of the initial feature extraction module in the first stage. Input the large-scale shallow features and the output of the supervised attention module in the second stage into the initial feature extraction module in the third stage respectively to obtain the output of the initial feature extraction module in the third stage. The specific calculation is as follows:
[0110] ,
[0111] ,
[0112] ,
[0113] in, The tensor concatenation method representing the channel, Represents the result of concatenating the tensors of the channels, Indicates The result of the attention group operation of the consecutive channels, Indicates Channel attention modules, represents the total number of channel attention modules, , Represents the large-scale initial features output by the initial feature extraction module in the third stage;
[0114] (4) Large-scale initial features Input the third stage image reconstruction module, select the scaling factor, and Contains multi-scale information to generate high-resolution medical images using optical magnification , where the image reconstruction module uses The method is used to reorganize pixels. The specific calculation is as follows:
[0115] .
[0116] S3 is as follows:
[0117] The loss function is calculated as follows:
[0118] ,
[0119] in, Represents the prediction results and the true value The loss function and prediction results That is, to generate high-resolution medical images based on the principle of optical magnification. represents the number of medical images in the training set, Indicates Prediction results of medical images, Indicates The true value of a medical image;
[0120] The peak signal-to-noise ratio PSNP and structural similarity SSIM are used as evaluation indicators to verify the generated optical magnification principle high-resolution medical image. The larger the PSNR and SSIM, the better the quality of the generated optical magnification principle high-resolution medical image.
[0121] The effects provided in the content of the invention are only the effects of the embodiments, rather than all the effects of the invention. The above technical solution has the following advantages or beneficial effects:
[0122] The present invention adopts a multi-stage progressive architecture design and combines multiple attention mechanisms for the reconstruction of relevant medical images. The present invention conducts continuous feature interaction under different-scale images through an attention mechanism with a variable window, so as to generate a higher-resolution medical image based on the principle of optical magnification from the input low-resolution medical image based on the principle of optical magnification under different scaling factors, thus being more in line with the diagnostic applications of clinical medicine. At the same time, the present invention proposes a new feature extraction method, a multi-convolution feature fusion module, which mainly consists of three parts: a channel reconstruction unit, a spatial reconstruction unit, and a standard convolution group. While balancing the order of spatial and channel feature extraction, it maximally improves the ability of the model to combine local detail features and global structural features, thereby generating a higher-quality medical image based on the principle of optical magnification and further improving the model performance. Therefore, the present invention overcomes the drawbacks of previous methods, can fully extract the deep features of the image, and at the same time adopts a multi-stage design concept to gradually repair details and refine the image, thereby generating a clearer medical image based on the principle of optical magnification. BRIEF DESCRIPTION OF THE DRAWINGS
[0123] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention.
[0124] Figure 1 It is a structural diagram of a multi-stage progressive multiple attention model in the present invention.
[0125] Figure 2 It is a comparative example of the processing result of skin disease images.
[0126] Figure 3 It is a comparative example of the processing result of eye disease images.
[0127] Figure 4 It is a comparative example of the processing result of ear disease images. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0128] In order to clearly illustrate the technical features of the present solution, the present invention will be described in detail below through specific embodiments and in conjunction with its drawings.
[0129] Embodiment 1
[0130] As Figure 1 shown, a multi-stage progressive super-resolution method for medical images based on the principle of optical magnification is as follows:
[0131] S1. Collect medical images for the optical magnification principle to form a dataset. The medical images include low-resolution images and original high-resolution images, and preprocess the low-resolution medical images in the dataset to obtain a preprocessed dataset. Then divide the dataset into a training set, a validation set, and a test set;
[0132] S2. Build a multi-stage progressive multi-attention network model. This model consists of three stages. The first stage and the second stage both include a shallow feature extraction module, an initial feature extraction module, and a supervised attention module. The third stage includes a shallow feature extraction module, an initial feature extraction module, and an image reconstruction module. Output the low-resolution image to this model for feature extraction, and finally generate a high-resolution medical image based on the optical magnification principle;
[0133] S3. Input the data in the training set, validation set, and test set into the multi-stage progressive multi-attention network model in turn. Input the data in the training set into the model and train the model in combination with a loss function. Then input the data in the validation set into the trained model and adjust the model parameters in combination with evaluation metrics. Finally, input the data in the test set into the adjusted model to test the model and output the final result.
[0134] S1 is specifically as follows:
[0135] S1.1. Build a dataset:
[0136] Select medical images according to the scaling factor and build a dataset according to the medical images , the medical images specifically include low-resolution medical images for the optical magnification principle and original high-resolution images. The dataset is represented as , where represents the th medical image, represents the th low-resolution medical image for the optical magnification principle, with a size of 64×64, represents the th original high-resolution picture for the optical magnification principle, with a size of 64S×64S, where S represents the scaling factor, represents the number of medical images in the dataset, represents the index of the medical image, ;
[0137] S1.2. Dataset preprocessing and division:
[0138] Perform horizontal flipping and rotation operations on the low-resolution medical images in the dataset for data augmentation, and then divide the preprocessed dataset into a training set, a validation set, and a test set according to a ratio.
[0139] S2 is as follows:
[0140] The low-resolution image goes through three stages in sequence. The low-resolution image is input into the first stage. The output of the first stage and the low-resolution image are used as the input of the second stage. Then, the output of the second stage and the low-resolution image are used as the output of the third stage. The third stage finally outputs a high-resolution medical image based on the optical magnification principle.
[0141] The process in the first stage of the multi-stage progressive multi-attention network model is as follows:
[0142] S2.1. First, the input low-resolution image is divided according to small scales, and then the division results are respectively input into the shallow feature extraction module of the first stage to obtain small-scale shallow features 、 、 and . Then, the results of the shallow feature extraction module are respectively input into the initial feature extraction module of the first stage for preliminary feature capture. The captured small-scale initial features are then stitched together in terms of image size to obtain initial features. Finally, the initial features and the low-resolution image are input into the supervised attention module of the first stage to obtain the final output of the first stage;
[0143] Among them, , represents the height, represents the width, represents the number of channels of the input low-resolution image, 、 、 、 , represents the number of channels of the features output by the shallow feature extraction module of the first stage,
[0144] (1) The small-scale division operation is as follows:
[0145] ,
[0146] Among them, represents the small-scale division operation, represents the result of the small-scale division;
[0147] (2) The shallow feature extraction module of the first stage consists of a convolution with a kernel size of 3 and a stride of 1. The calculation in the shallow feature extraction module of the first stage is as follows:
[0148] ,
[0149] Among them, Indicates a convolution operation;
[0150] (3) The initial feature extraction module in the first stage is essentially a variety of attention groups. The variety of attention groups include an overall channel attention module, a multiple convolution feature fusion module, and a dual attention module. The multiple convolution feature fusion module includes a channel reconstruction unit, a spatial reconstruction unit, and a standard convolution group. First, the small-scale shallow features are respectively input into the overall channel attention module to capture pixel information. Then, the output of the overall channel attention module is input into the multiple convolution feature fusion module for local and global information fusion. Finally, the output of the multiple convolution feature fusion module is input into the dual attention module to capture context information through different windows, obtaining small-size initial features ;
[0151] (4) The small-size initial features are concatenated, and then the concatenation result and the low-resolution image are jointly input into the supervised attention module of the first stage to obtain the final output of the first stage. The specific calculation is as follows:
[0152] ,
[0153] ,
[0154] ,
[0155] ,
[0156] ,
[0157] ,
[0158] ,
[0159] where, represents the tensor concatenation method of the scale, and represent two different concatenation results, represents the set of concatenation results, represents the supervised attention module, represents the intermediate output of the supervised attention module, and represent two different outputs of the supervised attention module in the first stage.
[0160] The specific content of the initial feature extraction module in the first stage is as follows:
[0161] (3-1) The specific calculation in the overall channel attention module is as follows:
[0162] ,
[0163] ,
[0164] ,
[0165] ,
[0166] Among them, represents the intermediate output of the overall channel attention module, represents the th small-scale shallow feature, , represents the convolution operation, represents the activation function operation, represents the output of the overall channel attention module, represents the operation of represents the convolution operation, represents the operation of the channel attention module, represents the activation function operation, represents the activation function operation, represents the average pooling operation;
[0167] The calculation in the (3 - 2) multi-convolution feature fusion module is as follows:
[0168] 1) Spatial reconstruction unit:
[0169]
[0170] ,
[0171] ,
[0172] ,
[0173] ,
[0174] Among them, represents the normalization operation, represents the result of processing, and respectively represent the trainable scaling parameter and translation parameter, and respectively represent the mean and standard deviation of A constant term for maintaining division stability Denotes the normalized correlation weight Denotes two different weights obtained after threshold processing Denotes threshold gating with the threshold set to 0.5 Denotes And The result obtained after the tensor cutting operation after multiplication Denotes And The result obtained after the tensor cutting operation after multiplication Denotes the tensor cutting operation Denotes the tensor concatenation operation Denotes the operation of the spatial reconstruction unit Denotes the output of the spatial reconstruction unit
[0175] 2) Channel reconstruction unit:
[0176] ,
[0177] ,
[0178] ,
[0179] ,
[0180] ,
[0181] ,
[0182] Among them, Denotes the result obtained after performing the tensor cutting operation on Denotes grouped convolution Denotes The result obtained after performing the grouped convolution operation on Denotes The result obtained after performing the tensor concatenation operation on Denotes the tensor concatenation operation Denotes And The result obtained after performing the tensor concatenation operation Denotes the normalized exponential function operation Denotes the weight of the normalized exponential function operation Denotes the intermediate feature of the fusion strategy process Denotes the operation of the channel reconstruction unit Denotes the output of the channel reconstruction unit
[0183] 3) Standard Convolution Group:
[0184] ,
[0185] ,
[0186] ,
[0187] ,
[0188] ,
[0189] ,
[0190] ,
[0191] ,
[0192] Among them, represents a spatial-channel operation, represents a channel-spatial operation, represents the result of the spatial-channel operation, represents the result of the channel-spatial operation, represents the result of the spatial-channel standard convolution, represents the result of the channel-spatial standard convolution, represents the input of the module, , represents the result of the nth standard convolution, represents the total number of standard convolutions, , represents the sum branch result, represents the difference branch result, represents the final output feature of the multi-convolution feature fusion module;
[0193] The calculation in the (3-3) dual attention module is as follows:
[0194] ,
[0195] ,
[0196] ,
[0197] ,
[0198] ,
[0199] The specific internal operations of the DAB module are as follows:
[0200] ,
[0201] ,
[0202] Among them, represents the layer normalization result, represents the deformable size window self-attention mechanism, represents the double routing attention mechanism, represents the intermediate output of the double attention module, represents the operation of the multi-layer perceptron, represents the output result of the double attention module, represents the overall channel attention, represents the multiple convolution feature fusion module, represents the double attention module, represents the th multiple attention group operation performed in the first stage, represents the number of multiple attention group operations, , represents the final feature output of the first-stage continuous multiple attention group module.
[0203] The process in the second stage of the multi-stage progressive multiple attention network model is as follows:
[0204] S2.2. First, the input low-resolution image is divided according to the medium scale, and then the division results are respectively input into the shallow feature extraction module of the second stage to obtain the medium-scale shallow features and . Then, the results of the shallow feature extraction module and the output of the supervised attention module of the first stage are respectively input into the initial feature extraction module of the second stage for preliminary feature capture. Then, the captured medium-scale initial features are stitched together in terms of image size to obtain the initial features. Finally, the initial features and the low-resolution image are input into the supervised attention module of the second stage to obtain the final output of the second stage;
[0205] Among them, 、 , represents the number of channels of the features output by the shallow feature extraction module of the second stage;
[0206] (1) The medium-scale division operation is as follows:
[0207] ,
[0208] Among them, represents the mesoscale division operation, and represents the result of mesoscale division;
[0209] (2) The shallow feature extraction module in the second stage consists of a convolution with a kernel size of 3 and a stride of 1. The calculations in the shallow feature extraction module in the second stage are as follows:
[0210] ,
[0211] Among them, represents the convolution operation;
[0212] (3) The structure of the initial feature extraction module in the second stage is the same as that in the first stage. The mesoscale shallow features and the output of the supervised attention module in the first stage are respectively input into the initial feature extraction module in the second stage to obtain the output of the initial feature extraction module in the second stage. The specific calculations are as follows:
[0213] ,
[0214] ,
[0215] ,
[0216] ,
[0217] Among them, represents the method of tensor concatenation of channels, and represent two different results of tensor concatenation of channels, and represent two different mesoscale initial features output by the initial feature extraction module in the second stage;
[0218] (4) The mesoscale initial features and are concatenated, and then the concatenation result and the low-resolution image are jointly input into the supervised attention module in the second stage to obtain the final output in the second stage. The specific calculations are as follows:
[0219] ,
[0220] ,
[0221] Among them, represents the method of tensor concatenation of scales, represents the operation of the supervised attention module, Represents the final output of the second stage.
[0222] The process in the third stage of the multi-stage progressive multi-attention network model is as follows:
[0223] S2.3. First, divide the input low-resolution image According to the large scale, and then input the division results into the shallow feature extraction module of the third stage respectively to obtain large-scale shallow features , then input the results of the shallow feature extraction module and the output of the supervised attention module of the second stage into the initial feature extraction module of the third stage for preliminary feature capture, and then splice the captured large-scale initial features in terms of image size to obtain the initial features. Finally, input the initial features and the low-resolution image into the supervised attention module of the third stage to obtain the final output of the third stage, that is, the high-resolution medical image based on the optical magnification principle;
[0224] Among them, , Represents the number of channels of the features output by the shallow feature extraction module of the third stage;
[0225] (1) The large-scale division operation is as follows:
[0226] ,
[0227] Among them, Represents the large-scale division operation, Represents the result of the large-scale division, that is, the input low-resolution image ;
[0228] (2) The shallow feature extraction module of the third stage consists of a convolution with a kernel size of 3 and a stride of 1. The calculation in the shallow feature extraction module of the third stage is as follows:
[0229] ,
[0230] Among them, Represents the convolution operation;
[0231] (3) The structure of the initial feature extraction module of the third stage is the same as that of the initial feature extraction module of the first stage. Input the large-scale shallow features and the output of the supervised attention module of the second stage into the initial feature extraction module of the third stage respectively to obtain the output of the initial feature extraction module of the third stage. The specific calculation is as follows:
[0232] ,
[0233] ,
[0234] ,
[0235] in, The tensor concatenation method representing the channel, Represents the result of concatenating the tensors of the channels, Indicates The result of the attention group operation of the consecutive channels, Indicates Channel attention modules, represents the total number of channel attention modules, , Represents the large-scale initial features output by the initial feature extraction module in the third stage;
[0236] (4) Large-scale initial features Input the third stage image reconstruction module, select the scaling factor, and Contains multi-scale information to generate high-resolution medical images using optical magnification , where the image reconstruction module uses The method is used to reorganize pixels. The specific calculation is as follows:
[0237] .
[0238] S3 is as follows:
[0239] The loss function calculation is as follows:
[0240] ,
[0241] in, Represents the prediction results and the true value The loss function and prediction results That is, to generate high-resolution medical images based on the principle of optical magnification. represents the number of medical images in the training set, Indicates Prediction results of medical images, Indicates The true value of a medical image;
[0242] The peak signal-to-noise ratio PSNP and structural similarity SSIM are used as evaluation indicators to verify the generated optical magnification principle high-resolution medical image. The larger the PSNR and SSIM, the better the quality of the generated optical magnification principle high-resolution medical image.
[0243] Example 2
[0244] like Figures 2 to 4As shown, they are the control group images of a dermoscope, glasses, and otoscope respectively. A dermoscope dataset, an ophthalmoscope dataset, and an otoscope dataset are constructed respectively. The specific process is as follows:
[0245] The dermoscope dataset contains a total of 10,000 high-resolution dermoscope images and 10,000 low-resolution dermoscope images. Images of important diagnostic categories in the field of skin lesions such as actinic keratosis, basal cell carcinoma, benign keratosis-like lesions (seborrheic keratosis and lichen planus-like keratosis), dermatofibroma, melanoma, melanocytic nevus, and vascular lesions (hemangioma, angiokeratoma) are collected. In the present invention, 7,000 dermoscope images are taken from 10,000 high-resolution dermoscope images and 10,000 low-resolution dermoscope images to form a training set, and then the remaining 3,000 high-resolution dermoscope images and 3,000 low-resolution dermoscope images are divided into three equal parts to form two validation sets and one test set;
[0246] As Figure 2 shown, seven kinds of skin disease images are selected from the dermoscope dataset. The low-resolution dermoscope images in the dermoscope dataset are processed by the method of the present invention to obtain the output of a multi-stage progressive multiple attention model. From Figure 2 it can be seen that the clarity of the images output by the model is significantly higher than that of the low-resolution dermoscope images, and the clarity of the images output by the model can reach the level of the original high-resolution dermoscope images.
[0247] An ophthalmoscope dataset is constructed. The ophthalmoscope dataset contains a total of 4,216 high-resolution dermoscope images and 4,216 low-resolution dermoscope images. Images of the left and right eyes of cataract, glaucoma, diabetic retinopathy, and normal conditions are collected. In the present invention, 3,316 dermoscope images are taken from 4,216 high-resolution dermoscope images and 4,216 low-resolution dermoscope images to form a training set, and then the remaining 900-resolution dermoscope images and 900-resolution dermoscope images are divided into three equal parts to form two validation sets and one test set;
[0248] As Figure 3 shown, four kinds of disease images are selected from the glasses dataset. The low-resolution glasses images in the glasses dataset are processed by the method of the present invention to obtain the output of a multi-stage progressive multiple attention model. From Figure 3 it can be seen that the clarity of the images output by the model is significantly higher than that of the low-resolution glasses images, and the clarity of the images output by the model can reach the level of the original high-resolution glasses images.
[0249] Construct an otoscope dataset. The otoscope dataset contains a total of 956 high-resolution dermoscopic images and 956 low-resolution dermoscopic images. Otoscopic images under acute otitis media, chronic suppurative otitis media, ear ventilation tubes, excessive earwax, foreign bodies in the ear, otitis externa, false tympanic membranes, tympanic membranes, and normal conditions are collected. In the present invention, 656 dermoscopic images are taken from 956 high-resolution dermoscopic images and 956 low-resolution dermoscopic images respectively to form a training set, and then the remaining 300 high-resolution dermoscopic images and 300 low-resolution dermoscopic images are divided into three equal parts to form two validation sets and one test set.
[0250] As Figure 4 shown, eight disease images and a group of normal images are selected from the otoscope dataset. The low-resolution otoscope images in the otoscope dataset are processed by the method of the present invention to obtain the output of a multi-stage progressive multi-attention model. As Figure 4 can be seen, the clarity of the images output by the model is significantly higher than that of the low-resolution otoscope images, and the clarity of the images output by the model can reach the level of the original high-resolution otoscope images.
[0251] In summary, it can be seen that the present invention can repair image details and refine medical images, and thus generate clearer medical images based on the principle of optical magnification.
[0252] Although the specific embodiments of the invention have been described above in conjunction with the accompanying drawings, it is not a limitation to the protection scope of the present invention. Based on the technical solutions of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.
Claims
1. A multi-stage progressive super-resolution method for medical images based on optical magnification principle, characterized in that: The following steps are involved: S1. Collect medical images used for optical magnification principle to form a data set, the medical images include low-resolution images and original high-resolution images, and preprocess the low-resolution medical images in the data set to obtain a preprocessed data set, and then divide the data set into a training set, a validation set and a test set; S2. Construct a multi-stage progressive multi-attention network model, which includes three stages. The first and second stages both include a shallow feature extraction module, an initial feature extraction module, and a supervised attention module. The third stage includes a shallow feature extraction module, an initial feature extraction module, and an image reconstruction module. The low-resolution image is output to the model for feature extraction, and finally a high-resolution medical image of the optical magnification principle is generated. Output the low-resolution image to the model for feature extraction: the low-resolution image goes through three stages in sequence. The low-resolution image is input to the first stage, the output of the first stage and the low-resolution image are used as input to the second stage, and the output of the second stage and the low-resolution image are used as input to the third stage. The third stage finally outputs a high-resolution medical image based on the optical magnification principle. First stage: First, the input low-resolution image Divide into small scales, and then input the division results into the shallow feature extraction module of the first stage to obtain small scale shallow features. , , and , and then input the results of the shallow feature extraction module into the initial feature extraction module of the first stage for preliminary feature capture, and then splice the captured small-scale initial features to obtain the initial features, and finally input the initial features and the low-resolution image into the supervised attention module of the first stage to obtain the final output of the first stage; in, , Indicates height, Indicates width, Indicates the number of channels of the input low-resolution image, , , , , Indicates the number of channels of features output by the shallow feature extraction module in the first stage ; The shallow feature extraction module in the first stage consists of a convolution with a kernel size of 3 and a stride of 1; The initial feature extraction module in the first stage is actually a variety of attention groups, which include the overall channel attention module, the multi-convolution feature fusion module and the dual attention module. The multi-convolution feature fusion module includes the channel reconstruction unit, the spatial reconstruction unit and the standard convolution group. Second stage: First, the input low-resolution image According to the mesoscale division, the division results are then input into the shallow feature extraction module of the second stage to obtain the mesoscale shallow feature and , and then input the results of the shallow feature extraction module and the output of the supervised attention module of the first stage into the initial feature extraction module of the second stage for preliminary feature capture, and then splice the captured mid-scale initial features to obtain the initial features, and finally input the initial features and the low-resolution image into the supervised attention module of the second stage to obtain the final output of the second stage; in, , , Indicates the number of channels of features output by the shallow feature extraction module in the second stage; The shallow feature extraction module in the second stage consists of a convolution with a kernel size of 3 and a stride of 1; The structure of the initial feature extraction module in the second stage is the same as that in the first stage; The third stage: First, the input low-resolution image According to the large-scale division, the division results are then input into the shallow feature extraction module of the third stage to obtain the large-scale shallow feature , and then the results of the shallow feature extraction module and the output of the second stage supervised attention module are respectively input into the initial feature extraction module of the third stage for preliminary feature capture, and then the captured large-scale initial features are spliced into image size to obtain initial features, and finally the initial features and low-resolution images are input into the third stage supervised attention module to obtain the final output of the third stage, i.e., the high-resolution medical image of the optical magnification principle; in, , Indicates the number of channels of features output by the shallow feature extraction module in the third stage; The shallow feature extraction module in the third stage consists of a convolution with a kernel size of 3 and a stride of 1; The structure of the initial feature extraction module in the third stage is the same as that in the first stage; S3. Input the data in the training set, validation set, and test set into the multi-stage progressive multiple attention network model in sequence. Input the data in the training set into the model and train the model with the loss function. Then input the data in the validation set into the trained model and adjust the model parameters with the evaluation index. Finally, input the data in the test set into the adjusted model to test the model and output the final result.
2. The multi-stage progressive super-resolution method for optical magnification principle medical images according to claim 1, characterized in that: S1 is as follows: S1.
1. Constructing the dataset: Select medical images based on scaling factors and build a dataset based on medical images , medical images specifically include low-resolution medical images used for optical magnification principles and original high-resolution images. The dataset is represented as ,in, Indicates Medical images, Indicates A low-resolution medical image of 64×64 pixels used for optical magnification. Indicates The original high-resolution image used for the optical magnification principle is 64S×64S in size, where S represents the scaling factor. represents the number of medical images in the dataset, represents the index of the medical image, ; S1.
2. Dataset preprocessing and partitioning: For the dataset The low-resolution medical images in the dataset are horizontally flipped and rotated for data augmentation, and then the preprocessed dataset is divided into training set, validation set and test set in proportion.
3. The multi-stage progressive super-resolution method for optical magnification principle medical images according to claim 2, characterized in that: S2.
1. The process in the first stage of the multi-stage progressive multiple attention network model is as follows: (1) The small-scale division operation is as follows: , in, represents the small-scale partitioning operation, Indicates the result of small-scale division; (2) The calculation in the shallow feature extraction module of the first stage is as follows: , in, Represents the convolution operation; (3) The initial feature extraction module in the first stage is actually a variety of attention groups. First, the small-scale shallow features are input into the overall channel attention module to capture pixel information. Then, the output of the overall channel attention module is input into the multi-convolution feature fusion module to fuse local information and global information. Finally, the output of the multi-convolution feature fusion module is input into the dual attention module to capture contextual information through different windows to obtain the small-scale initial features. ; (4) Small size initial features Stitching is then performed and the stitching result is combined with the low-resolution image They are input together into the supervised attention module of the first stage to obtain the final output of the first stage. The specific calculation is as follows: , , , , , , , in, A tensor concatenation method representing scale, and Represents two different splicing results, Represents the collection of splicing results, represents the supervised attention module, represents the intermediate output of the supervised attention module, and Represents two different outputs of the supervised attention module in the first stage.
4. The multi-stage progressive super-resolution method for optical magnification principle medical images according to claim 3, characterized in that: The initial feature extraction module of the first stage is as follows: (3-1) The calculation in the overall channel attention module is as follows: , , , , in, represents the intermediate output of the overall channel attention module, Indicates Small-scale shallow features, , represents the convolution operation, express Activation function operation, represents the output of the overall channel attention module, express Operation, represents the convolution operation, represents the operation of the channel attention module, express Activation function operation, express Activation function operation, Represents a tie pooling operation; (3-2) The calculation in the multiple convolution feature fusion module is as follows: 1) Space reconstruction unit: , , , , in, represents the normalization operation, express The result of the processing, and denote the trainable scaling and translation parameters, respectively. and Respectively The mean and standard deviation of represents the constant term that maintains the stability of division, represents the normalized correlation weight, Represents two different weights obtained after threshold processing, Indicates threshold gating, the threshold is set to 0.5, express and The result obtained after the tensor cutting operation after multiplication, express and The result obtained after the tensor cutting operation after multiplication, represents a tensor cutting operation, represents a tensor concatenation operation, represents the operation of the spatial reconstruction unit, represents the output of the spatial reconstruction unit; 2) Channel reconstruction unit: , , , , , , in, Express The result obtained after the tensor cutting operation, represents grouped convolution, Express The result obtained after performing group convolution operation, Express The result of tensor concatenation operation is: represents a tensor concatenation operation, express and The result obtained after tensor splicing operation, represents the normalized exponential function operation, represents the weight of the normalized exponential function operation, represents the intermediate features of the fusion strategy process, Represents the operation of the channel reconstruction unit, represents the output of the channel reconstruction unit; 3) Standard convolution group: , , , , , , , , in, represents a space-channel operation, represents channel-space operation, represents the result of the space-channel operation, represents the result of the channel-space operation, represents the spatial-channel standard convolution result, represents the channel-space standard convolution result, express The module input, , Indicates The result of substandard convolution, represents the total number of standard convolutions, represents the index of the number of standard convolutions, , Represents and branch results, represents the difference branch result, Represents the final output features of the multiple convolutional feature fusion module; (3-3) The calculation in the dual attention module is as follows: , , , , , The internal operations of the DAB module are as follows: , , in, Represents the normalized result of the layer, represents the self-attention mechanism of deformable size window, represents a two-layer routing attention mechanism, represents the intermediate output of the dual attention module, represents the operation of a multilayer perceptron, represents the output of the dual attention module, represents the overall channel attention, represents the multiple convolution feature fusion module, represents the dual attention module, Indicates the first stage of Multiple attention group operations, represents the number of operations of various attention groups, , Represents the final feature output of the first stage continuous multi-attention group module.
5. The multi-stage progressive super-resolution method for optical magnification principle medical images according to claim 4, characterized in that: S2.
2. The process in the second stage of the multi-stage progressive multiple attention network model is as follows: (1) The specific operation of mesoscale division is as follows: , in, represents the mesoscale partitioning operation, represents the result of mesoscale partitioning; (2) The shallow feature extraction module of the second stage is composed of a convolution with a convolution kernel size of 3 and a step size of 1. The calculation in the shallow feature extraction module of the second stage is as follows: , in, Represents the convolution operation; (3) The structure of the initial feature extraction module in the second stage is the same as that in the first stage. The mid-scale shallow features and the output of the supervised attention module in the first stage are respectively input into the initial feature extraction module in the second stage to obtain the output of the initial feature extraction module in the second stage. The specific calculation is as follows: , , , , in, The tensor concatenation method representing the channel, and Represents two different results of channel tensor concatenation, and Represents two different mesoscale initial features output by the initial feature extraction module of the second stage; (4) The initial mesoscale features and Stitching is then performed and the stitching result is combined with the low-resolution image They are input together into the supervised attention module of the second stage to obtain the final output of the second stage. The specific calculation is as follows: , , in, A tensor concatenation method representing scale, represents the operation of the supervised attention module, Represents the final output of the second stage.
6. The multi-stage progressive super-resolution method for optical magnification principle medical images according to claim 5, characterized in that: S2.
3. The process in the third stage of the multi-stage progressive multiple attention network model is as follows: (1) The large-scale division operation is as follows: , in, represents a large-scale partitioning operation, Represents the result of large-scale division, that is, the input low-resolution image ; (2) The calculation in the shallow feature extraction module of the third stage is as follows: , in, Represents the convolution operation; (3) The large-scale shallow features and the output of the supervised attention module in the second stage are respectively input into the initial feature extraction module in the third stage to obtain the output of the initial feature extraction module in the third stage. The specific calculation is as follows: , , , in, The tensor concatenation method representing the channel, Represents the result of concatenating the tensors of the channels, Indicates The result of the attention group operation of the consecutive channels, Indicates Channel attention modules, represents the total number of channel attention modules, , Represents the large-scale initial features output by the initial feature extraction module in the third stage; (4) Large-scale initial features Input the third stage image reconstruction module, select the scaling factor, and Contains multi-scale information to generate high-resolution medical images using optical magnification , where the image reconstruction module uses The method is used to reorganize pixels. The specific calculation is as follows: 。 7. The multi-stage progressive super-resolution method for optical magnification principle medical images according to claim 6, characterized in that S3 The details are as follows: The loss function calculation is as follows: , in, Represents the prediction results and the true value The loss function and prediction results That is, to generate high-resolution medical images based on the principle of optical magnification. represents the number of medical images in the training set, Indicates Prediction results of medical images, Indicates The true value of a medical image; The peak signal-to-noise ratio PSNP and structural similarity SSIM are used as evaluation indicators to verify the generated optical magnification principle high-resolution medical image. The larger the PSNR and SSIM, the better the quality of the generated optical magnification principle high-resolution medical image.
Citation Information
Patent Citations
Image super-resolution reconstruction method based on lightweight hybrid attention network
CN117745541A