Method, System and Device for Improving Super-Resolution of Laryngoscope Images Based on Deep Learning
Through a neural network system based on deep learning, the problems of low resolution and blurred details of laryngoscopy images are solved, and the reconstruction and quality of high-resolution images are achieved, meeting the needs of high-precision medical images.
Patent Information
- Application Number
- CN202510206299.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-02-25
AI Technical Summary
In the prior art, laryngoscopy images have low resolution and blurred details, making it difficult to meet the needs of high-precision medical images.
A neural network system based on deep learning is adopted, including image preprocessing module, shallow feature extraction module, deep feature extraction module and image reconstruction module, and high-resolution laryngoscope images are reconstructed through multi-scale feature fusion and deep feature capture.
It significantly improves the resolution and quality of laryngoscopic images, clearly presents detailed information of tissues and lesions, reduces artifacts and noise interference, and improves the visual quality of the image and the accuracy of diagnosis.
Smart Images

Figure CN119693371B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of medical image processing, artificial intelligence, and computer data processing, and particularly to a method, system, and device for enhancing the super-resolution of laryngoscope images based on deep learning. Background Art
[0002] With the rapid development of medical imaging technology, the medical field has an increasingly urgent need for high-definition images. Especially during diagnosis and surgical procedures, clear images can significantly improve the diagnostic accuracy of doctors and the success rate of surgeries. As an important medical device, the laryngoscope is widely used in the examination and treatment of laryngeal diseases. However, due to the special nature of the laryngoscope working environment, obtaining high-quality images has always been a challenge.
[0003] Traditional laryngoscope images usually have problems such as low resolution, high noise, and blurred details, which bring many inconveniences to subsequent use. To improve the clarity and resolution of images, traditional methods mainly rely on the upgrade and improvement of hardware, such as using high-resolution cameras and advanced optical systems. However, this method is not only costly but also limited by physical space and the size of the instrument, making it difficult to be widely promoted in practical applications. On the other hand, software methods have also been widely used in image super-resolution technology, such as algorithms like bilinear interpolation and bicubic interpolation. These algorithms improve the image resolution by inserting new pixels between existing image pixels. Although they can improve the image quality in some cases, since these methods rely on the global information of the image, they often result in blurred edges and lost details, making it difficult to restore high-quality detailed information. Especially at high magnification, problems such as artifacts, noise, or image blurring are likely to occur, making it difficult to meet the requirements of high-precision medical images.
[0004] In recent years, with the development of deep learning technology, especially the successful application of convolutional neural networks (CNNs) in the field of image processing, image super-resolution technology has achieved new breakthroughs. By constructing and training deep neural networks, high-resolution images can be reconstructed from low-resolution images, significantly improving the image quality. Deep learning methods can not only learn the global features of images but also capture the local details of images, thus performing excellently in image reconstruction.
[0005] Deep learning-based image super-resolution methods, such as SRCNN, SRResNet, and SRGAN, have demonstrated high reconstruction effects in natural image super-resolution tasks. SRCNN is a super-resolution model based on convolutional neural networks that achieves image reconstruction by establishing a direct mapping from low resolution to high resolution. However, the network structure of SRCNN is relatively shallow, making it difficult to capture high-order features of images, resulting in limited effects when processing high-precision scenarios such as medical images. SRResNet introduced the residual network (ResNet) on the basis of SRCNN, enhancing the network's learning ability by stacking residual modules and achieving remarkable results in the reconstruction of natural images. However, its applicability to medical images still needs to be further verified. SRGAN is a model that combines generative adversarial networks (GANs) for super-resolution reconstruction. By introducing adversarial loss, the generated images are made more natural and realistic. However, the network training of SRGAN is difficult, easily leading to unstable training, and may produce artifacts in medical images, affecting the observation of lesions. Summary of the Invention
[0006] To solve the above technical problems existing in the prior art, the present invention proposes a method, system, and device for enhancing the super-resolution of laryngoscope images based on a deep learning neural network, aiming to solve the problems of low image resolution and blurred details in the prior art. Specifically, for the characteristics of medical laryngoscope images, an efficient and practical solution is provided. Through this method, the resolution and quality of laryngoscope images can be significantly improved. Specifically, the present invention provides the following technical solutions:
[0007] On the one hand, the present invention provides a system for enhancing the super-resolution of laryngoscope images based on deep learning, which includes: an image preprocessing module, a shallow feature extraction module, a deep feature extraction module, and an image reconstruction module;
[0008] The image preprocessing module is used to preprocess the low-resolution laryngoscope image, enhance the locally blurred areas in the image, obtain the preprocessed laryngoscope image, and input the obtained preprocessed laryngoscope image into the shallow feature extraction module;
[0009] The shallow feature extraction module is used to extract shallow features from the low-resolution laryngoscope image, and through multi-scale feature fusion, input the obtained shallow features into the deep feature extraction module;
[0010] The deep feature extraction module is used to capture deep features from the shallow features, generate final features by fusing the shallow features with the captured deep features, and input these final features into the image reconstruction module;
[0011] The image reconstruction module is used to reconstruct a complete image from the learned final features, thereby outputting a high-resolution laryngoscope image;
[0012] The deep feature extraction module includes two paths, one is a deep feature extraction unit and an enhanced convolutional attention unit connected in sequence, wherein the deep feature extraction unit is provided with multiple parallel paths; the other is composed of a graph convolutional neural network unit and two parallel causal convolutional units, the two parallel causal convolutional units are both connected to the graph convolutional neural network unit, the graph convolutional neural network unit receives shallow features, and the two parallel causal convolutional units are both composed of causal convolutional blocks and activation functions.
[0013] Preferably, the shallow feature extraction module includes: three parallel 1×1 convolution and ReLU activation layers, 3×3 convolution and ReLU activation layers, and 5×5 convolution and ReLU activation layers; a primary feature fusion unit, and a 1×1 convolution layer connected to the primary feature fusion unit;
[0014] The low-resolution laryngoscope image is respectively input into a 1×1 convolution and ReLU activation layer, a 3×3 convolution and ReLU activation layer, and a 5×5 convolution and ReLU activation layer arranged in parallel, and three primary features are respectively extracted; the three primary features are fused by a primary feature fusion unit, and then passed through a 1×1 convolution layer to obtain shallow features.
[0015] Preferably, the deep feature extraction unit comprises five parallel paths, a feature fusion unit and a convolution layer with a convolution kernel of 1×1;
[0016] The five parallel paths are set as follows: the first path is a 1×1 convolution layer; the second path includes a 1×1 convolution layer and a 3×3 convolution layer connected in sequence; the third path includes a 1×1 convolution layer, a 3×3 convolution layer, and an expansion convolution layer with a dilation value of 3 and a convolution kernel size of 3×3 connected in sequence; the fourth path includes a 1×1 convolution layer, a 5×5 convolution layer, and an expansion convolution layer with a dilation value of 5 and a convolution kernel size of 3×3 connected in sequence; the fifth path is directly connected to the input.
[0017] Preferably, the feature fusion unit fuses the outputs of the first path, the second path, the third path, and the fourth path, and inputs the outputs to a convolution layer with a convolution kernel of 1×1;
[0018] The output of the convolution layer with a convolution kernel of 1×1 is added to the output of the fifth path.
[0019] Preferably, the image preprocessing module preprocesses the low-resolution laryngoscope image based on wavelet transform, specifically in the following manner:
[0020] Construct wavelet basis functions and set constraints on the wavelet basis functions;
[0021] Calculate the square of the wavelet coefficients and obtain them in ascending order , where s represents the square value of the corresponding wavelet coefficient and N represents the number of wavelet coefficients;
[0022] Calculate the risk value and threshold selection parameter of each wavelet coefficient, and retain the wavelet coefficients whose absolute value of the risk value is greater than the threshold selection parameter to obtain an optimized wavelet transform function;
[0023] Apply the optimized wavelet transform function to the low-resolution laryngoscope image for wavelet transform to obtain a preprocessed laryngoscope image.
[0024] Preferably, the constructed wavelet basis function is:
[0025] ;
[0026] where t represents the time variable of the wavelet transform.
[0027] Preferably, the kernel function of the wavelet transform is:
[0028] ;
[0029] where is the scale factor, is the translation factor, .
[0030] Preferably, the constraint condition is:
[0031] ;
[0032] where represents the Fourier transform form of , R represents the set of real numbers, that is, it means that the calculations and operations in the wavelet transform are carried out in the real number domain, and L() represents the wavelet transform.
[0033] Preferably, the risk value calculation method is: calculate the square of the wavelet coefficients and sort them in ascending order to obtain . According to the number of wavelet coefficients , calculate the risk value:
[0034] ;
[0035] Preferably, the calculation method of the threshold selection parameter is:
[0036] ;
[0037] where represents the mean square error, that is, calculate the mean square error of the wavelet coefficient as the risk value The screening conditions. If , then the wavelet coefficient is retained, otherwise it is deleted. Thus, the optimization of the wavelet transform function is completed. The optimized wavelet transform function is used to process the low-resolution laryngoscope image to improve the overall image quality.
[0038] Preferably, after wavelet transform, further screen whether there is a blurred area, and edge enhancement can be performed on the blurred area. The edge information can be enhanced by calculating the average gradient of the blurred area:
[0039] ;
[0040] where the size of the blurred area is M×N, represents the pixel gradient of the blurred area image, and the average gradient G obtained through calculation replaces the pixel value of the original area to achieve the effect of edge strengthening. It should be understood here that gradient enhancement is an optional additional method, and it is also possible not to perform gradient enhancement and only use the above optimized wavelet transform for processing.
[0041] Through the above calculation, after the wavelet transform is improved, the edge information of the local blurred area of the image can be accurately extracted, and the enhancement of the local image blur can be effectively completed using the improved wavelet transform, and the overall image quality can be improved. Subsequently, the result is input to the shallow feature extraction module.
[0042] Preferably, the blurred area can be, for example, an area where the image is unclear caused by excessive local exposure, etc. The blurred area can be delimited by manual division or other means. Or, the blurred area is determined by calculating the average contrast and image quality index of the image area. For example, when the average contrast and image quality index are respectively less than the preset thresholds, it is defined as a blurred area. When screening the blurred area, the image can be divided into areas of a certain unit area and processed block by block, for example, divided into blocks with a size of T pixels.
[0043] Preferably, the average contrast is used to measure the brightness difference in different areas of the image (the higher the average contrast, the more obvious the brightness difference), and the expression is:
[0044] ;
[0045] where represents the average value of the contrast, represents the intensity of the local blur variance of the image.
[0046] Preferably, the image quality index is the average degree of pixel brightness change within the image area, used to evaluate the quality of the image, and used to guide the enhancement of the local blurred area of the image. The higher the value, the higher the image quality. is calculated as:
[0047] ;
[0048] where x represents the gray value of pixels in the original low-resolution laryngoscope image, and y represents the gray value of pixels in the laryngoscope image after wavelet transform. and are respectively and the average values of and are respectively and the second-order moments of is and the covariance of
[0049] ;
[0050] ;
[0051] ;
[0052] ;
[0053] ;
[0054] Preferably, the enhanced convolutional attention unit includes a channel attention block and a spatial attention block connected in sequence.
[0055] The channel attention block performs convolutional processing on the shallow features input to the enhanced convolutional attention unit, multiplies them with the shallow features, and then inputs them to the spatial attention block.
[0056] Multiply the input of the spatial attention block with the output of the spatial attention block to obtain the final features output by the deep feature extraction module.
[0057] Preferably, the channel attention block is provided with two paths: the first path includes a 1×1 convolutional layer, a 3×3 convolutional layer, a 1×1 convolutional layer, a squeeze-and-excitation layer, a 1×1 convolutional layer, a batch normalization layer, and a ReLU activation layer connected in sequence; the second path directly connects to the input of the channel attention block.
[0058] The outputs of the first path and the second path are added to obtain the output of the channel attention block.
[0059] Preferably, the spatial attention block is provided with two paths: the first path sequentially passes through a 1×1 convolutional layer, a strided convolutional layer, a pooling layer, a convolutional groups layer, an upsampling layer, a 1×1 convolutional layer, and a Sigmoid activation layer; the second path is directly connected to the input of the spatial attention block;
[0060] The outputs of the first path and the second path are added together to obtain the output of the spatial attention block.
[0061] Preferably, two paths are provided between the input and the output of the squeeze-and-excitation layer: the first path is set to be sequentially connected with a global pooling layer, a first fully connected layer, a ReLU activation layer, a second fully connected layer, and a Sigmoid activation layer; the second path is directly connected to the input of the squeeze-and-excitation layer;
[0062] The output of the first path and the output of the second path are multiplied element by element to obtain the output of the squeeze-and-excitation layer.
[0063] Preferably, the calculation method of the graph convolutional neural network unit is as follows:
[0064] ;
[0065] where, and respectively represent the values of the -th layer and the -th layer, represents the activation function, represents the weight matrix;
[0066] ;
[0067] ;
[0068] The shallow features are represented by the graph G, , A represents the adjacency matrix, which is a set of real numbers, E represents the edge set, V represents the vertex set, I represents the identity matrix, represents the element value of the i-th row and j-th column of the matrix.
[0069] Preferably, the specific derivation process of the graph convolutional neural network unit is as follows: For a graph output by a shallow feature extraction module, where, represents the vertex set, that is, the pixels in the image, represents the edge set, is the adjacency matrix. The Laplacian operator is used to establish the dependency relationship between pixels, where, is the degree matrix, representing the degree of each node. Then, the Laplacian matrix is normalized:
[0070] ;
[0071] where is the identity matrix. Subsequently, a convolutional kernel is introduced, where . The convolution operation can be expressed as: where is the input signal, is the eigenvector matrix of is the eigenvalue of can be regarded as a function of the eigenvalue . To simplify the calculation, Chebyshev polynomials are used to approximate :
[0072] ;
[0073] where are Chebyshev polynomials.
[0074] ;
[0075] where
[0076] Based on the above formula, the calculation method of the graph convolutional neural network is as follows:
[0077] ;
[0078] where and respectively represent the values of the -th layer and the -th layer, represents the activation function, represents the weight matrix. Through the above rules, the graph convolutional neural network can aggregate the features of adjacent nodes, thereby establishing long-range dependencies between different nodes and significantly improving the ability of pixel points to combine spatial context information.
[0079] Preferably, the causal convolution block is implemented by right alignment to ensure that the convolution only acts on the current time point, focuses on past states, and prevents any information leakage from future states. In each convolutional layer, a gated activation unit is adopted, which consists of two non-linear functions: a Sigmoid function and a Tanh function. The output results of two parallel causal convolution units are fused in the following specific way:
[0080] ;
[0081] Among them, * represents the convolution operation, and the Tanh function is defined as , and the Sigmoid function is defined as . Among them, the Tanh term is used as a common activation function to constrain the output value within the range of [-1, 1], and the Sigmoid term is used as a gating mechanism to dynamically activate or deactivate unnecessary time points. The output result is fused with the output result of the enhanced convolutional attention unit.
[0082] Preferably, the image reconstruction module is composed of a first Conv unit, a sub-pixel convolutional layer, and a second Conv unit connected in sequence; both the first Conv unit and the second Conv unit are composed of a 1×1 convolutional layer and a ReLU activation layer.
[0083] In addition, the present invention also provides a method for enhancing the super-resolution of laryngoscope images based on deep learning. This method is applied to the system as described above, and this method includes:
[0084] S1. Obtain a low-resolution laryngoscope image . After preprocessing the low-resolution laryngoscope image, a preprocessed laryngoscope image is obtained, and the preprocessed laryngoscope image is input into shallow feature extraction;
[0085] S2. Extract shallow features from the preprocessed laryngoscope image through the shallow feature extraction module;
[0086] S3. Input the shallow features into the deep feature extraction module to extract deep features;
[0087] S4. Fuse the shallow features and the deep features to obtain the final features;
[0088] S5. Input the final features into the image reconstruction module for image reconstruction to obtain a high-resolution laryngoscope image .
[0089] On the other hand, the present invention also provides a device for enhancing the super-resolution of laryngoscope images based on deep learning. This device includes a processor and a memory. The processor calls the computer program stored in the memory to execute the method for enhancing the super-resolution of laryngoscope images based on deep learning as described above.
[0090] Compared with the prior art, the present solution has the following beneficial effects:
[0091] (1) A significant improvement in image resolution and details
[0092] Through an improved deep learning network structure, this solution achieves high-quality reconstruction of low-resolution laryngoscope images, enabling the generated high-resolution images to clearly present the detailed information of tissues and lesions. In particular, the texture and edge details of tiny lesions are effectively enhanced.
[0093] (2) Effectively reduce artifacts and noise interference and improve the visual quality of images
[0094] Traditional super-resolution methods often introduce artifacts and noise when magnifying images, affecting the diagnostic effect. By introducing a specific loss control mechanism in the network design, this solution significantly reduces the artifacts and noise interference in the reconstruction process, making the generated images more natural and the details more realistic.
[0095] (3) Improve the efficiency and stability of image processing
[0096] Through the optimization of the deep learning model, this solution enables the model to have higher computational efficiency while maintaining high-quality reconstruction, meeting the requirements of clinical real-time applications. At the same time, the stability of the model during training is improved, and it can maintain consistent reconstruction effects in different image samples and enhancement tasks, ensuring the quality consistency and diagnostic reliability of the generated images.
[0097] (4) Provide a convenient and efficient tool for medical image data analysis
[0098] This solution also has high adaptability and can be integrated into existing laryngoscope devices and image analysis systems, with simple and convenient operation. In addition, this solution can be extended to other medical imaging fields, such as endoscopes, ophthalmoscopes, and other low-resolution medical imaging scenarios, with great application potential. BRIEF DESCRIPTION OF THE DRAWINGS
[0099] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0100] Figure 1 It is a schematic diagram of the network architecture of the embodiment of the present invention;
[0101] Figure 2 It is a structural diagram of the shallow feature extraction module of the embodiment of the present invention;
[0102] Figure 3 It is a structural diagram of the MRDA unit of the embodiment of the present invention;
[0103] Figure 4The structural diagram of the Channel Attention Block (CAB) according to an embodiment of the present invention;
[0104] Figure 5 The structural diagram of the Squeeze-and-Excitation layer (SE layer) according to an embodiment of the present invention;
[0105] Figure 6 The structural diagram of the Spatial Attention Block (SAB) according to an embodiment of the present invention;
[0106] Figure 7 The structural diagram of the image reconstruction module according to an embodiment of the present invention. Detailed implementation manners
[0107] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be clear that the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0108] Those skilled in the art should be aware that the following specific embodiments or specific implementation manners are a series of optimized setting manners listed by the present invention to further explain the specific invention content, and these setting manners can be combined with each other or used in association with each other, unless it is clearly stated in the present invention that some or a specific embodiment or implementation manner cannot be associated or used jointly with other embodiments or implementation manners. At the same time, the following specific embodiments or implementation manners are only used as the optimized setting manners, rather than being used to limit the understanding of the protection scope of the present invention.
[0109] The objective of the present invention is to solve the deficiencies in the resolution and detail presentation of existing laryngoscope images, provide higher-quality image information for clinical use, and assist doctors in accurately observing and diagnosing laryngeal lesions. Due to the limitations in hardware imaging of laryngoscope devices, the obtained images often have low resolution and blurred details. Especially when observing tiny lesions, the image clarity directly affects the doctor's judgment, which may lead to missed diagnoses or misdiagnoses. Therefore, improving the clarity and detail restoration of laryngoscope images is an important requirement in this field. Based on deep learning super-resolution technology, the present invention proposes a novel laryngoscope image super-resolution improvement scheme, aiming to achieve the following technical objectives:
[0110] (1) Enhance the image resolution and improve the presentation of lesion details; (2) Reduce artifact and noise interference and improve the image quality; (3) Improve the sensitivity to lesions and highlight the features of medical images; (4) Have a stable and efficient training process for convenient clinical application.
[0111] The following elaborates on this solution in detail with specific embodiments: Refer to Figure 1As shown, the laryngoscope image super-resolution enhancement system proposed in this scheme includes three main parts: a shallow feature extraction module, a deep feature extraction module and an image reconstruction module.
[0112] The shallow feature extraction module is used to extract shallow feature data from the low-resolution laryngoscope image, and input the obtained shallow features into the deep feature extraction module through multi-scale feature fusion.
[0113] The deep feature extraction module is used to capture deep features from shallow features, generate richer feature information by fusing shallow features with captured deep features, form final features, and input these final feature data into the image reconstruction module;
[0114] The deep feature extraction module includes two paths, one of which is a deep feature extraction unit and an enhanced convolutional attention unit connected in sequence, wherein the deep feature extraction unit is provided with multiple parallel paths; the other is composed of a graph convolutional neural network unit and two parallel causal convolutional units, wherein the two parallel causal convolutional units are both connected to the graph convolutional neural network unit, wherein the graph convolutional neural network unit receives shallow features, and the two parallel causal convolutional units are both composed of a causal convolutional block and an activation function. After the features obtained by the two causal convolutional units are multiplied, they are fused with the output of the enhanced convolutional attention unit (i.e., concat operation), and then fused with the shallow features.
[0115] The image reconstruction module is used to reconstruct a complete image from the learned final feature data, thereby outputting a high-resolution laryngoscope image. It should be noted here that the above-mentioned feature fusion can be performed by, for example, a concat operation, which is a common method in the art and will not be described here.
[0116] The following is an elaboration of the structure of each key module.
[0117] 1. Combination Figure 2 As shown, the shallow feature extraction module is used to extract shallow feature data from low-resolution laryngoscope images. The module uses 1×1 convolution + ReLU activation layer, 3×3 convolution + ReLU activation layer and 5×5 convolution + ReLU activation layer to capture primary features, that is, the activation layers corresponding to the above three convolution kernels are arranged in parallel. The shallow features are obtained by fusing the primary features extracted at three scales and then passing through a 1x1 convolution, and the obtained shallow features are input into the deep feature extraction module. Here, the fusion operation of the primary features can be performed by, for example, the concat operation, that is, the primary features output by each channel are spliced.
[0118] 2. The deep feature extraction module consists of a deep feature extraction unit based on multi-receptive fields and dilated attention (Multi-Receptive Fields and Dilated Attention block, MRDA) and an enhanced convolutional block attention module (Enhanced Convolutional Block Attention Module, eCBAM).
[0119] As shown in Figure 3 , in a preferred real-time manner, five parallel paths are set between the input and output of the deep feature extraction unit (i.e., the MRDA unit). The first path passes through a 1×1 convolutional layer; the second path passes through a 1×1 convolutional layer and a 3×3 convolutional layer in sequence; the third path passes through a 1×1 convolutional layer, a 3×3 convolutional layer, and a dilated convolutional layer with a dilation value of 3 and a kernel size of 3×3 in sequence; the fourth path passes through a 1×1 convolutional layer, a 5×5 convolutional layer, and a dilated convolutional layer with a dilation value of 5 and a kernel size of 3×3 in sequence; the fifth path is directly connected to the input. After the outputs of the first, second, third, and fourth paths are fused (for example, fused using a concat operation), they pass through a 1×1 convolutional layer, and then are added to the output of the fifth path to obtain a fused result, which is input to the eCBAM unit. It should be noted here that in this embodiment, 5 paths are set for the deep feature extraction unit. Of course, those skilled in the art can appropriately adjust the number of paths and the kernel sizes of the corresponding paths based on the actual image processing requirements and the data conditions of the image objects. For example, the number of paths can be adjusted to 3, 2, etc. For example, in the case of setting 3 paths, the first path can pass through a 1×1 convolutional layer, a 3×3 convolutional layer, and a dilated convolutional layer with a dilation value of 3 and a kernel size of 3×3 in sequence, the second path can pass through a 1×1 convolutional layer, a 5×5 convolutional layer, and a dilated convolutional layer with a dilation value of 5 and a kernel size of 3×3 in sequence, and the third path can be directly connected to the input, etc. These setting methods should all be regarded as falling within the protection scope of the present invention.
[0120] The enhanced convolutional block attention module (i.e., the eCBAM unit) consists of two sub-modules: a channel attention block (Channel Attention Block, CAB) and a spatial attention block (Spatial Attention Block, SAB). As shown in Figure 4 , Figure 5 , Figure 6As shown, the input of the eCBAM unit is one-dimensionally convolved by the CAB unit, and the convolution result is multiplied by the output result of the MRDA unit and then sent as the input to the SAB unit for two-dimensional convolution. Then, the result of the two-dimensional convolution is multiplied by the output result of the CAB unit to obtain the final output. Finally, this output is sent to the image reconstruction part. It should be noted here that: two paths are set inside the CAB unit itself, and the fused result (for example, fused by using the concat operation) after the CAB two paths are processed is multiplied by the output result of the MRDA unit and then input to the SAB unit; similarly, two paths are also set inside the SAB unit, and the output results of the two paths of the SAB are fused (for example, fused by using the concat operation) and then multiplied by the input of the SAB unit (that is, the result after the CAB unit result is multiplied by the MRDA unit output) to obtain the output.
[0121] Combined with Figure 4 As shown, in this embodiment, two paths are set between the input and output of the channel attention block (i.e., CAB). The first path sequentially passes through a 1×1 convolutional layer, a 3×3 convolutional layer, a 1×1 convolutional layer, a squeeze and excitation layer (Squeeze and Excitation Block, SE), a 1×1 convolutional layer, and a batch normalization operation (Batch Normalization, BN)+ReLU activation layer; the second path is directly connected to the input. Finally, the output of the first path is added to the output of the second path to obtain the final output of the channel attention block.
[0122] More preferably, combined with Figure 5 As shown, two paths are set between the input and output of the squeeze and excitation layer (SE layer). The first path sequentially passes through a global pooling layer, a first fully connected layer (FullyConnected Layer, FC), a ReLU activation layer, a second fully connected layer, and a Sigmoid activation layer. Among them, the input feature map size of the global pooling layer is H×W×C, and through the global pooling operation, a tensor of size 1×1×C is obtained; then, the features are reduced in dimension to , where r is a scaling factor used to reduce the number of parameters and the amount of calculation; the dimension-reduced features pass through the ReLU activation layer, and the tensor size remains , then restore the feature dimension to 1×1×C through the second fully connected layer; finally, limit the output within the range of [0,1] through the Sigmoid activation layer to obtain a tensor of size 1×1×C, which represents the importance weights of each channel. The second path is directly connected to the input. Finally, multiply each channel output by the corresponding weight of the second path element-wise, that is, multiply the output of the first path by the output of the second path element-wise to adjust the feature response of each channel. This process can enhance the features of important channels while suppressing the features of unimportant channels, thus obtaining the final output of the SE layer.
[0123] Combined with Figure 6 As shown, in this embodiment, there are two paths between the input and output of the spatial attention block (i.e., SAB). The first path sequentially passes through a 1×1 convolutional layer, a strided convolutional layer (Strided Conv), a pooling layer (Pooling), a grouped convolutional layer (Conv Groups), an upsampling layer (Upsampling), a 1×1 convolutional layer, and a Sigmoid activation layer. Among them, the strided convolutional layer (Strided Conv) can reduce the spatial dimension of the feature map (i.e., downsampling) while extracting more abstract features; the pooling layer (Pooling) further reduces the spatial dimension of the feature map and aggregates global features; the grouped convolutional layer (Conv Groups) divides the channels of the feature map into several groups, and each group performs a convolutional operation separately to reduce the computational amount and capture local spatial information; the upsampling layer (Upsampling) restores the spatial resolution of the feature map to match the size of the original input feature map; the Sigmoid activation layer limits the result within the range of [0,1] to generate spatial attention weights. Finally, multiply the generated spatial attention weights by the output result of the CAB unit element-wise to highlight important spatial positions and suppress unimportant parts, thus obtaining the final output of the spatial attention block.
[0124] Combined again with Figure 1 As shown, in the deep feature extraction module, another branch is mainly composed of a graph convolutional neural network unit and two parallel causal convolutional units. The two parallel causal convolutional units are both connected to the graph convolutional neural network unit. The graph convolutional neural network unit receives shallow features, and both of the two parallel causal convolutional units are composed of a causal convolutional block and an activation function. The following elaborates on the structural design of this branch.
[0125] For a graph output by a shallow feature extraction module , where represents the vertex set, that is, the pixels in the image, represents the edge set, is the adjacency matrix. Use the Laplacian operator , establish the dependency relationship between pixels, where is the degree matrix, representing the degree of each node. Then normalize the Laplacian matrix:
[0126] ;
[0127] where is the identity matrix. Subsequently, introduce a convolutional kernel , where . The convolution operation can be expressed as: , where is the input signal, is 's eigenvector matrix, is 's eigenvalue, can be regarded as a function of the eigenvalue . To simplify the calculation, use Chebyshev polynomials to fit :
[0128] ;
[0129] where is the Chebyshev polynomial.
[0130] ;
[0131] where ; Based on the above formula, the calculation method of the graph convolutional neural network is as follows:
[0132]
[0133] where and respectively represent the values of the -th layer and the -th layer, represents the activation function, represents the weight matrix. Through the above rules, the graph convolutional neural network can aggregate the features of adjacent nodes, thereby establishing long-range dependency relationships between different nodes and significantly improving the ability of pixel points to combine spatial context information.
[0134] The causal convolution part is implemented by right alignment to ensure that the convolution only acts on the current time point, focuses on past states, and prevents any leakage of information from future states. In each convolutional layer, a gated activation unit is adopted, which consists of two non-linear functions: a Sigmoid function and a Tanh function. This unit dynamically adjusts the sequence as follows:
[0135] ;
[0136] where * represents the convolution operation. The Tanh function is defined as , and the Sigmoid function is defined as , where the Tanh term is used as a common activation function to constrain the output value within the range of [-1, 1], and the Sigmoid term is used as a gating mechanism to dynamically activate or deactivate unnecessary time points.
[0137] 3. In this embodiment, as shown in Figure 7 , the image reconstruction module is successively composed of a first Conv unit, a sub-pixel convolution layer, and a second Conv unit. Among them, the Conv unit is composed of a 1×1 convolution layer and a ReLU activation layer. After being processed by the image reconstruction module, a reconstructed high-resolution laryngoscope image is obtained .
[0138] 4. In a more preferred embodiment, the present solution can also perform enhancement processing on the original input low-resolution laryngoscope image in advance to facilitate subsequent further identification of image features and improve the effect of image feature recognition and reconstruction. In this setting mode, the present solution further includes an image preprocessing module, which is used to preprocess the low-resolution laryngoscope image, enhance the local blurred area in the image, obtain a preprocessed laryngoscope image, and input the obtained preprocessed laryngoscope image into the shallow feature extraction module.
[0139] Preferably, the image preprocessing module preprocesses the low-resolution laryngoscope image based on wavelet transform, and the specific method is as follows:
[0140] Construct a wavelet basis function and set the constraint conditions of the wavelet basis function;
[0141] Calculate the square of the wavelet coefficients and obtain them in ascending order , where s represents the square value of the corresponding wavelet coefficient, and N represents the number of wavelet coefficients;
[0142] Calculate the risk value and threshold selection parameter of each wavelet coefficient, and retain the wavelet coefficients whose absolute value of the risk value is greater than the threshold selection parameter to obtain an optimized wavelet transform function;
[0143] Apply the optimized wavelet transform function to the low-resolution laryngoscope image for wavelet transform to obtain a preprocessed laryngoscope image.
[0144] Preferably, the constructed wavelet basis function is:
[0145] ;
[0146] Among them, \(t\) represents the time variable of wavelet transform.
[0147] Among them, the set kernel function of wavelet transform is:
[0148] ;
[0149] Among them, is the scale factor, is the translation factor, . \(R\) represents the set of real numbers.
[0150] Preferably, the set constraint conditions of wavelet transform are:
[0151] ;
[0152] Among them, represents in the form of Fourier transform, \(R\) represents the set of real numbers, that is, it means that the calculations and operations in wavelet transform are carried out in the real number domain, and \(L()\) represents wavelet transform.
[0153] After that, the wavelet transform function is optimized by the risk value method. Calculate the square of the wavelet coefficients and sort them from small to large to get . According to the number of wavelet coefficients , calculate the risk value:
[0154] ;
[0155] Further calculate the threshold selection parameter, and the value of the threshold selection parameter is:
[0156] ;
[0157] Among them, represents the mean square error, that is, calculate the mean square error of the wavelet coefficient as the screening condition of the risk value . If , then the wavelet coefficient is retained, otherwise it is deleted. Thus, the optimization of the wavelet transform function is completed. Use the optimized wavelet transform function to process the low-resolution laryngoscope image to improve the overall quality of the image.
[0158] In addition, it can be further preferably that after wavelet transform, further screen whether there is a blurred area, and edge enhancement can be performed on the blurred area. The edge information can be enhanced by calculating the average gradient of the blurred area:
[0159] ;
[0160] Among them, the size of the blurred area is \(M\times N\), Indicates the pixel gradient of the blurred area image. After calculation, the average gradient G is used to replace the pixel values of the original area to achieve the effect of edge enhancement. It should be understood here that gradient enhancement is an optional additional method, and it is also possible not to perform gradient enhancement and only use the above-mentioned optimized wavelet transform for processing.
[0161] Through the above calculations, after the wavelet transform is improved, the edge information of the local blurred area of the image can be accurately extracted, and the improved wavelet transform is used to effectively complete the enhancement of local image blurring and improve the overall quality of the image. Subsequently, the result is input into the shallow feature extraction module.
[0162] In addition, it is further explained that the blurred area can be, for example, an area where the image is unclear caused by excessive local exposure, etc. The blurred area can be delimited by manual division or other means. Or, the blurred area is determined by calculating the average contrast and image quality index of the image area. For example, when the average contrast and image quality index are respectively less than the preset thresholds, it is defined as a blurred area. When screening the blurred area, the image can be divided into areas of a certain unit area and processed block by block. For example, it can be divided into blocks with a size containing T pixels. The following further explains with a specific example of calculating the average contrast and image quality index:
[0163] Preferably, the average contrast is used to measure the brightness difference between different regions in the image, that is, the higher the average contrast, the more obvious the brightness difference, and the better the graphic quality (clarity). The calculation method can be set as:
[0164] ;
[0165] Among them, represents the average value of the contrast, represents the intensity of the local blurred variance of the image.
[0166] Preferably, the image quality index is the average degree of pixel brightness change within the image area, used to evaluate the quality of the image and guide the enhancement of the local blurred area of the image. The higher the value, the higher the image quality (clarity). The calculation method is:
[0167] ;
[0168] Among them, x represents the gray value of the pixel in the original low-resolution laryngoscope image, y represents the gray value of the pixel in the laryngoscope image after wavelet transform, and are respectively and the average values, and are respectively and the second moments of is and the covariance of. The specific calculation formulas for each parameter are as follows:
[0169] ;
[0170] ;
[0171] ;
[0172] ;
[0173] ;
[0174] This solution improves the instability of traditional deep learning models during training, solves the problems of difficult training and unstable effects, enabling the model to maintain a high generation quality while being efficiently trained. This improvement helps to enhance the operability and stability of the model in practical applications, providing reliable technical support for clinical promotion. In addition, this solution has good adaptability and can achieve high-quality improvement of images under different devices and environments, with good generalization ability.
[0175] In summary, through the innovative deep learning super-resolution enhancement method, the present invention realizes the high-resolution reconstruction of laryngoscope images, aiming to enhance image details and clarity, reduce noise and artifact interference, enhance the sensitivity to lesions, and improve the stability and practicality of the model, thus assisting doctors in more accurately diagnosing and treating laryngeal diseases.
[0176] In another embodiment, based on the super-resolution enhancement system already disclosed in the above embodiment, the method for enhancing and reconstructing laryngoscope images in this solution is as follows:
[0177] S1. Obtain a low-resolution laryngoscope image , perform image preprocessing on the low-resolution laryngoscope image, and then input the obtained image into shallow feature extraction;
[0178] S2. Extract shallow features from the low-resolution laryngoscope image through the shallow feature extraction module;
[0179] S3. Input the shallow features into the deep feature extraction module to extract deep features;
[0180] S4. Fuse the shallow features and the deep features (for example, use the concat operation for fusion) to obtain the final features;
[0181] S5. Input the final features into the image reconstruction module for image reconstruction to obtain a high-resolution laryngoscope image .
[0182] In another implementation manner of this solution, it can be implemented in the form of a device. The device may include corresponding modules for executing each or several steps in the above-mentioned various implementation manners. Therefore, each step or several steps in the above-mentioned various implementation manners can be executed by the corresponding modules, and the electronic device may include one or more of these modules. The module may be one or more hardware modules specifically configured to execute the corresponding steps, or implemented by a processor configured to execute the corresponding steps, or stored in a computer-readable medium for implementation by the processor, or implemented through a certain combination. The device can be implemented using a bus architecture.
[0183] Any process or method description shown in the flowchart or described in other ways herein can be understood as representing a module, segment, or part of code including one or more executable instructions for implementing a specific logical function or process. The scope of the preferred implementation manner of this solution includes additional implementations, where the functions may be executed in a substantially simultaneous manner or in the reverse order according to the functions involved, rather than in the order shown or discussed. This should be understood by those skilled in the technical field to which the implementation manner of this solution belongs. The processor executes the various methods and processes described above. For example, the method implementation manner in this solution can be implemented as a software program tangibly contained in a machine-readable medium, such as a memory. In some implementation manners, part or all of the software program can be loaded and / or installed via the memory and / or communication interface. When the software program is loaded into the memory and executed by the processor, one or more steps in the method described above can be executed. Alternatively, in other implementation manners, the processor can be configured to execute one of the above methods in any other appropriate way (for example, by means of firmware).
[0184] As described above, the above are only specific implementation manners of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the technical field to which the present invention pertains within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A laryngoscope image super-resolution enhancement system based on deep learning, characterized in that: The system comprises: an image preprocessing module, a shallow feature extraction module, a deep feature extraction module and an image reconstruction module; The image preprocessing module is used to preprocess the low-resolution laryngoscope image, enhance the local blurred area in the image, obtain a preprocessed laryngoscope image, and input the obtained preprocessed laryngoscope image into the shallow feature extraction module; The shallow feature extraction module is used to extract shallow features from the preprocessed laryngoscope image, and input the obtained shallow features into the deep feature extraction module through a multi-scale feature concat operation; The deep feature extraction module is used to capture deep features from shallow features, generate final features by fusing the shallow features with the captured deep features, and input these final features into the image reconstruction module; The image reconstruction module is used to reconstruct a complete image from the learned final features, thereby outputting a high-resolution laryngoscope image; The deep feature extraction module includes two paths, one of which is a deep feature extraction unit and an enhanced convolutional attention unit connected in sequence, wherein the deep feature extraction unit is provided with a plurality of parallel paths; the other is composed of a graph convolutional neural network unit and two parallel causal convolutional units, wherein the two parallel causal convolutional units are both connected to the graph convolutional neural network unit, the graph convolutional neural network unit receives shallow features, and the two parallel causal convolutional units are both composed of a causal convolutional block and an activation function; The deep feature extraction unit includes five parallel paths, a feature fusion unit and a convolution layer with a convolution kernel of 1×1; The five parallel paths are set as follows: the first path is a 1×1 convolution layer; the second path includes a 1×1 convolution layer and a 3×3 convolution layer connected in sequence; the third path includes a 1×1 convolution layer, a 3×3 convolution layer, and an expansion convolution layer with a dilation value of 3 and a convolution kernel size of 3×3 connected in sequence; the fourth path includes a 1×1 convolution layer, a 5×5 convolution layer, and an expansion convolution layer with a dilation value of 5 and a convolution kernel size of 3×3 connected in sequence; the fifth path is directly connected to the input; The feature fusion unit fuses the outputs of the first path, the second path, the third path, and the fourth path, and inputs the outputs to a convolution layer with a convolution kernel of 1×1; the output of the convolution layer with a convolution kernel of 1×1 is added to the output of the fifth path; The enhanced convolutional attention unit includes a channel attention block and a spatial attention block connected in sequence; the channel attention block convolves the shallow features of the input enhanced convolutional attention unit, multiplies them with the shallow features, and then inputs them into the spatial attention block; the input of the spatial attention block is multiplied by the output of the spatial attention block to obtain the final features output by the deep feature extraction module.
2. The system according to claim 1, characterized in that The shallow feature extraction module includes: three parallel 1×1 convolution and ReLU activation layers, 3×3 convolution and ReLU activation layers, and 5×5 convolution and ReLU activation layers; a primary feature fusion unit, and a 1×1 convolution layer connected to the primary feature fusion unit; The preprocessed laryngoscope image is respectively input into a 1×1 convolution and ReLU activation layer, a 3×3 convolution and ReLU activation layer, and a 5×5 convolution and ReLU activation layer arranged in parallel to obtain three primary features respectively; the three primary features are fused through a primary feature fusion unit, and then through a 1×1 convolution layer to obtain shallow features.
3. The system according to claim 1, characterized in that The image preprocessing module preprocesses the low-resolution laryngoscope image based on wavelet transform, specifically in the following way: Construct wavelet basis functions and set constraints on the wavelet basis functions; Calculate the square of the wavelet coefficients and get them in order from small to large , where s represents the square value of the corresponding wavelet coefficient, and N represents the number of wavelet coefficients; Calculate the risk value and threshold selection parameter of each wavelet coefficient, and retain the wavelet coefficient whose absolute value of risk value is greater than the threshold selection parameter to obtain the optimized wavelet transform function; The low-resolution laryngoscope image is subjected to wavelet transform using the optimized wavelet transform function to obtain the preprocessed laryngoscope image.
4. The system according to claim 1, characterized in that The channel attention block is provided with two paths: the first path includes a 1×1 convolution layer, a 3×3 convolution layer, a 1×1 convolution layer, a compression and excitation layer, a 1×1 convolution layer, a batch normalization layer and a ReLU activation layer connected in sequence; the second path is directly connected to the input of the channel attention block; The outputs of the first and second paths are summed to obtain the output of the channel attention block.
5. The system according to claim 1, characterized in that The spatial attention block is provided with two paths: the first path includes a 1×1 convolution layer, a stride convolution layer, a pooling layer, a grouped convolution layer, an upsampling layer, a 1×1 convolution layer and a sigmoid activation layer connected in sequence; the second path is directly connected to the input of the spatial attention block; The outputs of the first and second pathways are summed to obtain the output of the spatial attention block.
6. A laryngoscope image super-resolution enhancement method based on deep learning, characterized in that: The method is applied to the system according to any one of claims 1 to 5, and the method comprises: S1, obtaining a low-resolution laryngoscope image, performing image preprocessing on the low-resolution laryngoscope image to obtain a preprocessed laryngoscope image, and inputting the preprocessed laryngoscope image into shallow feature extraction; S2, through the shallow feature extraction module, the shallow features of the pre-processed laryngoscope image are proposed; S3, inputting shallow features into a deep feature extraction module to extract deep features; S4, fuse shallow features with deep features to obtain the final features; S5. The final features are input into an image reconstruction module to perform image reconstruction to obtain a high-resolution laryngoscope image.
7. A laryngoscope image super-resolution enhancement device based on deep learning, characterized in that: The device includes a processor and a memory, and the processor calls a computer program stored in the memory to execute the deep learning-based laryngoscope image super-resolution enhancement method as described in claim 6.
Citation Information
Patent Citations
Mine image super-resolution reconstruction method and system based on multi-scale residual network
CN113592718A
Image super-resolution reconstruction method, device and equipment
CN115564649A