Deep learning based laryngeal structure and detail high-resolution image enhancement method
By employing deep learning methods and combining ResNet and U-Net++ models, and introducing anatomical attention and dynamic receptive field adjustment, the problems of feature weakening, motion artifacts, and loss of stereo correlation in laryngeal image processing were solved, achieving high-resolution enhancement of laryngeal images and improving diagnostic accuracy.
Patent Information
- Application Number
- CN202511142653.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-08-15
AI Technical Summary
Traditional laryngeal image processing techniques have shortcomings in feature capture, motion artifacts, loss of stereo correlation information, and super-resolution reconstruction accuracy, which affect the accuracy and reliability of diagnosis.
A deep learning-based approach is adopted, which combines ResNet and U-Net++ models with subpixel convolutional layers, introduces an anatomical attention module and dynamic receptive field adjustment technology, constructs a time-varying weight matrix, performs feature extraction and super-resolution reconstruction, enhances the feature representation of key regions and compensates for organ displacement.
It effectively enhances the characteristic features of key areas of the larynx, reduces motion artifacts, preserves three-dimensional spatial information, and improves the accuracy and reliability of diagnosis.
Smart Images

Figure CN120746872B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image enhancement technology, and more specifically, to a method for high-resolution image enhancement of throat structure and details based on deep learning. Background Technology
[0002] Traditional models struggle to effectively capture minute yet crucial features, such as microvessels in areas prone to cancer. Furthermore, the significant displacement of the larynx during respiration or vocalization, a dynamic organ, leads to motion artifacts in CT images, severely impacting the accuracy of clinical assessments such as tumor invasion depth. Simultaneously, standard isotropic 3D convolutional kernels result in feature loss or blurring at specific viewpoints. Conventional 3D feature extraction ignores the stereoscopic relationship between the laryngeal ventricle, vocal cords, and trachea, failing to support effective assessment of dynamic functional impairments such as vocal cord paralysis. In super-resolution reconstruction, existing techniques are insufficient in representing high-resolution image details, particularly edges and critical structures, and have limitations in the clarity and contrast of mucosal regions, all of which restrict diagnostic accuracy and reliability. Therefore, this paper proposes a deep learning-based method for high-resolution image enhancement of laryngeal structure and details. Summary of the Invention
[0003] The purpose of this invention is to provide a high-resolution image enhancement method for throat structure and details based on deep learning, so as to solve the problems of traditional throat image processing techniques mentioned in the background art in terms of feature capture, motion artifacts, loss of stereo correlation information and super-resolution reconstruction accuracy.
[0004] To achieve the above objectives, the present invention aims to provide a high-resolution image enhancement method for throat structure and details based on deep learning, comprising the following steps:
[0005] S1. Acquire low-resolution images from the laryngeal endoscope and corresponding three-dimensional image data of the larynx;
[0006] S2. Based on the ResNet model, features are extracted from low-resolution images. The 3D-ResNet model is used to extract features from three-dimensional image data. A time-varying weight matrix with sinusoidal periodic modulation is constructed in the 3D-ResNet model so that the image enhancement process can perceive the physiological motion state of the larynx in real time. The extracted features are then fused using cross-scale feature pyramid technology.
[0007] S3. Using the U-Net++ model that incorporates sub-pixel convolutional layers, super-resolution reconstruction is performed on images that have undergone feature fusion processing.
[0008] S4. The U-Net segmentation model based on deep learning performs targeted local enhancement of the mucous membrane region in the image.
[0009] As a further improvement to this technical solution, in step S2, features are extracted from low-resolution images based on the ResNet model, including the following steps:
[0010] S2.1 Use a pre-trained laryngeal key point detection model to automatically identify and mark key anatomical landmarks in the larynx, and generate a structural heatmap based on these landmarks;
[0011] S2.2 Adaptive gamma correction for the clinical region;
[0012] S2.3. Input the image processed in step S2.2 into the improved ResNet model that incorporates the anatomical attention module to begin the feature extraction process;
[0013] S2.4 Perform the first stage of ResNet feature extraction to obtain preliminary feature maps;
[0014] S2.5. Combine the feature map of the current stage with the structural thermogram generated in step S2.1. Figure 1 The input is fed into the anatomy attention module, where the feature information from both is combined through convolution operations to generate attention weights. These attention weights are then applied to weight the feature map.
[0015] S2.6. Dynamically adjust the convolution kernel parameters based on the spatial frequency of the current feature map;
[0016] S2.7 Repeat the above steps to obtain multi-level feature output.
[0017] As a further improvement to this technical solution, in step S2, features are extracted from three-dimensional image data based on the 3D-ResNet model, including the following steps:
[0018] S2.8, Collect data from the laryngeal motion sensor;
[0019] S2.9 Construct a time-varying weight matrix based on the acquired laryngeal motion sensor data;
[0020] S2.10. Decompose the 3D convolutional kernel of the 3D-ResNet model into convolutional kernels for specific directions of the throat, use a gating network to assign weights to the feature maps extracted from specific directions of the throat, and fuse the specific directions of the throat according to the weights.
[0021] S2.11 The feature maps obtained through the above steps in different directions are fed into the dynamic gating unit for weighted fusion to generate a comprehensive feature representation.
[0022] As a further improvement to this technical solution, in step S2.9, constructing a time-varying weight matrix based on the acquired throat motion sensor data includes the following steps:
[0023] S2.91. Extract key parameters from throat motion sensor data;
[0024] S2.92. Based on the above key parameters, classify the state and use the feature threshold judgment strategy to classify the current moment into one of the three states: vocalization phase, inhalation phase, and resting phase.
[0025] S2.93. Introduce the segmented organ mask in the baseline 3D image, and set the response weight factor of the organ structure according to the physiological state of different organ structures in each time phase.
[0026] S2.94. In a voxel space with the same resolution as the input image, initialize a three-dimensional voxel matrix and assign values to each organ structure voxel according to the organ mask information.
[0027] S2.95. Based on the physiological state of the current frame, apply a phase modulation factor to the weight of each organ voxel and introduce a periodic modulation term in the form of a sine function.
[0028] S2.96. Based on the vertical displacement and impedance changes, the 3D local displacement field at the current moment is reconstructed. To prevent artifacts from weight jumps, the modulated weight matrix is spatially smoothed.
[0029] S2.97, Output the dynamic time-varying weight matrix after spatial smoothing.
[0030] As a further improvement to this technical solution, in step S2, the extracted features are fused using cross-scale feature pyramid technology, including the following steps:
[0031] S2.12. Use the low-resolution 2D endoscopic image and 3D laryngeal image data obtained through feature extraction as input;
[0032] S2.13. For each scale, the feature map of the endoscope image is fused with the feature map of the three-dimensional image by stitching.
[0033] S2.14 Construct a feature pyramid based on the top-down path;
[0034] S2.15, Output cross-scale fusion feature pyramid.
[0035] As a further improvement to this technical solution, in step S3, the U-Net++ model, which incorporates sub-pixel convolutional layers, is used to perform super-resolution reconstruction on the image that has undergone feature fusion processing, including the following steps:
[0036] S3.1. The fused feature map output by the cross-scale feature pyramid is used as input, and the fused feature map is processed by a deep learning-based semantic segmentation network to generate a vocal cord mask.
[0037] S3.2 Construct the encoder part of the U-Net++ model. In the encoder part, feature extraction is performed sequentially through three encoding blocks and a bottleneck layer.
[0038] S3.3 Add a vocal cord attention gate module after each coding block to enhance the feature regions related to the free edge of the vocal cords by fusing vocal cord masks;
[0039] S3.4. The high-dimensional representation output by the encoder is fed into the bottleneck layer, and upsampling is performed step by step using multi-level subpixel convolution. A dynamic kernel strategy is introduced during the subpixel convolution process.
[0040] S3.5. The dense skip connection structure of the U-Net++ model is adopted to connect the feature map of each level in the encoder with the feature map of the corresponding scale in the decoder. A structure-aware weighting mechanism is introduced into the skip connection path of the U-Net++ model to assign weights to the skip connections according to the structural importance.
[0041] S3.6. Add an edge-guided loss term, use the Sobel operator to extract the vocal cord edge gradient, and use L1 loss for supervision.
[0042] As a further improvement to this technical solution, in step S3.4, multi-level sub-pixel convolution is used to perform upsampling step by step, and a dynamic kernel strategy is introduced during the sub-pixel convolution process, including the following steps:
[0043] S3.41. Initialize an empty list to store the upsampled feature map of each level's output, set the output feature map of the bottleneck layer of the U-Net++ model as the initial input, and set the number of upsampling times.
[0044] S3.42. Perform channel compression on the input feature map using 1×1 convolution;
[0045] S3.43 Calculate the gradient magnitude of the compressed feature map and count the proportion of high-frequency regions;
[0046] S3.44. Select the convolution kernel based on the high frequency ratio dynamic kernel, and use the selected convolution kernel to convolve the input feature map;
[0047] S3.45. Perform channel expansion convolution and use the PixelShuffle operation to upsample to obtain the upsampled feature map;
[0048] S3.46. Obtain the vocal cord mask at the current resolution and calculate the average gradient intensity of the corresponding vocal cord mask region in the upsampled feature map.
[0049] S3.47. Fuse the upsampled feature map with the same-scale encoded feature map at the next level;
[0050] S3.48. Save the feature map of the current upsampling output to an empty list and use the feature map as the input for the next stage of upsampling;
[0051] S3.49. Repeat steps S3.42 to S3.48 until all upsampling layers are completed, and finally output a high-resolution feature map.
[0052] As a further improvement to this technical solution, in S3.5, a structure-aware weighting mechanism is introduced into the skip connection path of the U-Net++ model, assigning weights to skip connections according to structural importance, including the following steps:
[0053] S3.51. Downsample the vocal cord mask to the same spatial size as the skip connection feature map to obtain the structural importance map;
[0054] S3.52. Apply the structural importance map to the skip connection feature map and perform pixel-by-pixel channel weighting.
[0055] S3.53. Concatenate the weighted feature map with the feature map of the corresponding scale in the decoder.
[0056] As a further improvement to this technical solution, in step S4, an attention map is generated based on the anatomical structure, and targeted local enhancement is performed on the mucosal region in the image, including the following steps:
[0057] S4.1. Use the super-resolution reconstructed image as input;
[0058] S4.2. Use the U-Net model for segmenting the mucosal region of the throat image to process the input image and output a binary mask;
[0059] In this context, the foreground pixels represent the location of the mucosal region, while the background pixels represent the non-mucosal region.
[0060] S4.3 Apply a local enhancement network based on dynamic filtering mechanism and adaptive weight adjustment to the segmented mucosal region;
[0061] S4.4 Integrate the locally enhanced mucosal region back into the original super-resolution reconstructed image;
[0062] S4.5 Generate the final fused image.
[0063] As a further improvement to this technical solution, in step S4.3, a local enhancement network based on dynamic filtering mechanism and adaptive weight adjustment is applied to the segmented mucosal region, including the following steps:
[0064] S4.31. Input the segmented mucosal region image into the input layer of the local enhancement network;
[0065] S4.32. Extract multi-scale features of the mucosal region through multiple convolutional layers;
[0066] S4.33. Adaptively generate filter kernel weights based on the extracted features, and introduce an attention mechanism to calculate the weight distribution of key textures and lesion features in the mucosal region;
[0067] S4.34. Combine the structural information of the mucosal region to dynamically adjust the weight allocation of different channels and spatial locations;
[0068] S4.35. After filtering and weight adjustment, a locally enhanced image of the mucosal region is generated.
[0069] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0070] 1. This deep learning-based high-resolution image enhancement method for laryngeal structure and details incorporates an anatomical attention module and dynamic receptive field adjustment technology. This method effectively enhances the features of key clinical areas (such as the free edge of the vocal cords and the anterior commissure), overcoming the feature weakening and multi-scale detail loss problems encountered by traditional ResNet models when processing complex anatomical structures. Furthermore, it utilizes a U-Net++ model with sub-pixel convolutional layers for super-resolution reconstruction, emphasizing the reconstruction quality of edge regions. The addition of an edge-guided loss term further improves the clarity of key anatomical structures, thereby helping physicians more accurately assess lesions and improve diagnostic accuracy.
[0071] 2. This deep learning-based high-resolution image enhancement method for laryngeal structure and details addresses the issue of motion artifacts in CT images caused by significant larynx displacement during respiration or phonation. The method compensates for organ displacement by constructing a time-varying weight matrix and employs three-way separable convolution and a dynamic gating network to preserve the anisotropic features of the larynx, enhancing feature representations at key perspectives (including the vocal cord level). This approach not only reduces image blurring and ghosting caused by organ movement but also preserves the three-dimensional spatial contextual information of the larynx through multi-directional feature gating fusion, supporting the assessment of dynamic functional disorders such as vocal cord paralysis and providing more accurate and reliable imaging evidence for clinical practice. Attached Figure Description
[0072] Figure 1 This is a flowchart of the overall method of the present invention. Detailed Implementation
[0073] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0074] Example: Please refer to Figure 1 As shown, this embodiment provides a deep learning-based method for high-resolution image enhancement of throat structure and details, including the following steps:
[0075] S1. Acquire low-resolution images from the laryngeal endoscope and corresponding three-dimensional images (CT images) of the larynx;
[0076] S2. Based on the ResNet model, features are extracted from low-resolution images. The 3D-ResNet model is used to extract features from three-dimensional image data. A time-varying weight matrix with sinusoidal periodic modulation is constructed in the 3D-ResNet model so that the image enhancement process can perceive the physiological motion state of the larynx in real time. The extracted features are then fused using cross-scale feature pyramid technology.
[0077] In this embodiment, features are extracted from low-resolution images based on the ResNet model, including the following steps:
[0078] The ResNet model is a classic convolutional neural network structure that can efficiently extract multi-level features and achieve deep network training stability. Its input is a throat image (single-channel grayscale CT / MRI image, or 3-channel fused image), and its output is an image feature vector or feature map. The core structure of ResNet is the residual block, which uses skip connections to ensure that features are not easily lost during deep network transmission. Each residual block includes two 3×3 convolutional layers and a shortcut branch.
[0079] Traditional ResNet feature extraction suffers from two major problems: weakening of anatomical structural features and loss of multi-scale details. The laryngeal anatomy is complex (including the free edge of the vocal cords and the anterior commissure), and traditional models process all pixels equally, causing key features such as microvessels (<0.1mm) in high-risk cancer areas to be submerged by background noise. Fixed-size convolutional kernels cannot simultaneously capture microscopic (mucosal capillaries) and macroscopic (arytenoid cartilage) structures, requiring multiple models to be cascaded, increasing computational latency and disrupting feature consistency. By using anatomical attention mechanisms and dynamic receptive field adjustment, targeted enhancement of clinically critical areas can be achieved.
[0080] The larynx, a complex region of the human body, contains several important anatomical structures, such as the free edge of the vocal cords and the anterior commissure. These areas are particularly important for diagnosing diseases such as cancer. However, traditional ResNet models, due to their equal treatment of all pixels, often fail to effectively distinguish between background noise and truly clinically significant key features. By introducing an anatomical attention module, these key areas can be automatically identified and highlighted, making lesion features more apparent and facilitating accurate diagnosis by doctors. The anatomical structures of the larynx encompass multiple scales, from micrometer-scale mucosal capillaries to millimeter-scale arytenoid cartilage. Fixed-size convolutional kernels struggle to capture information at such a wide scale simultaneously, easily leading to the loss of important details. The anatomical attention mechanism, combined with dynamic receptive field adjustment technology, can flexibly change the size of the convolutional kernel according to the specific needs of different locations, thereby more effectively extracting and representing multi-scale features and improving the overall image quality.
[0081] S2.1 Use a pre-trained laryngeal key point detection model to automatically identify and mark key anatomical landmarks in the larynx. These landmarks include, but are not limited to, the free edge of the vocal cords and the anterior commissure of the larynx. Generate a structural heatmap based on these landmarks.
[0082] S2.2 Adaptive gamma correction is applied to clinical areas (abnormal mucosa or lesion sites) to enhance the contrast and detail in these areas;
[0083] S2.3. Input the image processed in step S2.2 into the improved ResNet model that incorporates the anatomical attention module to begin the feature extraction process;
[0084] S2.4 Perform the first stage of ResNet feature extraction to obtain preliminary feature maps;
[0085] S2.5. Combine the feature map of the current stage with the structural thermogram generated in step S2.1. Figure 1 The input is fed into the anatomical attention module, and the feature information of both is combined through convolution operation to generate attention weights between 0 and 1. The attention weights are applied to weight the feature map, highlighting the features of important regions while suppressing information of non-critical regions.
[0086] S2.6. Dynamically adjust the convolution kernel parameters according to the spatial frequency of the current feature map to effectively capture features at different scales (use a smaller convolution kernel (3×3) for high-frequency features, while dilated convolution or a larger convolution kernel (5×5) may be used for low-frequency features).
[0087] S2.7 Repeat the above steps (feature extraction → attention weighting → dynamic receptive field adjustment) to obtain multi-level feature outputs, covering different levels of information from microvascular to macrocartilage structure.
[0088] Furthermore, features are extracted from 3D image data based on the 3D-ResNet model, including the following steps:
[0089] The 3D-ResNet model is a three-dimensional extension of ResNet. All convolution and pooling operations are performed in the time-space three-dimensional domain, and each residual block contains two or three convolutional layers.
[0090] The larynx undergoes significant displacement during respiration / vocalization. Traditional 3D-ResNet treats dynamic organs as static voxels, resulting in motion artifacts on CT images (including blurred vocal cord edges and cartilage ghosting), which severely affects the assessment of tumor invasion depth. By dynamically adjusting voxel values through a temporal weight matrix, the data from the laryngeal dynamic endoscope sensor is synchronized with the three-dimensional images to compensate for organ displacement during phonation / inspiration and reduce motion artifacts.
[0091] Clinical CT slice thickness is much greater than intra-slice resolution. Standard 3D convolution kernel isotropic processing leads to the loss of features in the coronal plane (a key perspective for assessing vocal cord symmetry) and blurring of details in the sagittal plane (for observing epiglottic mobility). By using three-dimensional separable convolution to specifically extract anatomical plane features, and by using a dynamic gating network to automatically enhance the weights of key clinical directions.
[0092] Conventional 3D feature extraction ignores the stereoscopic relationship between the laryngeal ventricle, vocal cords, and trachea (such as the 3D rotation angle of the arytenoid cartilage), and cannot support the assessment of dynamic functional disorders such as vocal cord paralysis. Multi-directional feature gating fusion preserves the stereoscopic spatial context of the larynx.
[0093] S2.8 Collect laryngeal motion sensor data (data obtained through laryngeal dynamic endoscope or accelerometer), which can reflect the positional changes of the larynx in different states;
[0094] S2.9. Construct a time-varying weight matrix based on the acquired laryngeal motion sensor data (this matrix adjusts the CT value according to the three stages of inhalation, phonation, and rest), dynamically adjust the CT value in the three-dimensional image to compensate for organ displacement caused by breathing and phonation. For image data with a slice thickness greater than 1 mm, use depth-sensing interpolation technology to interpolate along the Z-axis to improve the spatial resolution of the image.
[0095] The process of constructing a time-varying weight matrix based on the acquired laryngeal motion sensor data includes the following steps:
[0096] The larynx undergoes significant structural changes during different physiological states, including phonation, inspiration, and rest. In particular, the position and morphology of tissues such as the vocal cords and epiglottis change dynamically over a short period. Ignoring these changes and using only static images or fixed attention areas for image enhancement can lead to structural mismatches, blurring, and edge artifacts, severely impacting the clarity and diagnostic value of key anatomical regions. Furthermore, maintaining temporal consistency during image enhancement is difficult, easily causing inter-frame jumps and unstable enhancement results. Introducing a time-varying weight matrix based on laryngeal motion sensor data allows the image enhancement process to perceive the physiological motion of the larynx in real time and dynamically adjust the weight response of different anatomical regions, thereby accurately focusing on the detailed reconstruction of key structures (the free edge of the vocal cords). Simultaneously, by introducing smoothing mechanisms such as sinusoidal periodic modulation and spatial interpolation, inter-frame noise and weight jumps are effectively suppressed, improving the structural consistency and temporal stability of the enhanced images, significantly enhancing the overall accuracy and clinical applicability of super-resolution reconstruction.
[0097] S2.91. Extract key parameters from the laryngeal motion sensor data. Key parameters include the glottal impedance waveform (reflecting the phonation state). Vertical displacement of the throat Rate of change of air pressure ;
[0098] The sensor data stream is as follows: ;
[0099] S2.92. Based on the above key parameters, classify the state and use the feature threshold judgment strategy to classify the current moment into one of the three states: vocalization phase, inhalation phase, and resting phase.
[0100] The feature threshold determination strategy is as follows:
[0101] ;
[0102] Indicates the current time Physiological state categories; Indicates the glottal impedance threshold (used to determine whether sound is produced). This indicates the upper limit of the vertical displacement of the throat, in mm. A threshold representing the rate of change of air pressure, in mm; This represents the lower limit of the vertical displacement of the throat, in mm.
[0103] S2.93. Introduce segmented organ masks (including glottis, larynx, vocal cords, tracheal cartilage, etc.) from the baseline 3D image, and set response weight factors for organ structures according to the physiological state of different organ structures in each time phase.
[0104] S2.94. In a voxel space with the same resolution as the input image, initialize a three-dimensional voxel matrix with all values equal to 1. , ( The height of the image. The width of the image. (For the depth of the image), assign values to each organ structure voxel based on the organ mask information: , This represents the organ masking function (the organ number corresponding to the voxel position). Represents the three-dimensional coordinates in the image data, namely the position indexes in the width, height, and depth directions. This represents a set of static weight values defined based on different organ structure types;
[0105] S2.95. Apply a phase modulation factor to the weight of each organ voxel based on the physiological state of the current frame. (i.e., the amplification / diminishing coefficients of weights at different physiological stages), and introduce a periodic modulation term in the form of a sine function. (i.e., a time function that dynamically adjusts the weights) to more realistically simulate the effects of respiratory rhythm;
[0106] Among them, the periodic modulation term in the form of a sine function for:
[0107] ;
[0108] In the formula, Indicates the average respiratory rate. This indicates the amplitude adjustment factor (default 0.1~0.3).
[0109] S2.96. Based on vertical displacement and impedance changes, the 3D local displacement field at the current moment is reconstructed (to describe the direction and magnitude of motion of each voxel, in order to model the spatial deformation of the laryngeal anatomical structure in the real physiological motion process into the image enhancement process, so that the enhancement result has temporal consistency, structural continuity and functional responsiveness). In order to prevent artifacts caused by weight jumps, the modulated weight matrix is spatially smoothed, especially the organ boundary layer (2mm thickness) is transitioned by bilinear interpolation.
[0110] Among them, the 3D local displacement field at the current moment is reconstructed. for:
[0111] ;
[0112] In the formula, This is the organ displacement coefficient (range 0.8~1.2). This is the impedance-displacement coupling coefficient (value range 0.05~0.2). This represents the rate of change of glottal impedance, calculated using time-series data from a laryngeal motion sensor, and reflects the instantaneous state of glottal movement. Glottal impedance as a function of time The function of change (instantaneous air pressure difference across the glottis divided by instantaneous airflow through the glottis); This indicates the displacement along the patient's left and right axes (normal range is within). between, At that time, the vocal cords are fixed; (At times, laryngospasm); This indicates displacement along the patient's anterior-posterior axis (normal range is within). between, At that time, unilateral vocal cord paralysis occurred; (At that time, the tumor was compressing the tumor). Indicates displacement along the patient's superior and inferior axes (normal range within) between, At that time, the arytenoid cartilage dislocated; At that time, cricoarytenoid arthritis);
[0113] The formula for spatial smoothing the modulated weight matrix is as follows:
[0114] ;
[0115] ;
[0116] ;
[0117] In the formula, The final dynamic weights after smoothing The distance to the organ boundary, Represents the mixing coefficient, which is It changes linearly with distance; The weight values are obtained by combining dynamic weights with directional adjustments based on the displacement field and its gradient. This represents a dynamic weight value that combines organ type, phase, and respiratory rhythm. The thickness of the transition zone at the organ boundary. This is a dimensionless adjustment coefficient for the displacement field to the weight change, which controls the intensity of the influence of the anatomical structure displacement on the weight gradient.
[0118] S2.97, Output the smoothed dynamic time-varying weight matrix ,
[0119] In the formula, Indicates the height direction (row number); Indicates the width direction (number of columns); Indicates the depth direction (number of slices).
[0120] In this embodiment, the dynamic time-varying weight matrix output by S2.97 is... Corresponding 3D laryngeal image data Voxel-by-voxel multiplication is performed to obtain the corrected 3D image data. ,in Representing the Hadamard product, the corrected It is used as input to the 3D-ResNet model for subsequent feature extraction;
[0121] S2.10. The 3D convolution kernel of the 3D-ResNet model is decomposed into convolution kernels for specific directions of the larynx, including axial, coronal, and sagittal convolutions, to reduce computation and preserve the anisotropic features of the larynx. A learnable gating network is used to assign weights to the feature maps extracted from specific directions of the larynx, and the specific directions of the larynx are fused according to the weights. This process enhances the feature representation of key perspectives (including the vocal cord level).
[0122] S2.11 The feature maps obtained through the above steps in different directions are fed into the dynamic gating unit for weighted fusion to generate a comprehensive feature representation.
[0123] Furthermore, the extracted features are fused using cross-scale feature pyramid technology, including the following steps:
[0124] S2.12. Using the low-resolution 2D endoscopic image and 3D laryngeal image data obtained through feature extraction as input, extract the 2D feature map from the 3D image from a specific viewpoint and adjust its size to match the size of the 2D endoscopic image. This step ensures that images from two different sources can be compared and fused at the same spatial resolution.
[0125] S2.13. For each scale, the feature map of the adjusted endoscope image is fused with the feature map of the three-dimensional image by stitching to obtain the fused features of the three scales (denoted as F1, F2 and F3).
[0126] S2.14. Construct a feature pyramid based on a top-down path (starting from the smallest scale feature map F3, upsample (bilinear interpolation) to obtain a feature map, then add it to F2, and then upsample the added feature map and add it to F1). Before adding it to a larger scale feature map after each upsampling, ensure that the two feature maps are matched in size. If necessary, further adjust the size of the feature maps to ensure accurate pixel-level alignment.
[0127] S2.15 Output a cross-scale fusion feature pyramid, which contains information at different scales, from fine-grained to coarse-grained.
[0128] S3. Using the U-Net++ model that incorporates sub-pixel convolutional layers, super-resolution reconstruction is performed on images that have undergone feature fusion processing.
[0129] In this embodiment, U-Net++ is an enhanced version of the U-Net model that improves the accuracy of semantic segmentation through dense skip connections and multi-scale feature fusion. When processing images, U-Net is first used for preliminary semantic segmentation, generating features such as vocal cord masks to locate and identify key regions in the image. Traditional upsampling methods (including bilinear interpolation or transposed convolution) may lead to blurring or checkerboard effects. Subpixel convolution (also known as PixelShuffle) is a more efficient upsampling technique that intelligently rearranges the channel data of a low-resolution image to generate a high-resolution output, thus preserving more detail. This method can reduce computational costs while restoring details in high-resolution images. In the U-Net++ model, a dynamic kernel strategy is introduced, particularly during the upsampling process, adaptively adjusting the convolution kernel weights based on the current local structure. This means that for regions containing more high-frequency information (such as edges and textures), a smaller receptive field kernel can be selected to capture details; while for regions with smoother structures, dilated convolutions may be used to cover a larger area. This strategy helps enhance the detail representation capability during subpixel reconstruction.
[0130] Super-resolution reconstruction of images that have undergone feature fusion processing is performed using the U-Net++ model, which incorporates sub-pixel convolutional layers. The steps include:
[0131] S3.1. The fused feature map output by the cross-scale feature pyramid is used as input. The fused feature map is processed by a deep learning-based semantic segmentation network (U-Net) to generate a vocal cord mask (the backbone network (ResNet-34) is pre-trained on ImageNet, and the segmentation head is fine-tuned on the laryngoscope dataset (Kvasir-SEG). This mask is used to locate the region in the image that is related to the free edge of the vocal cord.
[0132] S3.2 Construct the encoder part of the U-Net++ model. In the encoder part, feature extraction is performed sequentially through three coding blocks and one bottleneck layer. The encoder contains three coding blocks and one bottleneck layer (the number of channels are 64 / 128 / 256 respectively, and the number of output channels of the bottleneck layer is 512). Each coding block includes two convolutional layers and a ReLU activation function to progressively extract the structural features of the image and perform spatial downsampling to capture feature information at different levels.
[0133] S3.3. After each coding block, a vocal cord attention gate module is added to enhance the feature regions related to the free edge of the vocal cord by fusing vocal cord masks (by fusing pre-generated vocal cord masks to emphasize the feature regions related to the free edge of the vocal cord, making these regions, which are crucial for diagnosis, more prominent. This method helps overcome the problem of traditional convolutional neural networks treating all pixels equally when processing complex anatomical structures, ensuring that key details are not drowned out by background noise; this module can selectively increase the weights of important regions such as the vocal cords during network training, making the information of these regions easier for subsequent layers to capture, and more effectively propagating gradient information backward, promoting the model to better learn the feature representations of these regions), thereby improving the weights of key regions and their gradient propagation ability, which helps to improve the reconstruction quality, especially the clarity of important anatomical structures;
[0134] S3.4. The high-dimensional representation output by the encoder is fed into the bottleneck layer, which has an output dimension of (20×15×2048). Multi-level subpixel convolution is used to gradually upsample (to restore high-definition image details). A dynamic kernel strategy is introduced in the subpixel convolution process to adaptively generate convolution kernel weights according to the current local structure in order to enhance the detail expression ability during subpixel reconstruction.
[0135] The method involves progressive upsampling using multi-level subpixel convolution (to restore high-resolution image details) and introducing a dynamic kernel strategy during the subpixel convolution process, including the following steps:
[0136] S3.41. Initialize an empty list to store the upsampled feature map of each level output, set the output feature map of the bottleneck layer of the U-Net++ model as the initial input, and set the number of upsampling times (3 times).
[0137] S3.42. Perform channel compression on the input feature map using 1×1 convolution to reduce computational cost and standardize the feature distribution;
[0138] S3.43 Calculate the gradient magnitude of the compressed feature map (using the Sobel operator) and count the proportion of high-frequency regions;
[0139] S3.44. Select the convolution kernel dynamically based on the high frequency ratio (if the high frequency ratio is greater than a (a represents the threshold of the high frequency region, with a value range of [0.2, 0.4]), it indicates that the high frequency information is rich, and a convolution kernel with a smaller receptive field is selected; if the high frequency ratio is less than or equal to a, it indicates that the structure is smooth, and dilated convolution is selected), and use the selected convolution kernel to convolve the input feature map.
[0140] S3.45. Perform channel expansion convolution and use the PixelShuffle operation to upsample (the PixelShuffle operation achieves efficient image upsampling by rearranging the channel dimensions in the feature map into spatial dimensions), to obtain the upsampled feature map;
[0141] S3.46. Obtain the vocal cord mask at the current resolution, and calculate the average gradient intensity of the corresponding vocal cord mask region in the upsampled feature map (by extracting the gradient magnitude of the corresponding region of the vocal cord mask in the upsampled feature map and averaging its pixel gradient to measure the detail intensity of the region), which is used for the structural weighting of subsequent skip connections.
[0142] S3.47. Fuse the upsampled feature map with the same-scale encoded feature map at an interval of one level (such as the previous layer) (first adjust the number of channels and spatial size of the same-scale encoded feature map, calculate the structure weighting factor (calculate the attention based on the average gradient strength and the gradient of the whole image), and then perform weighted fusion).
[0143] S3.48. Save the feature map of the current upsampling output to an empty list and use the feature map as the input for the next stage of upsampling;
[0144] S3.49. Repeat steps S3.42 to S3.48 until all upsampling layers (N=3) are completed, and finally output a high-resolution feature map.
[0145] S3.5. Employing the dense skip connection structure of the U-Net++ model (a dense skip connection structure is a mechanism designed to enhance information flow and feature reuse. Through dense skip connections, U-Net++ can better combine information from different levels, including both detailed low-level features and abstract high-level features, which is crucial for accurately segmenting complex and intricate anatomical structures), the feature maps of each level in the encoder are combined. ( Indicates network hierarchy, The feature maps (representing the depth of dense connections) are connected to the feature maps of the corresponding scale in the decoder to achieve the fusion of semantic and spatial information. This helps control the transmission of redundant information and focus on key structures. Furthermore, a structure-aware weighting mechanism (a technique used in deep learning models to optimize feature transmission and enhance the recognition of specific anatomical structures) is introduced into the skip connection path of the U-Net++ model. Weights are assigned to skip connections based on structural importance (whether they belong to the vocal cords), enabling more targeted reconstruction and adjustment.
[0146] Furthermore, a structure-aware weighting mechanism is introduced into the skip connection paths of the U-Net++ model, assigning weights to skip connections based on structural importance (whether they belong to the vocal cords), including the following steps:
[0147] S3.51, Downsample the vocal cord mask to the feature map connected to the jump. For the same spatial dimensions, a structural importance diagram is obtained. Each location value reflects whether the location is within the vocal cord region or near the boundary;
[0148] Structural importance diagram Represented as:
[0149] ;
[0150] In the formula, ; Indicates the spatial coordinate position in the image; This indicates that the entire vocal cord area is present; Indicates the weight of non-vocal cord regions;
[0151] S3.52, Structure Importance Diagram Applied to skip connection feature maps Pixel-wise channel weighting is performed, which preserves the feature intensity of the vocal cord region and suppresses the transmission of redundant features in non-critical regions.
[0152] The specific steps for performing pixel-by-pixel channel weighting are as follows:
[0153] ;
[0154] In the formula, This is the weighted feature map. Channels representing feature maps;
[0155] S3.53, Weighted feature map By stitching the feature maps of the corresponding scale in the decoder, the decoder can focus more on restoring key structural regions when recovering image details.
[0156] S3.6 To particularly emphasize and improve the reconstruction quality of edge regions, an edge guidance loss term is added. The Sobel operator is used to extract the vocal cord edge gradient, and L1 loss is used for supervision (the Sobel operator is used to calculate the gradient map of the predicted image and the real label image respectively, the pixel-by-pixel difference of the two gradient maps is calculated, the average value of the difference is calculated using the L1 loss function, and the loss is added to the total loss as an edge guidance term to optimize the model's edge reconstruction capability).
[0157] S4. The U-Net segmentation model based on deep learning performs targeted local enhancement of the mucous membrane region in the image;
[0158] In this embodiment, an attention map is generated based on the anatomical structure, and targeted local enhancement is performed on the mucosal region in the image, including the following steps:
[0159] S4.1. Use the super-resolution reconstructed image as input;
[0160] S4.2. Use the U-Net model for segmenting the mucosal region of the throat image to process the input image and output a binary mask;
[0161] In this context, the foreground pixels represent the location of the mucosal region, while the background pixels represent the non-mucosal region.
[0162] S4.3 Apply a local enhancement network based on dynamic filtering mechanism and adaptive weight adjustment to the segmented mucosal region;
[0163] Furthermore, the dynamic filtering mechanism can automatically adjust filter parameters according to the specific characteristics of different mucosal regions to achieve the best image enhancement effect. This mechanism can effectively remove noise interference while preserving or even enhancing important tissue structure information. It also allows for different filtering strategies to be used for different types of mucosa (such as inflammation, ulcers, etc.), thereby improving the visualization of specific lesion features. Adaptive weight adjustment technology can automatically assign different weight values according to different features within the mucosal region during processing, ensuring that key areas receive more attention. This technology can enhance sensitivity to subtle structural changes, including minute morphological changes in early cancer, making these changes easier to detect and further improving diagnostic efficiency.
[0164] Applying a local enhancement network based on dynamic filtering and adaptive weight adjustment to the segmented mucosal region includes the following steps:
[0165] S4.31. Input the segmented mucosal region image into the input layer of the local enhancement network as the high-interest region to be processed;
[0166] S4.32. Multi-scale features of the mucosal region are extracted through multiple convolutional layers to capture the texture features, edge information and microstructure of the mucosal region, ensuring multi-level expression of detailed information.
[0167] S4.33. Based on the extracted features, adaptively generate filter kernel weights. The filter kernel will be dynamically adjusted according to the different local structures to enhance details rather than simply smooth them. An attention mechanism is introduced to calculate the weight distribution of key textures and lesion features in the mucosal region, strengthen the attention to important regions, and improve the response intensity of these regions.
[0168] S4.34. Combining the structural information of the mucosal region (including edge strength and texture complexity), dynamically adjust the weight allocation of different channels and spatial locations to achieve differentiated processing of different locations and channels, highlighting important details.
[0169] S4.35. After filtering and weight adjustment, a locally enhanced image of the mucosal region is generated.
[0170] S4.4 Integrate the locally enhanced mucosal region back into the original super-resolution reconstructed image;
[0171] S4.5 Generate the final fused image.
[0172] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention.
Claims
1. A high-resolution image enhancement method for throat structure and details based on deep learning, characterized in that, Includes the following steps: S1. Acquire low-resolution images from the laryngeal endoscope and corresponding three-dimensional image data of the larynx; S2. Based on the ResNet model, features are extracted from low-resolution images. The 3D-ResNet model is used to extract features from three-dimensional image data. A time-varying weight matrix with sinusoidal periodic modulation is constructed in the 3D-ResNet model so that the image enhancement process can perceive the physiological motion state of the larynx in real time. The extracted features are then fused using cross-scale feature pyramid technology. S3. Using the U-Net++ model that incorporates sub-pixel convolutional layers, super-resolution reconstruction is performed on images that have undergone feature fusion processing. S4. The U-Net segmentation model based on deep learning performs targeted local enhancement of the mucous membrane region in the image.
2. The method for high-resolution image enhancement of throat structure and details based on deep learning according to claim 1, characterized in that: In step S2, features are extracted from low-resolution images based on the ResNet model, including the following steps: S2.1 Use a pre-trained laryngeal key point detection model to automatically identify and mark key anatomical landmarks in the larynx, and generate a structural heatmap based on these landmarks; S2.2 Adaptive gamma correction for the clinical region; S2.
3. Input the image processed in step S2.2 into the improved ResNet model that incorporates the anatomical attention module to begin the feature extraction process; S2.4 Perform the first stage of ResNet feature extraction to obtain preliminary feature maps; S2.
5. Input the feature map of the current stage together with the structural heatmap generated in step S2.1 into the anatomy attention module, combine the feature information of the two through convolution operation, generate attention weights, and apply the attention weights to weight the feature map. S2.
6. Dynamically adjust the convolution kernel parameters based on the spatial frequency of the current feature map; S2.7 Repeat the above steps to obtain multi-level feature output.
3. The method for high-resolution image enhancement of throat structure and details based on deep learning according to claim 2, characterized in that: In step S2, features are extracted from 3D image data based on the 3D-ResNet model, including the following steps: S2.8, Collect data from the laryngeal motion sensor; S2.9 Construct a time-varying weight matrix based on the acquired laryngeal motion sensor data; S2.
10. Decompose the 3D convolutional kernel of the 3D-ResNet model into convolutional kernels for specific directions of the throat, use a gating network to assign weights to the feature maps extracted from specific directions of the throat, and fuse the specific directions of the throat according to the weights. S2.11 The feature maps obtained through the above steps in different directions are fed into the dynamic gating unit for weighted fusion to generate a comprehensive feature representation.
4. The method for high-resolution image enhancement of throat structure and details based on deep learning according to claim 3, characterized in that: In step S2.9, constructing a time-varying weight matrix based on the acquired throat motion sensor data includes the following steps: S2.
91. Extract key parameters from throat motion sensor data; S2.
92. Based on the above key parameters, classify the state and use the feature threshold judgment strategy to classify the current moment into one of the three states: vocalization phase, inhalation phase, and resting phase. S2.
93. Introduce the segmented organ mask in the baseline 3D image, and set the response weight factor of the organ structure according to the physiological state of different organ structures in each time phase. S2.
94. In a voxel space with the same resolution as the input image, initialize a three-dimensional voxel matrix and assign values to each organ structure voxel according to the organ mask information. S2.
95. Based on the physiological state of the current frame, apply a phase modulation factor to the weight of each organ voxel and introduce a periodic modulation term in the form of a sine function. S2.
96. Based on the vertical displacement and impedance changes, the 3D local displacement field at the current moment is reconstructed. To prevent artifacts from weight jumps, the modulated weight matrix is spatially smoothed. S2.97, Output the dynamic time-varying weight matrix after spatial smoothing.
5. The method for high-resolution image enhancement of throat structure and details based on deep learning according to claim 3, characterized in that: In step S2, the extracted features are fused using cross-scale feature pyramid technology, including the following steps: S2.
12. Use the low-resolution 2D endoscopic image and 3D laryngeal image data obtained through feature extraction as input; S2.
13. For each scale, the feature map of the endoscope image is fused with the feature map of the three-dimensional image by stitching. S2.14 Construct a feature pyramid based on the top-down path; S2.15, Output cross-scale fusion feature pyramid.
6. The method for high-resolution image enhancement of throat structure and details based on deep learning according to claim 1, characterized in that: In step S3, the U-Net++ model, which incorporates sub-pixel convolutional layers, is used to perform super-resolution reconstruction on the image that has already undergone feature fusion processing. This includes the following steps: S3.
1. The fused feature map output by the cross-scale feature pyramid is used as input, and the fused feature map is processed by a deep learning-based semantic segmentation network to generate a vocal cord mask. S3.2 Construct the encoder part of the U-Net++ model. In the encoder part, feature extraction is performed sequentially through three encoding blocks and a bottleneck layer. S3.3 Add a vocal cord attention gate module after each coding block to enhance the feature regions related to the free edge of the vocal cords by fusing vocal cord masks; S3.
4. The high-dimensional representation output by the encoder is fed into the bottleneck layer, and upsampling is performed step by step using multi-level subpixel convolution. A dynamic kernel strategy is introduced during the subpixel convolution process. S3.
5. The dense skip connection structure of the U-Net++ model is adopted to connect the feature map of each level in the encoder with the feature map of the corresponding scale in the decoder. A structure-aware weighting mechanism is introduced into the skip connection path of the U-Net++ model to assign weights to the skip connections according to the structural importance. S3.
6. Add an edge-guided loss term, use the Sobel operator to extract the vocal cord edge gradient, and use L1 loss for supervision.
7. The method for high-resolution image enhancement of throat structure and details based on deep learning according to claim 6, characterized in that: In step S3.4, upsampling is performed step by step using multi-level sub-pixel convolution, and a dynamic kernel strategy is introduced during the sub-pixel convolution process, including the following steps: S3.
41. Initialize an empty list to store the upsampled feature map of each level output, set the output feature map of the bottleneck layer of the U-Net++ model as the initial input, and set the number of upsampling times. S3.
42. Perform channel compression on the input feature map using 1×1 convolution; S3.43 Calculate the gradient magnitude of the compressed feature map and count the proportion of high-frequency regions; S3.
44. Select the convolution kernel based on the high frequency ratio dynamic kernel, and use the selected convolution kernel to convolve the input feature map; S3.
45. Perform channel expansion convolution and use the PixelShuffle operation to upsample to obtain the upsampled feature map; S3.
46. Obtain the vocal cord mask at the current resolution and calculate the average gradient intensity of the corresponding vocal cord mask region in the upsampled feature map. S3.
47. Fuse the upsampled feature map with the same-scale encoded feature map at the next level; S3.
48. Save the feature map of the current upsampling output to an empty list and use the feature map as the input for the next stage of upsampling; S3.
49. Repeat steps S3.42 to S3.48 until all upsampling layers are completed, and finally output a high-resolution feature map.
8. The method for high-resolution image enhancement of throat structure and details based on deep learning according to claim 6, characterized in that: In S3.5, a structure-aware weighting mechanism is introduced into the skip connection path of the U-Net++ model, assigning weights to skip connections based on structural importance, including the following steps: S3.
51. Downsample the vocal cord mask to the same spatial size as the skip connection feature map to obtain the structural importance map; S3.
52. Apply the structural importance map to the skip connection feature map and perform pixel-by-pixel channel weighting. S3.
53. Concatenate the weighted feature map with the feature map of the corresponding scale in the decoder.
9. The method for high-resolution image enhancement of throat structure and details based on deep learning according to claim 1, characterized in that: In step S4, an attention map is generated based on the anatomical structure, and targeted local enhancement is performed on the mucosal region in the image, including the following steps: S4.
1. Use the super-resolution reconstructed image as input; S4.
2. Use the U-Net model for segmenting the mucosal region of the throat image to process the input image and output a binary mask; In this context, the foreground pixels represent the location of the mucosal region, while the background pixels represent the non-mucosal region. S4.3 Apply a local enhancement network based on dynamic filtering mechanism and adaptive weight adjustment to the segmented mucosal region; S4.4 Integrate the locally enhanced mucosal region back into the original super-resolution reconstructed image; S4.5 Generate the final fused image.
10. The method for high-resolution image enhancement of throat structure and details based on deep learning according to claim 9, characterized in that: In step S4.3, a local enhancement network based on dynamic filtering mechanism and adaptive weight adjustment is applied to the segmented mucosal region, including the following steps: S4.
31. Input the segmented mucosal region image into the input layer of the local enhancement network; S4.
32. Extract multi-scale features of the mucosal region through multiple convolutional layers; S4.
33. Adaptively generate filter kernel weights based on the extracted features, and introduce an attention mechanism to calculate the weight distribution of key textures and lesion features in the mucosal region; S4.
34. Combine the structural information of the mucosal region to dynamically adjust the weight allocation of different channels and spatial locations; S4.
35. After filtering and weight adjustment, a locally enhanced image of the mucosal region is generated.
Citation Information
Patent Citations
Throat image enhancement method capable of sensing structure and details
CN115131230A
Junk image denoising method based on multi-dimensional image information fusion
CN116543168A