A fundus OCT image reconstruction system, method, device and storage medium imitating the top cover of Amaris
The eagle-eye inspired system addresses the limitations of existing deep learning methods in OCT image super-resolution by using vertical and horizontal feature extraction and fusion techniques to enhance high-frequency details and contrast, resulting in clearer and more detailed eye fundus images.
Patent Information
- Application Number
- CN202210553987.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-19
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-05-19
AI Technical Summary
The existing single-frame image super-resolution reconstruction method based on deep learning has problems in OCT images with slow network operation speed, unobvious image edge contour information, loss of detail information, and poor contrast. It is difficult to effectively improve image resolution and restore fundus health.
The fundus OCT image reconstruction system is adopted to imitate the eagle vision top cover. Through feature extraction blocks, information processing blocks and reconstruction blocks, the hollow convolution, dense connections, channel attention and spatial attention mechanisms are used to mine high-frequency features step by step from the vertical and horizontal dimensions, highlight significant information, and perform image reconstruction.
It has achieved efficient improvement of OCT image resolution, clearly presenting the details and contour characteristics of the lesion area, improving image contrast, truly restoring the health of the fundus, and improving the visual effect and diagnosis and treatment accuracy of the image.
Smart Images

Figure CN114820325B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of image reconstruction, and relates to an eagle-eye tectum-like fundus OCT image reconstruction system, method, device and storage medium. Background Art
[0002] Conducting fundus disease screening, proposing diagnosis and prevention plans as early as possible, and avoiding patients from losing their ability to work are of great significance to promoting social development. Optical Coherence Tomography (OCT) is an important means of examining retinal abnormalities such as glaucoma, macular edema, and diabetic retinopathy. It collects multiple images continuously and quickly after fixing the patient's eye position, and then aligns and averages them. In actual applications, due to the shortcomings of the equipment itself and the unintentional shaking of the patient's body during shooting, the images obtained are often blurred and the visual effect is poor, which has a great impact on the analysis of the disease and the proposal of diagnosis and treatment plans. At present, without improving the hardware equipment, super-resolution reconstruction provides a new solution to improve the clarity and contrast of OCT images.
[0003] With the rise of the deep learning craze, the idea of using convolutional neural networks for image super-resolution reconstruction has received extensive attention from scholars. Dong et al. first applied deep learning knowledge to the reconstruction technology and proposed SRCNN, which avoided the artificial design of feature extraction methods and achieved the learning of the image itself, thus realizing image reconstruction. For details, see "C. Dong, C. C. Loy, K. He, X. Tang, Image super-resolution using deep convolutional networks, in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), 2014, pp. 184-199". Ledig et al. proposed the reconstruction model SRGAN based on generative adversarial networks, which completed image reconstruction through the competition between the generator and the discriminator. For details, see "C. Ledig et al., "Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network," 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 105-114". Zheng et al. proposed the Information Multi-distillation Network (IMDN), which improved the reconstruction speed without increasing the model complexity by mining deep features. For details, see "H. Zheng, X. Gao, Y. Yang, and X. Wang, "Lightweight Image Super-Resolution with Information Multi-distillation Network," Proceedings of the 27th ACM International Conference on Multimedia (ACM), 2019, pp. 2024-2032". Liu et al. proposed the Residual Feature Distillation Network (RFDN) with channel separation feature distillation structure, which achieved speed improvement by reducing the number of parameters. For details, see "J. Liu, J. Tang, and G. Wu, "Residual Feature Distillation Network for Lightweight Image Super-Resolution," Computer Vision-ECCV 2020 Workshops, 2020, pp. 41-55".Zhu et al. proposed a lightweight image super-resolution reconstruction network EMASRN, which uses the expectation-maximization attention mechanism to improve the network speed. For details, see "X. Zhu, K. Guo, S. Ren, B. Hu, M. Hu and H. Fang, "Lightweight Image Super-Resolution With Expectation-Maximization Attention Mechanism," in IEEE Transactions on Circuits and Systems for Video Technology, vol. 32, no. 3, 2022, pp. 1273-1284". Although the above learning-based methods can reconstruct high-resolution images, there is a lack of corresponding connections between convolutional layers, and significant features cannot be highlighted spatially. They perform poorly in OCT images with little change in gray level, unclear hierarchical features, and low contrast.
[0004] The eagle-eye technology has significant advantages in image recognition. For example, Fu et al. proposed a lightweight vision system for target detection and recognition by imitating the eagle's vision system. For details, see "Q. Fu, S. T. Wang, J. Wang, S. N. Liu and Y. B. Sun, "A Lightweight Eagle-Eye-Based Vision System for Target Detection and Recognition," in IEEE Sensors Journal, vol. 21, no. 22, 2021, pp. 26140-26148". Liu et al. proposed an eagle-eye multi-task CNN model for aerial image classification to distinguish subtle differences between different aerial images. For details, see "Y. Liu, Z. Han, C. Chen, L. Ding and Y. Liu, "Eagle-Eyed Multitask CNNs for Aerial Image Retrieval and Scene Classification," in IEEE Transactions on Geoscience and Remote Sensing, vol. 58, no. 9, 2020, pp. 6699-6721". In the eagle-eye vision system, the visual attention mechanism is used to analyze the information in the image scene, making it easier to select specific regions of interest, which can be used to solve the problems of blurred edges and unclear detailed information in the lesion areas of OCT images.
[0005] The inventors have found through research that the reconstruction effect of existing single-image super-resolution reconstruction methods based on deep learning is not good, and there are mainly the following defects: (1) Most existing networks pursue good effects by increasing the depth of convolutional layers, resulting in slow network operation speed; (2) The complementarity of convolutional kernels is insufficient, the edge contour information of the reconstructed image is not obvious, and there is a loss of detail information; (3) The color of the reconstructed image is darker than that of the original image, the contrast is poor, and the visual effect is not good. How to obtain more useful information from low-resolution fundus OCT images, efficiently improve the image resolution, restore the real scene of the fundus health condition, maximize the presentation of the information contained in the image, and promote the intelligence of fundus disease screening has become an urgent problem to be solved. Summary of the Invention
[0006] To solve the above problems, the present invention provides a fundus OCT image reconstruction system imitating the eagle optic tectum, which has high image contrast, rich detail features in the lesion area, can enrich the content of the picture while ensuring the image clarity, and solves the problems existing in the prior art.
[0007] The second object of the present invention is to provide a fundus OCT image reconstruction method imitating the eagle optic tectum.
[0008] The third object of the present invention is to provide an electronic device.
[0009] The fourth object of the present invention is to provide a computer storage medium.
[0010] The technical solution adopted by the present invention is a fundus OCT image reconstruction system imitating the eagle optic tectum, which includes a feature extraction block, an information processing block, and a reconstruction block;
[0011] The feature extraction block is used to extract shallow features from low-resolution OCT images to obtain useful information;
[0012] The information processing block extracts high-frequency features by gradually expanding the receptive field from two dimensions, vertical and horizontal, imitating the information processing mechanism of the eagle optic tectum;
[0013] The reconstruction block is used to perform preliminary reconstruction on the extracted shallow features and high-frequency features, and then perform upsampling operations and then deep reconstruction to obtain the reconstructed high-resolution OCT image.
[0014] Further, the vertical dimension of the information processing block is composed of 6 eagle optic tectum imitation blocks (EVB), which are responsible for mining high-frequency features;
[0015] The EVB in the vertical dimension of the information processing block is responsible for extracting high-frequency features from the feature maps of each channel, and using the channel attention mechanism to promote the fusion between different channels;
[0016] The EVB module includes two parts: information deep extraction and information fusion;
[0017] The information deep extraction part is composed of four dilated convolutions with a convolution kernel size of 3×3 and dilation rates of 1, 1, 2, and 3 respectively. When each convolution kernel outputs, the number of channels of the original features is halved to reduce the number of parameters. At the same time, the output of the previous convolution kernel is propagated layer by layer to the subsequent convolution kernel to form dense connections, facilitating the deep learning of features by the subsequent convolution kernel;
[0018] The information fusion part fuses the outputs of the four convolution kernels through concat, uses a 1×1 convolution kernel for channel number transformation, and introduces a channel attention mechanism to mine significant information;
[0019] The channel attention mechanism is divided into three parts: global context embedding, channel normalization, and gate adaptation;
[0020] The global context embedding uses the l2 criterion for normalization operations and introduces a training parameter α to assist the adaptive output of l2 normalization;
[0021] The channel normalization uses the l2 criterion for cross-channel normalization between features to reduce the number of parameters;
[0022] The gate adaptation controls the competition and cooperation relationship between neurons through weights γ and biases β.
[0023] Furthermore, the horizontal dimension of the information processing block is responsible for fusing the high-frequency features extracted from the vertical dimension, deeply mining significant information, using a spatial attention mechanism to strengthen the deep features, clearly expressing details such as the edema contour of the lesion area, and transmitting the extracted high-frequency features to the reconstruction module;
[0024] The spatial attention mechanism divides the extracted features into 64 subspaces in space. Each subspace first uses depthwise separable convolution to extract new features for each group of features to reduce redundancy within the channel range. After max pooling, pointwise convolution is used to reduce redundancy within the spatial range. Finally, the softmax function is used to scale the feature map in the H dimension to ensure that the weights sum to 1. After fusion with the original features, 64 subspace features are obtained;
[0025] The different subspaces are connected through concat to obtain the final output.
[0026] Furthermore, the reconstruction block uses a 3×3 convolution kernel for one layer to obtain a basic reconstructed image. After upsampling operations, a 3×3 convolution kernel is used for deep reconstruction to obtain a clear fundus OCT reconstructed image.
[0027] A method for reconstructing fundus OCT images imitating the optic tectum of an eagle is carried out for reconstruction according to the following steps:
[0028] S1. Input the low-resolution fundus OCT image into the feature extraction block to complete the extraction of shallow features and send them to the information processing block;
[0029] S2. The vertical dimension of the information processing block is responsible for extracting high-frequency features. After the horizontal dimension of the information processing block fuses the high-frequency features extracted by the vertical dimension and highlights the significant information, it is sent to the reconstruction block;
[0030] S3. The reconstruction block uses the shallow features and high-frequency features extracted by the feature extraction block and the information processing block to complete image reconstruction and obtain a clear fundus OCT image.
[0031] Further, the information processing block in S2 extracts high-frequency features in the vertical dimension according to the following formula
[0032] F1 = C 3×3 (X i-1 )
[0033] F2 = C 3×3 (C 3×3 (F1))
[0034]
[0035] F5 = C 1×1 ({F1,F2,F3,F4})
[0036]
[0037] where, F k represents the output of the k (k = {1,2,3,4})-th convolutional kernel, C 3×3 represents a convolutional kernel of size 3×3, represents a dilated convolution with a convolutional kernel size of 3×3 and a dilation rate of 2, represents a dilated convolution with a convolutional kernel size of 3×3 and a dilation rate of 3, {} represents a concatenation operation, X i-1 and X i are respectively the input and output of the i (i = 1,...,n)-th EVB module, F n represents the output of the n (n = {1,2,3,4})-th convolutional kernel, F5 and are respectively the input and output of channel attention, R c represents the output of global context embedding, R = R1,R2,...,R c-1 ,R c represents R c decomposed into C (C = 64) feature maps, represents the output of channel normalization, represents for A scalar normalized by scale, α, β, and γ represent trainable parameters, ε represents a constant, and f tanh represents the tanh activation function.
[0038] Furthermore, the information processing block of S2 fuses and enhances deep features in the horizontal dimension according to the following formula
[0039] X C = C 1×1 ({X1, X2,..., X n-1 , X n})
[0040]
[0041] where X C and respectively represent the input and output of the channel attention mechanism, X i (i = 1,..., n) represents the output of the i-th EVB module, and respectively represent the input and output of the g-th (g = 1,..., 64) sub-feature space, C 1×1 represents a convolutional kernel of size 1×1, PW 1×1 represents a depthwise separable convolution with a convolutional kernel size of 1×1, DW 1×1 represents a pointwise convolution with a convolutional kernel size of 1×1, f maxpool represents a max pooling operation, f softmax represents the softmax activation function;
[0042] The mathematical model of the information processing block in S2 is:
[0043]
[0044] X0 = C 3×3 (C 3×3 (X))
[0045] where X represents the low-resolution image input to the feature extraction block, C 3×3 represents a convolutional kernel of size 3×3, X0 represents the shallow features obtained by the feature extraction block, where X out represents the output of the information processing block.
[0046] Furthermore, S3 completes the reconstruction according to the following formula
[0047] Y = C 3×3 (f upsampler (C 3×3 (X out )))
[0048] Where Y represents the output of the reconstruction module, f upsampler Represents an upsampling operation.
[0049] An electronic device adopts the above method to realize image reconstruction.
[0050] A computer storage medium stores at least one program instruction, and the at least one program instruction is loaded and executed by a processor to implement the above-mentioned image reconstruction method.
[0051] The beneficial effects of the present invention are:
[0052] Aiming at the problems of poor contrast of fundus OCT images and blurred edges of lesion areas, the present invention draws on the information processing mechanism of the eagle vision system and proposes a super-resolution reconstruction method EOTRN for a single OCT image. It imitates the idea of gradually expanding the receptive field of the eagle's visual tectum to gradually mine high-frequency features from the vertical and horizontal dimensions. In the vertical dimension, with the help of dilated convolution, dense connection, and channel attention, the receptive field is gradually expanded, the features of different network layers are propagated, and the "competition" and "cooperation" between the features of different channels are achieved, and the preliminary extraction of high-frequency features is completed. In the horizontal dimension, with the help of 64 feature subspaces, redundant information in high-frequency features is eliminated, significant information is corrected and highlighted, and the texture and contour features of the lesion area are enhanced. Finally, the shallow features and high-frequency features are upsampled and deeply reconstructed to obtain a high-definition OCT image. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0054] Figure 1 It is a schematic diagram of the structure of a reconstruction system according to an embodiment of the present invention.
[0055] Figure 2 It is a schematic diagram of the structure of the EVB module in the reconstruction system according to an embodiment of the present invention.
[0056] Figure 3 It is a structural diagram of the channel attention mechanism module in the reconstruction system of an embodiment of the present invention.
[0057] Figure 4 It is a schematic diagram of the structure of the spatial attention module in the reconstruction system of an embodiment of the present invention.
[0058] Figure 5This is a comparison chart of the reconstruction effects of the reconstruction method of an embodiment of the present invention and other algorithms on patients with cystoid macular edema.
[0059] Figure 6 This is a comparison chart of the reconstruction effects of the reconstruction method of an embodiment of the present invention and other algorithms on patients with serous retinal detachment. DETAILED DESCRIPTION
[0060] The following will be combined with the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0061] Embodiment 1,
[0062] An eagle-like visual tectum-like fundus OCT image reconstruction system, the structure of which is as follows Figure 1 As shown, it includes a feature extraction module, an information processing module, and a reconstruction module connected in sequence;
[0063] Among them, the feature extraction module uses two 3×3 convolution kernels in series to extract features from low-resolution images. Compared with the convolution kernel parallel method, this series method ensures sufficient feature extraction without reducing the complexity of the model.
[0064] The information processing module is used to extract high-frequency information such as the contour of the lesion area in the low-resolution OCT image that is easily ignored by the feature extraction module. It is designed to deeply extract high-frequency features from the vertical and horizontal dimensions, imitating the information processing mechanism of the eagle-eye tectum. The vertical dimension is responsible for extracting high-frequency features from the feature maps of each channel, and the horizontal dimension is responsible for fusing the high-frequency features extracted from the vertical dimension to deeply mine significant information.
[0065] like Figure 2As shown in the figure, in the vertical dimension, the eagle-like tectal information block (EVB) uses dilated convolution to gradually expand the receptive field without increasing the complexity, and uses dense connections to obtain more deep features; the introduction of channel attention promotes the "competition" and "cooperation" between different channel features, and uses 6 EVBs to extract high-frequency features from the feature maps of each channel. The EVB module includes two parts: deep information extraction and information fusion. Inspired by the eagle-like tectal to gradually expand the receptive field of visual information for multi-layer processing, the deep information extraction part uses four convolution kernels to extract features in depth, and the last two layers use dilated convolution to expand the receptive field and mine more effective information. When each convolution kernel outputs, the number of channels of the original feature is halved to reduce the number of parameters. At the same time, the output of the previous convolution kernel is propagated layer by layer to the convolution kernel of the next layer to form dense connections, so as to facilitate the deep learning of the features of the next layer convolution kernel. The information fusion part fuses the outputs of the four convolution kernels through concat, and the fused features may cause partial redundancy. In order to eliminate redundancy, promote the fusion of different channels, and highlight significant information such as the contour of the lesion area, a 1×1 convolution kernel is used to transform the number of channels, and a channel attention mechanism is used to mine significant information. Figure 3 As shown in Figure 1, it is divided into three parts: global context embedding, channel normalization, and gate adaptation. Global context embedding uses the l2 criterion for normalization, and introduces a training parameter α to assist the adaptive output of l2 normalization; channel normalization uses the l2 criterion to perform cross-channel normalization between features to reduce the number of parameters, and gate adaptation controls the competition and synergy between neurons through weights γ and bias β.
[0066] like Figure 1 As shown in the figure, in the horizontal dimension, 6 EVB blocks are cascaded to fuse the features learned by different blocks, spatial attention is used to calibrate the features spatially, and the high-frequency features extracted in the vertical dimension are fused to deeply mine significant information.
[0067] like Figure 4 As shown in the figure, spatial attention divides the extracted features into 64 subspaces in space. Each subspace first uses a depth-separable convolution to extract new features for each group of features to reduce redundancy within the channel range. After maximum pooling, point-by-point convolution is used to reduce redundancy within the spatial range. Finally, the softmax function is used to scale the feature map to H dimensions to ensure that the weight sum is 1. After fusion with the original features, 64 subspace features are obtained. Different subspaces are connected through concat to obtain the final output.
[0068] The reconstruction module is used to reconstruct the deep and shallow layer features extracted by the feature extraction module and the information processing module to restore the true contour and detail features of the fundus OCT image.
[0069] like Figure 1As shown, the reconstruction module first performs basic reconstruction using a convolutional kernel of size 3×3. After the upsampling operation, it then uses a convolutional kernel of size 3×3 for deep reconstruction to obtain the reconstructed fundus OCT image.
[0070] The design of network models based on deep learning is based on the application background. For example, image classification and image segmentation focus on the analysis and understanding of image content. In model design, they tend to focus on the recognition of targets, and separate target information from the whole for identification and classification. Image super-resolution reconstruction focuses on the nonlinear mapping relationship between low-resolution images and high-resolution images. By extracting each feature information in the image, the original weak contour features are enhanced, and the details and textures are improved. Using low-resolution images to infer all missing high-frequency details is the key to reconstruction. In order to fully extract feature information from low-resolution images and minimize high-frequency details, an embodiment of the present invention proposes a fundus OCT image reconstruction method that simulates the eagle's eye tectum. The method uses two 3×3 convolution kernels to extract underlying features from low-resolution OCT images to obtain useful information. The information processing block imitates the eagle's eye tectum information processing mechanism to deeply extract high-frequency features from both vertical and horizontal dimensions. In the vertical dimension, the eagle-like tectal information block (EVB) uses dilated convolution to gradually expand the receptive field without increasing the complexity, and uses dense connections to obtain more deep features; the introduction of channel attention promotes the "competition" and "cooperation" between different channel features, and uses 6 EVBs to extract high-frequency features from the feature maps of each channel. The EVB module includes two parts: deep information extraction and information fusion. The deep information extraction part uses four convolution kernels to extract features in depth, and the last two layers use dilated convolution to expand the receptive field and mine more effective information. When each convolution kernel outputs, the number of channels of the original feature is halved to reduce the number of parameters. At the same time, the output of the previous convolution kernel is propagated layer by layer to the next layer convolution kernel to form dense connections, so as to facilitate the deep learning of the next layer convolution kernel for features. The information fusion part fuses the outputs of the four convolution kernels through concat, uses a 1×1 convolution kernel to transform the number of channels, and uses the channel attention mechanism to eliminate redundancy and highlight significant information. The channel attention mechanism is divided into three parts: global context embedding, channel normalization, and gate adaptation. The global context embedding uses the l2 criterion for normalization, and introduces the training parameter α to assist the adaptive output of l2 normalization; channel normalization uses the l2 criterion to perform cross-channel normalization between features to reduce the number of parameters, and gate adaptation controls the competition and synergy between neurons through weights γ and bias β. In the horizontal dimension, 6 EVB blocks are cascaded to fuse the features learned by different blocks, and spatial attention is used to calibrate the features spatially to highlight significant information. The extracted features are divided into 64 subspaces in space. Each subspace first uses a depth-separable convolution to extract new features from each group of features to reduce redundancy within the channel range, and then uses point-by-point convolution after maximum pooling to reduce redundancy within the spatial range. Finally, the softmax function is used to scale the feature map to H dimensions to ensure that the weight sum is 1. After fusion with the original features, 64 subspace features are obtained. Different subspaces are connected by concat to obtain the final output.The reconstruction module first performs basic reconstruction operations using 3×3 convolutions, and then performs deep reconstruction using 3×3 convolutions after upsampling operations to obtain a clear image.
[0071] Example 2
[0072] A fundus OCT image reconstruction method imitating the eagle vision tectum, characterized in that it is carried out according to the following steps:
[0073] S1. Input the low-resolution fundus OCT image into the feature extraction block to complete the extraction of shallow features and send them to the information processing block;
[0074] S2. The longitudinal dimension of the information processing block is responsible for extracting high-frequency features. After the transverse dimension of the information processing block fuses the high-frequency features extracted by the longitudinal dimension and highlights the significant information, it is sent to the reconstruction block;
[0075] S3. The reconstruction block uses the shallow features and high-frequency features extracted by the feature extraction block and the information processing block to complete image reconstruction and obtain a clear fundus OCT image.
[0076] Furthermore, the information processing block in S2 extracts high-frequency features according to the following formula in the longitudinal dimension
[0077] F1 = C 3×3 (X i-1 )
[0078] F2 = C 3×3 (C 3×3 (F1))
[0079]
[0080] F5 = C 1×1 ({F1,F2,F3,F4})
[0081]
[0082] where F k represents the output of the k-th (k = {1,2,3,4}) layer convolution kernel, C 3×3 represents a convolution kernel of size 3×3, represents a dilated convolution with a convolution kernel size of 3×3 and a dilation rate of 2, represents a dilated convolution with a convolution kernel size of 3×3 and a dilation rate of 3, {} represents a concatenation operation, X i-1 and X i are the input and output of the i-th (i = 1,...,n) EVB module respectively, F n represents the output of the n-th (n = {1,2,3,4}) layer convolution kernel, F5 and respectively represent the input and output of the channel attention, R c represents the output of the global context embedding, R = R1, R2, ..., R c-1 , R c represents R c C (C = 64) feature maps after decomposition, represents the output of the channel normalization, represents scalar for normalizing the scale, α, β, γ represent trainable parameters, ε represents a constant, f tanh represents the tanh activation function.
[0083] Furthermore, the information processing block of the S2 fuses and enhances the deep features in the horizontal dimension according to the following formula,
[0084] X C = C 1×1 ({X1, X2, ..., X n-1 , X n})
[0085]
[0086] where, X C and respectively represent the input and output of the channel attention mechanism, X i (i = 1, ..., n) represents the output of the i-th EVB module, and respectively represent the input and output of the g-th (g = 1, ..., 64) sub-feature space, C 1×1 represents a convolutional kernel of size 1×1, PW 1×1 represents a depthwise separable convolution with a convolutional kernel size of 1×1, DW 1×1 represents a pointwise convolution with a convolutional kernel size of 1×1, f maxpool represents the max pooling operation, f softmax represents the softmax activation function;
[0087] The mathematical model of the information processing block in the S2 is:
[0088]
[0089] X0 = C 3×3 (C 3×3 (X))
[0090] where, X represents the low-resolution image input to the feature extraction block, C 3×3 represents a convolutional kernel of size 3×3, X0 represents the shallow features obtained by the feature extraction block, where, X outRepresents the output of the information processing block.
[0091] Furthermore, the S3 completes the reconstruction according to the following formula
[0092] Y = C 3×3 (f upsampler (C 3×3 (X out )))
[0093] where Y represents the output of the reconstruction module, and f upsampler represents the upsampling operation.
[0094] The ultimate goal of constructing the super-resolution reconstruction network is to construct the mapping function between X and X HR For a given training dataset The mapping relationship between X and X is established through the following formula, making HR reach the minimum value: value reaches the minimum:
[0095]
[0096] where θ = {W1, W2, W3... W m , b1, b2, b3... b m}, represents the weights and biases of the m-layer neural network, represents the least squares solution of θ, F θ () represents the mapping relationship after adding the weights and biases of the m-layer neural network, X i and respectively represent the low-resolution image and the high-resolution image of the i-th pair of training sets, represents the loss function used to minimize the difference between Y i and ; this mapping, that is, the upsampling operation, maps the X feature to the X HR feature.
[0097] To avoid introducing unnecessary training techniques and accelerate the convergence speed, we finally choose the L1 loss function. The loss function is defined as:
[0098]
[0099] where, ||||1 represents the operation of solving the 1-norm.
[0100] Use the training set to train the fundus OCT image reconstruction system of the imitation eagle vision tectum to minimize the loss and find the optimal parameters.
[0101] To verify the effectiveness of the fundus OCT image reconstruction method imitating the eagle eye top cover in the embodiments of the present invention, fundus OCT images are selected as the test set, and the algorithms of Keys (R. Keys, Cubic convolution interpolation for digital image processing, in IEEE Transactions on Acoustics, Speech, and Signal Processing, vol. 29, December 1981, pp. 1153-1160.); Dong (C. Dong, C. C. Loy, K. He, X. Tang, Image super-resolution using deep convolutional networks, in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), 2014, pp. 184-199.); Ledig (C. Ledig et al., Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network, 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 105-114); Zheng (H. Zheng, X. Gao, Y. Yang, and X. Wang, Lightweight Image Super-Resolution with Information Multi-distillation Network, Proceedings of the 27th ACM International Conference on Multimedia (ACM MM), 2019, pp. 2024-2032); Liu (J. Liu, J. Tang, and G. Wu, Residual Feature Distillation Network for Lightweight Image Super-Resolution, Computer Vision-ECCV 2020 Workshops, 2020, pp. 41-55); Zhu (X. Zhu, K. Guo, S. Ren, B. Hu, M. Hu and H.Fang, "Lightweight Image Super-Resolution With Expectation-Maximization Attention Mechanism", in IEEE Transactions on Circuits and Systems for Video Technology, vol. 32, no. 3, 2022, pp. 1273-1284) and the experimental results of the present invention are compared and analyzed and verified from both subjective and objective aspects.
[0102] As Figure 5 shown, it is the experimental effect diagram of the fundus OCT image reconstruction method imitating the eagle vision tectum provided by the embodiment of the present invention and other algorithms for reconstructing the OCT images of patients with cystoid macular edema, and a local comparison of the lesions is carried out. Among them, Figure 5 (a) is the original image corresponding to the lesion area, Figure 5 (b) is the reconstruction result diagram of Keys' bicubic method, Figure 5 (c) is the reconstruction result diagram of Dong's SRCNN method, Figure 5 (d) is the reconstruction result diagram of Ledig's SRGAN method, Figure 5 (e) is the reconstruction result diagram of Zheng's IMDN method, Figure 5 (f) is the reconstruction result diagram of Liu's RFDN method, Figure 5 (g) is the reconstruction result diagram of Zhu's EMASRN method, Figure 5 (h) is the reconstruction result diagram of the method of the embodiment of the present invention. Cystoid macular edema is manifested as low-signal intraretinal cystoid spaces and high-signal septa. By observing the enlarged view of the cystoid edema area, it can be seen that the fundus OCT image reconstructed by the method of the embodiment of the present invention has clear contours and strong contrast, and can relatively clearly distinguish layers such as the retinal pigment epithelium layer, the inner and outer photoreceptor layers, and the outer limiting membrane, and can observe the separation of the cystoid arms between different cysts, truly restoring the fundus lesion condition; while the images reconstructed by other methods have the phenomena of visual blurring, poor contrast, and unclear boundaries between different layers. Therefore, the method of the embodiment of the present invention effectively restores the real fundus health condition and improves the contrast of the image.
[0103] As Figure 6 shown, it is the comparison diagram of the reconstruction effects of the fundus OCT image reconstruction method imitating the eagle vision tectum provided by the embodiment of the present invention and other algorithms for patients with serous retinal detachment, and a local comparison of the lesions is carried out. Among them, Figure 6 (a) is the original image corresponding to the lesion area, Figure 6 (b) is the reconstruction result diagram of Keys' bicubic method, Figure 6(c) Reconstruction result graph of Dong's SRCNN method, Figure 6 (d) Reconstruction result graph of Ledig's SRGAN method, Figure 6 (e) Reconstruction result graph of Zheng's IMDN method, Figure 6 (f) Reconstruction result graph of Liu's RFDN method, Figure 6 (g) Reconstruction result graph of Zhu's EMASRN method, Figure 6 (h) Reconstruction result graph of the method of the embodiment of the present invention. Serous retinal detachment is manifested as a shallow detachment of the retina and an optically transparent area between the neuroepithelial layer and the pigment epithelial layer. It can be seen from the observed magnified local pictures that the hierarchical features of the pictures reconstructed by the bicubic, SRCNN, and SRGAN methods are not obvious and the visual effect is not good; although the IMDN, RFDN, and EMASRN methods improve the contrast, the edges of the cysts are relatively blurred and the smaller cysts are difficult to observe. The method of the embodiment of the present invention gradually expands the receptive field from two dimensions of vertical and horizontal to ensure the extraction of high-frequency features, and the small cysts in the diseased part of the reconstructed picture are clearer and the diseased levels are more distinct.
[0104] In this example, to avoid the deviation caused by qualitative analysis, two objective indicators, peak signal-to-noise ratio (PSNR) and structural similarity (SSIM), are used for quantitative evaluation. Through the reconstruction and restoration comparison of two different magnification factors of 2 and 4 times on four test data sets, the results are shown in Table 1 and Table 2 respectively:
[0105] Table 1 Comparison data of PSNR / SSIM results of different methods on different data sets at ×2 magnification
[0106]
[0107] Table 2 Comparison data of PSNR / SSIM results of different methods on different data sets at ×4 magnification
[0108]
[0109] As can be seen from Table 1, at a magnification of ×2, the method of the embodiment of the present invention achieved the optimal PSNR value on three test sets. Compared with EMASRN, the PSNR values of the method of the embodiment of the present invention (our) on the four test sets increased by 0.15 dB, 0.11 dB, 0.12 dB, and 0.11 dB respectively; especially on Test Set 1, the SSIM value of the method of the embodiment of the present invention (our) increased by 2.73% compared with IMDN. As can be seen from Table 2, at a magnification of ×4, the method of the embodiment of the present invention (our) achieved the optimal indicators on all four test sets, and the sub-optimal indicators were obtained by RFDN or EMASRN. Compared with RFDN, the SSIM values of the method of the embodiment of the present invention (our) on the four test sets increased by 0.04%, 0.33%, 0.36%, and 0.52% respectively. Compared with EMASRN, the PSNR values of the method of the embodiment of the present invention (our) on the four test sets increased by 0.24 dB, 0.26 dB, 0.29 dB, and 0.27 dB respectively.
[0110] From the data in Table 1 and Table 2, it can be seen that for PSNR and SSIM, the higher the value, the more similar the result is to the real image, and the higher the image quality. Table 1 and Table 2 clearly show the average scores of the test data of different image data sets under different indicators. Therefore, the method of the embodiment of the present invention has a significant improvement in the peak signal-to-noise ratio and structural similarity of the reconstructed image, the visual quality of the reconstructed image is improved, and the detail features are richer.
[0111] If the image reconstruction method described in the embodiment of the present invention is implemented in the form of software functional modules and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the image reconstruction method described in the embodiment of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0112] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are all included in the protection scope of the present invention.
Claims
1. An ophthalmic OCT image reconstruction system imitating the Eagle Eye top cover, characterized in that, It includes a feature extraction block, an information processing block and a reconstruction block in sequence; The feature extraction block is used to extract shallow features from low-resolution OCT images to obtain useful information; The information processing block imitates the eagle vision tectum information processing mechanism to gradually expand the receptive field from the vertical and horizontal dimensions to extract high-frequency features; the vertical dimension is composed of 6 eagle vision tectum blocks EVB, which are responsible for mining high-frequency features; The horizontal dimension is responsible for fusing the high-frequency features extracted by the vertical dimension, deeply mining significant information, using the spatial attention mechanism to strengthen the deep features, clearly expressing the details such as the edema contour of the lesion area, and transmitting the extracted high-frequency features to the reconstruction module; The reconstruction block is used to perform preliminary reconstruction on the extracted shallow features and high-frequency features, and then perform deep reconstruction after upsampling operation to obtain a reconstructed high-resolution OCT image.
2. The fundus OCT image reconstruction system imitating the eagle eye top cover according to claim 1, wherein The EVB in the vertical dimension of the information processing block is responsible for extracting high-frequency features from the feature maps of each channel and promoting the fusion of different channels by using the channel attention mechanism; The EVB module includes two parts: information deep extraction and information fusion; The deep information extraction part is composed of four convolution kernels with a size of 3×3 and dilated convolutions with dilated rates of 1, 1, 2, and 3 respectively. When each convolution kernel outputs, the number of channels of the original feature is halved to reduce the number of parameters. At the same time, the output of the previous convolution kernel is propagated layer by layer to the subsequent convolution kernel to form a dense connection, so as to facilitate the deep learning of the features by the subsequent convolution kernel. The information fusion part fuses the outputs of the four convolution kernels through concat, uses a 1×1 convolution kernel to transform the number of channels, and introduces a channel attention mechanism to mine significant information; The channel attention mechanism is divided into three parts: global context embedding, channel normalization, and gate adaptation; The global context embedding is normalized using the l2 criterion, and a training parameter α is introduced to facilitate the adaptive output of the l2 normalization; The channel normalization uses the l2 criterion to perform cross-channel normalization between features to reduce the number of parameters; The gate adaptation controls the competition and cooperation relationship between neurons through the weight γ and bias β.
3. The fundus OCT image reconstruction system imitating the eagle eye top cover according to claim 1, characterized in that The spatial attention mechanism divides the extracted features into 64 subspaces in space. Each subspace first uses a depthwise separable convolution to extract new features from each group of features to reduce redundancy within the channel range. After maximum pooling, point-by-point convolution is used to reduce redundancy within the spatial range. Finally, the softmax function is used to scale the feature map in H dimensions to ensure that the sum of the weights is 1. After being fused with the original features, 64 subspace features are obtained. Different subspaces are connected through concat to get the final output.
4. The fundus OCT image reconstruction system imitating the EagleEye top cover according to claim 1, characterized in that, The reconstruction block uses a layer of 3×3 convolution kernels to obtain a basic reconstructed image, and after an upsampling operation, uses a 3×3 convolution kernel to perform deep reconstruction to obtain a clear fundus OCT reconstructed image.
5. A method for reconstructing fundus OCT images imitating the top cover of a hawk's eye, characterized in that, Follow these steps: S1, input the low-resolution fundus OCT image into the feature extraction block, complete the extraction of shallow features, and send it to the information processing block; S2. Information processing block, which extracts high-frequency features by gradually expanding the receptive field from the vertical and horizontal dimensions following the information processing mechanism of the eagle-eyed tectum; the vertical dimension consists of 6 eagle-eyed tectum-like blocks EVB, responsible for mining high-frequency features; The horizontal dimension is responsible for fusing the high-frequency features extracted in the vertical dimension, deeply mining significant information, enhancing deep features using the spatial attention mechanism, clearly expressing detailed information such as the edema contour of the lesion area, and transmitting the extracted high-frequency features to the reconstruction module; S3. Reconstruction block, which is used to preliminarily reconstruct the extracted shallow features and high-frequency features, and then perform upsampling operations and deep reconstruction to obtain the reconstructed high-resolution OCT image.
6. The method for reconstructing fundus OCT images imitating the eagle-eye top cover according to claim 5, wherein The information processing block of S2 extracts high-frequency features according to the following formula in the vertical dimension, F1 = C 3×3 (X i-1 ) F2 = C 3×3 (C 3×3 (F1)) Among them, F k represents the output of the convolutional kernel of the k-th (k = {1, 2, 3, 4}) layer, C 3×3 represents a convolutional kernel of size 3×3, represents a dilated convolution with a convolutional kernel size of 3×3 and a dilation rate of 2, represents a dilated convolution with a convolutional kernel size of 3×3 and a dilation rate of 3, {} represents the concatenation operation, X i-1 and X i are respectively the input and output of the i-th (i = 1,..., n) EVB module, F n represents the output of the convolutional kernel of the n-th (n = {1, 2, 3, 4}) layer, F5 and are respectively the input and output of channel attention, R c represents the output of global context embedding, R = R1, R2,..., R c-1 , R c represents R c decomposed into C (C = 64) feature maps, represents the output of channel normalization, represents the scalar for normalizing the scale, α, β, γ represent trainable parameters, ε represents a constant, f tanh represents the tanh activation function.
7. A fundus OCT image reconstruction method imitating the eagle eye top cover according to claim 5, characterized in that The information processing block of S2 fuses and enhances deep features according to the following formula in the horizontal dimension, X C = C 1×1 ({X1, X2,..., X n-1 , X n [[ID=8}]}) Among them, X C and respectively represent the input and output of the channel attention mechanism, and X i (i = 1,..., n) represents the output of the i-th EVB module, and respectively represent the input and output of the g-th (g = 1,..., 64) sub-feature space, and C 1×1 represents a convolutional kernel of size 1×1, PW 1×1 represents a depthwise separable convolution with a convolutional kernel size of 1×1, DW 1×1 represents a pointwise convolution with a convolutional kernel size of 1×1, f maxpool represents a max pooling operation, and f softmax represents a softmax activation function; The mathematical model of the information processing block in S2 is: X0 = C 3×3 (C 3×3 (X)) Among them, X represents the low-resolution image input to the feature extraction block, and C 3×3 represents a convolutional kernel of size 3×3, and X0 represents the shallow features obtained by the feature extraction block. Among them, X out represents the output of the information processing block.
8. The method for reconstructing fundus OCT images imitating the top cover of a hawk's eye according to claim 5, wherein S3 completes the reconstruction according to the following formula, Y = C 3×3 (f upsampler (C 3×3 (X out ))) Among them, Y represents the output of the reconstruction module, and f upsampler represents the upsampling operation.
9. An electronic device, characterized in that, The image reconstruction is realized by using the method according to any one of claims 5 to 8.
10. A computer storage medium, characterized in that, At least one program instruction is stored in the storage medium, and the at least one program instruction is loaded and executed by the processor to realize the image reconstruction method according to any one of claims 5 to 8.
Citation Information
Patent Citations
Super-resolution multi-scale residual fusion model of single image and restoration method thereof
CN111861961A
Adaptive Interface for a Medical Imaging System
US20140177935A1