Medical video super-resolution video frame extraction method and system
Through the video super-resolution model based on recursive backprojection network, combined with optical flow algorithms and deep cascade technology, the problems of large calculation volume, low efficiency and poor detail reconstruction in the existing technology are solved, and efficient and clear super-resolution reconstruction of medical images are achieved.
Patent Information
- Application Number
- CN202011431866.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-10
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2040-12-10
AI Technical Summary
The existing video super-resolution reconstruction model has a large amount of calculation and low efficiency, and the reconstruction effect of medical image edges, textures and other details is not clear enough, which affects doctors' diagnosis.
The video super-resolution model based on recursive backprojection network is adopted, and the backprojection module and recursive structure are constructed, and the optical flow algorithm is used to perform motion estimation and motion compensation. Combined with deep cascade, projection error correction and dense connection technology, high-quality super-resolution video frames are reconstructed.
It improves the efficiency and quality of video super-resolution reconstruction, significantly improves the reconstruction clarity of medical image edges and texture details, and is suitable for telemedicine video super-resolution reconstruction.
Smart Images

Figure CN114626980B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of smart medical care, and in particular to a method for extracting super-resolution video frames from medical videos. Background Art
[0002] With the development of deep learning, convolutional neural networks (CNNs) have proven to be a powerful tool for computer vision. A CNN-based auxiliary diagnosis model can classify skin disease images produced by confocal laser scanning microscopy almost as accurately as dermatologists [1]. Based on CNN, AT&T has developed telemedicine, where doctors can communicate with patients in need through online video communication and provide them with diagnosis and treatment options; the telemedicine system jointly developed by Peking University People's Hospital and Guangzhou Institute of Respiratory Health provides expert medical advice and diagnosis from top American medical institutions to patients with major diseases such as cancer, cardiovascular disease, and family genetic diseases in my country. In the field of telemedicine, high-quality videos are an important basis for providing patients with accurate treatment options. However, due to the current limitations of hardware manufacturing processes, hardware costs, available storage space, and system transmission conditions, the resolution of some medical videos is low. Limited by low-resolution (LR) videos, doctors cannot clearly observe the patient's lesions and make accurate diagnoses. Therefore, it is very important to reconstruct high-resolution videos from low-resolution videos. Video super-resolution technology can restore the corresponding high-resolution video from low-resolution videos, which is of great significance for doctors to diagnose patients' conditions through telemedicine systems.
[0003] Most of the existing video super-resolution methods are developed based on image super-resolution methods. At present, the most advanced image SR methods are based on deep learning, such as SRCNN, VDSR, and EDSR[2][3][4]. These deep learning-based image super-resolution methods have achieved excellent results in single image super-resolution reconstruction, but they are not suitable for direct use in video super-resolution because these methods ignore the relationship between frames and cannot utilize the useful information of adjacent frames in the same video scene. Subsequently, multi-frame video super-resolution methods such as video super-resolution network (VSRnet), motion compensation and residual network (MCResNet), sub-pixel motion compensation (SPMC), and residual recurrent convolution network (RRCN) were proposed. VSRnet adopts a method combining local-global and total variation (CLG-TV) to reconstruct the central frame with 5 consecutive input frames to improve the reconstruction quality and reduce the training time[5]. ESPCN adopts a new CNN structure to extract feature maps in the LR space and maps the LR features to high-resolution output through an efficient sub-pixel convolution layer to reconstruct the SR video[6]. MCResNet uses an optical flow algorithm for motion estimation and motion compensation as a preprocessing step, and uses a new deep residual convolutional neural network (CNN) to predict high-resolution images using multiple motion compensated observations [7]. These methods use the existing optical flow algorithm to perform motion estimation and motion compensation on multiple input frames, and use the compensated frames and the center frame as input to reconstruct a super-resolution image of the center frame, effectively improving the effect of video super-resolution reconstruction. A panoramic video interpolation method, device, and corresponding storage medium provide a panoramic video interpolation method, which uses an optical flow map to interpolate panoramic videos to improve the video frame rate, without super-resolution reconstruction of video frames [8].
[0004] In summary, existing methods have the following problems:
[0005] (1) Existing video super-resolution reconstruction models are mostly feed-forward network structures that improve the reconstruction effect by stacking network layers. The model has large computational complexity and low efficiency, and is not suitable for telemedicine video super-resolution reconstruction.
[0006] (2) The existing video super-resolution reconstruction model is not clear enough in reconstructing detailed information such as edges and textures of medical images, which is not conducive to doctors' diagnosis. Summary of the invention
[0007] The technical problem to be solved by the present invention is to provide a medical video super-resolution video frame extraction method in view of the deficiencies in the prior art, so as to solve the problems of large computational complexity and low efficiency of the existing video super-resolution reconstruction model.
[0008] In order to solve the above technical problems, the technical solution adopted by the present invention is: a medical video super-resolution video frame reconstruction method, comprising the following steps:
[0009] S1. Extract central frames from medical videos sequentially and the two adjacent frames of the central frame calculate and Optical flow graph and Where t = 2,3…N-1, N is the total number of frames of the medical video;
[0010] S2, two adjacent frames of the central frame Perform motion compensation to obtain compensated frame Where ω represents the Warp operation;
[0011] S3. As the input of the convolutional neural network, reconstruct the central frame of the video Super-resolution video frames
[0012] S4. Reconstruct all super-resolution video frames Get super-resolution medical videos.
[0013] The super-resolution video frames reconstructed by the present invention can comprehensively utilize the spatial and temporal information of low-resolution video frames to improve video clarity, thereby solving the problems of large computational complexity and low efficiency of existing video super-resolution reconstruction models.
[0014] The specific implementation process of step S3 includes:
[0015] 1) Fusion characteristics, and obtain
[0016] 2) Set the number of iterations i = 1;
[0017] 3) Use upsampling units to transform features Zoom in to the specified multiple and get Then Reduce to original size and Subtract and get the residual Will Zoom in to the specified multiple and Add together and get
[0018] 4) Use the downsampling unit to Reduce to original size Will Zoom in to the specified multiple and Subtract the residual Will Scale back to original size and Add together and get
[0019] 5) As the input of the up-sampling unit, and let the value of i increase by 1, return to step 3), when the number of iterations i=K, stop the iteration; 6)
[0021] 1) Concatenate the output vectors of the upsampling units in the 1st to Kth iterations, and use the output convolution layer to reduce the dimension of the concatenated vectors to three-dimensional space to output the super-resolution video frame.
[0022] The present invention uses the connection layer to fuse the output of the upsampling unit, and can utilize the upsampling information of different depths of the neural network to improve the quality of video frame reconstruction.
[0023] After step S3 and before step S4, the process further includes: using the super-resolution video frame as the input of the convolution layer to obtain a super-resolution image of the central frame, and using the L1 norm As the loss function, the gradient descent method is used to make the super-resolution reconstruction result of the central frame of the medical video consistent with the real image. Stay consistent; Represents a super-resolution frame reconstructed using a convolutional neural network.
[0024] The L1 norm can constrain the pixels of super-resolution video frames to be consistent with the real image, which can improve the convergence speed and stability during network training.
[0025] The upsampling unit includes two deconvolution layers and a first convolution layer; wherein,
[0026] The first deconvolution layer is used to transform the features Zoom in to the specified multiple and get
[0027] The first convolutional layer is used to Reduce to original size and compare with input Subtract and get the residual
[0028] The second deconvolution layer is used to Zoom in to the specified multiple and Add and get the output
[0029] The upsampling unit uses the residual to correct the output of the first deconvolution layer, which can reduce the amount of calculation and improve the upsampling quality.
[0030] The down sampling unit comprises:
[0031] The second convolutional layer is used to convert the input Scaling back to the original size gives
[0032] The third deconvolution layer is used to Zoom in to the specified multiple and compare with the input Subtract the residual
[0033] The fourth convolutional layer is used to Scale back to original size and Add and get the output
[0034] The downsampling unit uses the residual to correct the output of the second convolutional layer, which can reduce the amount of calculation and improve the downsampling quality.
[0035] A medical video super-resolution video frame reconstruction system, comprising:
[0036] Feature extraction unit, used to sequentially extract central frames from medical videos and the two adjacent frames of the central frame calculate and Optical flow graph and For the two adjacent frames of the central frame Perform motion compensation to obtain compensated frame Where ω represents the Warp operation; where t = 2,3…N-1, N is the total number of frames of the medical video;
[0037] Convolutional neural network, As input, reconstruct the center frame of the video Super-resolution video frames
[0038] Reconstruction unit, used to reconstruct all super-resolution video frames Get super-resolution medical videos.
[0039] The reconstructed super-resolution video frames can comprehensively utilize the spatial and temporal information of the low-resolution video frames to improve the video clarity.
[0040] The feature extraction unit comprises:
[0041] The third convolutional layer is used to extract Features;
[0042] The fourth convolutional layer is used to extract Features;
[0043] The fifth convolutional layer is used to extract Features;
[0044] The first connection layer is used to fuse the output results of the third convolutional layer, the fourth convolutional layer, and the fifth convolutional layer;
[0045] The sixth convolutional layer is used to extract the features of the output result
[0046] The third, fourth, and fifth convolutional layers are used to extract features of low-resolution video frames, respectively, so that detailed information of low-resolution video frames can be extracted, so that the fusion features extracted by the sixth convolutional layer have rich feature information.
[0047] The convolutional neural network comprises:
[0048] A plurality of stacking units connected in series, each of the stacking units comprising an up-sampling unit and a down-sampling unit connected in sequence;
[0049] Among them, the upsampling unit of the first stacking unit is used to use the upsampling unit to convert the feature Zoom in to the specified multiple and get Then Reduce to original size and Subtract and get the residual Will Zoom in to the specified multiple and Add together and get
[0050] The downsampling unit of the first stacking unit is used to use the downsampling unit to Reduce to original size Will Zoom in to the specified multiple and Subtract the residual Will Scale back to original size and Add together and get
[0051] The output end of the down-sampling unit of the last stacking unit is connected to the input end of the up-sampling unit of the second stacking unit;
[0052] The upsampling units of all stacked units are connected to the output upsampling unit; the output upsampling unit is connected to the seventh convolutional layer through the second connection layer, and the seventh convolutional layer outputs the video center frame Super-resolution video frames
[0053] The second connection layer cascades the outputs of upsampling units at different depths, which can mine super-resolution information at different levels of the neural network, making the reconstructed super-resolution video frames rich in high-frequency information and clear in visual perception.
[0054] The first stacking unit comprises:
[0055] The upsampling unit includes a first deconvolution layer, a second deconvolution layer, and a first convolution layer; wherein,
[0056] The first deconvolution layer is used to transform the features Zoom in to the specified multiple and get
[0057] The first convolutional layer is used to Reduce to original size and compare with input Subtract and get the residual
[0058] The second deconvolution layer is used to Zoom in to the specified multiple and Add and get the output
[0059] The downsampling unit includes a third deconvolution layer, a fourth deconvolution layer, and a second convolution layer; wherein,
[0060] The second convolutional layer is used to convert the input Scaling back to the original size gives
[0061] The third deconvolution layer is used to Zoom in to the specified multiple and compare with the input Subtract the residual
[0062] The fourth deconvolution layer is used to Scale back to original size and Add and get the output
[0063] The stacking unit enables the neural network to learn the reconstruction information from low-resolution video frames to super-resolution video frames. Combined with the loop structure designed in the present invention, it can learn the reconstruction information from low-resolution video frames to super-resolution video frames multiple times, thereby improving the quality of re-resolution reconstruction.
[0064] The optimization unit is also included, which is used to use the super-resolution video frame as the input of the convolution layer to obtain a super-resolution image of the central frame, and use the L1 norm As the loss function, the gradient descent method is used to make the super-resolution reconstruction result of the central frame of the medical video consistent with the real image. Stay consistent; Represents a super-resolution frame reconstructed using a convolutional neural network.
[0065] Compared with the prior art, the present invention has the following beneficial effects: in view of the problem that the existing video super-resolution reconstruction model has a large amount of calculation and low efficiency, the present invention proposes a video super-resolution model based on a recursive back-projection network (a network structure with multiple stacked units connected in series), which fully exploits the nonlinear mapping between low-resolution video frames and super-resolution video frames by constructing a back-projection module (alternating up- and down-sampling layers), avoids using the stacked network layer method to exploit the nonlinear mapping between low-resolution video frames and super-resolution video frames, and further reduces model parameters and improves model efficiency through a recursive structure to adapt to telemedicine video super-resolution reconstruction. In view of the problem that the existing video super-resolution reconstruction model does not have a clear effect on the reconstruction of medical image edge, texture and other detail information, the present invention proposes a deep cascade, projection error correction and dense connection technology based on a recursive back-projection network, which fuses the super-resolution images of different depths of the recursive back-projection network, and uses the input video frame to correct the output video frame through a shortcut connection and dense connection technology to improve the edge, texture and other detail information of the medical image, so that the reconstructed super-resolution image is clearer, which can effectively assist doctors in diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 This is a schematic diagram of the method of an embodiment of the present invention;
[0067] Figure 2 Comparison of reconstruction effects between gastroscopy video frames and retinopathy video frames. DETAILED DESCRIPTION
[0068] The video super-resolution method includes three steps: motion estimation, motion compensation and super-resolution reconstruction. Motion estimation is used to estimate the motion between low-resolution image frames. Motion compensation is to align the adjacent frame images with the current frame in the same coordinate system, and predict and compensate the current frame. Super-resolution reconstruction uses multiple frames after motion compensation as input to reconstruct the super-resolution image of the central frame. In order to train an end-to-end medical video super-resolution model, this application uses the CNN-based optical flow algorithm FlowNet2 for motion estimation and motion compensation [9]
[10] . The FlowNet2 optical flow algorithm can improve the speed of optical flow estimation while ensuring accuracy, and can solve the problems of large displacement optical flow estimation and small displacement optical flow estimation at the same time, and its prediction ability can be continuously improved through learning. First, FlowNet2 is used to estimate the motion of the two adjacent frames of the central frame, and then the Warp operation based on FlowNet2 is used to perform motion compensation on the adjacent frames. Finally, the compensated multiple frames are fused, and the fused multiple frames are used as input to super-resolve the central frame image.
[0069] In order to improve the efficiency and quality of medical video super-resolution reconstruction, this application proposes a recursive back-projection super-resolution network structure RBPN (Recursive Back-Projection Networks) based on video frame motion estimation and motion compensation. RBPN includes three parts: feature fusion and extraction, nonlinear mapping and super-resolution reconstruction. First, three convolutional layers (conv) are used to extract the features of the central frame and two adjacent compensation frames respectively. After the extracted features are fused using concat, a convolutional layer is used to extract the features of the fused frame as input; secondly, the nonlinear mapping from low-resolution frames to super-resolution frames is learned by constructing a back-projection structure in which up- and down-sampling units are placed alternately. This application proposes to alternately place 4 up-sampling units and 3 down-sampling units, and set a recursive loop for the second and third pairs of up- and down-sampling units (the number of loops is 8), which can improve the super-resolution effect without increasing the parameters of the network model; finally, after the outputs of all up-sampling units are concat fused, a convolutional layer is used to reconstruct the super-resolution image of the central frame.
[0070] The steps of motion estimation, motion compensation, multi-frame fusion and super-resolution in medical video super-resolution reconstruction are as follows:
[0071] Step 1: Extract central frames from medical videos sequentially and two adjacent frames Calculated using FlowNet2 optical flow algorithm and Optical flow graph and Where t = 2, 3…N-1, N is the total number of video frames
[10] .
[0072] Step 2: Based on the adjacent frame optical flow map generated in the first step, use the Warp operation to transform the adjacent frames.
[0073] Perform motion compensation to obtain compensated frame and where ω represents the Warp operation
[10] .
[0074] Step 3: Use 3 convolutional layers (2DConv) to extract The proposed features are fused using the Concat operation, and the fused features are extracted using a convolutional layer (2DConv)
[0075] Step 4: Construct upsampling and downsampling units. The upsampling unit includes 2 deconvolution layers (Deconv) and 1 convolution layer. The upsampling unit uses a shortcut connection to connect the input Subtract the output of the middle convolutional layer to get the output of the projection error correction upsampling unit The downsampling unit consists of 2 convolutional layers and 1 deconvolutional layer. The downsampling unit uses a shortcut connection to connect the input Subtract the output of the intermediate deconvolution layer to get the projection error correction upsampling unit output
[0076] Step 5: Construct a back-projection network structure with up-sampling and down-sampling units stacked in sequence (such as Figure 1 As shown), 4 up-sampling units and 3 down-sampling units are placed to realize nonlinear mapping from low-resolution video frames to super-resolution video frames.
[0077] Step 6: Set a recursive loop structure for the second and third groups of up and down sampling units, and set the number of loops to 8. Through the recursive loop structure, a deeper level of super-resolution reconstruction features can be learned (the recursive loop structure expands the 2 groups of up and down sampling units to 16 groups of up and down sampling units) without changing the number of parameters.
[0078] Step 7: Output of all upsampling units The Concat operation is used for fusion, and the upsampled outputs at different depths are used to reconstruct the super-resolution video frame of the central frame.
[0079] Step 8: Based on the fusion features obtained in step 7, use a convolutional layer to obtain a super-resolution image of the central frame and use the L1 norm As the loss function (where represents a high-resolution frame, represents the super-resolution frame reconstructed using the back-projection network), optimizing the super-resolution reconstruction results of the central frame of the medical video.
[0080] The present invention adopts the deep learning method in artificial intelligence, combined with the video super-resolution technology, and invents a medical video super-resolution method based on recursive back-projection network. It can use shallow neural networks (4 upsampling, 3 downsampling layers) to reconstruct low-resolution medical videos into super-resolution videos with clear edges and detail information (relative to interpolation methods, methods based on residual networks, etc.), and can improve the efficiency of medical video super-resolution (the time to reconstruct 500 frames is about 3.53 seconds). Experiments have proved that the medical video super-resolution method MABPN based on FlowNet2 optical flow algorithm and recursive back-projection network of the present invention can improve the quality and efficiency of medical super-resolution videos. The objective indicators of the video (peak signal-to-noise ratio PSNR and structural similarity SSIM) and reconstruction time are better than the comparison methods (SRCNN, EDSR, Video Enhancer). The experimental comparison results are shown in Table 1 and Figure 2As shown in Table 1, the medical video super-resolution method MRBPN proposed in this paper is superior to the comparative methods EDSR and Video Enhancer in terms of the objective evaluation index PSNR. PSNR is the most widely used objective image evaluation index, which is an error-sensitive image quality evaluation index based on corresponding pixels. MRBPN has the highest PSNR value, indicating that the reconstructed super-resolution image is closer to the real image at the pixel level. Figure 2 It can be seen from the above that the medical images reconstructed by MRBPN have no structural deformation and no artifacts compared with EDSR and SRCNN, and are closer to the real image. In addition, it can be seen from Table 1 that MRBPN has the highest reconstruction efficiency, and it takes 3.53 seconds to reconstruct 500 frames of gastroscopy video, which is better than the comparison method EDSR's 224.45 seconds and Video Enhancer's 10 seconds.
[0081] Table 1: Comparison of objective indicators of medical video reconstruction
[0082]
[0083] References
[0084] [1]Guo K, Li T, Huang R, et al.DDA: A deep neural network-based cognitive system for IoT-aided dermatosis discrimination. Ad Hoc Networks, 2018,80: 95-103.018
[0085] [2]Dong C, Loy CC, He K, et al. Learning a deep convolutional network for image super-resolution. / / European Conference on Computer Vision, Springer, Cham, 2014:184-199.
[0086] [3] Kim J, Kwon Lee J, Mu Lee K. Accurate image super-resolution using very deep convolutional networks. / / Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016: 1646-1654.
[0087] [4]Bee Lim, Sanghyun Son, Heewon Kim, Seungjun Nah, and Kyoung MuLee. Enhanced Deep Residual Networks for Single Image SR. 10.1109 / CVPRW.2017.151:136-144.
[0088] [5]Kappeler A, Yoo S, Dai Q, et al. Video super-resolution with convolutional neural networks. IEEE Transactions on Computational Imaging, 2016, 2(2):109-122.
[0089] [6]Wenzhe Shi,Jose Caballero,FerencHuszar,Johannes Totz,AndrewP.Aitken,Rob Bishop,Daniel Rueckert,Zehan Wang.Real-Time Single Image andVideo SR Using an Efficient Sub-Pixel Convolutional Neural Network.InCVPR2016:1874-1883.
[0090] [7]Li D,Wang Z.Video Super-Resolution via Motion Compensation and Deep Residual Learning.IEEE Transactions on Computational Imaging, 2017,1(1):749-762.
[0091] [8] Chen Dan, Zhang Yuyao, Tan Zhigang. Panoramic video frame insertion method, device and corresponding storage medium: China, 111372087A[P / OL]. 2020-07-03.
[0092] [9]Dosovitskiy A,Fischer P,Ilg E, P,Hazirbas C,Golkov C,SmagtP, Cremers D,and Brox T.Flownet:Learning optical flow with convolutionalnetworks. / / In IEEE International Conference on ComputerVision(ICCV),2015:2758-2766.
[0093]
[10] IlgE,Mayer N,Saikia T,et al.FlowNet 2:Evolution of Optical FlowEstimation with Deep Networks. / / Proceedings ofthe IEEE Conference on ComputerVision and Pattern Recognition,2017,1647-1655.
Claims
1. A medical video super-resolution video frame reconstruction method, characterized in that: The following steps are involved: S1. Extract central frames from medical videos sequentially and the two adjacent frames of the central frame calculate and Optical flow graph and Where t = 2, 3…N-1, N is the total number of frames of the medical video; S2, two adjacent frames of the central frame Perform motion compensation to obtain compensated frame Where ω represents the Warp operation; S3. As the input of the convolutional neural network, reconstruct the central frame of the video Super-resolution video frames S4. Reconstruct all super-resolution video frames Obtain super-resolution medical video; The specific implementation process of step S3 includes: 1) Fusion characteristics, and obtain 2) Set the number of iterations i = 1; 3) Use upsampling units to transform features Zoom in to the specified multiple and get Then Reduce to original size and Subtract and get the residual Will Zoom in to the specified multiple and Add together and get 4) Use the downsampling unit to Reduce to original size Will Zoom in to the specified multiple and Subtract the residual Will Scale back to original size and Add together and get 5) As the input of the up-sampling unit, and let the value of i increase by 1, return to step 3), when the number of iterations i=K, stop the iteration; 6) Concatenate the output vectors of the upsampling units in the 1st to Kth iterations, and use the output convolution layer to reduce the dimension of the concatenated vectors to three-dimensional space, and output the super-resolution video frame; After step S3 and before step S4, the method further includes: The super-resolution video frame is used as the input of the convolutional layer to obtain the super-resolution image of the central frame, and the L1 norm is used to As the loss function, the gradient descent method is used to make the super-resolution reconstruction result of the central frame of the medical video consistent with the real image. Stay consistent; Represents a super-resolution frame reconstructed using a convolutional neural network.
2. The medical video super-resolution video frame reconstruction method according to claim 1, characterized in that: The upsampling unit includes two deconvolution layers and a first convolution layer; wherein, The first deconvolution layer is used to transform the features Zoom in to the specified multiple and get The first convolutional layer is used to Reduce to original size and compare with input Subtract and get the residual The second deconvolution layer is used to Zoom in to the specified multiple and Add and get the output 3. The medical video super-resolution video frame reconstruction method according to claim 1, characterized in that: The down sampling unit comprises: The second convolutional layer is used to transform the input Scaling back to the original size gives The third deconvolution layer is used to Zoom in to the specified multiple and compare with the input Subtract the residual The fourth deconvolution layer is used to Scale back to original size and Add and get the output 4. A medical video super-resolution video frame reconstruction system, characterized in that: include: Feature extraction unit, used to sequentially extract central frames from medical videos and the two adjacent frames of the central frame calculate and Optical flow graph and For the two adjacent frames of the central frame Perform motion compensation to obtain compensated frame Where ω represents the Warp operation; where t = 2,3…N-1, N is the total number of frames of the medical video; Convolutional neural network, As input, reconstruct the center frame of the video Super-resolution video frames Reconstruction unit, used to reconstruct all super-resolution video frames Obtain super-resolution medical video; The feature extraction unit comprises: The third convolutional layer is used to extract Features; The fourth convolutional layer is used to extract Features; The fifth convolutional layer is used to extract Features; The first connection layer is used to fuse the output results of the third convolutional layer, the fourth convolutional layer, and the fifth convolutional layer; The sixth convolutional layer is used to extract the features of the output result The convolutional neural network comprises: A plurality of stacking units connected in series, each of the stacking units comprising an up-sampling unit and a down-sampling unit connected in sequence; Among them, the upsampling unit of the first stacking unit is used to use the upsampling unit to convert the feature Zoom in to the specified multiple and get Then Reduce to original size and Subtract and get the residual Will Zoom in to the specified multiple and Add together and get The downsampling unit of the first stacking unit is used to use the downsampling unit to Reduce to original size Will Zoom in to the specified multiple and Subtract the residual Will Scale back to original size and Add together and get The output end of the down-sampling unit of the last stacking unit is connected to the input end of the up-sampling unit of the second stacking unit; The upsampling units of all stacked units are connected to the output upsampling unit; the output upsampling unit is connected to the seventh convolutional layer through the second connection layer, and the seventh convolutional layer outputs the video center frame Super-resolution video frames The first stacking unit includes: The upsampling unit includes a first deconvolution layer, a second deconvolution layer, and a first convolution layer; wherein the first deconvolution layer is used to convert the feature Zoom in to the specified multiple and get The first convolutional layer is used to Reduce to original size and compare with input Subtract and get the residual The second deconvolution layer is used to Zoom in to the specified multiple and Add and get the output The downsampling unit includes a third deconvolution layer, a fourth deconvolution layer, and a second convolution layer; wherein the second convolution layer is used to convert the input Scaling back to the original size gives The third deconvolution layer is used to Zoom in to the specified multiple and compare with the input Subtract the residual The fourth deconvolution layer is used to Scale back to original size and Add and get the output The optimization unit is also included, which is used to use the super-resolution video frame as the input of the convolution layer to obtain a super-resolution image of the central frame, and use the L1 norm As the loss function, the gradient descent method is used to make the super-resolution reconstruction result of the central frame of the medical video consistent with the real image. Stay consistent; Represents a super-resolution frame reconstructed using a convolutional neural network.
Citation Information
Patent Citations
Panoramic video frame insertion method and device and corresponding storage medium
CN111372087A
Blurred video super-resolution method and system based on deep learning
CN110458756A
Video super-resolution reconstruction method based on multi-frame fusion optical flow
CN111311490A