Video processing method, device, computer equipment and medium
By acquiring feature information and difference information between video frames and determining the conversion parameters, the problem of inaccurate acquisition of super-resolution information in the prior art is solved, and the accuracy of super-resolution reconstruction and video processing is improved.
Patent Information
- Application Number
- CN202210273317.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-18
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2042-03-18
AI Technical Summary
The existing super-resolution reconstruction technology is difficult to accurately acquire super-resolution information of images, resulting in reduced accuracy of super-resolution reconstruction and video processing.
By obtaining image feature information and difference information of the front and back frame images of the current frame image, the conversion parameters are determined using these information, and then super-resolved information of the current frame image is obtained.
The accuracy of super-resolution reconstruction and video processing are improved, and the super-resolution information of the current frame image can be accurately obtained.
Smart Images

Figure CN114612841B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of video technology, and in particular to a video processing method, apparatus, computer equipment and medium. Background Art
[0002] In the field of video technology, super-resolution reconstruction technology has a wide range of applications and research significance, and with the development of deep learning, super-resolution reconstruction technology based on convolutional neural networks has achieved rapid development. Among them, super-resolution reconstruction technology uses low-resolution images to reconstruct high-resolution images with higher pixel density and more complete details at the corresponding moment.
[0003] At present, video super-resolution technology based on convolutional neural networks usually involves: using two-dimensional convolution, three-dimensional convolution or other types of convolution to construct a convolutional neural network, extracting super-resolution information of multiple frames of images included in the video based on the convolutional neural network, so as to obtain the detailed features required to reconstruct high-resolution images, thereby converting the multiple frames of low-resolution images included in the video into multiple frames of high-resolution images.
[0004] However, the currently used super-resolution reconstruction technology still has difficulty in accurately acquiring super-resolution information of images, which reduces the accuracy of super-resolution reconstruction and the accuracy of video processing. Summary of the invention
[0005] The present disclosure provides a video processing method, device, computer equipment and medium, which can accurately obtain super-resolution information of the current frame image, improve the accuracy of super-resolution reconstruction, and improve the accuracy of video processing. The technical solution of the present disclosure is as follows:
[0006] According to a first aspect of an embodiment of the present disclosure, a video processing method is provided, the method comprising:
[0007] For the i-th frame image in the video, image feature information and first difference information of the i-1-th frame image and image feature information and second difference information of the i+1-th frame image are obtained, where i is a positive integer greater than 1, the image feature information represents the detail feature of the corresponding image, the first difference information represents the difference between the corresponding image and a subsequent frame image of the image in the detail feature, and the second difference information represents the difference between the corresponding image and a previous frame image of the image in the detail feature;
[0008] Based on the image feature information of the i-1th frame image and the first difference information, first conversion information is determined; based on the image feature information of the i+1th frame image and the second difference information, second conversion information is determined, the first conversion information indicating the parameters required when the i-1th frame image is converted to the i-th frame image, and the second conversion information indicating the parameters required when the i+1th frame image is converted to the i-th frame image;
[0009] Determine super-resolution information of the i-th frame of image based on the image feature information of the i-th frame of image, the first conversion information, and the second conversion information;
[0010] A super-resolution video is obtained based on super-resolution information of multiple frames of images in the video.
[0011] In the disclosed embodiment, for the i-th frame in the video, by obtaining the image feature information and the first difference information of the i-1th frame image and the image feature information and the second difference information of the i+1th frame image, and then using the image feature information and the first difference information of the i-1th frame image, the feature information of the i-1th frame image is converted to the i-th frame image to obtain the first conversion information, and using the image feature information and the second difference information of the i+1th frame image, the feature information of the i+1th frame image is converted to the i-th frame image to obtain the second conversion information, and then using the image feature information, the first conversion information and the second conversion information of the i-th frame image to obtain the super-resolution information of the i-th frame image. In this way, when determining the super-resolution information of the i-th frame image, not only the conversion information of the previous frame image to the current frame image is referenced, but also the conversion information of the next frame image to the current frame is referenced, which increases the amount of referenced information, can accurately obtain the super-resolution information of the current frame image, improves the accuracy of super-resolution reconstruction, and then based on multiple frames in the video, can accurately obtain super-resolution video, improves the accuracy of video processing.
[0012] In some embodiments, the process of acquiring the image feature information and the first difference information of the i-1th frame image includes:
[0013] Inputting the i-1th frame image and the adjacent frame image of the i-1th frame image into a feature extraction network, and extracting the hidden layer features of the i-1th frame image based on the i-1th frame image and the adjacent frame image of the i-1th frame image through the feature extraction network, wherein the feature extraction network is trained based on at least one sample frame image, the adjacent frame image of the at least one sample frame image, and the hidden layer features of the at least one sample frame image;
[0014] Based on the hidden layer features of the i-1th frame image, image feature information of the i-1th frame image is determined; based on the hidden layer features of the i-1th frame image, the i-th frame image and the i-1th frame image, first difference information of the i-1th frame image is determined.
[0015] In the disclosed embodiment, for the i-1th frame image, the hidden layer features of the i-1th frame image are extracted through a feature extraction network, that is, the detail features of the i-1th frame image are extracted, and then the hidden layer features of the i-1th frame image are used to determine the image feature information and the first difference information of the i-1th frame image, thereby improving the accuracy of determining the image feature information and the first difference information.
[0016] In some embodiments, when i is a positive integer greater than 2, the method further includes:
[0017] The hidden layer features of the i-1th frame image, the adjacent frame images of the i-1th frame image, and the i-2th frame image are input into the feature extraction network. The hidden layer features of the i-1th frame image are extracted based on the hidden layer features of the i-1th frame image, the adjacent frame images of the i-1th frame image, and the i-2th frame image through the feature extraction network. The feature extraction network is trained based on the hidden layer features of at least one frame sample image, the adjacent frame images of the at least one frame sample image, the previous frame image of the at least one frame sample image, and the hidden layer features of the at least one frame sample image.
[0018] In the disclosed embodiment, when determining the hidden layer features of the i-1th frame image, the hidden layer features of the i-2th frame image are also referenced. In this way, the hidden layer features of the previous frame image are referenced to determine the hidden layer features of the current frame image, thereby increasing the amount of referenced information, being able to accurately obtain the hidden layer features of the current frame image, and improving the accuracy of obtaining the hidden layer features.
[0019] In some embodiments, based on the hidden layer features of the i-1th frame image, determining the image feature information of the i-1th frame image, and based on the i-1th frame image, the i-th frame image and the hidden layer features of the i-1th frame image, determining the first difference information of the i-1th frame image includes:
[0020] Inputting the hidden layer features of the i-1th frame image into a first feature extraction subnetwork, extracting image feature information of the i-1th frame image based on the hidden layer features of the i-1th frame image through the first feature extraction subnetwork, wherein the first feature extraction subnetwork is trained based on the hidden layer features of at least one frame of sample image and the image feature information of the at least one frame of sample image;
[0021] The i-1th frame image, the i-th frame image and the hidden layer features of the i-1th frame image are input into a second feature extraction subnetwork, and the first difference information of the i-1th frame image is extracted based on the i-1th frame image, the i-th frame image and the hidden layer features of the i-1th frame image through the second feature extraction subnetwork. The second feature extraction subnetwork is trained based on at least one frame sample image, the subsequent frame image of the at least one frame sample image, the hidden layer features of the at least one frame sample image and the first difference information of the at least one frame sample image.
[0022] In the disclosed embodiment, for the i-1th frame image, the image feature information of the i-1th frame image can be quickly extracted through the first feature extraction subnetwork, thereby improving the accuracy of obtaining the image feature information, and, through the second feature extraction subnetwork, the first difference information of the i-1th frame image can be quickly extracted, thereby improving the accuracy of obtaining the first difference information.
[0023] In some embodiments, the process of acquiring the image feature information and the second difference information of the (i+1)th frame image includes:
[0024] Inputting the i+1th frame image and the adjacent frame image of the i+1th frame image into a feature extraction network, extracting hidden layer features of the i+1th frame image based on the i+1th frame image and the adjacent frame image of the i+1th frame image through the feature extraction network, wherein the feature extraction network is trained based on at least one sample frame image, the adjacent frame image of the at least one sample frame image, and the hidden layer features of the at least one sample frame image;
[0025] Based on the hidden layer features of the i+1th frame image, the image feature information of the i+1th frame image is determined; based on the hidden layer features of the i+1th frame image, the i-th frame image and the i+1th frame image, the second difference information of the i+1th frame image is determined.
[0026] In the disclosed embodiment, for the i+1th frame image, the hidden layer features of the i+1th frame image are extracted through a feature extraction network, that is, the detail features of the i+1th frame image are extracted, and then the hidden layer features of the i+1th frame image are used to determine the image feature information and the second difference information of the i+1th frame image, thereby improving the accuracy of determining the image feature information and the second difference information.
[0027] In some embodiments, the method further comprises:
[0028] The i+1th frame image, the adjacent frame images of the i+1th frame image, and the hidden layer features of the i+1th frame image are input into the feature extraction network. The hidden layer features of the i+1th frame image are extracted based on the hidden layer features of the i+1th frame image, the adjacent frame images of the i+1th frame image, and the i-th frame image through the feature extraction network. The feature extraction network is trained based on the hidden layer features of at least one frame sample image, the adjacent frame images of the at least one frame sample image, the previous frame image of the at least one frame sample image, and the hidden layer features of the at least one frame sample image.
[0029] In the disclosed embodiment, when determining the hidden layer features of the i+1th frame image, the hidden layer features of the i-th frame image are also referenced. In this way, the hidden layer features of the previous frame image are referenced to determine the hidden layer features of the current frame image, thereby increasing the amount of referenced information, being able to accurately obtain the hidden layer features of the current frame image, and improving the accuracy of obtaining the hidden layer features.
[0030] In some embodiments, based on the hidden layer features of the i+1th frame image, determining the image feature information of the i+1th frame image, and based on the hidden layer features of the i+1th frame image, the i-th frame image and the i+1th frame image, determining the second difference information of the i+1th frame image includes:
[0031] Inputting the hidden layer features of the i+1th frame image into a first feature extraction subnetwork, extracting image feature information of the i+1th frame image based on the hidden layer features of the i+1th frame image through the first feature extraction subnetwork, wherein the first feature extraction subnetwork is trained based on the hidden layer features of at least one frame of sample image and the image feature information of the at least one frame of sample image;
[0032] The i+1th frame image, the i-th frame image and the hidden layer features of the i+1th frame image are input into a third feature extraction subnetwork, and the second difference information of the i+1th frame image is extracted based on the hidden layer features of the i+1th frame image, the i-th frame image and the i+1th frame image through the third feature extraction subnetwork. The third feature extraction subnetwork is trained based on at least one frame sample image, the previous frame image of the at least one frame sample image, the hidden layer features of the at least one frame sample image and the second difference information of the at least one frame sample image.
[0033] In the disclosed embodiment, for the i+1th frame image, the image feature information of the i+1th frame image can be quickly extracted through the first feature extraction subnetwork, thereby improving the accuracy of obtaining the image feature information, and, through the third feature extraction subnetwork, the second difference information of the i+1th frame image can be quickly extracted, thereby improving the accuracy of obtaining the second difference information.
[0034] In some embodiments, the feature extraction network is composed of multiple residual modules, wherein a residual module includes a first two-dimensional convolutional layer, an activation function connected to the first two-dimensional convolutional layer, and a second two-dimensional convolutional layer connected to the activation function, and the activation function is used to indicate the functional mapping relationship between the corresponding image, the adjacent frame images of the image and the hidden layer features of the image.
[0035] In the disclosed embodiment, by setting a residual module, the problem of gradient disappearance when the hidden layer is deepened can be avoided.
[0036] In some embodiments, the process of acquiring the image feature information and the first difference information of the i-1th frame image includes:
[0037] Based on the i-1th frame image and the i-th frame image, determine the optical flow feature information of the i-1th frame image and the first optical flow information, the optical flow feature information represents the optical flow feature of the corresponding image, and the first optical flow information represents the pixel movement between the corresponding image and a subsequent frame image of the image;
[0038] Interpolation processing is performed on the optical flow feature information of the i-1th frame image and the first optical flow information respectively, and the optical flow feature information and the first optical flow information after the interpolation processing are determined as the image feature information and the first difference information of the i-1th frame image.
[0039] In the disclosed embodiment, for the i-1th frame image, by extracting the optical flow feature information of the i-1th frame image and then performing interpolation processing on the optical flow feature information, the detail features of the i-1th frame image can be quickly determined, that is, the image feature information of the i-1th frame image is determined, and, by extracting the first optical flow information of the i-1th frame image, the pixel movement between the i-1th frame image and the i-th frame image can be quickly determined, and then the first optical flow information is interpolated, the difference in detail features between the i-1th frame image and the i-th frame image can be quickly determined, that is, the first difference information of the i-1th frame image is determined, thereby improving the efficiency and accuracy of determining the image feature information and the first difference information.
[0040] In some embodiments, the process of acquiring the image feature information and the second difference information of the (i+1)th frame image includes:
[0041] Based on the (i+1)th frame image and the (i)th frame image, determining optical flow feature information of the (i+1)th frame image and second optical flow information, the optical flow feature information representing the optical flow feature of the corresponding image, and the second optical flow information representing the pixel movement between the corresponding image and a previous frame image of the image;
[0042] Interpolation processing is performed on the optical flow feature information of the (i+1)th frame image and the second optical flow information respectively, and the optical flow feature information and the second optical flow information after the interpolation processing are determined as the image feature information of the (i+1)th frame image and the second difference information.
[0043] In the disclosed embodiment, for the i+1th frame image, by extracting the optical flow feature information of the i+1th frame image and then performing interpolation processing on the optical flow feature information, the detail features of the i+1th frame image can be quickly determined, that is, the image feature information of the i+1th frame image is determined, and, by extracting the first optical flow information of the i+1th frame image, the pixel movement between the i+1th frame image and the i-th frame image can be quickly determined, and then the first optical flow information is interpolated, so that the difference in detail features between the i+1th frame image and the i-th frame image can be quickly determined, that is, the second difference information of the i+1th frame image is determined, thereby improving the efficiency and accuracy of determining the image feature information and the second difference information.
[0044] In some embodiments, determining the first conversion information based on the image feature information of the (i-1)th frame image and the first difference information, and determining the second conversion information based on the image feature information of the (i+1)th frame image and the second difference information includes:
[0045] Determine a difference between the image feature information of the (i-1)th frame image and the first difference information of the (i-1)th frame image as the first conversion information;
[0046] A difference between the image feature information of the (i+1)th frame image and the second difference information of the (i+1)th frame image is determined as the second conversion information.
[0047] In the disclosed embodiment, by taking the difference method, the feature information of the i-1th frame image can be quickly converted to the i-th frame image, and the feature information of the i+1th frame image can be converted to the i-th frame image, so that the conversion information of the previous frame image to the current frame image and the conversion information of the next frame image to the current frame can be used to determine the super-resolution information of the current frame image, thereby improving the accuracy of super-resolution reconstruction.
[0048] In some embodiments, determining the super-resolution information of the i-th frame of image based on the image feature information of the i-th frame of image, the first conversion information, and the second conversion information includes:
[0049] The image feature information of the i-th frame image, the first conversion information and the second conversion information are input into a temporal convolutional network, and convolution processing is performed based on the image feature information of the i-th frame image, the first conversion information and the second conversion information through the temporal convolutional network to obtain super-resolution information of the i-th frame image. The temporal convolutional network is trained based on the image feature information of at least one frame of sample image, the first conversion information, the second conversion information and the super-resolution information of the at least one frame of sample image.
[0050] In the disclosed embodiment, for the i-th frame image, the image feature information of the i-th frame image, the first conversion information and the second conversion information are convolved through a temporal convolutional network, so that the super-resolution information of the i-th frame image can be quickly obtained, thereby improving the efficiency of obtaining super-resolution information.
[0051] In some embodiments, based on super-resolution information of multiple frames of images in the video, obtaining a super-resolution video includes:
[0052] Based on the super-resolution information of the multiple frames of images in the video, sub-pixel rearrangement processing is performed to obtain sub-pixel rearrangement results of the multiple frames of images;
[0053] Performing upsampling processing on the multiple frames of images to obtain upsampling results of the multiple frames of images;
[0054] The super-resolution video is generated based on the sub-pixel rearrangement results of the multiple-frame images and the up-sampling results of the multiple-frame images.
[0055] In the disclosed embodiment, the corresponding super-resolution image can be quickly generated by utilizing the sub-pixel rearrangement result of the image and the upsampling result of the image, and then the super-resolution video can be quickly obtained by utilizing the super-resolution images corresponding to multiple frames of images, thereby improving the efficiency of obtaining the super-resolution video.
[0056] In some embodiments, the image feature information, the first difference information, and the second difference information are all in the form of residual maps. The residual map corresponding to the image feature information is used to represent the distribution of sub-pixels in the corresponding image; the residual map corresponding to the first difference information is used to represent the sub-pixel difference between the image and the next frame image of the image; the residual map corresponding to the second difference information is used to represent the sub-pixel difference between the image and the previous frame image of the image.
[0057] According to a second aspect of an embodiment of the present disclosure, a video processing device is provided, the device comprising:
[0058] an information acquisition unit, configured to acquire, for an i-th frame image in a video, image feature information and first difference information of an i-1th frame image, and image feature information and second difference information of an i+1th frame image, where i is a positive integer greater than 1, the image feature information represents a detail feature of a corresponding image, the first difference information represents a difference between the corresponding image and a subsequent frame image of the image in the detail feature, and the second difference information represents a difference between the corresponding image and a previous frame image of the image in the detail feature;
[0059] a conversion information determination unit, configured to determine first conversion information based on the image feature information of the i-1th frame image and the first difference information, and determine second conversion information based on the image feature information of the i+1th frame image and the second difference information, wherein the first conversion information indicates parameters required when the i-1th frame image is converted to the i-th frame image, and the second conversion information indicates parameters required when the i+1th frame image is converted to the i-th frame image;
[0060] a super-resolution information determining unit, configured to determine the super-resolution information of the i-th frame of image based on the image feature information of the i-th frame of image, the first conversion information and the second conversion information;
[0061] The video acquisition unit is configured to execute super-resolution information based on multiple frame images in the video to acquire a super-resolution video.
[0062] In some embodiments, the information acquisition unit includes:
[0063] A feature extraction subunit is configured to input the i-1th frame image and the adjacent frame images of the i-1th frame image into a feature extraction network, and extract hidden layer features of the i-1th frame image based on the i-1th frame image and the adjacent frame images of the i-1th frame image through the feature extraction network, wherein the feature extraction network is trained based on at least one sample frame image, the adjacent frame images of the at least one sample frame image, and the hidden layer features of the at least one sample frame image;
[0064] A determination subunit is configured to determine image feature information of the i-1th frame image based on hidden layer features of the i-1th frame image, and to determine first difference information of the i-1th frame image based on hidden layer features of the i-1th frame image, the i-th frame image and the i-1th frame image.
[0065] In some embodiments, when i is a positive integer greater than 2, the feature extraction subunit is further configured to perform:
[0066] The hidden layer features of the i-1th frame image, the adjacent frame images of the i-1th frame image, and the i-2th frame image are input into the feature extraction network. The hidden layer features of the i-1th frame image are extracted based on the hidden layer features of the i-1th frame image, the adjacent frame images of the i-1th frame image, and the i-2th frame image through the feature extraction network. The feature extraction network is trained based on the hidden layer features of at least one frame sample image, the adjacent frame images of the at least one frame sample image, the previous frame image of the at least one frame sample image, and the hidden layer features of the at least one frame sample image.
[0067] In some embodiments, the determining subunit is configured to execute:
[0068] Inputting the hidden layer features of the i-1th frame image into a first feature extraction subnetwork, extracting image feature information of the i-1th frame image based on the hidden layer features of the i-1th frame image through the first feature extraction subnetwork, wherein the first feature extraction subnetwork is trained based on the hidden layer features of at least one frame of sample image and the image feature information of the at least one frame of sample image;
[0069] The i-1th frame image, the i-th frame image and the hidden layer features of the i-1th frame image are input into a second feature extraction subnetwork, and the first difference information of the i-1th frame image is extracted based on the i-1th frame image, the i-th frame image and the hidden layer features of the i-1th frame image through the second feature extraction subnetwork. The second feature extraction subnetwork is trained based on at least one frame sample image, the subsequent frame image of the at least one frame sample image, the hidden layer features of the at least one frame sample image and the first difference information of the at least one frame sample image.
[0070] In some embodiments, the information acquisition unit includes:
[0071] A feature extraction subunit is configured to input the i+1th frame image and the adjacent frame images of the i+1th frame image into a feature extraction network, and extract hidden layer features of the i+1th frame image based on the i+1th frame image and the adjacent frame images of the i+1th frame image through the feature extraction network, wherein the feature extraction network is trained based on at least one sample frame image, the adjacent frame images of the at least one sample frame image, and the hidden layer features of the at least one sample frame image;
[0072] The determination subunit is configured to determine the image feature information of the i+1th frame image based on the hidden layer features of the i+1th frame image, and determine the second difference information of the i+1th frame image based on the hidden layer features of the i+1th frame image, the i-th frame image and the i+1th frame image.
[0073] In some embodiments, the feature extraction subunit is further configured to perform:
[0074] The i+1th frame image, the adjacent frame images of the i+1th frame image, and the hidden layer features of the i+1th frame image are input into the feature extraction network. The hidden layer features of the i+1th frame image are extracted based on the hidden layer features of the i+1th frame image, the adjacent frame images of the i+1th frame image, and the i-th frame image through the feature extraction network. The feature extraction network is trained based on the hidden layer features of at least one frame sample image, the adjacent frame images of the at least one frame sample image, the previous frame image of the at least one frame sample image, and the hidden layer features of the at least one frame sample image.
[0075] In some embodiments, the determining subunit is configured to execute:
[0076] Inputting the hidden layer features of the i+1th frame image into a first feature extraction subnetwork, extracting image feature information of the i+1th frame image based on the hidden layer features of the i+1th frame image through the first feature extraction subnetwork, wherein the first feature extraction subnetwork is trained based on the hidden layer features of at least one frame of sample image and the image feature information of the at least one frame of sample image;
[0077] The i+1th frame image, the i-th frame image and the hidden layer features of the i+1th frame image are input into a third feature extraction subnetwork, and the second difference information of the i+1th frame image is extracted based on the hidden layer features of the i+1th frame image, the i-th frame image and the i+1th frame image through the third feature extraction subnetwork. The third feature extraction subnetwork is trained based on at least one frame sample image, the previous frame image of the at least one frame sample image, the hidden layer features of the at least one frame sample image and the second difference information of the at least one frame sample image.
[0078] In some embodiments, the feature extraction network is composed of multiple residual modules, wherein a residual module includes a first two-dimensional convolutional layer, an activation function connected to the first two-dimensional convolutional layer, and a second two-dimensional convolutional layer connected to the activation function, and the activation function is used to indicate the functional mapping relationship between the corresponding image, the adjacent frame images of the image and the hidden layer features of the image.
[0079] In some embodiments, the information acquisition unit includes:
[0080] A determination subunit is configured to determine, based on the i-1th frame image and the i-th frame image, optical flow feature information of the i-1th frame image and first optical flow information, the optical flow feature information representing the optical flow feature of the corresponding image, and the first optical flow information representing the pixel movement between the corresponding image and a subsequent frame image of the image;
[0081] The processing subunit is configured to perform interpolation processing on the optical flow feature information of the i-1th frame image and the first optical flow information respectively, and determine the optical flow feature information and the first optical flow information after the interpolation processing as the image feature information and the first difference information of the i-1th frame image.
[0082] In some embodiments, the information acquisition unit includes:
[0083] a determination subunit configured to determine, based on the (i+1)th frame image and the (i)th frame image, optical flow feature information of the (i+1)th frame image and second optical flow information, the optical flow feature information representing the optical flow feature of the corresponding image, and the second optical flow information representing the pixel movement between the corresponding image and a previous frame image of the image;
[0084] The processing subunit is configured to perform interpolation processing on the optical flow feature information of the i+1th frame image and the second optical flow information respectively, and determine the optical flow feature information and the second optical flow information after the interpolation processing as the image feature information of the i+1th frame image and the second difference information.
[0085] In some embodiments, the conversion information determining unit is configured to perform:
[0086] Determine a difference between the image feature information of the (i-1)th frame image and the first difference information of the (i-1)th frame image as the first conversion information;
[0087] A difference between the image feature information of the (i+1)th frame image and the second difference information of the (i+1)th frame image is determined as the second conversion information.
[0088] In some embodiments, the super-resolution information determining unit is configured to perform:
[0089] The image feature information of the i-th frame image, the first conversion information and the second conversion information are input into a temporal convolutional network, and convolution processing is performed based on the image feature information of the i-th frame image, the first conversion information and the second conversion information through the temporal convolutional network to obtain super-resolution information of the i-th frame image. The temporal convolutional network is trained based on the image feature information of at least one frame of sample image, the first conversion information, the second conversion information and the super-resolution information of the at least one frame of sample image.
[0090] In some embodiments, the video acquisition unit is configured to perform:
[0091] Based on the super-resolution information of the multiple frames of images in the video, sub-pixel rearrangement processing is performed to obtain sub-pixel rearrangement results of the multiple frames of images;
[0092] Performing upsampling processing on the multiple frames of images to obtain upsampling results of the multiple frames of images;
[0093] The super-resolution video is generated based on the sub-pixel rearrangement results of the multiple-frame images and the up-sampling results of the multiple-frame images.
[0094] In some embodiments, the image feature information, the first difference information, and the second difference information are all in the form of residual maps. The residual map corresponding to the image feature information is used to represent the distribution of sub-pixels in the corresponding image; the residual map corresponding to the first difference information is used to represent the sub-pixel difference between the image and the next frame image of the image; the residual map corresponding to the second difference information is used to represent the sub-pixel difference between the image and the previous frame image of the image.
[0095] According to a third aspect of an embodiment of the present disclosure, a computer device is provided, the computer device comprising:
[0096] one or more processors;
[0097] A memory for storing program codes executable by the processor;
[0098] The processor is configured to execute the program code to implement the above-mentioned video processing method.
[0099] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, which includes: when the program code in the computer-readable storage medium is executed by a processor of a computer device, the computer device is enabled to execute the above-mentioned video processing method.
[0100] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, including a computer program, which implements the above-mentioned video processing method when executed by a processor.
[0101] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0102] The drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute improper limitations on the present disclosure.
[0103] Figure 1 is a schematic diagram of an implementation environment of a video processing method according to an exemplary embodiment;
[0104] Figure 2 is a flow chart of a video processing method according to an exemplary embodiment;
[0105] Figure 3 is a flow chart of a video processing method according to an exemplary embodiment;
[0106] Figure 4 is a structural schematic diagram of a super-resolution model according to an exemplary embodiment;
[0107] Figure 5 is a schematic structural diagram of a residual module according to an exemplary embodiment;
[0108] Figure 6 is a schematic diagram showing a super-resolution test result according to an exemplary embodiment;
[0109] Figure 7 is a block diagram of a video processing device according to an exemplary embodiment;
[0110] Figure 8 is a block diagram of a terminal according to an exemplary embodiment;
[0111] Fig. 9 It is a block diagram of a server according to an exemplary embodiment. DETAILED DESCRIPTION
[0112] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings.
[0113] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0114] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in the embodiments of the present disclosure are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions. For example, the image feature information, first difference information or second difference information involved in the embodiments of the present disclosure are all obtained with full authorization. In some embodiments, a permission inquiry page is provided in the embodiments of the present disclosure, and the permission inquiry page is used to inquire whether to grant permission to obtain the above information. In the permission inquiry page, an authorization consent control and an authorization rejection control are displayed. When a trigger operation on the authorization consent control is detected, the video processing method provided in the embodiments of the present disclosure is used to obtain the above information, thereby realizing super-resolution reconstruction of the video.
[0115] The video processing method provided by the embodiments of the present disclosure can be applied to the super-resolution reconstruction scenario of videos, for example, super-resolution reconstruction of surveillance videos, super-resolution reconstruction of medical videos, super-resolution reconstruction of filmed videos, etc. Among them, super-resolution reconstruction is to use low-resolution images to reconstruct high-resolution images with higher pixel density and more complete details at the corresponding moment. Correspondingly, super-resolution reconstruction of videos is to use low-resolution videos to reconstruct high-resolution videos with higher pixel density and more complete details at the corresponding moment. It should be understood that a high-resolution image can provide richer information, and it is easier to further mine and utilize the information therein than a low-resolution image.
[0116] Figure 1 is a schematic diagram of an implementation environment of a video processing method according to an exemplary embodiment. Figure 1 , the implementation environment includes: a terminal 101 and a server 102.
[0117] The terminal 101 may be at least one of a smart phone, a smart watch, a desktop computer, a laptop, a virtual reality terminal, an augmented reality terminal, a wireless terminal, and a laptop computer. The terminal 101 has a communication function and can access a wired network or a wireless network. The terminal 101 may generally refer to one of a plurality of terminals. This embodiment is only illustrated by taking the terminal 101 as an example. Those skilled in the art may know that the number of the above terminals may be more or less.
[0118] The server 102 may be an independent physical server, or a server cluster or distributed file system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. Optionally, the number of the above-mentioned servers 102 may be more or less, and the embodiments of the present disclosure are not limited to this. Of course, the server 102 may also include other functional servers to provide more comprehensive and diversified services.
[0119] In some embodiments, the video processing method provided in the embodiments of the present disclosure is executed by the terminal 101. For example, after the terminal 101 detects a super-resolution reconstruction operation for a video, the video processing method provided in the embodiments of the present disclosure is used to obtain a super-resolution video of the video. In other embodiments, the video processing method provided in the embodiments of the present disclosure is executed by the server 102. For example, after the server 102 receives a super-resolution reconstruction request for a video, the video processing method provided in the embodiments of the present disclosure is used to obtain a super-resolution video of the video. In some embodiments, the server 102 is directly or indirectly connected to the terminal 101 through a wired or wireless communication method, which is not limited in the embodiments of the present disclosure. Accordingly, in some embodiments, if the terminal 101 detects a super-resolution reconstruction operation for a video, a super-resolution reconstruction request for the video is sent to the server 102 to request the server 102 to use the video processing method provided in the embodiments of the present disclosure to obtain a super-resolution video of the video. In the embodiments of the present disclosure, a computer device is used to refer to the terminal 101 or the server 102.
[0120] Figure 2 is a flowchart of a video processing method according to an exemplary embodiment. Figure 2 As shown, the method is executed by a computer device, which can be provided as the above Figure 1 The terminal or server shown schematically comprises the following steps:
[0121] In step 201, the computer device obtains image feature information and first difference information of the i-1th frame image and image feature information and second difference information of the i+1th frame image for the i-th frame image in the video, where i is a positive integer greater than 1, the image feature information represents the detail feature of the corresponding image, the first difference information represents the difference between the corresponding image and the next frame image of the image in the detail feature, and the second difference information represents the difference between the corresponding image and the previous frame image of the image in the detail feature.
[0122] In step 202, the computer device determines first conversion information based on the image feature information of the i-1th frame image and the first difference information, and determines second conversion information based on the image feature information of the i+1th frame image and the second difference information. The first conversion information represents the parameters required for converting the i-1th frame image to the i-th frame image, and the second conversion information represents the parameters required for converting the i+1th frame image to the i-th frame image.
[0123] In step 203, the computer device determines super-resolution information of the i-th frame of image based on the image feature information of the i-th frame of image, the first conversion information, and the second conversion information.
[0124] In step 204, the computer device obtains a super-resolution video based on super-resolution information of multiple frames of images in the video.
[0125] The technical solution provided by the embodiment of the present disclosure is for the i-th frame in the video. By obtaining the image feature information and the first difference information of the i-1-th frame image and the image feature information and the second difference information of the i+1-th frame image, the feature information of the i-1-th frame image is converted to the i-th frame image by using the image feature information and the first difference information, and the feature information of the i+1-th frame image is converted to the i-th frame image by using the image feature information and the second difference information, and the second conversion information is obtained. Then, the image feature information, the first conversion information and the second conversion information of the i-th frame image are used to obtain the super-resolution information of the i-th frame image. In this way, when determining the super-resolution information of the i-th frame image, not only the conversion information of the previous frame image to the current frame image is referred to, but also the conversion information of the next frame image to the current frame is referred to, which increases the amount of referenced information, can accurately obtain the super-resolution information of the current frame image, improves the accuracy of super-resolution reconstruction, and then based on multiple frames in the video, can accurately obtain the super-resolution video, and improves the accuracy of video processing.
[0126] In some embodiments, the process of acquiring the image feature information and the first difference information of the i-1th frame image includes:
[0127] Inputting the i-1th frame image and the adjacent frame image of the i-1th frame image into a feature extraction network, and extracting the hidden layer features of the i-1th frame image based on the i-1th frame image and the adjacent frame image of the i-1th frame image through the feature extraction network, wherein the feature extraction network is trained based on at least one sample frame image, the adjacent frame image of the at least one sample frame image, and the hidden layer features of the at least one sample frame image;
[0128] Based on the hidden layer features of the i-1th frame image, image feature information of the i-1th frame image is determined; based on the hidden layer features of the i-1th frame image, the i-th frame image and the i-1th frame image, first difference information of the i-1th frame image is determined.
[0129] In some embodiments, when i is a positive integer greater than 2, the method further includes:
[0130] The hidden layer features of the i-1th frame image, the adjacent frame images of the i-1th frame image, and the i-2th frame image are input into the feature extraction network. The hidden layer features of the i-1th frame image are extracted based on the hidden layer features of the i-1th frame image, the adjacent frame images of the i-1th frame image, and the i-2th frame image through the feature extraction network. The feature extraction network is trained based on the hidden layer features of at least one frame sample image, the adjacent frame images of the at least one frame sample image, the previous frame image of the at least one frame sample image, and the hidden layer features of the at least one frame sample image.
[0131] In some embodiments, based on the hidden layer features of the i-1th frame image, determining the image feature information of the i-1th frame image, and based on the i-1th frame image, the i-th frame image and the hidden layer features of the i-1th frame image, determining the first difference information of the i-1th frame image includes:
[0132] Inputting the hidden layer features of the i-1th frame image into a first feature extraction subnetwork, extracting image feature information of the i-1th frame image based on the hidden layer features of the i-1th frame image through the first feature extraction subnetwork, wherein the first feature extraction subnetwork is trained based on the hidden layer features of at least one frame of sample image and the image feature information of the at least one frame of sample image;
[0133] The i-1th frame image, the i-th frame image and the hidden layer features of the i-1th frame image are input into a second feature extraction subnetwork, and the first difference information of the i-1th frame image is extracted based on the i-1th frame image, the i-th frame image and the hidden layer features of the i-1th frame image through the second feature extraction subnetwork. The second feature extraction subnetwork is trained based on at least one frame sample image, the subsequent frame image of the at least one frame sample image, the hidden layer features of the at least one frame sample image and the first difference information of the at least one frame sample image.
[0134] In some embodiments, the process of acquiring the image feature information and the second difference information of the (i+1)th frame image includes:
[0135] Inputting the i+1th frame image and the adjacent frame image of the i+1th frame image into a feature extraction network, extracting hidden layer features of the i+1th frame image based on the i+1th frame image and the adjacent frame image of the i+1th frame image through the feature extraction network, wherein the feature extraction network is trained based on at least one sample frame image, the adjacent frame image of the at least one sample frame image, and the hidden layer features of the at least one sample frame image;
[0136] Based on the hidden layer features of the i+1th frame image, the image feature information of the i+1th frame image is determined; based on the hidden layer features of the i+1th frame image, the i-th frame image and the i+1th frame image, the second difference information of the i+1th frame image is determined.
[0137] In some embodiments, the method further comprises:
[0138] The i+1th frame image, the adjacent frame images of the i+1th frame image, and the hidden layer features of the i+1th frame image are input into the feature extraction network. The hidden layer features of the i+1th frame image are extracted based on the hidden layer features of the i+1th frame image, the adjacent frame images of the i+1th frame image, and the i-th frame image through the feature extraction network. The feature extraction network is trained based on the hidden layer features of at least one frame sample image, the adjacent frame images of the at least one frame sample image, the previous frame image of the at least one frame sample image, and the hidden layer features of the at least one frame sample image.
[0139] In some embodiments, based on the hidden layer features of the i+1th frame image, determining the image feature information of the i+1th frame image, and based on the hidden layer features of the i+1th frame image, the i-th frame image and the i+1th frame image, determining the second difference information of the i+1th frame image includes:
[0140] Inputting the hidden layer features of the i+1th frame image into a first feature extraction subnetwork, extracting image feature information of the i+1th frame image based on the hidden layer features of the i+1th frame image through the first feature extraction subnetwork, wherein the first feature extraction subnetwork is trained based on the hidden layer features of at least one frame of sample image and the image feature information of the at least one frame of sample image;
[0141] The i+1th frame image, the i-th frame image and the hidden layer features of the i+1th frame image are input into a third feature extraction subnetwork, and the second difference information of the i+1th frame image is extracted based on the hidden layer features of the i+1th frame image, the i-th frame image and the i+1th frame image through the third feature extraction subnetwork. The third feature extraction subnetwork is trained based on at least one frame sample image, the previous frame image of the at least one frame sample image, the hidden layer features of the at least one frame sample image and the second difference information of the at least one frame sample image.
[0142] In some embodiments, the feature extraction network is composed of multiple residual modules, wherein a residual module includes a first two-dimensional convolutional layer, an activation function connected to the first two-dimensional convolutional layer, and a second two-dimensional convolutional layer connected to the activation function, and the activation function is used to indicate the functional mapping relationship between the corresponding image, the adjacent frame images of the image and the hidden layer features of the image.
[0143] In some embodiments, the process of acquiring the image feature information and the first difference information of the i-1th frame image includes:
[0144] Based on the i-1th frame image and the i-th frame image, determine the optical flow feature information of the i-1th frame image and the first optical flow information, the optical flow feature information represents the optical flow feature of the corresponding image, and the first optical flow information represents the pixel movement between the corresponding image and a subsequent frame image of the image;
[0145] Interpolation processing is performed on the optical flow feature information of the i-1th frame image and the first optical flow information respectively, and the optical flow feature information and the first optical flow information after the interpolation processing are determined as the image feature information and the first difference information of the i-1th frame image.
[0146] In some embodiments, the process of acquiring the image feature information and the second difference information of the (i+1)th frame image includes:
[0147] Based on the (i+1)th frame image and the (i)th frame image, determining optical flow feature information of the (i+1)th frame image and second optical flow information, the optical flow feature information representing the optical flow feature of the corresponding image, and the second optical flow information representing the pixel movement between the corresponding image and a previous frame image of the image;
[0148] Interpolation processing is performed on the optical flow feature information of the (i+1)th frame image and the second optical flow information respectively, and the optical flow feature information and the second optical flow information after the interpolation processing are determined as the image feature information of the (i+1)th frame image and the second difference information.
[0149] In some embodiments, determining the first conversion information based on the image feature information of the (i-1)th frame image and the first difference information, and determining the second conversion information based on the image feature information of the (i+1)th frame image and the second difference information includes:
[0150] Determine a difference between the image feature information of the (i-1)th frame image and the first difference information of the (i-1)th frame image as the first conversion information;
[0151] A difference between the image feature information of the (i+1)th frame image and the second difference information of the (i+1)th frame image is determined as the second conversion information.
[0152] In some embodiments, determining the super-resolution information of the i-th frame of image based on the image feature information of the i-th frame of image, the first conversion information, and the second conversion information includes:
[0153] The image feature information of the i-th frame image, the first conversion information and the second conversion information are input into a temporal convolutional network, and convolution processing is performed based on the image feature information of the i-th frame image, the first conversion information and the second conversion information through the temporal convolutional network to obtain super-resolution information of the i-th frame image. The temporal convolutional network is trained based on the image feature information of at least one frame of sample image, the first conversion information, the second conversion information and the super-resolution information of the at least one frame of sample image.
[0154] In some embodiments, based on super-resolution information of multiple frames of images in the video, obtaining a super-resolution video includes:
[0155] Based on the super-resolution information of the multiple frames of images in the video, sub-pixel rearrangement processing is performed to obtain sub-pixel rearrangement results of the multiple frames of images;
[0156] Performing upsampling processing on the multiple frames of images to obtain upsampling results of the multiple frames of images;
[0157] The super-resolution video is generated based on the sub-pixel rearrangement results of the multiple-frame images and the up-sampling results of the multiple-frame images.
[0158] In some embodiments, the image feature information, the first difference information, and the second difference information are all in the form of residual maps. The residual map corresponding to the image feature information is used to represent the distribution of sub-pixels in the corresponding image; the residual map corresponding to the first difference information is used to represent the sub-pixel difference between the image and the next frame image of the image; the residual map corresponding to the second difference information is used to represent the sub-pixel difference between the image and the previous frame image of the image.
[0159] Above Figure 2 The following is only a basic process of the present disclosure. The solution provided by the present disclosure is further described based on a specific implementation method. Figure 3 is a flowchart of a video processing method according to an exemplary embodiment. Figure 3 , the method comprising:
[0160] In step 301, the computer device obtains image feature information and first difference information of the i-1th frame image for the i-th frame image in the video, where i is a positive integer greater than 1. The image feature information represents the detail feature of the corresponding image, and the first difference information represents the difference between the corresponding image and the next frame image of the image in the detail feature.
[0161] Among them, the computer device can be provided as a terminal or a server, and the computer device provides a function of super-resolution reconstruction of the video. In the embodiment of the present disclosure, the video refers to the video to be processed, that is, the video to be super-resolution reconstructed. Among them, super-resolution reconstruction is to use a low-resolution image to reconstruct a high-resolution image with a higher pixel density and more complete details at the corresponding moment. Correspondingly, super-resolution reconstruction of the video is to use a low-resolution video to reconstruct a high-resolution video with a higher pixel density and more complete details at the corresponding moment. In some embodiments, the video is a video stored locally in the terminal, or the video is a video stored in the server, or the video is a video stored in a video library associated with the server, etc., and the embodiments of the present disclosure do not limit this.
[0162] The i-th frame image refers to the image to be super-resolution reconstructed in the video, and the i-th frame image represents any frame image in the video. Correspondingly, the i-1-th frame image is also the previous frame image of any frame image. The image feature information of the i-1-th frame image represents the detail feature of the i-1-th frame image, and the first difference information of the i-1-th frame image represents the difference between the i-1-th frame image and the i-th frame image in the detail feature. Among them, the detail feature is used to characterize the texture detail information in the corresponding image. In some embodiments, the detail feature is in the form of a feature vector. It should be noted that the detail feature is the information required to reconstruct a high-resolution image. It can be understood that the image feature information and the first difference information are both predicted high-resolution information.
[0163] In some embodiments, the image feature information is in the form of a residual map, and the residual map corresponding to the image feature information is used to represent the distribution of sub-pixels in the corresponding image. For example, the image feature information is in the form of a spatial residual map, which is used to represent the distribution of sub-pixels in the image in the spatial dimension. In some embodiments, the first difference information is in the form of a residual map, and the residual map corresponding to the first difference information is used to represent the difference in sub-pixels between the image and the next frame of the image. For example, the first difference information is in the form of a time series residual map, which is used to represent the difference in the distribution of sub-pixels between the image and the next frame of the image in the time series dimension. Wherein, sub-pixels refer to the detail information between two pixels. In this way, by adopting the form of a residual map, the detail features in the form of a vector are converted into detail features in the form of a picture, so as to use the residual map of the detail features in the time series to perform the subsequent super-resolution reconstruction process.
[0164] In some embodiments, a computer device inputs the i-1th frame image and the adjacent frame images of the i-1th frame image into a feature extraction network, extracts hidden layer features of the i-1th frame image based on the i-1th frame image and the adjacent frame images of the i-1th frame image through the feature extraction network, determines image feature information of the i-1th frame image based on the hidden layer features of the i-1th frame image, and determines first difference information of the i-1th frame image based on the hidden layer features of the i-1th frame image, the i-th frame image, and the i-1th frame image.
[0165] Among them, the adjacent frame images of the i-1th frame image are also the previous frame image and the next frame image of the i-1th frame image. Of course, for the first frame image of the video, the adjacent frame images of the first frame image are also the next frame image of the first frame image, and for the last frame image of the video, the adjacent frame images of the last frame image are also the previous frame image of the last frame image. For example, taking i as a positive integer greater than 2, the computer device inputs the i-1th frame image, the i-2th frame image and the i-th frame image into the feature extraction network, so as to subsequently use the feature extraction network to extract the hidden layer features of the i-1th frame image. In some embodiments, the computer device inputs the i-1th frame image, the i-2th frame image and the i-th frame image into a feature extraction network, extracts the image features of the i-1th frame image, the i-2th frame image and the i-th frame image respectively through the feature extraction layer of the feature extraction network, and splices the image features of the i-1th frame image, the i-2th frame image and the i-th frame image in the color dimension of the image to obtain the spliced image features, inputs the spliced image features into the hidden layer of the feature extraction network, and performs convolution processing on the spliced image features through the hidden layer of the feature extraction network to obtain the hidden layer features of the i-1th frame image. It should be noted that the image features are three-dimensional features, such as features in three dimensions: height, width and color channel.
[0166] In this embodiment, for the i-1th frame image, the hidden layer features of the i-1th frame image are extracted through the feature extraction network, that is, the detail features of the i-1th frame image are extracted, and then the hidden layer features of the i-1th frame image are used to determine the image feature information and the first difference information of the i-1th frame image, thereby improving the accuracy of determining the image feature information and the first difference information.
[0167] In some embodiments, the feature extraction network is obtained by training based on at least one sample image, adjacent frame images of the at least one sample image, and hidden layer features of the at least one sample image. Accordingly, the training process of the feature extraction network is: the computer device performs model training based on at least one sample image, adjacent frame images of the at least one sample image, and hidden layer features of the at least one sample image, to obtain the feature extraction network. Specifically, in some embodiments, during the mth iteration of model training, the server inputs the at least one sample image and the adjacent frame images of the at least one sample image into the feature extraction network determined by the m-1th iteration process to obtain the hidden layer features extracted by the mth iteration process, wherein m is a positive integer greater than 1; based on the hidden layer features extracted by the mth iteration process and the hidden layer features of the at least one sample image, the model parameters of the feature extraction network determined by the m-1th iteration process are adjusted, and the m+1th iteration process is performed based on the adjusted model parameters, and the above training iteration process is repeated until the training meets the target conditions.
[0168] In some embodiments, the target condition satisfied by the training is that the number of training iterations of the model reaches the target number, which is a preset number of training iterations, such as 1000 times; or, the target condition satisfied by the training is that the loss value meets the target threshold condition, such as the loss value is less than 0.00001. The present disclosure embodiment does not limit the setting of the target condition.
[0169] In this way, through iterative training, the network model with better model parameters is obtained as the feature extraction network, so as to obtain a feature extraction network with better extraction ability, thereby improving the extraction accuracy of the feature extraction network.
[0170] Regarding the process of extracting hidden layer features by the above-mentioned computer device using the feature extraction network, in some embodiments, when i is a positive integer greater than 2, the computer device inputs the hidden layer features of the i-1th frame image, the adjacent frame image of the i-1th frame image, and the i-2th frame image into the feature extraction network, and through the feature extraction network, the hidden layer features of the i-1th frame image are extracted based on the hidden layer features of the i-1th frame image, the adjacent frame image of the i-1th frame image, and the i-2th frame image. In this embodiment, the feature extraction network is trained based on the hidden layer features of at least one frame sample image, the adjacent frame image of the at least one frame sample image, the previous frame image of the at least one frame sample image, and the hidden layer features of the at least one frame sample image. Among them, the model training process of the feature extraction network is similar to the model training process of the above-mentioned feature extraction network, and will not be repeated.
[0171] In this embodiment, when determining the hidden layer features of the i-1th frame image, the hidden layer features of the i-2th frame image are also referenced. In this way, the hidden layer features of the previous frame image are referenced to determine the hidden layer features of the current frame image, thereby increasing the amount of referenced information, being able to accurately obtain the hidden layer features of the current frame image, and improving the accuracy of obtaining the hidden layer features.
[0172] Regarding the process of determining the image feature information of the i-1th frame image based on the hidden layer features by the computer device, in some embodiments, the computer device inputs the hidden layer features of the i-1th frame image into a first feature extraction subnetwork, and extracts the image feature information of the i-1th frame image based on the hidden layer features of the i-1th frame image through the first feature extraction subnetwork. In this embodiment, for the i-1th frame image, the image feature information of the i-1th frame image can be quickly extracted through the first feature extraction subnetwork, thereby improving the accuracy of obtaining the image feature information.
[0173] In some embodiments, the first feature extraction subnetwork is trained based on the hidden layer features of at least one frame of sample image and the image feature information of the at least one frame of sample image. Accordingly, the training process of the first feature extraction subnetwork is: the computer device performs model training based on the hidden layer features of at least one frame of sample image and the image feature information of the at least one frame of sample image to obtain the first feature extraction subnetwork. Specifically, in some embodiments, during the mth iteration of model training, the server inputs the hidden layer features of the at least one frame of sample image into the first feature extraction subnetwork determined by the m-1th iteration process to obtain the image feature information extracted by the mth iteration process; based on the image feature information extracted by the mth iteration process and the image feature information of the at least one frame of sample image, the model parameters of the first feature extraction subnetwork determined by the m-1th iteration process are adjusted, and the m+1th iteration process is performed based on the adjusted model parameters, and the above training iteration process is repeated until the training meets the target conditions. In this way, through iterative training, the network model with better model parameters is obtained as the first feature extraction sub-network, so as to obtain the first feature extraction sub-network with better extraction ability, thereby improving the extraction accuracy of the first feature extraction sub-network.
[0174] Regarding the process of determining the first difference information of the i-1th frame image based on the hidden layer features by the computer device, in some embodiments, the computer device inputs the i-1th frame image, the i-th frame image, and the hidden layer features of the i-1th frame image into a second feature extraction subnetwork, and extracts the first difference information of the i-1th frame image based on the hidden layer features of the i-1th frame image, the i-th frame image, and the i-1th frame image through the second feature extraction subnetwork. In this embodiment, the first difference information of the i-1th frame image can be quickly extracted through the second feature extraction subnetwork, thereby improving the accuracy of obtaining the first difference information.
[0175] In some embodiments, the second feature extraction subnetwork is trained based on at least one frame of sample image, a frame image after the at least one frame of sample image, hidden layer features of the at least one frame of sample image, and first difference information of the at least one frame of sample image. Accordingly, the training process of the second feature extraction subnetwork is: the computer device performs model training based on at least one frame of sample image, a frame image after the at least one frame of sample image, hidden layer features of the at least one frame of sample image, and first difference information of the at least one frame of sample image to obtain the second feature extraction subnetwork. Specifically, in some embodiments, during the mth iteration of model training, the server inputs the at least one sample image, the next frame of the at least one sample image, and the hidden layer features of the at least one sample image into the second feature extraction subnetwork determined by the m-1th iteration, and obtains the first difference information extracted by the mth iteration; based on the first difference information extracted by the mth iteration and the first difference information of the at least one sample image, the model parameters of the second feature extraction subnetwork determined by the m-1th iteration are adjusted, and the m+1th iteration is performed based on the adjusted model parameters, and the iteration process of the above training is repeated until the training meets the target conditions. In this way, through iterative training, the network model with better model parameters is obtained as the second feature extraction subnetwork, so as to obtain the second feature extraction subnetwork with better extraction ability, thereby improving the extraction accuracy of the second feature extraction subnetwork.
[0176] In the above embodiment, a method is provided for obtaining the image feature information and the first difference information of the i-1th frame image based on the hidden layer features. In other embodiments, the computer device can also obtain the image feature information and the first difference information of the i-1th frame image based on the optical flow features, and the corresponding process is: based on the i-1th frame image and the i-th frame image, determine the optical flow feature information and the first optical flow information of the i-1th frame image; interpolate the optical flow feature information and the first optical flow information of the i-1th frame image respectively, and determine the optical flow feature information and the first optical flow information after the interpolation as the image feature information and the first difference information of the i-1th frame image.
[0177] Among them, the optical flow feature information represents the optical flow feature of the corresponding image. The first optical flow information represents the pixel movement between the corresponding image and the next frame of the image. Correspondingly, the first optical flow information of the i-1th frame image represents the pixel movement between the i-1th frame image and the i-th frame image. It should be noted that the optical flow feature is a feature extracted based on a low-resolution image. It can be understood that the optical flow feature information and the first optical flow information are both extracted low-resolution information. Furthermore, the optical flow feature information and the first optical flow information obtained by interpolation processing are also high-resolution information.
[0178] With respect to the above-mentioned process of extracting optical flow features, in some embodiments, the computer device uses an optical flow prediction algorithm to extract the optical flow feature information and the first optical flow information of the i-1th frame image. In some embodiments, the optical flow prediction algorithm is any one of a sparse optical flow prediction algorithm, a dense optical flow prediction algorithm, and a deep learning optical flow prediction algorithm. With respect to the above-mentioned process of interpolation processing, in some embodiments, the computer device uses an interpolation algorithm to interpolate the optical flow feature information and the first optical flow information of the i-1th frame image. In some embodiments, the interpolation algorithm is any one of a nearest neighbor algorithm, a bilinear interpolation algorithm, and a cubic interpolation algorithm.
[0179] In this embodiment, for the i-1th frame image, by extracting the optical flow feature information of the i-1th frame image and then performing interpolation processing on the optical flow feature information, the detail features of the i-1th frame image can be quickly determined, that is, the image feature information of the i-1th frame image is determined, and, by extracting the first optical flow information of the i-1th frame image, the pixel movement between the i-1th frame image and the i-th frame image can be quickly determined, and then the first optical flow information is interpolated, so that the difference in detail features between the i-1th frame image and the i-th frame image can be quickly determined, that is, the first difference information of the i-1th frame image is determined, thereby improving the efficiency and accuracy of determining the image feature information and the first difference information.
[0180] In step 302, the computer device obtains image feature information and second difference information of the (i+1)th frame image, where the second difference information indicates the difference between the corresponding image and the previous frame image of the image in the detail feature.
[0181] The i-th frame image represents any frame image in the video, and accordingly, the i+1-th frame image is the next frame image of any frame image. The image feature information of the i+1-th frame image represents the detail feature of the i+1-th frame image, and the second difference information of the i+1-th frame image represents the difference between the i+1-th frame image and the i-th frame image in the detail feature. In some embodiments, the second difference information is in the form of a residual map, and the residual map corresponding to the second difference information is used to represent the sub-pixel difference between the image and the previous frame image of the image. For example, the second difference information is in the form of a time series residual map, which is used to represent the distribution difference of sub-pixels between the image and the previous frame image of the image in the time series dimension.
[0182] In some embodiments, a computer device inputs the i+1th frame image and the adjacent frame images of the i+1th frame image into a feature extraction network, extracts hidden layer features of the i+1th frame image based on the i+1th frame image and the adjacent frame images of the i+1th frame image through the feature extraction network, determines image feature information of the i+1th frame image based on the hidden layer features of the i+1th frame image, and determines second difference information of the i+1th frame image based on the hidden layer features of the i+1th frame image, the i-th frame image, and the i+1th frame image.
[0183] The adjacent frame image of the i+1th frame image is at least one of the previous frame image and the next frame image of the i+1th frame image. For example, the computer device inputs the i+1th frame image, the i-th frame image and the i+2th frame image into a feature extraction network, so that the feature extraction network is subsequently used to extract the hidden layer features of the i+1th frame image.
[0184] In some embodiments, a computer device inputs the i+1th frame image, the i-th frame image and the i+2th frame image into a feature extraction network, extracts image features of the i+1th frame image, the i-th frame image and the i+2th frame image respectively through the feature extraction layer of the feature extraction network, and splices the image features of the i+1th frame image, the i-th frame image and the i+2th frame image in the color dimension of the image to obtain spliced image features, inputs the spliced image features into the hidden layer of the feature extraction network, performs convolution processing on the spliced image features through the hidden layer of the feature extraction network to obtain the hidden layer features of the i+1th frame image.
[0185] In this embodiment, for the i+1th frame image, the hidden layer features of the i+1th frame image are extracted through the feature extraction network, that is, the detail features of the i+1th frame image are extracted, and then the hidden layer features of the i+1th frame image are used to determine the image feature information and the second difference information of the i+1th frame image, thereby improving the accuracy of determining the image feature information and the second difference information.
[0186] In some embodiments, the feature extraction network is trained based on at least one sample image, adjacent frames of the at least one sample image, and hidden features of the at least one sample image. The model training process of the feature extraction network refers to the model training process for the feature extraction network in step 301, which will not be repeated here.
[0187] Regarding the process of extracting hidden features using a feature extraction network by the above-mentioned computer device, in some embodiments, the computer device inputs the i+1th frame image, the adjacent frame image of the i+1th frame image, and the hidden features of the i-th frame image into the feature extraction network, and extracts the hidden features of the i+1th frame image based on the i+1th frame image, the adjacent frame image of the i+1th frame image, and the hidden features of the i-th frame image through the feature extraction network. In this embodiment, the feature extraction network is trained based on the hidden features of at least one frame sample image, the adjacent frame image of the at least one frame sample image, the previous frame image of the at least one frame sample image, and the hidden features of the at least one frame sample image. Among them, the model training process of the feature extraction network is similar to the model training process of the feature extraction network in step 301, and will not be repeated.
[0188] In this embodiment, when determining the hidden layer features of the i+1th frame image, the hidden layer features of the i-th frame image are also referenced. In this way, the hidden layer features of the previous frame image are referenced to determine the hidden layer features of the current frame image, thereby increasing the amount of referenced information, being able to accurately obtain the hidden layer features of the current frame image, and improving the accuracy of obtaining the hidden layer features.
[0189] Regarding the process of determining the image feature information of the i+1th frame image based on the hidden layer features by the computer device, in some embodiments, the computer device inputs the hidden layer features of the i+1th frame image into the first feature extraction subnetwork, and extracts the image feature information of the i+1th frame image based on the hidden layer features of the i+1th frame image through the first feature extraction subnetwork. In this embodiment, for the i+1th frame image, the image feature information of the i+1th frame image can be quickly extracted through the first feature extraction subnetwork, thereby improving the accuracy of obtaining the image feature information.
[0190] In some embodiments, the first feature extraction subnetwork is trained based on the hidden layer features of at least one frame of sample image and the image feature information of the at least one frame of sample image. The model training process of the first feature extraction subnetwork refers to the model training process of the first feature extraction subnetwork shown in step 301, which will not be repeated here.
[0191] Regarding the process of determining the second difference information of the i+1th frame image based on the hidden layer features by the computer device, in some embodiments, the computer device inputs the hidden layer features of the i+1th frame image, the i-th frame image, and the i+1th frame image into a third feature extraction subnetwork, and extracts the second difference information of the i+1th frame image based on the hidden layer features of the i+1th frame image, the i-th frame image, and the i+1th frame image through the third feature extraction subnetwork. In this embodiment, the second difference information of the i+1th frame image can be quickly extracted through the third feature extraction subnetwork, thereby improving the accuracy of obtaining the second difference information.
[0192] In some embodiments, the third feature extraction subnetwork is trained based on at least one sample image, a previous frame of the at least one sample image, hidden features of the at least one sample image, and second difference information of the at least one sample image. Accordingly, the training process of the third feature extraction subnetwork is: the computer device performs model training based on at least one sample image, a previous frame of the at least one sample image, hidden features of the at least one sample image, and second difference information of the at least one sample image to obtain the third feature extraction subnetwork. Specifically, in some embodiments, during the mth iteration of model training, the server inputs the at least one sample image, the previous image of the at least one sample image, and the hidden features of the at least one sample image into the third feature extraction subnetwork determined by the m-1th iteration, and obtains the second difference information extracted by the mth iteration; based on the second difference information extracted by the mth iteration and the second difference information of the at least one sample image, the model parameters of the third feature extraction subnetwork determined by the m-1th iteration are adjusted, and the m+1th iteration is performed based on the adjusted model parameters, and the above training iteration process is repeated until the training meets the target conditions. In this way, through iterative training, the network model with better model parameters is obtained as the third feature extraction subnetwork, so as to obtain the third feature extraction subnetwork with better extraction ability, thereby improving the extraction accuracy of the third feature extraction subnetwork.
[0193] For the process of using hidden layer features to obtain image feature information, first difference information and second difference information in step 301 and step 302, the embodiment of the present disclosure also provides a super-resolution model, which is provided with the above-mentioned feature extraction network, the first feature extraction sub-network, the second feature extraction sub-network and the third feature extraction sub-network. For example, Figure 4 is a schematic diagram of a super-resolution model according to an exemplary embodiment. Figure 4 , Figure 4In the example, the It frame image is the current frame image, the It-1 frame image is the previous frame image, and the It+1 frame image is the next frame image. When obtaining the above information corresponding to the It frame image, the It frame image, the It-1 frame image, and the It+1 frame image are input as follows: Figure 4 The super-resolution model shown in FIG. 1 firstly extracts the hidden layer features of the It-th frame image through the feature extraction network in the super-resolution model, wherein the feature extraction network can be Figure 4 Furthermore, in some embodiments, the hidden layer features output by the feature extraction network are input into the feature extraction network at the next moment, that is, the input Figure 4 The "Ht+1 network" shown in the figure is used so that the "Ht+1 network" uses the hidden features of the It frame image to determine the hidden features of the It+1 frame image; in other embodiments, the hidden features output by the feature extraction network are input into the first feature extraction subnetwork, and the image feature information of the It frame image is extracted through the first feature extraction subnetwork, the hidden features output by the feature extraction network are input into the second feature extraction subnetwork, and the first difference information of the It frame image is extracted through the second feature extraction subnetwork, and the hidden features output by the feature extraction network are input into the third feature extraction subnetwork, and the second difference information of the It frame image is extracted through the third feature extraction subnetwork, wherein the first feature extraction subnetwork can be Figure 4 The “network” 401 shown, accordingly, the image feature information predicted by the first feature extraction subnetwork is Figure 4 St shown; the second feature extraction subnetwork can be Figure 4 The “network” 402 shown, accordingly, the first difference information predicted by the second feature extraction subnetwork is Figure 4 The third feature extraction subnetwork can be Figure 4 The “network” 403 shown, accordingly, the second difference information predicted by the third feature extraction subnetwork is Figure 4 The Pt shown.
[0194] With respect to the feature extraction network shown in the above steps 301 and 302, in some embodiments, the feature extraction network is based on a plurality of residual modules, and further, the hidden layer in the feature extraction network is a cascaded architecture of a plurality of residual modules. Among them, a residual module includes a first two-dimensional convolutional layer, an activation function connected to the first two-dimensional convolutional layer, and a second two-dimensional convolutional layer connected to the activation function, and the activation function is used to indicate the functional mapping relationship between the corresponding image, the adjacent frame image of the image, and the hidden layer features of the image. In some embodiments, both the first two-dimensional convolutional layer and the second two-dimensional convolutional layer use a 3*3 convolution kernel.
[0195] For example, Figure 5 is a schematic diagram of a structure of a residual module according to an exemplary embodiment. Figure 5 , the first two-dimensional convolutional layer is Figure 5 The "2D convolutional layer" 501 shown has an activation function of Figure 5 The "ReLU" 502 shown, the second two-dimensional convolutional layer is Figure 5 As shown in the "2D convolution layer" 503, it can be found that in a residual module, after the feature is input into the first two-dimensional convolution layer, the convolution processing is performed through the first two-dimensional convolution layer, and the feature after the convolution processing is output, and the feature output by the first two-dimensional convolution layer is used as the input of the activation function, and the input feature is operated through the function mapping relationship indicated by the activation function, and the operation result of the activation function is output, and the output of the activation function is used as the input of the second two-dimensional convolution layer, and the operation result output by the activation function is convoluted through the second two-dimensional convolution layer, and the feature after the second convolution processing is output, and then the feature after the second convolution processing and the input feature of the residual module are input into the next residual module. In the disclosed embodiment, by setting the residual module, the effect of gradient feedback can be achieved, thereby avoiding the problem of gradient disappearance when the hidden layer is deepened.
[0196] In the above embodiment, a method for obtaining the image feature information and the second difference information of the i+1th frame image based on the hidden layer features is provided. In other embodiments, the computer device can also obtain the image feature information and the second difference information of the i+1th frame image based on the optical flow features, and the corresponding process is: based on the i+1th frame image and the i-th frame image, determine the optical flow feature information and the second optical flow information of the i+1th frame image; interpolate the optical flow feature information and the second optical flow information of the i+1th frame image respectively, and determine the optical flow feature information and the second optical flow information after the interpolation process as the image feature information and the second difference information of the i+1th frame image.
[0197] The second optical flow information represents the pixel movement between the corresponding image and the previous frame image of the image. Correspondingly, the second optical flow information of the i+1th frame image represents the pixel movement between the i+1th frame image and the i-th frame image.
[0198] In some embodiments, the computer device uses an optical flow estimation method to extract the optical flow feature information and the second optical flow information of the i+1th frame image. In some embodiments, the computer device uses an interpolation algorithm to perform interpolation processing on the optical flow feature information and the second optical flow information of the i+1th frame image.
[0199] In this embodiment, for the i+1th frame image, by extracting the optical flow feature information of the i+1th frame image and then performing interpolation processing on the optical flow feature information, the detail features of the i+1th frame image can be quickly determined, that is, the image feature information of the i+1th frame image is determined, and, by extracting the first optical flow information of the i+1th frame image, the pixel movement between the i+1th frame image and the i-th frame image can be quickly determined, and then the first optical flow information is interpolated to quickly determine the difference in detail features between the i+1th frame image and the i-th frame image, that is, the second difference information of the i+1th frame image is determined, thereby improving the efficiency and accuracy of determining the image feature information and the second difference information.
[0200] In step 303, the computer device determines first conversion information based on the image feature information of the i-1th frame image and the first difference information, where the first conversion information represents the parameters required for converting the i-1th frame image to the i-th frame image.
[0201] The first conversion information represents the parameters required to convert the i-1th frame image to the i-th frame image in the time dimension, and further, the parameters required to convert the information predicted at the time of the i-1th frame image to the time of the i-th frame image. It should be understood that for any frame image of the video, the previous frame image of the image is the image at the previous moment, and the next frame image of the image is the image at the next moment.
[0202] In some embodiments, the computer device determines a difference between the image feature information of the (i-1)th frame image and the first difference information of the (i-1)th frame image as the first conversion information.
[0203] For example, see Figure 4 ,against Figure 4 At the time t-1 shown, the image corresponding to the time t-1 is the i-1 frame image. The image feature information predicted at the time t-1 (i.e. Figure 4 St-1 in the Figure 4 By subtracting Ft-1 from the previous moment, the information predicted at the t-1th moment can be converted to the tth moment. The first converted information is Figure 4 Shown
[0204] In step 304, the computer device determines second conversion information based on the image feature information of the (i+1)th frame image and the second difference information, where the second conversion information represents the parameters required for converting the (i+1)th frame image to the (i)th frame image.
[0205] Among them, the second conversion information represents the parameters required to convert the i+1th frame image to the i-th frame image in the time dimension. Further, it means converting the information predicted at the i+1th frame image to the parameters required at the i-th frame image.
[0206] In some embodiments, the computer device determines a difference between the image feature information of the (i+1)th frame image and the second difference information of the (i+1)th frame image as the second conversion information.
[0207] For example, see Figure 4 ,against Figure 4 At the time t+1 shown, the image corresponding to the time t+1 is the i+1th frame image. The image feature information predicted at the time t+1 (i.e. Figure 4 St+1 in the Figure 4 By subtracting Pt+1 from the time t+1, the information predicted at time t+1 can be converted to time t. The second converted information obtained is Figure 4 Shown
[0208] In the above steps 303 to 304, by taking the difference method, the feature information of the i-1th frame image can be quickly converted to the i-th frame image, and the feature information of the i+1th frame image can be converted to the i-th frame image, so that the conversion information of the previous frame image to the current frame image and the conversion information of the next frame image to the current frame can be used to determine the super-resolution information of the current frame image, thereby improving the accuracy of super-resolution reconstruction.
[0209] In step 305 , the computer device determines super-resolution information of the i-th frame of image based on the image feature information of the i-th frame of image, the first conversion information, and the second conversion information.
[0210] In some embodiments, a computer device inputs the image feature information of the i-th frame image, the first conversion information, and the second conversion information into a temporal convolutional network, and performs convolution processing based on the image feature information of the i-th frame image, the first conversion information, and the second conversion information through the temporal convolutional network to obtain super-resolution information of the i-th frame image. In this embodiment, for the i-th frame image, the image feature information of the i-th frame image, the first conversion information, and the second conversion information are convolved through the temporal convolutional network, so that the super-resolution information of the i-th frame image can be quickly obtained, thereby improving the efficiency of obtaining super-resolution information.
[0211] For example, see Figure 4 ,against Figure 4At the time t shown, the image corresponding to the time t is the i-th frame image, and the image feature information of the i-th frame image predicted at the time t (i.e. St), the first conversion information (i.e. ) and the second conversion information (ie ) to obtain the spliced features, and input the spliced features into Figure 4 The temporal convolutional network shown is used to convolve the image feature information of the i-th frame image, the first conversion information and the second conversion information to obtain super-resolution information of the i-th frame image. In some embodiments, the temporal convolutional network is a structure of multiple residual modules cascaded.
[0212] In some embodiments, the temporal convolutional network is trained based on the image feature information of at least one frame of sample image, the first conversion information, the second conversion information and the super-resolution information of the at least one frame of sample image. Accordingly, the training process of the temporal convolutional network is: the computer device performs model training based on the image feature information of at least one frame of sample image, the first conversion information, the second conversion information and the super-resolution information of the at least one frame of sample image, to obtain the temporal convolutional network. Specifically, in some embodiments, during the mth iteration of model training, the server inputs the image feature information, the first conversion information and the second conversion information of the at least one frame of sample image into the temporal convolutional network determined by the m-1th iteration process to obtain the super-resolution information extracted by the mth iteration process; based on the super-resolution information extracted by the mth iteration process and the super-resolution information of the at least one frame of sample image, the model parameters of the temporal convolutional network determined by the m-1th iteration process are adjusted, and the m+1th iteration process is performed based on the adjusted model parameters, and the above training iteration process is repeated until the training meets the target conditions. In this way, through iterative training, the network model with better model parameters is obtained as a temporal convolutional network, so as to obtain a temporal convolutional network with better extraction capability, thereby improving the extraction accuracy of the temporal convolutional network.
[0213] In step 306, the computer device obtains a super-resolution video based on super-resolution information of multiple frames of images in the video.
[0214] In some embodiments, the computer device performs sub-pixel rearrangement processing based on super-resolution information of multiple frames of images in the video to obtain sub-pixel rearrangement results of the multiple frames of images, performs upsampling processing on the multiple frames of images to obtain upsampling results of the multiple frames of images, and generates the super-resolution video based on the sub-pixel rearrangement results of the multiple frames of images and the upsampling results of the multiple frames of images.
[0215] Regarding the above-mentioned upsampling process, in some embodiments, the computer device performs the above-mentioned upsampling process based on a linear interpolation upsampling method, or the computer device performs the above-mentioned upsampling process based on a deep learning upsampling method (such as deconvolution). The embodiments of the present disclosure do not limit the content of the upsampling process.
[0216] In some embodiments, a computer device inputs the sub-pixel rearrangement results of the multi-frame images and the upsampling results of the multi-frame images into an adder, and the adder adds the sub-pixel rearrangement results of the multi-frame images and the upsampling results of the multi-frame images respectively to obtain the super-resolution images corresponding to the multi-frame images, and combines the super-resolution images corresponding to the multi-frame images in the time order of the multi-frame images to obtain the super-resolution video.
[0217] In this embodiment, the corresponding super-resolution image can be quickly generated by utilizing the sub-pixel rearrangement result of the image and the upsampling result of the image, and then the super-resolution video can be quickly obtained by utilizing the super-resolution images corresponding to multiple frames of images, thereby improving the efficiency of obtaining the super-resolution video.
[0218] In this way, the feature information of the past and future moments is converted to the current moment by utilizing the difference in timing of the predicted detail features, and the result of the current moment is further optimized by utilizing the converted information. Compared with the related art that directly uses a convolutional neural network to output a single prediction result, the solution provided by the embodiment of the present disclosure can utilize the high-resolution prediction results of the future and past moments for further optimization, and provides a time-series round-trip optimization solution under high resolution, which can more fully utilize the feature information of the past and future moments, thereby obtaining better super-resolution results, and producing richer details and accurate structures in the super-resolution results.
[0219] The technical solution provided by the embodiment of the present disclosure is for the i-th frame in the video. By obtaining the image feature information and the first difference information of the i-1-th frame image and the image feature information and the second difference information of the i+1-th frame image, the feature information of the i-1-th frame image is converted to the i-th frame image by using the image feature information and the first difference information, and the feature information of the i+1-th frame image is converted to the i-th frame image by using the image feature information and the second difference information, and the second conversion information is obtained. Then, the image feature information, the first conversion information and the second conversion information of the i-th frame image are used to obtain the super-resolution information of the i-th frame image. In this way, when determining the super-resolution information of the i-th frame image, not only the conversion information of the previous frame image to the current frame image is referred to, but also the conversion information of the next frame image to the current frame is referred to, which increases the amount of referenced information, can accurately obtain the super-resolution information of the current frame image, improves the accuracy of super-resolution reconstruction, and then based on multiple frames in the video, can accurately obtain the super-resolution video, and improves the accuracy of video processing.
[0220] In the embodiment of the present disclosure, super-resolution reconstruction experiments were also conducted using the Vid4 dataset and the UDM10 dataset, where both the Vid4 test set and the UDM10 test set are video super-resolution test sets. In this experiment, super-resolution reconstruction of the video was conducted under the conditions of N=0 and N=1, respectively, where N=0 means that the time series round-trip optimization method is not used, and the unidirectional recurrent convolutional network in the related technology is directly used for super-resolution reconstruction; N=1 means that the time series round-trip optimization method shown in the embodiment of the present disclosure is used, and the conversion information of the previous frame image and the conversion information of the next frame image are used to obtain the super-resolution information of the current frame image. See Table 1. In this experiment, PSNR (peak signal-to-noise ratio) is used as the test index. Under the condition of N=0, the signal-to-noise ratio of the high-resolution image reconstructed using the Vid4 dataset is 28.04db, and the signal-to-noise ratio of the high-resolution image reconstructed using the UDM10 dataset is 39.68db. Under the condition of N=1, the signal-to-noise ratio of the high-resolution image reconstructed using the Vid4 dataset is 28.21db, and the signal-to-noise ratio of the high-resolution image reconstructed using the UDM10 dataset is 39.80db. In this way, the reconstructed high-resolution image can express richer levels and contain richer colors. In addition, for the optical flow method and the time series round-trip method shown in the embodiments of the present disclosure, super-resolution reconstruction experiments were also carried out using the above-mentioned Vid4 dataset and UDM10 dataset. Obviously, the images obtained by super-resolution reconstruction based on the optical flow method and the images obtained by super-resolution reconstruction based on the time series round-trip method have higher signal-to-noise ratios than the images obtained by super-resolution reconstruction based on the relevant technology.
[0221] Table 1
[0222]
[0223] For example, Figure 6 is a schematic diagram showing a super-resolution test result according to an exemplary embodiment, see Figure 6 , Figure 6 The first column of images shown is a high-resolution image obtained using a traditional upsampling interpolation algorithm. Figure 6 The second column of images shown is an image obtained by super-resolution reconstruction using a convolutional neural network in the related art. Figure 6 The third column of images shown is an image obtained by using the single-step timing round-trip optimization method provided by the embodiment of the present disclosure. Figure 6 The fourth column of images shown is the real image. It can be found that Figure 6 The texture details in the second column of images shown are not clear enough and the numbers are blurry. Figure 6 The resolution of the images in the third column is obviously improved and is close to the real images in the fourth column.
[0224] Figure 7 is a block diagram of a video processing device according to an exemplary embodiment. Figure 7 The device includes an information acquisition unit 701, a conversion information determination unit 702, a super-resolution information determination unit 703 and a video acquisition unit 704.
[0225] The information acquisition unit 701 is configured to execute, for the i-th frame image in the video, acquiring the image feature information and the first difference information of the i-1th frame image and the image feature information and the second difference information of the i+1th frame image, where i is a positive integer greater than 1, the image feature information represents the detail feature of the corresponding image, the first difference information represents the difference between the corresponding image and the next frame image of the image in the detail feature, and the second difference information represents the difference between the corresponding image and the previous frame image of the image in the detail feature;
[0226] The conversion information determining unit 702 is configured to determine first conversion information based on the image feature information of the i-1th frame image and the first difference information, and determine second conversion information based on the image feature information of the i+1th frame image and the second difference information, wherein the first conversion information indicates parameters required when the i-1th frame image is converted to the i-th frame image, and the second conversion information indicates parameters required when the i+1th frame image is converted to the i-th frame image;
[0227] The super-resolution information determining unit 703 is configured to determine the super-resolution information of the i-th frame of image based on the image feature information of the i-th frame of image, the first conversion information and the second conversion information;
[0228] The video acquisition unit 704 is configured to acquire a super-resolution video based on super-resolution information of multiple frames of images in the video.
[0229] The technical solution provided by the embodiment of the present disclosure is for the i-th frame in the video. By obtaining the image feature information and the first difference information of the i-1-th frame image and the image feature information and the second difference information of the i+1-th frame image, the feature information of the i-1-th frame image is converted to the i-th frame image by using the image feature information and the first difference information, and the feature information of the i+1-th frame image is converted to the i-th frame image by using the image feature information and the second difference information, and the second conversion information is obtained. Then, the image feature information, the first conversion information and the second conversion information of the i-th frame image are used to obtain the super-resolution information of the i-th frame image. In this way, when determining the super-resolution information of the i-th frame image, not only the conversion information of the previous frame image to the current frame image is referred to, but also the conversion information of the next frame image to the current frame is referred to, which increases the amount of referenced information, can accurately obtain the super-resolution information of the current frame image, improves the accuracy of super-resolution reconstruction, and then based on multiple frames in the video, can accurately obtain the super-resolution video, and improves the accuracy of video processing.
[0230] In some embodiments, the information acquisition unit 701 includes:
[0231] A feature extraction subunit is configured to input the i-1th frame image and the adjacent frame images of the i-1th frame image into a feature extraction network, and extract hidden layer features of the i-1th frame image based on the i-1th frame image and the adjacent frame images of the i-1th frame image through the feature extraction network, wherein the feature extraction network is trained based on at least one sample frame image, the adjacent frame images of the at least one sample frame image, and the hidden layer features of the at least one sample frame image;
[0232] A determination subunit is configured to determine image feature information of the i-1th frame image based on hidden layer features of the i-1th frame image, and to determine first difference information of the i-1th frame image based on hidden layer features of the i-1th frame image, the i-th frame image and the i-1th frame image.
[0233] In some embodiments, when i is a positive integer greater than 2, the feature extraction subunit is further configured to perform:
[0234] The hidden layer features of the i-1th frame image, the adjacent frame images of the i-1th frame image, and the i-2th frame image are input into the feature extraction network. The hidden layer features of the i-1th frame image are extracted based on the hidden layer features of the i-1th frame image, the adjacent frame images of the i-1th frame image, and the i-2th frame image through the feature extraction network. The feature extraction network is trained based on the hidden layer features of at least one frame sample image, the adjacent frame images of the at least one frame sample image, the previous frame image of the at least one frame sample image, and the hidden layer features of the at least one frame sample image.
[0235] In some embodiments, the determining subunit is configured to execute:
[0236] Inputting the hidden layer features of the i-1th frame image into a first feature extraction subnetwork, extracting image feature information of the i-1th frame image based on the hidden layer features of the i-1th frame image through the first feature extraction subnetwork, wherein the first feature extraction subnetwork is trained based on the hidden layer features of at least one frame of sample image and the image feature information of the at least one frame of sample image;
[0237] The i-1th frame image, the i-th frame image and the hidden layer features of the i-1th frame image are input into a second feature extraction subnetwork, and the first difference information of the i-1th frame image is extracted based on the i-1th frame image, the i-th frame image and the hidden layer features of the i-1th frame image through the second feature extraction subnetwork. The second feature extraction subnetwork is trained based on at least one frame sample image, the subsequent frame image of the at least one frame sample image, the hidden layer features of the at least one frame sample image and the first difference information of the at least one frame sample image.
[0238] In some embodiments, the information acquisition unit 701 includes:
[0239] A feature extraction subunit is configured to input the i+1th frame image and the adjacent frame images of the i+1th frame image into a feature extraction network, and extract hidden layer features of the i+1th frame image based on the i+1th frame image and the adjacent frame images of the i+1th frame image through the feature extraction network, wherein the feature extraction network is trained based on at least one sample frame image, the adjacent frame images of the at least one sample frame image, and the hidden layer features of the at least one sample frame image;
[0240] The determination subunit is configured to determine the image feature information of the i+1th frame image based on the hidden layer features of the i+1th frame image, and determine the second difference information of the i+1th frame image based on the hidden layer features of the i+1th frame image, the i-th frame image and the i+1th frame image.
[0241] In some embodiments, the feature extraction subunit is further configured to perform:
[0242] The i+1th frame image, the adjacent frame images of the i+1th frame image, and the hidden layer features of the i+1th frame image are input into the feature extraction network. The hidden layer features of the i+1th frame image are extracted based on the hidden layer features of the i+1th frame image, the adjacent frame images of the i+1th frame image, and the i-th frame image through the feature extraction network. The feature extraction network is trained based on the hidden layer features of at least one frame sample image, the adjacent frame images of the at least one frame sample image, the previous frame image of the at least one frame sample image, and the hidden layer features of the at least one frame sample image.
[0243] In some embodiments, the determining subunit is configured to execute:
[0244] Inputting the hidden layer features of the i+1th frame image into a first feature extraction subnetwork, extracting image feature information of the i+1th frame image based on the hidden layer features of the i+1th frame image through the first feature extraction subnetwork, wherein the first feature extraction subnetwork is trained based on the hidden layer features of at least one frame of sample image and the image feature information of the at least one frame of sample image;
[0245] The i+1th frame image, the i-th frame image and the hidden layer features of the i+1th frame image are input into a third feature extraction subnetwork, and the second difference information of the i+1th frame image is extracted based on the hidden layer features of the i+1th frame image, the i-th frame image and the i+1th frame image through the third feature extraction subnetwork. The third feature extraction subnetwork is trained based on at least one frame sample image, the previous frame image of the at least one frame sample image, the hidden layer features of the at least one frame sample image and the second difference information of the at least one frame sample image.
[0246] In some embodiments, the feature extraction network is composed of multiple residual modules, wherein a residual module includes a first two-dimensional convolutional layer, an activation function connected to the first two-dimensional convolutional layer, and a second two-dimensional convolutional layer connected to the activation function, and the activation function is used to indicate the functional mapping relationship between the corresponding image, the adjacent frame images of the image and the hidden layer features of the image.
[0247] In some embodiments, the information acquisition unit 701 includes:
[0248] A determination subunit is configured to determine, based on the i-1th frame image and the i-th frame image, optical flow feature information of the i-1th frame image and first optical flow information, the optical flow feature information representing the optical flow feature of the corresponding image, and the first optical flow information representing the pixel movement between the corresponding image and a subsequent frame image of the image;
[0249] The processing subunit is configured to perform interpolation processing on the optical flow feature information of the i-1th frame image and the first optical flow information respectively, and determine the optical flow feature information and the first optical flow information after the interpolation processing as the image feature information and the first difference information of the i-1th frame image.
[0250] In some embodiments, the information acquisition unit 701 includes:
[0251] a determination subunit configured to determine, based on the (i+1)th frame image and the (i)th frame image, optical flow feature information of the (i+1)th frame image and second optical flow information, the optical flow feature information representing the optical flow feature of the corresponding image, and the second optical flow information representing the pixel movement between the corresponding image and a previous frame image of the image;
[0252] The processing subunit is configured to perform interpolation processing on the optical flow feature information of the i+1th frame image and the second optical flow information respectively, and determine the optical flow feature information and the second optical flow information after the interpolation processing as the image feature information of the i+1th frame image and the second difference information.
[0253] In some embodiments, the conversion information determining unit 702 is configured to perform:
[0254] Determine a difference between the image feature information of the (i-1)th frame image and the first difference information of the (i-1)th frame image as the first conversion information;
[0255] A difference between the image feature information of the (i+1)th frame image and the second difference information of the (i+1)th frame image is determined as the second conversion information.
[0256] In some embodiments, the super-resolution information determining unit 703 is configured to perform:
[0257] The image feature information of the i-th frame image, the first conversion information and the second conversion information are input into a temporal convolutional network, and convolution processing is performed based on the image feature information of the i-th frame image, the first conversion information and the second conversion information through the temporal convolutional network to obtain super-resolution information of the i-th frame image. The temporal convolutional network is trained based on the image feature information of at least one frame of sample image, the first conversion information, the second conversion information and the super-resolution information of the at least one frame of sample image.
[0258] In some embodiments, the video acquisition unit 704 is configured to perform:
[0259] Based on the super-resolution information of the multiple frames of images in the video, sub-pixel rearrangement processing is performed to obtain sub-pixel rearrangement results of the multiple frames of images;
[0260] Performing upsampling processing on the multiple frames of images to obtain upsampling results of the multiple frames of images;
[0261] The super-resolution video is generated based on the sub-pixel rearrangement results of the multiple-frame images and the up-sampling results of the multiple-frame images.
[0262] In some embodiments, the image feature information, the first difference information, and the second difference information are all in the form of residual maps. The residual map corresponding to the image feature information is used to represent the distribution of sub-pixels in the corresponding image; the residual map corresponding to the first difference information is used to represent the sub-pixel difference between the image and the next frame image of the image; the residual map corresponding to the second difference information is used to represent the sub-pixel difference between the image and the previous frame image of the image.
[0263] It should be noted that: the video processing device provided in the above embodiment only uses the division of the above functional modules as an example to illustrate the video processing. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the video processing device provided in the above embodiment and the video processing method embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.
[0264] The computer device mentioned in the embodiments of the present disclosure may be provided as a terminal. Figure 8 The structure block diagram of a terminal 800 provided by an exemplary embodiment of the present disclosure is shown. The terminal 800 may be: a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer or a desktop computer. The terminal 800 may also be called a user device, a portable terminal, a laptop terminal, a desktop terminal or other names.
[0265] Typically, the terminal 800 includes a processor 801 and a memory 802 .
[0266] The processor 801 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 801 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 801 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 801 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 801 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.
[0267] The memory 802 may include one or more computer-readable storage media, which may be non-transitory. The memory 802 may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 802 is used to store at least one program code, which is used to be executed by the processor 801 to implement the process executed by the terminal in the video processing method provided in the method embodiment of the present disclosure.
[0268] In some embodiments, the terminal 800 may further optionally include: a peripheral device interface 803 and at least one peripheral device. The processor 801, the memory 802 and the peripheral device interface 803 may be connected via a bus or a signal line. Each peripheral device may be connected to the peripheral device interface 803 via a bus, a signal line or a circuit board. Specifically, the peripheral device includes: at least one of a radio frequency circuit 804, a display screen 805, a camera assembly 806, an audio circuit 807, a positioning assembly 808 and a power supply 809.
[0269] The peripheral device interface 803 may be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 801 and the memory 802. In some embodiments, the processor 801, the memory 802, and the peripheral device interface 803 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 801, the memory 802, and the peripheral device interface 803 may be implemented on a separate chip or circuit board, which is not limited in this embodiment.
[0270] The radio frequency circuit 804 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 804 communicates with communication networks and other communication devices through electromagnetic signals. The radio frequency circuit 804 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals. Optionally, the radio frequency circuit 804 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, and the like. The radio frequency circuit 804 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes, but is not limited to: a metropolitan area network, various generations of mobile communication networks (2G, 3G, 4G and 5G), a wireless local area network and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 804 may also include circuits related to NFC (Near Field Communication), which is not limited in the present disclosure.
[0271] The display screen 805 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 805 is a touch display screen, the display screen 805 also has the ability to collect touch signals on the surface or above the surface of the display screen 805. The touch signal can be input to the processor 801 as a control signal for processing. At this time, the display screen 805 can also be used to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards. In some embodiments, the display screen 805 can be one, set on the front panel of the terminal 800; in other embodiments, the display screen 805 can be at least two, respectively set on different surfaces of the terminal 800 or in a folding design; in other embodiments, the display screen 805 can be a flexible display screen, set on the curved surface or folding surface of the terminal 800. Even, the display screen 805 can also be set to a non-rectangular irregular shape, that is, a special-shaped screen. The display screen 805 can be made of materials such as LCD (Liquid Crystal Display), OLED (Organic Light-Emitting Diode, organic light-emitting diode).
[0272] The camera assembly 806 is used to capture images or videos. Optionally, the camera assembly 806 includes a front camera and a rear camera. Typically, the front camera is arranged on the front panel of the terminal, and the rear camera is arranged on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth of field camera, a wide-angle camera, and a telephoto camera, so as to realize the fusion of the main camera and the depth of field camera to realize the background blur function, the fusion of the main camera and the wide-angle camera to realize panoramic shooting and VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera assembly 806 may also include a flash. The flash can be a monochrome temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation at different color temperatures.
[0273] The audio circuit 807 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, and convert the sound waves into electrical signals and input them into the processor 801 for processing, or input them into the radio frequency circuit 804 to achieve voice communication. For the purpose of stereo acquisition or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the terminal 800. The microphone may also be an array microphone or an omnidirectional acquisition microphone. The speaker is used to convert the electrical signal from the processor 801 or the radio frequency circuit 804 into sound waves. The speaker may be a traditional film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for purposes such as ranging. In some embodiments, the audio circuit 807 may also include a headphone jack.
[0274] The positioning component 808 is used to locate the current geographical location of the terminal 800 to implement navigation or LBS (Location Based Service).
[0275] The power supply 809 is used to power various components in the terminal 800. The power supply 809 can be an alternating current, a direct current, a disposable battery, or a rechargeable battery. When the power supply 809 includes a rechargeable battery, the rechargeable battery can support wired charging or wireless charging. The rechargeable battery can also be used to support fast charging technology.
[0276] In some embodiments, the terminal 800 further includes one or more sensors 810 , including but not limited to: an acceleration sensor 811 , a gyroscope sensor 812 , a pressure sensor 813 , a fingerprint sensor 814 , an optical sensor 815 , and a proximity sensor 816 .
[0277] The acceleration sensor 811 can detect the magnitude of acceleration on the three coordinate axes of the coordinate system established by the terminal 800. For example, the acceleration sensor 811 can be used to detect the components of gravity acceleration on the three coordinate axes. The processor 801 can control the display screen 805 to display the user interface in a horizontal view or a vertical view according to the gravity acceleration signal collected by the acceleration sensor 811. The acceleration sensor 811 can also be used to collect game or user motion data.
[0278] The gyroscope sensor 812 can detect the body direction and rotation angle of the terminal 800, and the gyroscope sensor 812 can cooperate with the acceleration sensor 811 to collect the user's 3D actions on the terminal 800. The processor 801 can implement the following functions based on the data collected by the gyroscope sensor 812: motion sensing (such as changing the UI according to the user's tilt operation), image stabilization during shooting, game control, and inertial navigation.
[0279] The pressure sensor 813 can be set on the side frame of the terminal 800 and / or the lower layer of the display screen 805. When the pressure sensor 813 is set on the side frame of the terminal 800, it can detect the user's holding signal of the terminal 800, and the processor 801 performs left and right hand recognition or shortcut operation according to the holding signal collected by the pressure sensor 813. When the pressure sensor 813 is set on the lower layer of the display screen 805, the processor 801 controls the operability controls on the UI interface according to the user's pressure operation on the display screen 805. The operability controls include at least one of a button control, a scroll bar control, an icon control, and a menu control.
[0280] The fingerprint sensor 814 is used to collect the user's fingerprint, and the processor 801 identifies the user's identity based on the fingerprint collected by the fingerprint sensor 814, or the fingerprint sensor 814 identifies the user's identity based on the collected fingerprint. When the user's identity is identified as a trusted identity, the processor 801 authorizes the user to perform relevant sensitive operations, including unlocking the screen, viewing encrypted information, downloading software, paying, and changing settings. The fingerprint sensor 814 can be set on the front, back, or side of the terminal 800. When a physical button or a manufacturer logo is set on the terminal 800, the fingerprint sensor 814 can be integrated with the physical button or the manufacturer logo.
[0281] The optical sensor 815 is used to collect the ambient light intensity. In one embodiment, the processor 801 can control the display brightness of the display screen 805 according to the ambient light intensity collected by the optical sensor 815. Specifically, when the ambient light intensity is high, the display brightness of the display screen 805 is increased; when the ambient light intensity is low, the display brightness of the display screen 805 is reduced. In another embodiment, the processor 801 can also dynamically adjust the shooting parameters of the camera assembly 806 according to the ambient light intensity collected by the optical sensor 815.
[0282] The proximity sensor 816, also called a distance sensor, is usually disposed on the front panel of the terminal 800. The proximity sensor 816 is used to collect the distance between the user and the front of the terminal 800. In one embodiment, when the proximity sensor 816 detects that the distance between the user and the front of the terminal 800 is gradually decreasing, the processor 801 controls the display screen 805 to switch from the screen-on state to the screen-off state; when the proximity sensor 816 detects that the distance between the user and the front of the terminal 800 is gradually increasing, the processor 801 controls the display screen 805 to switch from the screen-off state to the screen-on state.
[0283] Those skilled in the art will understand that Figure 8The structure shown in the figure does not constitute a limitation on the terminal 800, and the terminal 800 may include more or less components than those shown in the figure, or combine some components, or adopt a different component arrangement.
[0284] The computer device mentioned in the embodiments of the present disclosure may be provided as a server. Fig. 9 It is a block diagram of a server according to an exemplary embodiment. The server 900 may have relatively large differences due to different configurations or performances, and may include one or more processors (Central Processing Units, CPU) 901 and one or more memories 902, wherein the one or more memories 902 store at least one program code, and the at least one program code is loaded and executed by the one or more processors 901 to implement the process executed by the server in the video processing method provided by the above-mentioned various method embodiments. Of course, the server 900 may also have components such as a wired or wireless network interface, a keyboard, and an input and output interface for input and output. The server 900 may also include other components for implementing device functions, which will not be described in detail here.
[0285] In an exemplary embodiment, a computer-readable storage medium including a program code is also provided, such as a memory 802 or a memory 902 including the program code, and the program code can be executed by the processor 801 of the terminal 800 or the processor 901 of the server 900 to complete the above-mentioned video processing method. In some embodiments, the computer-readable storage medium can be a ROM (Read-Only Memory), a RAM (Random Access Memory), a CD-ROM (Compact-Disc Read-Only Memory), a magnetic tape, a floppy disk, an optical data storage device, etc.
[0286] In an exemplary embodiment, a computer program product is also provided, including a computer program, and when the computer program is executed by a processor, the above-mentioned video processing method is implemented.
[0287] In some embodiments, the computer program involved in the embodiments of the present disclosure may be deployed and executed on a computer device, or on multiple computer devices located at one location, or on multiple computer devices distributed at multiple locations and interconnected through a communication network. Multiple computer devices distributed at multiple locations and interconnected through a communication network may constitute a blockchain system.
[0288] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The description and examples are to be considered exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.
[0289] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A video processing method, characterized in that: The method comprises: For the i-th frame image in the video, image feature information and first difference information of the i-1-th frame image and image feature information and second difference information of the i+1-th frame image are obtained, where i is a positive integer greater than 1, the image feature information represents detail features of the corresponding image, the first difference information represents the difference between the corresponding image and a subsequent frame image of the image in the detail features, and the second difference information represents the difference between the corresponding image and a previous frame image of the image in the detail features; Determine the difference between the image feature information of the i-1th frame image and the first difference information of the i-1th frame image as the first conversion information; determine the difference between the image feature information of the i+1th frame image and the second difference information of the i+1th frame image as the second conversion information; the first conversion information represents the parameters required when the i-1th frame image is converted to the i-th frame image, and the second conversion information represents the parameters required when the i+1th frame image is converted to the i-th frame image; Determining super-resolution information of the i-th frame of image based on image feature information of the i-th frame of image, the first conversion information, and the second conversion information; Based on the super-resolution information of multiple frames of images in the video, a super-resolution video is obtained.
2. The video processing method according to claim 1, characterized in that: The process of acquiring the image feature information and the first difference information of the i-1th frame image includes: Inputting the i-1th frame image and the adjacent frame images of the i-1th frame image into a feature extraction network, extracting hidden layer features of the i-1th frame image based on the i-1th frame image and the adjacent frame images of the i-1th frame image through the feature extraction network, wherein the feature extraction network is trained based on at least one sample frame image, the adjacent frame images of the at least one sample frame image, and the hidden layer features of the at least one sample frame image; Based on the hidden layer features of the i-1th frame image, image feature information of the i-1th frame image is determined; based on the hidden layer features of the i-1th frame image, the i-th frame image and the i-1th frame image, first difference information of the i-1th frame image is determined.
3. The video processing method according to claim 2, characterized in that: In the case where i is a positive integer greater than 2, the method further includes: The hidden layer features of the i-1th frame image, the adjacent frame images of the i-1th frame image, and the i-2th frame image are input into the feature extraction network. The hidden layer features of the i-1th frame image are extracted based on the hidden layer features of the i-1th frame image, the adjacent frame images of the i-1th frame image, and the i-2th frame image through the feature extraction network. The feature extraction network is trained based on the hidden layer features of at least one frame sample image, the adjacent frame images of the at least one frame sample image, the previous frame image of the at least one frame sample image, and the hidden layer features of the at least one frame sample image.
4. The video processing method according to claim 2, characterized in that: The determining of image feature information of the i-1th frame image based on the hidden layer features of the i-1th frame image, and the determining of first difference information of the i-1th frame image based on the i-1th frame image, the i-th frame image and the hidden layer features of the i-1th frame image include: Inputting the hidden layer features of the i-1th frame image into a first feature extraction subnetwork, extracting image feature information of the i-1th frame image based on the hidden layer features of the i-1th frame image through the first feature extraction subnetwork, wherein the first feature extraction subnetwork is trained based on the hidden layer features of at least one frame of sample image and the image feature information of the at least one frame of sample image; The i-1th frame image, the i-th frame image and the hidden layer features of the i-1th frame image are input into a second feature extraction subnetwork, and the first difference information of the i-1th frame image is extracted based on the i-1th frame image, the i-th frame image and the hidden layer features of the i-1th frame image through the second feature extraction subnetwork, and the second feature extraction subnetwork is trained based on at least one frame of sample image, the subsequent frame of the at least one frame of sample image, the hidden layer features of the at least one frame of sample image and the first difference information of the at least one frame of sample image.
5. The video processing method according to claim 1, characterized in that: The process of acquiring the image feature information and the second difference information of the i+1th frame image includes: Inputting the i+1th frame image and the adjacent frame image of the i+1th frame image into a feature extraction network, extracting hidden layer features of the i+1th frame image based on the i+1th frame image and the adjacent frame image of the i+1th frame image through the feature extraction network, wherein the feature extraction network is trained based on at least one sample frame image, the adjacent frame image of the at least one sample frame image, and the hidden layer features of the at least one sample frame image; Based on the hidden layer features of the i+1th frame image, the image feature information of the i+1th frame image is determined; based on the hidden layer features of the i+1th frame image, the i-th frame image and the i+1th frame image, the second difference information of the i+1th frame image is determined.
6. The video processing method according to claim 5, characterized in that: The method further comprises: The i+1th frame image, the adjacent frame images of the i+1th frame image, and the hidden layer features of the i+1th frame image are input into the feature extraction network. The hidden layer features of the i+1th frame image are extracted based on the hidden layer features of the i+1th frame image, the adjacent frame images of the i+1th frame image, and the i-th frame image through the feature extraction network. The feature extraction network is trained based on the hidden layer features of at least one frame sample image, the adjacent frame images of the at least one frame sample image, the previous frame image of the at least one frame sample image, and the hidden layer features of the at least one frame sample image.
7. The video processing method according to claim 5, characterized in that: The determining of image feature information of the i+1th frame image based on the hidden layer features of the i+1th frame image, and the determining of second difference information of the i+1th frame image based on the hidden layer features of the i+1th frame image, the i-th frame image and the i+1th frame image comprises: Inputting the hidden layer features of the i+1th frame image into a first feature extraction subnetwork, extracting image feature information of the i+1th frame image based on the hidden layer features of the i+1th frame image through the first feature extraction subnetwork, wherein the first feature extraction subnetwork is trained based on the hidden layer features of at least one frame of sample image and the image feature information of the at least one frame of sample image; The i+1th frame image, the i-th frame image and the hidden layer features of the i+1th frame image are input into a third feature extraction subnetwork, and the second difference information of the i+1th frame image is extracted based on the hidden layer features of the i+1th frame image, the i-th frame image and the i+1th frame image through the third feature extraction subnetwork. The third feature extraction subnetwork is trained based on at least one frame of sample image, the previous frame of the at least one frame of sample image, the hidden layer features of the at least one frame of sample image and the second difference information of the at least one frame of sample image.
8. The video processing method according to any one of claims 2, 3, 5 and 6, characterized in that: The feature extraction network is composed of multiple residual modules, wherein a residual module includes a first two-dimensional convolutional layer, an activation function connected to the first two-dimensional convolutional layer, and a second two-dimensional convolutional layer connected to the activation function, and the activation function is used to indicate the functional mapping relationship between the corresponding image, the adjacent frame images of the image and the hidden layer features of the image.
9. The video processing method according to claim 1, characterized in that: The determining the super-resolution information of the i-th frame image based on the image feature information of the i-th frame image, the first conversion information, and the second conversion information includes: The image feature information of the i-th frame image, the first conversion information and the second conversion information are input into a temporal convolutional network, and convolution processing is performed based on the image feature information of the i-th frame image, the first conversion information and the second conversion information through the temporal convolutional network to obtain super-resolution information of the i-th frame image. The temporal convolutional network is trained based on the image feature information of at least one frame of sample image, the first conversion information, the second conversion information and the super-resolution information of the at least one frame of sample image.
10. The video processing method according to claim 1, characterized in that: The step of obtaining a super-resolution video based on super-resolution information of multiple frames of images in the video includes: Based on the super-resolution information of the multiple frames of images in the video, sub-pixel rearrangement processing is performed to obtain sub-pixel rearrangement results of the multiple frames of images; Performing upsampling processing on the multiple frames of images to obtain upsampling results of the multiple frames of images; The super-resolution video is generated based on the sub-pixel rearrangement results of the multiple frames of images and the up-sampling results of the multiple frames of images.
11. The video processing method according to any one of claims 1 to 7, 9 to 10, characterized in that: The image feature information, the first difference information, and the second difference information are all in the form of residual maps. The residual map corresponding to the image feature information is used to represent the distribution of sub-pixels in the corresponding image; the residual map corresponding to the first difference information is used to represent the sub-pixel difference between the image and the next frame image of the image; the residual map corresponding to the second difference information is used to represent the sub-pixel difference between the image and the previous frame image of the image.
12. The video processing method according to claim 8, characterized in that: The image feature information, the first difference information, and the second difference information are all in the form of residual maps. The residual map corresponding to the image feature information is used to represent the distribution of sub-pixels in the corresponding image; the residual map corresponding to the first difference information is used to represent the sub-pixel difference between the image and the next frame image of the image; the residual map corresponding to the second difference information is used to represent the sub-pixel difference between the image and the previous frame image of the image.
13. A video processing device, characterized in that: The device comprises: an information acquisition unit, configured to acquire, for an i-th frame image in a video, image feature information and first difference information of an i-1th frame image, and image feature information and second difference information of an i+1th frame image, wherein i is a positive integer greater than 1, the image feature information represents detail features of a corresponding image, the first difference information represents a difference between a corresponding image and a subsequent frame image of the image in the detail features, and the second difference information represents a difference between a corresponding image and a previous frame image of the image in the detail features; a conversion information determination unit, configured to determine a difference between image feature information of the i-1th frame image and first difference information of the i-1th frame image as first conversion information, and determine a difference between image feature information of the i+1th frame image and second difference information of the i+1th frame image as second conversion information; the first conversion information represents parameters required when the i-1th frame image is converted to the i-th frame image, and the second conversion information represents parameters required when the i+1th frame image is converted to the i-th frame image; a super-resolution information determining unit, configured to determine the super-resolution information of the i-th frame image based on the image feature information of the i-th frame image, the first conversion information, and the second conversion information; The video acquisition unit is configured to execute super-resolution information based on multiple frames of images in the video to acquire a super-resolution video.
14. The video processing device according to claim 13, characterized in that: The information acquisition unit comprises: a feature extraction subunit, configured to input the i-1th frame image and the adjacent frame images of the i-1th frame image into a feature extraction network, and extract hidden layer features of the i-1th frame image based on the i-1th frame image and the adjacent frame images of the i-1th frame image through the feature extraction network, wherein the feature extraction network is trained based on at least one sample frame image, the adjacent frame images of the at least one sample frame image, and the hidden layer features of the at least one sample frame image; A determination subunit is configured to determine image feature information of the i-1th frame image based on hidden layer features of the i-1th frame image, and to determine first difference information of the i-1th frame image based on hidden layer features of the i-1th frame image, the i-th frame image and the i-1th frame image.
15. The video processing device according to claim 14, characterized in that: In the case where i is a positive integer greater than 2, the feature extraction subunit is further configured to perform: The hidden layer features of the i-1th frame image, the adjacent frame images of the i-1th frame image, and the i-2th frame image are input into the feature extraction network. The hidden layer features of the i-1th frame image are extracted based on the hidden layer features of the i-1th frame image, the adjacent frame images of the i-1th frame image, and the i-2th frame image through the feature extraction network. The feature extraction network is trained based on the hidden layer features of at least one frame sample image, the adjacent frame images of the at least one frame sample image, the previous frame image of the at least one frame sample image, and the hidden layer features of the at least one frame sample image.
16. The video processing device according to claim 14, characterized in that: The determining subunit is configured to execute: Inputting the hidden layer features of the i-1th frame image into a first feature extraction subnetwork, extracting image feature information of the i-1th frame image based on the hidden layer features of the i-1th frame image through the first feature extraction subnetwork, wherein the first feature extraction subnetwork is trained based on the hidden layer features of at least one frame of sample image and the image feature information of the at least one frame of sample image; The i-1th frame image, the i-th frame image and the hidden layer features of the i-1th frame image are input into a second feature extraction subnetwork, and the first difference information of the i-1th frame image is extracted based on the i-1th frame image, the i-th frame image and the hidden layer features of the i-1th frame image through the second feature extraction subnetwork, and the second feature extraction subnetwork is trained based on at least one frame of sample image, the subsequent frame of the at least one frame of sample image, the hidden layer features of the at least one frame of sample image and the first difference information of the at least one frame of sample image.
17. The video processing device according to claim 13, characterized in that: The information acquisition unit comprises: a feature extraction subunit, configured to input the i+1th frame image and the adjacent frame image of the i+1th frame image into a feature extraction network, and extract the hidden layer features of the i+1th frame image based on the i+1th frame image and the adjacent frame image of the i+1th frame image through the feature extraction network, wherein the feature extraction network is trained based on at least one sample frame image, the adjacent frame image of the at least one sample frame image, and the hidden layer features of the at least one sample frame image; The determination subunit is configured to determine the image feature information of the i+1th frame image based on the hidden layer features of the i+1th frame image, and determine the second difference information of the i+1th frame image based on the hidden layer features of the i+1th frame image, the i-th frame image and the i+1th frame image.
18. The video processing device according to claim 17, characterized in that: The feature extraction subunit is further configured to perform: The i+1th frame image, the adjacent frame images of the i+1th frame image, and the hidden layer features of the i+1th frame image are input into the feature extraction network. The hidden layer features of the i+1th frame image are extracted based on the hidden layer features of the i+1th frame image, the adjacent frame images of the i+1th frame image, and the i-th frame image through the feature extraction network. The feature extraction network is trained based on the hidden layer features of at least one frame sample image, the adjacent frame images of the at least one frame sample image, the previous frame image of the at least one frame sample image, and the hidden layer features of the at least one frame sample image.
19. The video processing device according to claim 17, characterized in that: The determining subunit is configured to execute: Inputting the hidden layer features of the i+1th frame image into a first feature extraction subnetwork, extracting image feature information of the i+1th frame image based on the hidden layer features of the i+1th frame image through the first feature extraction subnetwork, wherein the first feature extraction subnetwork is trained based on the hidden layer features of at least one frame of sample image and the image feature information of the at least one frame of sample image; The i+1th frame image, the i-th frame image and the hidden layer features of the i+1th frame image are input into a third feature extraction subnetwork, and the second difference information of the i+1th frame image is extracted based on the hidden layer features of the i+1th frame image, the i-th frame image and the i+1th frame image through the third feature extraction subnetwork. The third feature extraction subnetwork is trained based on at least one frame of sample image, the previous frame of the at least one frame of sample image, the hidden layer features of the at least one frame of sample image and the second difference information of the at least one frame of sample image.
20. The video processing device according to any one of claims 14, 15, 17 and 18, characterized in that: The feature extraction network is composed of multiple residual modules, wherein a residual module includes a first two-dimensional convolutional layer, an activation function connected to the first two-dimensional convolutional layer, and a second two-dimensional convolutional layer connected to the activation function, and the activation function is used to indicate the functional mapping relationship between the corresponding image, the adjacent frame images of the image and the hidden layer features of the image.
21. The video processing device according to claim 13, characterized in that: The super-resolution information determining unit is configured to execute: The image feature information of the i-th frame image, the first conversion information and the second conversion information are input into a temporal convolutional network, and convolution processing is performed based on the image feature information of the i-th frame image, the first conversion information and the second conversion information through the temporal convolutional network to obtain super-resolution information of the i-th frame image. The temporal convolutional network is trained based on the image feature information of at least one frame of sample image, the first conversion information, the second conversion information and the super-resolution information of the at least one frame of sample image.
22. The video processing device according to claim 13, characterized in that: The video acquisition unit is configured to execute: Based on the super-resolution information of the multiple frames of images in the video, sub-pixel rearrangement processing is performed to obtain sub-pixel rearrangement results of the multiple frames of images; Performing upsampling processing on the multiple frames of images to obtain upsampling results of the multiple frames of images; The super-resolution video is generated based on the sub-pixel rearrangement results of the multiple frames of images and the up-sampling results of the multiple frames of images.
23. The video processing device according to any one of claims 13 to 19, 21 to 22, characterized in that: The image feature information, the first difference information, and the second difference information are all in the form of residual maps. The residual map corresponding to the image feature information is used to represent the distribution of sub-pixels in the corresponding image; the residual map corresponding to the first difference information is used to represent the sub-pixel difference between the image and the next frame image of the image; the residual map corresponding to the second difference information is used to represent the sub-pixel difference between the image and the previous frame image of the image.
24. The video processing device according to claim 20, characterized in that: The image feature information, the first difference information, and the second difference information are all in the form of residual maps. The residual map corresponding to the image feature information is used to represent the distribution of sub-pixels in the corresponding image; the residual map corresponding to the first difference information is used to represent the sub-pixel difference between the image and the next frame image of the image; the residual map corresponding to the second difference information is used to represent the sub-pixel difference between the image and the previous frame image of the image.
25. A computer device, characterized in that: The computer device comprises: one or more processors; a memory for storing program code executable by the processor; The processor is configured to execute the program code to implement the video processing method according to any one of claims 1 to 12.
26. A computer-readable storage medium, characterized in that: When the program code in the computer-readable storage medium is executed by a processor of a computer device, the computer device is enabled to execute the video processing method according to any one of claims 1 to 12.
27. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the video processing method according to any one of claims 1 to 12 is implemented.