Video Transmission System Based on Adaptive Key Frame Selection Strategy for Video Super-Resolution at Low Bit Rates

By adopting an adaptive keyframe selection strategy in the video transmission system and selecting high-resolution keyframes based on the difference in frames within the time window, the problems of low reconstruction quality and unreduced data volume in traditional systems are solved, and the effects of high-quality reconstruction and data volume reduction are achieved.

CN119031187BActive Publication Date: 2025-08-05TSINGHUA UNIVERSITY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410947739.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-16
Publication Date
2025-08-05
Estimated Expiration
2044-07-16

AI Technical Summary

Technical Problem

In traditional video transmission systems, the unified keyframe selection strategy fails to fully utilize the specificity of video content, resulting in low quality of reconstructed videos and ineffective data transmission.

Method used

Adaptive keyframe selection strategy based on inter-frame differences within the time window is adopted to select high-resolution keyframes from high-resolution video sequences and downsample them at the terminal, send low-resolution video and high-resolution keyframes to the cloud for reconstruction, and use the advantages of cloud computing to reduce the amount of data.

Benefits of technology

Through adaptive selection of keyframes, the quality of video reconstruction is significantly improved, and the amount of data transmission is greatly reduced, making full use of the specificity of video content and cloud computing capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119031187B_ABST
    Figure CN119031187B_ABST
Patent Text Reader

Abstract

The present invention provides a video transmission system based on an adaptive key frame selection strategy for video super-resolution at low bit rates, which relates to the technical field of video transmission. In the embodiments of the present invention, only high-resolution key frames and low-resolution videos are transmitted between the terminal and the cloud, which can greatly reduce the amount of data during the transmission process. In addition, in the embodiments of the present invention, after obtaining the high-resolution key frames and low-resolution videos, image compression standards and video compression standards can be used for compression respectively, further reducing the amount of data during the transmission process. In the embodiments of the present invention, key frames are determined based on the differences between frames within a time window, and the positions of the selected key frames vary with the specific content of the video, which can improve the effect of video super-resolution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the technical field of video transmission, and in particular, to a video transmission system based on an adaptive key frame selection strategy for video super-resolution at low bitrates. Background Art

[0002] With the continuous improvement of social governance requirements and the development of people's security awareness, more and more surveillance cameras are deployed in various locations, including supermarkets, airports, villages, streets, and private houses. As a result, a large amount of video data is generated, and this video data needs to be sent to a cloud server for subsequent analysis, storage, and video content retrieval. In the traditional video field, the popularity of video stream live broadcast, video on demand, and people's demand for higher video quality have also contributed to the explosion of video data volume.

[0003] In order to reduce the amount of data during video transmission, the traditional method is to compress and encode the video content using various methods before transmitting the video content. The goal of video compression is to minimize video distortion at a certain bitrate, mainly by using the intra-frame and inter-frame correlations of video frames to remove temporal and spatial redundancies, significantly reducing the amount of data while maintaining the perceptual quality.

[0004] However, traditional video compression systems often have difficulty in comprehensively optimizing the system. The powerful end-to-end learning ability of deep learning enables the implementation of video compression systems based on deep learning. Compared with traditional methods, video compression systems based on deep learning show superior performance. Video compression systems based on deep learning can learn video compression in an end-to-end manner.

[0005] In addition to video compression systems aimed at reducing the amount of transmitted data, better end-cloud collaboration solutions can further reduce the amount of data during video transmission. In related technologies, a surveillance video data transmission system based on end-cloud computing is proposed. This system can significantly reduce the bitrate of transmitted data while maintaining the reconstructed video quality. Specifically, the captured high-resolution (HR) video is downsampled at the terminal and then sent to the cloud server together with sparsely selected key frames. On the cloud server, the received low-resolution (LR) video and HR key frame data stream are decoded, and the video is super-resolved using a video super-resolution model to obtain the final HR video. Since video resolution significantly affects the amount of data during transmission, reducing video resolution can effectively reduce the amount of video data.

[0006] The edge-cloud collaborative system can further reduce the amount of data during transmission based on existing video compression methods. During transmission, LR videos lack high-frequency information such as details and textures, and the amount of data is significantly reduced. While HR key frames have more high-frequency details and can provide rich prior knowledge for the video reconstruction model to reconstruct the video. Therefore, the key frame selection strategy will significantly affect the quality of the reconstructed video. However, in traditional edge-cloud collaborative systems, a unified key frame selection strategy is usually adopted, that is, key frames are selected at fixed intervals, and the quality of the finally reconstructed video based on these key frames is not high, not fully utilizing the specific content of the video, and there is room for further improvement. Summary of the Invention

[0007] Embodiments of the present invention provide a video transmission system with an adaptive key frame selection strategy based on video super-resolution at low bitrates to at least partially solve the problems existing in existing video transmission systems.

[0008] In a first aspect of an embodiment of the present invention, a video transmission system with an adaptive key frame selection strategy based on video super-resolution at low bitrates is provided. The system includes:

[0009] An edge device that acquires a high-resolution video, performs downsampling, and converts it into a low-resolution video with the same frame rate; uses the difference between frames within a time window as an index to measure the quality of key frames, selects high-resolution key frames from the high-resolution video sequence, and sends the low-resolution video and high-resolution key frames to the cloud;

[0010] The cloud reconstructs based on the received low-resolution video and the received high-resolution key frames to obtain the reconstructed high-resolution video.

[0011] Optionally, using the difference between frames within a time window as an index to measure the quality of key frames and selecting high-resolution key frames from the high-resolution video sequence includes:

[0012] Determine the length of the time window and the selection range of high-resolution key frames;

[0013] Based on the selection range of high-resolution key frames, determine multiple center frames and multiple surrounding frames for each center frame. The total number of frames of a center frame and its multiple surrounding frames is the length of the time window;

[0014] Calculate the average value of the PSNR of each center frame for its multiple surrounding frames. The PSNR of a center frame for its surrounding frames is used to measure the ability of the center frame to provide high-frequency information for its surrounding frames;

[0015] Determine the center frame with the largest average value of PSNR among the multiple center frames as the high-resolution key frame.

[0016] Optionally, taking the difference between frames within a time window as an indicator to measure the quality of key frames, and selecting high-resolution key frames from a high-resolution video sequence, including:

[0017] Taking the first frame and the last frame of the high-resolution video as key frames;

[0018] The selection range of the high-resolution key frames is from a to b, where a > 1 and b < m, and m is the total number of frames of the high-resolution video.

[0019] Optionally, determining the selection range of high-resolution key frames, including:

[0020] For each video in the validation set, using high-resolution video frames with different frame numbers as key frames, calculating the change in the average value of PSNR corresponding to each video, and the change in the average value of PSNR is used to characterize the change in the reconstructed high-resolution video;

[0021] According to the change in the average value of PSNR corresponding to each video, determining the frame number range where the average value of PSNR is not lower than the preset threshold as the selection range of high-resolution key frames.

[0022] Optionally, determining the length of the time window, including:

[0023] For each video in the validation set, using high-resolution video frames with different frame numbers as high-resolution key frames, calculating the change in PSNR corresponding to different high-resolution key frames, and the change in PSNR is used to characterize the change in the reconstructed high-resolution video;

[0024] According to the change in PSNR corresponding to different high-resolution key frames, determining the propagation range of high-frequency information provided by different high-resolution key frames;

[0025] According to the propagation range of high-frequency information provided by different high-resolution key frames, determining the length of the time window.

[0026] Optionally, according to the propagation range of high-frequency information provided by different high-resolution key frames, determining the length of the time window, including:

[0027] Comparing the propagation ranges of high-frequency information provided by different high-resolution key frames to determine the overlapping propagation range;

[0028] Based on the overlapping propagation range, determining the length of the time window.

[0029] Optionally, when multiple high-resolution key frames are determined, calculating the average interval of the multiple high-resolution key frames;

[0030] In the case where the average interval does not meet the sparse parameter value, re-determine the length of the time window and the selection range of high-resolution key frames, and re-determine multiple high-resolution key frames until the average interval of the determined multiple high-resolution key frames meets the sparse parameter value.

[0031] In a second aspect of the embodiments of the present invention, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes, it implements the video transmission system of the adaptive key frame selection strategy based on video super-resolution at low bitrates as described in the first aspect of the present invention.

[0032] In a third aspect of the embodiments of the present invention, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, it implements the video transmission system of the adaptive key frame selection strategy based on video super-resolution at low bitrates as described in the first aspect of the present invention.

[0033] In a fourth aspect of the embodiments of the present invention, a computer program product is provided, including a computer program / instructions. When the computer program / instructions are implemented by a processor, it implements the video transmission system of the adaptive key frame selection strategy based on video super-resolution at low bitrates as described in the first aspect of the present invention.

[0034] In the embodiments of the present invention, only high-resolution key frames and low-resolution videos are transmitted between the terminal and the cloud, which can greatly reduce the amount of data during the transmission process. In addition, in the embodiments of the present invention, after obtaining high-resolution key frames and low-resolution videos, image compression standards and video compression standards can be used for compression respectively, further reducing the amount of data during the transmission process. In the embodiments of the present invention, key frames are determined based on the differences between frames within the time window. Compared with the key frames determined at a fixed interval, the positions of the selected key frames change with the specific content of the video, which can improve the effect of video super-resolution. Description of the Drawings

[0035] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0036] Figure 1 It is a schematic diagram of the architecture of the video transmission system of the adaptive key frame selection strategy based on video super-resolution at low bitrates provided by the embodiments of the present invention;

[0037] Figure 2It is an exemplary schematic diagram for determining key frames in a video transmission system based on an adaptive key frame selection strategy for video super-resolution at a low bit rate provided by an embodiment of the present invention;

[0038] Figure 3 It is an exemplary schematic diagram of the change in the average value of PSNR corresponding to different reconstructed videos when frames with different frame numbers are used as high-resolution key frames in a video transmission system based on an adaptive key frame selection strategy for video super-resolution at a low bit rate provided by an embodiment of the present invention;

[0039] Figure 4 It is an exemplary schematic diagram of PSNR corresponding to high-resolution video frames with different frame numbers as high-resolution key frames in a video transmission system based on an adaptive key frame selection strategy for video super-resolution at a low bit rate provided by an embodiment of the present invention;

[0040] Figure 5 It is a schematic flow diagram of the steps of a video transmission method based on an adaptive key frame selection strategy for video super-resolution at a low bit rate provided by an embodiment of the present invention. Detailed implementation manners

[0041] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners.

[0042] In a traditional end-cloud collaboration system, a unified key frame selection strategy is adopted, which selects key frames at a fixed interval, ignoring the correlation between the picture contents of each frame and also ignoring the context information of each frame in a longer time dimension.

[0043] Based on this, an embodiment of the present invention proposes to use the difference between frames within a time window as an index to measure the quality of key frames, select high-resolution key frames from a high-resolution video sequence, and send the low-resolution video and high-resolution key frames to the cloud. Thus, key frames can be determined based on the differences between each frame within a time window. This selection strategy takes into account the differences between all frames within a time window, can consider the context information of each frame in a longer time dimension, and, by measuring the quality of key frames based on the difference between frames, can measure the ability of key frames to transmit high-frequency information to other frames. Therefore, a reconstructed video with better quality can be obtained based on the selected optimal key frames.

[0044] Specifically, in an embodiment of the present invention, a video transmission system based on an adaptive key frame selection strategy for video super-resolution at a low bit rate is proposed. As Figure 1 shown, it shows an architecture schematic diagram of a video transmission system based on an adaptive key frame selection strategy for video super-resolution at a low bit rate provided by an embodiment of the present invention. The system includes:

[0045] The terminal acquires a high-resolution video, downsamples it, and converts it into a low-resolution video with the same frame rate; uses the difference between frames within a time window as an index to measure the quality of key frames, selects high-resolution key frames from the high-resolution video sequence, and sends the low-resolution video and the high-resolution key frames to the cloud;

[0046] The cloud reconstructs based on the received low-resolution video and the received high-resolution key frames to obtain a reconstructed high-resolution video.

[0047] In the embodiments of the present invention, the terminal refers to a terminal that processes the acquired video on the video acquisition side, which can be any terminal with data processing capabilities, including a video acquisition end with data processing capabilities. The cloud refers to a cloud server that analyzes, stores, and retrieves video content of the videos uploaded by the terminal, which can be any server with data processing capabilities.

[0048] In the embodiments of the present invention, in order to make full use of the advantages of cloud computing and reduce the amount of data in the transmission process from the terminal to the cloud, the terminal downsamples the high-resolution video acquired by the video acquisition device, reduces the resolution, and converts it into a low-resolution video with the same frame rate.

[0049] Specifically, the high-resolution video can be downsampled multiple times to obtain a low-resolution video that meets the low bit rate requirements in the transmission process.

[0050] In the embodiments of the present invention, after the terminal processes to obtain the low-resolution video, it also needs to select high-resolution key frames from the high-resolution video sequence, and upload both the low-resolution video and the high-resolution key frames to the cloud, so that the cloud restores the low-resolution video based on the high-frequency information provided by the high-resolution key frames to obtain a reconstructed resolution video.

[0051] Specifically, the terminal can compress the low-resolution video using a video compression standard to obtain video compression data, compress the high-resolution key frames using an image compression standard to obtain key frame compression data, and send the video compression data and the key frame compression data to the cloud. The cloud decompresses the received data, and inputs the decompressed low-resolution video and high-resolution key frames into a video reconstruction model based on deep learning to obtain a reconstructed resolution video.

[0052] In the embodiments of the present invention, during the reconstruction process, the video reconstruction model utilizes temporal correlation to obtain a reconstructed resolution video based on the low-resolution video and the high-resolution key frames.

[0053] Specifically, in the embodiments of the present invention, the video reconstruction model can directly transfer the features of the high-resolution key frames to the features of each low-resolution video frame to improve the quality of video reconstruction.

[0054] Specifically, the video reconstruction model extracts features from the input high-resolution key frames and low-resolution video frames through multiple residual blocks, and then feeds the extracted features into a recurrent neural network for bidirectional propagation. The propagation layer is divided into four layers, namely two forward propagation layers and two backward propagation layers. Due to the motion between different frames, optical flow-guided deformable convolution is used for feature alignment, and multiple residual blocks are used to extract features from the low-resolution video frames and high-resolution key frames. g t denotes the features extracted from the low-resolution video frame L t and K j (where j represents the j-th key frame) denotes the features extracted from the high-resolution key frame. is the aggregated feature calculated at the t-th time step in the l-th propagation layer (where ).

[0055] To obtain the aggregated feature in forward propagation First, calculate the optical flow between the t-th low-resolution video frame and the (t - 1) low-resolution video frame, and the optical flow between the t-th low-resolution video frame and the j-th high-resolution key frame, which are respectively denoted as and Use the calculated optical flow to transform and K j as follows:

[0056]

[0057] where W represents the spatial transformation operation. Then calculate the optical flow residual to obtain the offsets o t→t-1 , o t→j and the modulation masks m t→t-1 , m t→j as follows:

[0058]

[0059] where C {o,m} represents the stack of convolutions, and σ represents the sigmoid function. Connect the separately calculated offsets and modulation masks to obtain the overall offset and modulation mask. Then use DCN to transform and K j as follows:

[0060] o t = c(o t→t-1 , o t→j )

[0061] m t = c(m t→t-1 , m t→j )

[0062]

[0063] Connect and and transfer them into the residual block stack to obtain aggregated features

[0064] For each low-resolution video frame, each layer has a feature output. These features are input into a filter based on an attention mechanism, as shown in the following formula:

[0065]

[0066] Where is the score of the attention between the low-resolution video input features and the aggregated output features of each layer.

[0067] Finally, a pixel-shuffle layer is used to obtain the video reconstruction result.

[0068] In the embodiments of the present invention, only high-resolution key frames and low-resolution videos are transmitted between the terminal and the cloud, which can greatly reduce the amount of data during the transmission process. In addition, in the embodiments of the present invention, after obtaining the high-resolution key frames and low-resolution videos, image compression standards and video compression standards can be used for compression respectively, further reducing the amount of data during the transmission process. In the embodiments of the present invention, key frames are determined based on the differences between frames within a time window. Compared with the key frames determined at a fixed interval, the positions of the selected key frames change with the specific content of the video, which can improve the effect of video super-resolution.

[0069] In an optional embodiment, taking the differences between frames within a time window as an index to measure the quality of key frames, and selecting high-resolution key frames from a high-resolution video sequence may include the following steps:

[0070] S101, determine the length of the time window and the selection range of high-resolution key frames;

[0071] S102, based on the selection range of high-resolution key frames, determine multiple center frames and multiple surrounding frames for each center frame. The total number of frames of a center frame and its multiple surrounding frames is the length of the time window;

[0072] S103, calculate the average value of the PSNR of each center frame for its multiple surrounding frames. The PSNR of a center frame for its surrounding frames is used to measure the ability of the center frame to provide high-frequency information for its surrounding frames;

[0073] In S104, the central frame with the largest average PSNR among multiple central frames is determined as the high-resolution key frame.

[0074] In the embodiments of the present invention, the above steps S101 to S104 are executed by the terminal.

[0075] In the embodiments of the present invention, for each video frame within the selection range of the high-resolution key frame, it is used as the central frame, and according to the length of the time window, with this central frame as the center, other surrounding frames within one time window of this central frame are determined, and then the similarity between the central frame and the surrounding frames within one time window is calculated. In the embodiments of the present invention, in order to reduce the complexity of similarity calculation, PSNR (Peak Signal-to-Noise Ratio) is selected to measure the inter-frame similarity. In the embodiments of the present invention, the average PSNR between the central frame and the surrounding frames within one time window is calculated, and the ability of the central frame to provide information for other frames is measured based on PSNR. Among the selection range of the high-resolution key frame, the central frame with the highest average PSNR is taken as the final key frame. In the embodiments of the present invention, the PSNR between two frames can be calculated based on the mean square error (MSE) between the two frames. MSE is obtained by calculating the square of the difference between the corresponding pixels of two images and then taking the average value. PSNR is the logarithmic reciprocal of MSE and is represented in logarithmic scale. A higher PSNR value indicates a smaller difference between two frame images. On the contrary, a lower PSNR value indicates a larger difference between the two frames.

[0076] Specifically, in the embodiments of the present invention, the key frame selection strategy can be implemented based on the key frame selection module of the terminal. Specifically, the low-resolution video sequence is input into the key frame selection module to select key frames. Let the length of the time window be k = 2n + 1, and the central frame at this time is L i . Calculate the average PSNR of the central frame for the surrounding frames L i-n , L i-n+1 , …, L i-1 , L i+1 , …, L i+n as follows:

[0077]

[0078] The position of the finally determined key frame Ik is:

[0079]

[0080] Exemplarily, as Figure 2As shown, it shows an exemplary schematic diagram for determining key frames in a video transmission system based on an adaptive key frame selection strategy for video super-resolution at a low bit rate provided by an embodiment of the present invention. Among them, the dashed box represents a time window, and the image frame within the red box represents the central frame for this calculation. Taking this central frame as the center, the total length of the time window is the number of frames included within the time window, and then the average PSNR between this central frame and other surrounding frames is calculated.

[0081] In an optional embodiment, the difference between frames within the time window is used as an index to measure the quality of key frames, and high-resolution key frames are selected from a high-resolution video sequence, including:

[0082] Taking the first frame and the last frame of the high-resolution video as key frames;

[0083] The selection range of the high-resolution key frames is from a to b, where a > 1 and b < m, and m is the total number of frames of the high-resolution video.

[0084] In the embodiment of the present invention, taking the first frame and the last frame of the high-resolution video as key frames to provide high-frequency information at the first frame and the last frame is beneficial for video reconstruction.

[0085] In the embodiment of the present invention, during the video transmission between the terminal and the cloud, in the form of a video stream, the terminal transmits a small number of video frame sequences to the cloud each time. For example, the number of frames of the video frame sequence transmitted each time can be 50 - 100 frames. In the case of transmitting a small number of video frames each time, the selection range of the high-resolution key frames can be determined as the second frame to the (m - 1)-th frame, so as to select one more frame as a key frame from the frames other than the first frame and the last frame, and finally transmit three high-resolution key frames to the cloud.

[0086] In an optional embodiment, determining the selection range of high-resolution key frames includes:

[0087] S201, for each video in the validation set, taking high-resolution video frames with different frame numbers as key frames, calculating the change of the average value of PSNR corresponding to each video, and the change of the average value of PSNR is used to characterize the change of the reconstructed high-resolution video;

[0088] S202, according to the change of the average value of PSNR corresponding to each video, determining the frame number range where the average value of PSNR is not lower than the preset threshold as the selection range of high-resolution key frames.

[0089] In the embodiment of the present invention, several video data can be obtained from a public video dataset as the validation set, and each video in the validation set is experimentally analyzed to determine the selection range of high-resolution key frames.

[0090] Specifically, video data with a video sequence length of 67 is selected as the validation set. Assume the positions of the selected high-resolution key frames are I k ={k1, k2, k3}, then the video sequence will be split into two parts. The first part is the low-resolution video frames L k1 to L k2 , which will be input into the video reconstruction model together with the high-resolution key frames H k1 and H k2 . The second part is the low-resolution video frames L k2 to L k3 , which will be input into the video reconstruction model together with the high-resolution key frames H k2 and H k3 . Since the video sequence needs to be split according to the key frames during the video reconstruction process, in the embodiments of the present invention, the first frame and the last frame of the video sequence are used as key frames. In order to reduce the bit rate while maintaining the quality of video reconstruction during transmission, it is necessary to determine a frame as a key frame among the intermediate video frames to optimize the video reconstruction effect. In the embodiments of the present invention, the relationship between the selection range of high-resolution key frames and the video reconstruction result is determined based on experiments on the validation set. For the four videos (000, 001, 002, 0) in the validation set, change the positions of the key frames, use high-resolution video frames with different frame numbers as key frames to reconstruct the video, and determine the average value of the corresponding reconstructed video PSNR, so as to obtain the change of the average value of PSNR corresponding to each video, as shown in Figure 3 shown. Figure 3 Shows the change of the average value of PSNR corresponding to each video. Among them, the upper left, upper right, lower left and lower right are respectively the schematic diagrams of the change of the average value of PSNR corresponding to the four videos (000, 001, 002, 0) during the change of the key frame position. Among them, the abscissa represents the frame number, the ordinate represents the PSNR value, and each point in the figure represents using the high-resolution video frame with the frame number corresponding to this point as the key frame to reconstruct the video and determining the average value of the corresponding reconstructed video PSNR.

[0091] Based on Figure 3 it can be seen that the optimal key frame position is usually near the center. Placing the key frame too close to the edge of the sequence will result in uneven segmentation, and the length of one segment will be too short. Such a short sequence cannot effectively utilize the long-term information aggregation, resulting in a decline in the video reconstruction effect.

[0092] In the embodiments of the present invention, according to the change of the average value of PSNR corresponding to each video, the frame number range where the average value of PSNR is not lower than the preset threshold can be determined as the selection range of high-resolution key frames.

[0093] Specifically, the frame number range with the average value of PSNR not lower than the preset threshold can be determined respectively based on the change conditions corresponding to each video. Finally, the frame number ranges of up to multiple videos are summarized, and the union is taken to obtain the selection range of the final high-resolution key frames.

[0094] In the embodiments of the present invention, the above steps S201 to S202 can be executed by the terminal or by the cloud, and the cloud sends the determined selection range of the high-resolution key frames to the terminal.

[0095] In an optional implementation manner, determining the length of the time window includes:

[0096] S301, for each video in the validation set, taking the high-resolution video frames with different frame numbers as high-resolution key frames, and calculating the change conditions of the PSNR corresponding to different high-resolution key frames, where the change conditions of the PSNR are used to characterize the change conditions of the reconstructed high-resolution video;

[0097] S302, according to the change conditions of the PSNR corresponding to different high-resolution key frames, determining the propagation range of the high-frequency information provided by different resolution key frames;

[0098] S303, according to the propagation range of the high-frequency information provided by different resolution key frames, determining the length of the time window.

[0099] In the embodiments of the present invention, for each video in the validation set, the change conditions of the PSNR of different video frames in the video reconstructed with different high-resolution key frames can be calculated in the video sequence, so as to obtain the change conditions of the PSNR corresponding to different high-resolution key frames. Finally, the change conditions of the PSNR corresponding to different high-resolution key frames respectively corresponding to multiple videos are analyzed to determine the length of the time window.

[0100] Specifically, as Figure 4 shown, it shows the change conditions of the PSNR corresponding to the high-resolution video frames with different frame numbers as high-resolution key frames for video 000 in the validation set. The straight line represents the change conditions of the PSNR between each frame and the 10th frame in the video reconstructed with the 10th frame as the key frame, the long dashed line represents the change conditions of the PSNR between each frame and the 33rd frame in the video reconstructed with the 33rd frame as the key frame, and the short dashed line represents the change conditions of the PSNR between each frame and the 55th frame in the video reconstructed with the 55th frame as the key frame. The abscissa represents the frame number and the ordinate represents the PSNR value.

[0101] By Figure 4It can be seen that as the frames move away from the key frame, the PSNR of the reconstructed video gradually decreases. This trend indicates that the frames closer to the key frame obtain more information from it. The high similarity in content and structure between these frames and the key frame makes it possible to use more valuable information. Therefore, the degree of similarity between the key frame and the frames can be used as a metric to measure the amount of information transmitted from the key frame to the frames. The information brought by the key frame can only spread a limited distance. When the distance between the surrounding frames and the key frame exceeds a certain distance, the key frame cannot provide more information to the surrounding frames. This may be because when the position of the surrounding frames exceeds a certain distance from the key frame, the motion between the key frame and the surrounding frames is too large to be effectively aligned by the deformable convolution guided by optical flow. That is to say, the key frame can usually provide more information to the surrounding frames only within a time window. Therefore, embodiments of the present invention propose that the length of the time window can be determined by comparing the propagation ranges of the high-frequency information provided by different high-resolution key frames.

[0102] Specifically, step S303 above includes:

[0103] S3031, comparing the propagation ranges of the high-frequency information provided by different high-resolution key frames to determine the overlapping propagation range;

[0104] S3032, determining the length of the time window based on the overlapping propagation range.

[0105] In embodiments of the present invention, steps S301 to S302 above can be executed by the terminal or by the cloud, and the cloud sends the determined length of the time window to the terminal.

[0106] For example Figure 4 as an example for illustrative purposes, in Figure 4 it can be seen that the overlapping part between the propagation ranges of the high-frequency information provided by different high-resolution key frames is 18 - 48. Based on this, the length of the time window can be set to 31.

[0107] In embodiments of the present invention, the key frame is selected within a suitable position range to avoid the position of the key frame being too close to the edge of the sequence, resulting in uneven sequence segmentation and inability to utilize long-term information aggregation. In addition, in embodiments of the present invention, the key frame is determined based on the inter-frame difference within a suitable time window. In embodiments of the present invention, selecting PSNR as a metric can characterize the amount of information that the key frame can bring to the surrounding frames.

[0108] In an alternative embodiment, when multiple high-resolution key frames are determined, the terminal further needs to calculate the average interval of the multiple high-resolution key frames; when the average interval does not meet the sparse parameter value, the length of the time window and the selection range of the high-resolution key frames are re-determined, and the multiple high-resolution key frames are re-determined until the average interval of the determined multiple high-resolution key frames meets the sparse parameter value.

[0109] In the embodiments of the present invention, in order to greatly reduce the amount of data generated by the selected high-resolution key frames, an evaluation of the sparsity of the determined high-resolution key frames is also proposed to select sufficiently sparse key frames.

[0110] In the embodiments of the present invention, a comparative embodiment is also provided to compare the influence of different key frame selection strategies on the video reconstruction quality. In the embodiments of the present invention, a training set is obtained based on the public dataset REDS, and a video reconstruction model is trained based on the training set. The batch size during training is set to 1. The initial learning rate of the main network during training is 1×10 -4 , and the initial learning rate of the optical flow network is 2.5×10 -5 . The patch size during training is set to 64×64. The model parameters are initialized using pre-trained parameters, and the model is trained for 100 rounds of iteration. Then, a validation set is obtained based on the public dataset, specifically as shown by the video numbers in Table 1. For each data in the validation set, two different key frame selection strategies are used to transmit and reconstruct the video, and the PSNR of the reconstructed video is calculated. During the experiment, the length of the time window in the key frame selection strategy provided by the embodiments of the present invention is 27. The experimental results are shown in Table 1 and Table 2. Table 1 shows the influence of different key frame selection strategies on the reconstruction results of each video in the validation set. It can be seen that the key frame selection strategy based on the time window can effectively improve the video reconstruction performance. Table 2 shows the influence of the key frame selection strategy on PSNR at different bitrates. It can be seen that as the bitrate of the low-resolution video decreases, video reconstruction becomes more and more difficult. This trend stems from the fact that as the bitrate decreases, the degradation of video frames becomes more and more serious.

[0111] Table 1 Influence of key frame selection strategy on PSNR

[0112]

[0113]

[0114] Table 2 Influence of key frame selection strategy on PSNR at different bitrates

[0115]

[0116] Based on the same inventive concept, an embodiment of the present invention further provides a video transmission method based on an adaptive key frame selection strategy for video super-resolution at a low bit rate, as follows Figure 5 shown, which shows a schematic flowchart of the steps of the video transmission method based on the adaptive key frame selection strategy for video super-resolution at a low bit rate provided by an embodiment of the present invention. Specifically, the method includes:

[0117] S501, the terminal acquires a high-resolution video, performs downsampling, and converts it into a low-resolution video with the same frame rate; uses the difference between frames within a time window as an index to measure the quality of key frames, selects high-resolution key frames from the high-resolution video sequence, and sends the low-resolution video and the high-resolution key frames to the cloud;

[0118] S502, the cloud reconstructs based on the received low-resolution video and the received high-resolution key frames to obtain a reconstructed high-resolution video.

[0119] Optionally, using the difference between frames within a time window as an index to measure the quality of key frames and selecting high-resolution key frames from the high-resolution video sequence includes:

[0120] Determine the length of the time window and the selection range of high-resolution key frames;

[0121] Based on the selection range of high-resolution key frames, determine multiple center frames and multiple surrounding frames for each center frame. The total number of frames of a center frame and its multiple surrounding frames is the length of the time window;

[0122] Calculate the average value of the PSNR of each center frame for its multiple surrounding frames. The PSNR of a center frame for its surrounding frames is used to measure the ability of the center frame to provide high-frequency information for its surrounding frames;

[0123] Determine the center frame with the largest average value of PSNR among the multiple center frames as the high-resolution key frame.

[0124] Optionally, using the difference between frames within a time window as an index to measure the quality of key frames and selecting high-resolution key frames from the high-resolution video sequence includes:

[0125] Use the first frame and the last frame of the high-resolution video as key frames;

[0126] The selection range of the high-resolution key frames is from a to b, where a>1 and b<m, and m is the total number of frames of the high-resolution video.

[0127] Optionally, determining the selection range of high-resolution key frames includes:

[0128] For each video in the validation set, high-resolution video frames with different frame numbers are used as key frames, and the change in the average value of the PSNR corresponding to each video is calculated. The change in the average value of the PSNR is used to characterize the change in the reconstructed high-resolution video;

[0129] According to the change in the average value of the PSNR corresponding to each video, the frame number range where the average value of the PSNR is not lower than the preset threshold is determined as the selection range of high-resolution key frames.

[0130] Optionally, determining the length of the time window includes:

[0131] For each video in the validation set, high-resolution video frames with different frame numbers are used as high-resolution key frames, and the change in the PSNR corresponding to different high-resolution key frames is calculated. The change in the PSNR is used to characterize the change in the reconstructed high-resolution video;

[0132] According to the change in the PSNR corresponding to different high-resolution key frames, the propagation range of the high-frequency information provided by different resolution key frames is determined;

[0133] According to the propagation range of the high-frequency information provided by different high-resolution key frames, the length of the time window is determined.

[0134] Optionally, according to the propagation range of the high-frequency information provided by different high-resolution key frames, determining the length of the time window includes:

[0135] Compare the propagation ranges of the high-frequency information provided by different high-resolution key frames to determine the overlapping propagation range;

[0136] Based on the overlapping propagation range, the length of the time window is determined.

[0137] Optionally, in the case where multiple high-resolution key frames are determined, calculate the average interval of the multiple high-resolution key frames;

[0138] In the case where the average interval does not meet the sparse parameter value, re-determine the length of the time window and the selection range of high-resolution key frames, and re-determine multiple high-resolution key frames until the average interval of the determined multiple high-resolution key frames meets the sparse parameter value.

[0139] Based on the same inventive concept, an embodiment of the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes, it implements the steps in the video transmission system of the adaptive key frame selection strategy based on video super-resolution at low bit rates as described in any of the above embodiments.

[0140] Based on the same inventive concept, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps in the video transmission system of the adaptive key-frame selection strategy based on video super-resolution at a low bit rate described in any of the above embodiments are implemented.

[0141] Based on the same inventive concept, an embodiment of the present invention provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, the steps in the video transmission system of the adaptive key-frame selection strategy based on video super-resolution at a low bit rate described in any of the above embodiments are implemented.

[0142] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is the difference from other embodiments. The same or similar parts among the embodiments can be referred to each other.

[0143] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a device, or a computer program product. Therefore, the embodiments of the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0144] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (devices), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable terminal devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable terminal devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0145] These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable terminal devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0146] These computer program instructions can also be loaded onto a computer or other programmable terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in one process or multiple processes and / or blocks. Figure 1 one process or multiple processes and / or blocks Figure 1 steps for implementing the functions specified in one block or multiple blocks.

[0147] Although the preferred embodiments of the embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they know the basic creative concept. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present invention.

[0148] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising the element.

[0149] The above has introduced in detail a video transmission system with an adaptive key frame selection strategy based on video super-resolution at a low code rate. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A video transmission system based on an adaptive key frame selection strategy for video super-resolution at low bit rate, characterized in that: The system comprises: The terminal obtains the high-resolution video, downsamples it, and converts it into a low-resolution video with the same frame rate; The terminal uses the difference between frames in the time window as an indicator to measure the quality of the key frame, and selects a high-resolution key frame from the high-resolution video sequence, including: determining the length of the time window k=2n+1 and the selection range of the high-resolution key frame; based on the selection range of the high-resolution key frame, determining multiple center frames and multiple surrounding frames of each center frame, and the total number of frames of a center frame and multiple surrounding frames of the center frame is the length of the time window; calculating the average PSNR of each center frame with respect to multiple surrounding frames of the center frame, including: for the i-th center frame L i , calculate the i-th central frame L i For each surrounding frame L in the 2n surrounding frames i−n , L i−n+1 ,…, L i−1 , L i+1 ,…,L i+n The average PSNR of a central frame to its surrounding frames is used to measure: the ability of the central frame to provide high-frequency information to its surrounding frames; the central frame with the largest average PSNR among multiple central frames is determined as the high-resolution key frame; The terminal sends the low-resolution video and the high-resolution key frame to the cloud; In the cloud, reconstruction is performed based on the received low-resolution video and the received high-resolution key frames to obtain a reconstructed high-resolution video; The terminal determines the length of the time window, including: For the videos in the validation set, high-resolution video frames with different frame numbers are used as high-resolution key frames. The changes in PSNR corresponding to different high-resolution key frames are calculated. The changes in PSNR are used to represent the changes in the reconstructed high-resolution video. According to the changes in PSNR corresponding to different high-resolution key frames, the propagation range of high-frequency information provided by different high-resolution key frames is determined; The length of the time window is determined according to the propagation range of the high-frequency information provided by different high-resolution key frames.

2. The video transmission system based on the adaptive key frame selection strategy for video super-resolution at low bit rate according to claim 1, characterized in that: The difference between frames within the time window is used as an indicator to measure the quality of key frames. High-resolution key frames are selected from high-resolution video sequences, including: Using the first frame and the last frame of the high-resolution video as key frames; The selection range of the high-resolution key frame is from a to b, where a is greater than 1 and b is less than m, and m is the total number of frames of the high-resolution video.

3. The video transmission system based on the adaptive key frame selection strategy for video super-resolution at low bit rate according to claim 1, characterized in that: Determine the selection range for high-resolution keyframes, including: For each video in the validation set, high-resolution video frames with different frame numbers are used as key frames. The change in the average PSNR corresponding to each video is calculated. The change in the average PSNR is used to represent the change in the reconstructed high-resolution video. According to the change of the average value of PSNR corresponding to each video, the frame number range in which the average value of PSNR is not lower than a preset threshold is determined as the selection range of high-resolution key frames.

4. The video transmission system based on the adaptive key frame selection strategy for video super-resolution at low bit rate according to claim 1, characterized in that: The length of the time window is determined based on the propagation range of the high-frequency information provided by different high-resolution keyframes, including: Compare the propagation range of high-frequency information provided by different high-resolution keyframes and determine the overlapping propagation range; Based on the overlapping propagation ranges, the length of the time window is determined.

5. The video transmission system based on the adaptive key frame selection strategy for video super-resolution at low bit rate according to any one of claims 1 to 4, characterized in that: When multiple high-resolution key frames are determined, calculating an average interval between the multiple high-resolution key frames; When the average interval does not satisfy the sparse parameter value, the length of the time window and the selection range of the high-resolution key frames are re-determined, and multiple high-resolution key frames are re-determined until the average interval of the multiple high-resolution key frames determined satisfies the sparse parameter value.

6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the video transmission system based on the adaptive key frame selection strategy for video super-resolution at low bit rate according to any one of claims 1 to 5 is implemented.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the video transmission system based on the adaptive key frame selection strategy of video super-resolution at low bit rate according to any one of claims 1 to 5 is realized.

8. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the video transmission system based on the adaptive key frame selection strategy of video super-resolution at low bit rate as described in any one of claims 1-5 is implemented.

Citation Information

Patent Citations

  • End-cloud combined super-resolution video reconstruction method and system based on key frame

    CN116523758A