Intelligent frame interpolation and fluency improvement system and method for video content

Through the intelligent frame interpolation and fluency improvement system, deep learning and generative adversarial network technology are used to solve the problems of low fluency of video frame interpolation, inaccurate feature extraction, motion estimation deviation and insufficient optimization, and high-quality and fluency of high-definition video are achieved.

CN119967112AInactive Publication Date: 2025-05-09GUANGZHOU WEIBO NETWORK INFORMATION TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202411870763.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-18
Publication Date
2025-05-09
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art has low fluency in video frame interpolation, inaccurate feature extraction, motion estimation deviation and insufficient optimization, making it difficult to meet the high standard requirements for video quality in the high-definition video era.

Method used

A smart frame interpolation and fluency improvement system for video content is designed, and a pre-processing module is used to perform denoising, color correction and brightness equalization. A deep convolutional neural network is used to extract multi-dimensional feature vectors, combined with optical flow algorithms and space-time attention mechanisms for motion estimation, and high-quality interpolated frames are generated by generating adversarial networks.

Benefits of technology

It significantly improves the fluency and nature of interpolated frames, improves the accuracy of feature extraction and the accuracy of motion estimation, realizes continuous optimization of video quality, and meets the high standards requirements of high-definition video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119967112A_ABST
    Figure CN119967112A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent frame interpolation and fluency improvement system and method for video content, and relates to the technical field of video processing, the system comprises a preprocessing module, a frame extraction module, a frame analysis module, a frame interpolation module, a fusion module, a post-processing module, a fluency optimization module and the like, and the preprocessing module carries out denoising, color correction and brightness equalization on a video; the frame extraction module extracts frame sequence features by using a deep convolutional neural network; the frame analysis module calculates inter-frame similarity and a motion vector; the frame interpolation module generates an interpolation frame based on deep learning; the fusion module fuses the interpolation frame and the original frame; the post-processing module performs edge smoothing and frame rate adjustment; according to the system, the video fluency can be effectively improved, the sawtooth phenomenon is reduced, the frame rate is dynamically adjusted according to the equipment performance and the user requirement, and the high-quality video playing experience is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video processing technology, and more specifically, to a system and method for intelligent frame interpolation and fluency improvement of video content. Background Art

[0002] In today's digital multimedia era, the quality and fluency of video content are crucial to user experience. With the continuous expansion of video application scenarios, such as high-definition video playback, video conferencing, virtual reality, etc., people have put forward higher requirements for the visual effects of videos. However, existing technologies have many shortcomings in video processing and are difficult to meet these requirements.

[0003] Traditional video frame interpolation methods often use simple linear interpolation or fixed template-based algorithms, which results in poor accuracy and authenticity of interpolated frames when processing complex scenes and fast-moving objects, and is prone to problems such as blurred images and deformation of moving objects, seriously affecting the smoothness of the video and viewing experience. Moreover, in the feature extraction stage, previous technologies usually use a single-scale feature extraction method, which cannot fully capture the rich semantics, texture and edge information in the video frame, making the subsequent frame interpolation and processing lack sufficiently accurate image feature basis. In addition, in terms of motion estimation, most methods only rely on basic optical flow algorithms, which are difficult to cope with complex motion patterns and scene changes. The motion vector calculation of the details and key areas of moving objects is not accurate enough, resulting in an unnatural connection between the interpolated frame and the original frame. At the same time, there is a lack of systematicity and efficiency in video quality assessment and parameter optimization, and it is impossible to dynamically adjust the processing parameters according to the real-time changes in the video content, making it difficult to achieve continuous performance improvement.

[0004] Therefore, the existing technology has the problems of low frame interpolation smoothness, inaccurate feature extraction, motion estimation deviation and insufficient optimization. Summary of the invention

[0005] In order to overcome the problems of low frame interpolation fluency, inaccurate feature extraction, motion estimation deviation and insufficient optimization in the prior art, the present invention designs an intelligent frame interpolation and fluency improvement system and method for video content, which can effectively solve the above technical problems.

[0006] In order to solve the above technical problems, the technical solution of the present invention is as follows:

[0007] A system for intelligent frame interpolation and fluency improvement of video content, comprising:

[0008] A preprocessing module is used to obtain the original video sequence and perform denoising, color correction and brightness equalization preprocessing operations on it;

[0009] A frame extraction module receives the preprocessed video sequence and decomposes it into a frame sequence, extracts the semantic features, texture features and edge features of the video image using a deep convolutional neural network for each frame sequence, and combines these features into a multi-dimensional feature vector;

[0010] A frame analysis module, which calculates the similarity between adjacent frame sequences using cosine similarity or a metric method based on feature distance according to the multidimensional feature vector, and uses an optical flow algorithm to preliminarily estimate the motion vector between the frame sequences and analyze the motion information between the frame sequences, and determines the scene change points, motion intensity and key frame positions in the video image based on the similarity and motion information, and generates a video frame sequence analysis report;

[0011] A frame interpolation module, comprising a feature extraction layer, a motion estimation layer and an image generation layer, wherein the preprocessed video sequence is input into a frame interpolation model based on deep learning according to the video frame sequence analysis report generated by the frame analysis module to generate an interpolated frame, wherein the feature extraction layer uses a multi-scale convolutional neural network to extract a feature map of the frame sequence, the motion estimation layer uses an optical flow algorithm combined with a spatiotemporal attention mechanism to calculate a motion vector between adjacent frame sequences, and the image generation layer generates the interpolated frame based on the feature map and the motion vector through a generative adversarial network;

[0012] A fusion module, used for fusing the generated interpolation frames with the original video sequence frames in time order to form a new video sequence;

[0013] A post-processing module, used for performing edge smoothing and frame rate adjustment processing operations on the new video sequence to improve the fluency of the video;

[0014] The fluency optimization module is used to evaluate the quality of the new video sequence and optimize the parameters of the frame interpolation module and the post-processing module according to the evaluation result.

[0015] Preferably, the multi-scale convolutional neural network includes a dilated convolution unit and a grouped convolution unit;

[0016] The atrous convolution unit captures the frame sequence feature map at different scales by applying an atrous convolution kernel with a atrous rate to skip a certain number of pixels and adjust the step size, and obtains a multi-scale global representation of the frame sequence feature map by combining the outputs of convolution layers with different atrous rates;

[0017] The grouped convolution unit groups the input frame sequence feature maps into groups, each group containing a specific number of channels, performs a convolution operation on each group to extract local features, and recombines the processed frame sequence feature maps between different groups through cross-group connections or feature map fusion to obtain richer feature representation.

[0018] Preferably, the motion estimation layer includes an optical flow calculation unit and a spatiotemporal attention unit;

[0019] The optical flow calculation unit preliminarily calculates the pixel motion vectors between adjacent frame sequences using an optical flow algorithm;

[0020] The spatiotemporal attention unit is used to perform weighted processing on the pixel motion vector through a spatiotemporal attention mechanism, focus on the area of ​​the moving object in the video sequence, and improve the accuracy of the motion vector;

[0021] The spatiotemporal attention unit includes a time attention subunit and a space attention subunit. The time attention subunit assigns weights according to the time sequence of the frame sequence, and the space attention subunit assigns weights according to pixel positions and feature values.

[0022] Preferably, the generative adversarial network in the image generation layer includes a generator unit and a discriminator unit;

[0023] The generator unit uses a residual structured U-Net network to gradually restore detail information of the interpolation frame through upsampling and downsampling operations;

[0024] The discriminator unit uses a multi-layer convolutional neural network to extract features from the input interpolated frames and real video sequence frames, and outputs the judgment results through a fully connected layer. According to the feedback information of the discriminator unit, the generator unit continuously adjusts the parameters for generating the interpolated frames to improve the quality of the generated interpolated frames.

[0025] Preferably, the fusion module adopts a weighted averaging method to fuse the interpolated frames and the original video sequence frames, and adaptively adjusts the weights of the interpolated frames and the original video sequence frames according to the intensity of motion of the video content. For areas with slow motion, the weight of the original video sequence frames is increased, and for areas with intense motion, the weight of the interpolated frames is increased.

[0026] Preferably, the post-processing module includes an edge smoothing unit and a frame rate adjustment unit;

[0027] The edge smoothing unit adopts a bilateral filtering algorithm to smooth the pixel values ​​of the video sequence frames while maintaining the edge information of the new video sequence, so as to reduce the edge jagged phenomenon generated during the interpolation process;

[0028] The frame rate adjustment unit is used to perform dynamic adjustment according to the performance of the video playback device and the viewing needs of the user, and is achieved by copying or deleting part of the video sequence frames.

[0029] Preferably, the fluency optimization module includes a quality assessment unit and a parameter optimization unit;

[0030] The quality assessment unit is used to quantitatively assess the video quality using a peak signal-to-noise ratio, a structural similarity index, and a video multi-method assessment fusion index;

[0031] The parameter optimization unit is used to feed back the evaluation results to the frame interpolation model and the post-processing module, and adjust the frame interpolation model parameters and the post-processing parameters through the back propagation algorithm to improve the intelligent frame interpolation and fluency improvement effects of the video content.

[0032] A method for intelligent frame interpolation and fluency improvement of video content, comprising the following steps:

[0033] Acquire an original video sequence, and preprocess the video sequence, including denoising, color correction, and brightness equalization;

[0034] Perform frame decomposition on the preprocessed video sequence to split it into multiple frame sequences, perform feature extraction on each frame sequence, use a deep convolutional neural network to extract semantic features, texture features and edge features of the video image, and combine these features into a multi-dimensional feature vector;

[0035] According to the multidimensional feature vector, the similarity between adjacent frame sequences is calculated by using cosine similarity or a measurement method based on feature distance, and the motion vector between frame sequences is preliminarily estimated by using an optical flow algorithm, and the motion information between the frame sequences is analyzed, and the scene change points, motion intensity and key frame positions in the video image are determined based on the calculated similarity and the analyzed motion information, and a video frame sequence analysis report is generated;

[0036] According to the generated video frame sequence analysis report, the preprocessed video sequence is input into a frame interpolation model based on deep learning, wherein the frame interpolation model includes a feature extraction layer, a motion estimation layer and an image generation layer, wherein the feature extraction layer uses a multi-scale convolutional neural network to extract the feature map of the frame sequence, the motion estimation layer uses an optical flow algorithm combined with a spatiotemporal attention mechanism to calculate the motion vector between adjacent frame sequences, and the image generation layer generates an interpolated frame based on the feature map and the motion vector through a generative adversarial network;

[0037] The generated interpolated frames are merged with the original video sequence frames in time order to form a new video sequence;

[0038] Post-processing the new video sequence, including edge smoothing and frame rate adjustment to improve the smoothness of the video;

[0039] The quality of the new video sequence after post-processing is evaluated, and the model parameters of the frame interpolation process and the parameters of the post-processing process are optimized and adjusted according to the evaluation results.

[0040] Preferably, an electronic device comprises:

[0041] A memory storing executable program code;

[0042] a processor coupled to the memory;

[0043] The processor calls the executable program code stored in the memory to execute the intelligent frame interpolation and fluency improvement method of video content as described above.

[0044] Preferably, a computer storage medium stores computer instructions, and when the computer instructions are called, they are used to execute the intelligent frame interpolation and fluency improvement method of the video content as described above.

[0045] Compared with the prior art, the present invention has the following beneficial effects: the present invention performs denoising, color correction and brightness balance on the original video sequence through the preprocessing module, thereby ensuring the quality of the video content before frame interpolation; the frame extraction module uses a deep convolutional neural network to extract the semantic, texture and edge features of the video image and generate a multi-dimensional feature vector, which improves the accuracy of feature extraction; the frame analysis module uses cosine similarity and optical flow algorithms to analyze the similarity and motion information between frames, accurately judges the scene change points and key frame positions, and provides accurate motion vectors and scene change information for frame interpolation; the frame interpolation module uses a multi-scale convolutional neural network and a generative adversarial network to generate high-quality interpolated frames, which significantly improves the fluency and naturalness of the interpolated frames; the fusion module uses a multi-scale convolutional neural network and a generative adversarial network to generate high-quality interpolated frames, thereby significantly improving the fluency and naturalness of the interpolated frames; the fusion module uses a deep convolutional neural network to extract the semantic, texture and edge features of the video image and generate a multi-dimensional feature vector, which improves the accuracy of feature extraction; the frame analysis module uses a cosine similarity and optical flow algorithm to analyze the similarity and motion information between frames and accurately judges the scene change points and key frame positions, thereby providing accurate motion vectors and scene change information for frame interpolation; the frame interpolation module uses a multi-scale convolutional neural network and a generative adversarial network to generate high-quality interpolated frames, thereby significantly improving the fluency and naturalness of the interpolated frames; the fusion module uses a multi-scale convolutional neural network and a generative adversarial network to generate high-quality ... The weights of the interpolated frames and the original frames are adaptively adjusted to further improve the video fluency. The post-processing module further optimizes the video quality through edge smoothing and frame rate adjustment, reduces jagged phenomena, and improves the viewing experience. The fluency optimization module realizes dynamic adjustment of the frame interpolation model and post-processing parameters through quality evaluation and parameter optimization, ensuring the continuous optimization of video fluency. In summary, this scheme comprehensively improves the accuracy, fluency and naturalness of video frame interpolation through advanced technologies such as deep learning, multi-scale feature extraction, spatiotemporal attention mechanism and generative adversarial network. At the same time, it realizes the continuous improvement of video fluency through quality evaluation and parameter optimization, which meets the high standards for video quality in the era of high-definition video and has significant technical innovation and practical value. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings in the following description are only exemplary, and for ordinary technicians in this field, other implementation drawings can be derived based on the provided drawings without paying any creative work.

[0047] Figure 1 A structure diagram of a system for intelligent frame interpolation and fluency improvement of video content;

[0048] Figure 2 A step-by-step diagram of a method for intelligent frame interpolation and smoothness improvement of video content. DETAILED DESCRIPTION

[0049] The drawings are for illustrative purposes only and should not be construed as limiting the present patent;

[0050] In order to better illustrate the present embodiment, some parts in the drawings may be omitted, enlarged or reduced, and do not represent the size of the actual product;

[0051] It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.

[0052] The technical solution of the present invention is further described below in conjunction with the accompanying drawings and embodiments.

[0053] Example 1

[0054] Choose a workstation with a high-performance GPU (such as NVIDIA GeForce RTX 3090), install Python 3.8 and TensorFlow 2.5 deep learning framework, and ensure that there is sufficient memory and storage space to process video data.

[0055] An intelligent frame interpolation and fluency improvement system for video content, such as Figure 1 As shown, including:

[0056] A preprocessing module is used to obtain the original video sequence and perform denoising, color correction and brightness equalization preprocessing operations on it;

[0057] An original video sequence file in AVI format is selected from a local storage device. The video sequence contains traffic scenes of city streets, including vehicles driving, pedestrians walking, and light changes in different time periods.

[0058] Use Python's cv2 library (OpenCV) to read the video file and convert it into a sequence of video frames.

[0059] Gaussian filtering is used to denoise the video frames. The Gaussian kernel size is set to 5x5 and the standard deviation is 1.5. Each pixel of each frame is traversed and the pixel neighborhood is weighted averaged according to the Gaussian distribution to remove the Gaussian noise in the image and make the image smoother.

[0060] Using automatic color balancing algorithms, such as CLAHE-contrast limited adaptive histogram equalization, the image is divided into multiple small blocks, such as 8x8 blocks, and histogram equalization is performed in each small block. At the same time, excessive contrast enhancement is limited to avoid overly bright or dark areas in the image, correct color deviation caused by shooting equipment or lighting conditions, and make the color of the video picture more natural and vivid.

[0061] Calculate the brightness histogram of the entire video frame sequence, and stretch the brightness distribution to a more appropriate range by adjusting the brightness gain of the image. For example, for videos with low brightness, increase the brightness gain appropriately to make the dark details more clearly visible; for videos with high brightness, reduce the brightness gain to prevent the image from being overexposed, so that the brightness of the entire video is more uniform.

[0062] A frame extraction module receives the preprocessed video sequence and decomposes it into a frame sequence, extracts the semantic features, texture features and edge features of the video image using a deep convolutional neural network (CNN) for each frame sequence, and combines these features into a multi-dimensional feature vector;

[0063] The preprocessed video frame sequence is extracted frame by frame to build a frame list; a deep convolutional neural network model is constructed, which is modified based on the ResNet architecture. The network contains multiple convolutional layers, residual blocks and fully connected layers. First, the low-level features of the image, such as edge and texture information, are extracted through the convolutional layer; then, the deep semantic features are further extracted through the residual block; finally, the features at different levels are mapped to a feature vector space of fixed dimension through the fully connected layer, and the semantic features, texture features and edge feature vectors of each frame sequence are obtained respectively; these feature vectors are spliced ​​and combined in dimension to form a multi-dimensional feature vector, which is used to characterize the comprehensive feature information of each video frame.

[0064] A frame analysis module, which calculates the similarity between adjacent frame sequences using cosine similarity or a metric method based on feature distance according to the multidimensional feature vector, and uses an optical flow algorithm to preliminarily estimate the motion vector between the frame sequences and analyze the motion information between the frame sequences, and determines the scene change points, motion intensity and key frame positions in the video image based on the similarity and motion information, and generates a video frame sequence analysis report;

[0065] The cosine similarity is used to calculate the similarity between the multidimensional feature vectors of adjacent frame sequences. For two adjacent frame feature vectors A and B, the formula for calculating their cosine similarity is: where A·B is the dot product of the vectors, | A | and | B |are the moduli of vectors A and B respectively. The similarity value obtained by calculation is used to determine the similarity between the two frames.

[0066] The Farneback optical flow algorithm is used to estimate the motion vector between frame sequences. Two adjacent frames of images are used as input to calculate the displacement of each pixel on the image plane to obtain a motion vector field, which represents the direction and speed of the pixel's movement. The intensity of the movement in the video image is analyzed based on the size and direction of the motion vector. For example, if the amplitude of the motion vector is large and the direction is more dispersed, it indicates that the area is moving violently; conversely, if the amplitude of the motion vector is small and the direction is more consistent, the movement is relatively slow.

[0067] The scene change points in the video image are determined by combining the similarity value and motion information. When the similarity between adjacent frames is lower than the set threshold, such as 0.3, and the motion vector changes significantly, it is considered that a scene switch may have occurred, and the frame is marked as a scene change point. At the same time, key frames are selected at scene change points, points where the intensity of motion changes significantly, or at certain time intervals, such as every 5 seconds, to generate a video frame sequence analysis report, which records in detail the similarity, motion information, scene change points, and key frame position information of each frame.

[0068] A frame interpolation module, comprising a feature extraction layer, a motion estimation layer and an image generation layer, wherein the preprocessed video sequence is input into a frame interpolation model based on deep learning according to the video frame sequence analysis report generated by the frame analysis module to generate an interpolated frame, wherein the feature extraction layer uses a multi-scale convolutional neural network to extract a feature map of the frame sequence, the motion estimation layer uses an optical flow algorithm combined with a spatiotemporal attention mechanism to calculate a motion vector between adjacent frame sequences, and the image generation layer generates the interpolated frame based on the feature map and the motion vector through a generative adversarial network;

[0069] Feature extraction layer, dilated convolution unit: construct a multi-scale convolutional neural network, set the dilated convolution kernel's dilation rate to 1, 3, and 5 respectively. For the convolution kernel with a dilation rate of 1, perform convolution operation on the frame sequence feature map in the conventional way to extract local features; for the convolution kernel with a dilation rate of 3, perform convolution every 2 pixels to capture feature information in a wider range; for the convolution kernel with a dilation rate of 5, perform convolution every 4 pixels to obtain a more macro feature representation, and fuse the outputs of convolution layers with different dilation rates to obtain a multi-scale global frame sequence feature map representation, so that it can cover image features of different scales and enhance the feature expression ability.

[0070] Grouped convolution unit: The input frame sequence feature map is evenly divided into 4 groups according to the number of channels. Each group contains a specific number of channels, for example, 16 channels per group. Convolution operation is performed on each group separately, using a 3x3 convolution kernel, a step size of 1, and a padding of 1 to extract local features. Then, the features of different groups are spliced ​​through cross-group connections to form a new feature map, so that the features of different groups can complement each other and obtain richer feature representation for subsequent motion estimation and image generation.

[0071] Motion estimation layer, optical flow calculation unit: The Lucas-Kanade optical flow algorithm is used to preliminarily calculate the pixel motion vector between adjacent frame sequences. Two adjacent frame images are used as input. The motion vector of each pixel is solved by minimizing the brightness difference of the pixel points between the two frames. A motion vector field is obtained to represent the displacement of the pixel on the image plane.

[0072] Spatiotemporal attention unit: The temporal attention subunit assigns weights according to the time sequence of the frame sequence. For example, for the current frame, a higher weight, such as 0.4, is given to the adjacent previous and next frames, while a lower weight, such as 0.1, is given to the frames farther away from the current frame, so that motion estimation pays more attention to the frame features that are close to the current frame in time and have strong motion coherence. The spatial attention subunit assigns weights according to pixel position and eigenvalue. A higher weight, such as 0.6, is given to areas with obvious features such as edges and textures in the image (judged by eigenvalues) and areas where moving objects are located (judged by motion vector amplitude), while a lower weight, such as 0.2, is given to relatively stable areas such as the background. In this way, the pixel motion vector is weighted through the spatiotemporal attention mechanism to improve the accuracy of the motion vector and make it more focused on the area of ​​moving objects in the video sequence.

[0073] Image generation layer, generator unit: A U-Net network with a residual structure is used to input the feature map extracted previously and the motion vector information processed by the motion estimation layer. In the downsampling part of the network, the resolution of the feature map is gradually reduced through the convolution layer and the pooling layer to extract high-level semantic features; in the upsampling part, the transposed convolution layer is used to gradually restore the resolution of the image, and the low-level features extracted in the downsampling process are fused with the high-level features in the upsampling process through the jump connection to supplement the detailed information of the image and finally generate an interpolated frame.

[0074] Discriminator unit: A 5-layer convolutional neural network is used to extract features from the input interpolated frames and the real video sequence frames obtained from the original video. The sizes of the convolution kernels are 7x7, 5x5, 3x3, 3x3, and 3x3, respectively. The number of channels in each layer gradually increases, such as 32, 64, 128, 256, and 512, with a step size of 2 and a padding of 1. Finally, the judgment result is output through the fully connected layer to determine whether the input frame is a real frame or a generated interpolated frame. According to the feedback information of the discriminator unit, the generator unit continuously adjusts the parameters of the generated interpolated frames through the back propagation algorithm, such as the weight and bias of the convolution kernel, so that the generated interpolated frames are closer to the real frames in terms of visual effects and feature performance, thereby improving the quality of the generated interpolated frames.

[0075] A fusion module, used for fusing the generated interpolation frames with the original video sequence frames in time order to form a new video sequence;

[0076] The generated interpolated frames are fused with the original video sequence frames in chronological order using the weighted averaging method to form a new video sequence. According to the video content motion intensity information obtained by the frame analysis module, the weight of the original video sequence frames is increased for slow-moving areas, such as background areas or areas where pedestrians walk slowly. The weight of the original frames is set to 0.7 and the weight of the interpolated frames is set to 0.3. For areas with intense motion, such as areas where vehicles are driving fast or people are moving quickly, the weight of the interpolated frames is increased. The weight of the original frames is set to 0.4 and the weight of the interpolated frames is set to 0.6. Through this adaptive weight adjustment, the fused video can maintain good visual effects in different motion states, which not only retains the authenticity of the original video, but also uses the interpolated frames to improve the smoothness of the motion area.

[0077] A post-processing module, used for performing edge smoothing and frame rate adjustment processing operations on the new video sequence to improve the fluency of the video;

[0078] Edge smoothing unit: Use bilateral filtering algorithm to perform edge smoothing on the new video sequence. For each frame of image, set the spatial domain standard deviation to 3 and the value domain standard deviation to 0.1. Bilateral filtering not only considers the proximity of pixel spatial positions, but also considers the similarity of pixel values. It can smooth the pixel values ​​of video sequence frames while maintaining the edge information of the video sequence. It determines the edge by detecting the places where the pixel gradient changes greatly in the image, reduces the edge jagged phenomenon generated during the interpolation process, makes the image edge more natural and smooth, and improves the visual quality of the video.

[0079] Frame rate adjustment unit: Dynamically adjust according to the performance of the video playback device, such as when it is detected that the playback device is a low-configuration mobile device and the user's viewing needs, such as when the user hopes to have a smoother viewing experience when watching sports event videos. If the playback device performance is low, the frame rate can be appropriately reduced to ensure smooth playback, which can be achieved by deleting some video sequence frames. For example, the original frame rate of 30 frames / second is reduced to 20 frames / second, and one frame is deleted every certain number of frames, such as every 3 frames. If the user wants a smoother viewing experience, the frame rate can be appropriately increased by copying some frames, for example, increasing the frame rate from 30 frames / second to 60 frames / second, and inserting copied frames between adjacent frames to make the video playback smoother and meet the needs of different users and devices.

[0080] The fluency optimization module is used to evaluate the quality of the new video sequence and optimize the parameters of the frame interpolation module and the post-processing module according to the evaluation result.

[0081] Quality assessment unit: Peak signal-to-noise ratio (PSNR), structural similarity index (SSIM) and video multi-method assessment fusion index (VMAF) are used to quantitatively evaluate video quality. For the entire processed video sequence, the values ​​of these indicators are calculated frame by frame. PSNR measures image quality by comparing the processed frame with the original frame. If there is a difference in pixel values ​​of the original reference frame, the higher the value, the better the quality; SSIM comprehensively evaluates the similarity of two frames of images from multiple dimensions such as brightness, contrast, and structure. The value range is between -1 and 1. The closer to 1, the more similar the two frames are, and the better the quality; VMAF evaluates video quality from a perspective that is more in line with human visual perception, and comprehensively considers multiple aspects of the video, such as clarity, color accuracy, motion smoothness, etc., to obtain a more comprehensive video quality assessment result.

[0082] Parameter optimization unit: Feedback the evaluation results obtained by the quality evaluation unit to the frame interpolation model and post-processing module. For the frame interpolation model, adjust the parameters in the model through the back propagation algorithm, such as the weights, biases, the hole rate of the hole convolution kernel, the number of groups in the group convolution, etc. in the neural network to improve the generation quality of the interpolated frame. For the post-processing module, adjust the spatial domain and range standard deviation of the bilateral filter, the strategy parameters of the frame rate adjustment, such as the number of interval frames for deleting or copying frames, etc., to continuously optimize the performance of the system, so that the next time the video is processed, a higher quality and better fluency result can be obtained, thereby realizing the adaptive optimization of the intelligent frame interpolation and fluency improvement system of the entire video content.

[0083] The multi-scale convolutional neural network includes a hole convolution unit and a grouped convolution unit;

[0084] The atrous convolution unit captures the frame sequence feature map at different scales by applying an atrous convolution kernel with a atrous rate to skip a certain number of pixels and adjust the step size, and obtains a multi-scale global representation of the frame sequence feature map by combining the outputs of convolution layers with different atrous rates;

[0085] The grouped convolution unit groups the input frame sequence feature maps into groups, each group containing a specific number of channels, performs a convolution operation on each group to extract local features, and recombines the processed frame sequence feature maps between different groups through cross-group connections or feature map fusion to obtain richer feature representation.

[0086] The motion estimation layer includes an optical flow calculation unit and a spatiotemporal attention unit;

[0087] The optical flow calculation unit preliminarily calculates the pixel motion vectors between adjacent frame sequences using an optical flow algorithm;

[0088] The spatiotemporal attention unit is used to perform weighted processing on the pixel motion vector through a spatiotemporal attention mechanism, focus on the area of ​​the moving object in the video sequence, and improve the accuracy of the motion vector;

[0089] The spatiotemporal attention unit includes a time attention subunit and a space attention subunit. The time attention subunit assigns weights according to the time sequence of the frame sequence, and the space attention subunit assigns weights according to pixel positions and feature values.

[0090] The generative adversarial network in the image generation layer includes a generator unit and a discriminator unit;

[0091] The generator unit uses a residual structured U-Net network to gradually restore detail information of the interpolation frame through upsampling and downsampling operations;

[0092] The discriminator unit uses a multi-layer convolutional neural network to extract features from the input interpolated frames and real video sequence frames, and outputs the judgment results through a fully connected layer. According to the feedback information of the discriminator unit, the generator unit continuously adjusts the parameters for generating the interpolated frames to improve the quality of the generated interpolated frames.

[0093] The fusion module fuses the interpolated frames and the original video sequence frames using a weighted average method, and adaptively adjusts the weights of the interpolated frames and the original video sequence frames according to the intensity of motion of the video content. For areas with slow motion, the weight of the original video sequence frames is increased, and for areas with intense motion, the weight of the interpolated frames is increased.

[0094] The post-processing module includes an edge smoothing unit and a frame rate adjustment unit;

[0095] The edge smoothing unit adopts a bilateral filtering algorithm to smooth the pixel values ​​of the video sequence frames while maintaining the edge information of the new video sequence, so as to reduce the edge jagged phenomenon generated during the interpolation process;

[0096] The frame rate adjustment unit is used to dynamically adjust according to the performance of the video playback device and the viewing needs of the user, by copying or deleting some video sequence frames. (When the performance of the video playback device is low or the user prefers a lower frame rate, some interpolated frames are appropriately deleted to reduce the frame rate. When the performance of the video playback device is high and the user needs a high frame rate viewing experience, some key frames or interpolated frames are copied to increase the frame rate.

[0097] The fluency optimization module includes a quality assessment unit and a parameter optimization unit;

[0098] The quality assessment unit is used to quantitatively assess the video quality using peak signal-to-noise ratio (PSNR), structural similarity index (SSIM) and video multi-method assessment fusion (VMAF) indicators;

[0099] The parameter optimization unit is used to feed back the evaluation results to the frame interpolation model and the post-processing module, and adjust the frame interpolation model parameters and the post-processing parameters through the back propagation algorithm to improve the intelligent frame interpolation and fluency improvement effects of the video content.

[0100] Example 2

[0101] A method for intelligent frame interpolation and fluency improvement of video content, such as Figure 2 As shown, the following steps are included:

[0102] Acquire an original video sequence, and preprocess the video sequence, including denoising, color correction, and brightness equalization;

[0103] Perform frame decomposition on the preprocessed video sequence to split it into multiple frame sequences, perform feature extraction on each frame sequence, use a deep convolutional neural network to extract semantic features, texture features and edge features of the video image, and combine these features into a multi-dimensional feature vector;

[0104] According to the multidimensional feature vector, the similarity between adjacent frame sequences is calculated by using cosine similarity or a measurement method based on feature distance, and the motion vector between frame sequences is preliminarily estimated by using an optical flow algorithm, and the motion information between the frame sequences is analyzed, and the scene change points, motion intensity and key frame positions in the video image are determined based on the calculated similarity and the analyzed motion information, and a video frame sequence analysis report is generated;

[0105] According to the generated video frame sequence analysis report, the preprocessed video sequence is input into a frame interpolation model based on deep learning, wherein the frame interpolation model includes a feature extraction layer, a motion estimation layer and an image generation layer, wherein the feature extraction layer uses a multi-scale convolutional neural network to extract the feature map of the frame sequence, the motion estimation layer uses an optical flow algorithm combined with a spatiotemporal attention mechanism to calculate the motion vector between adjacent frame sequences, and the image generation layer generates an interpolated frame based on the feature map and the motion vector through a generative adversarial network;

[0106] The generated interpolated frames are merged with the original video sequence frames in time order to form a new video sequence;

[0107] Post-processing the new video sequence, including edge smoothing and frame rate adjustment to improve the smoothness of the video;

[0108] The quality of the new video sequence after post-processing is evaluated, and the model parameters of the frame interpolation process and the parameters of the post-processing process are optimized and adjusted according to the evaluation results.

[0109] An electronic device, comprising:

[0110] A memory storing executable program code;

[0111] a processor coupled to the memory;

[0112] The processor calls the executable program code stored in the memory to execute the intelligent frame interpolation and fluency improvement method of video content as described above.

[0113] A computer storage medium stores computer instructions, which, when called, are used to execute the above-mentioned method for intelligent frame interpolation and fluency improvement of video content.

[0114] In a specific implementation, a 10-minute character dance video with a resolution of 1280x720 is selected from a local video folder as an original video sequence.

[0115] The median filter algorithm is used for denoising, which effectively removes the salt and pepper noise in the video and makes the picture clearer.

[0116] The color correction method based on white point balance is used to make the character's skin color more natural and realistic, and the overall color more harmonious.

[0117] The brightness equalization algorithm based on local contrast enhancement is adopted to highlight the details of the characters and background in the video and improve the visual effect.

[0118] The preprocessed video sequence is decomposed into groups of 3 frames to obtain a total of multiple frame sequences.

[0119] For each frame sequence, a deep convolutional neural network with 10 convolutional layers is used for feature extraction. The network has been trained with a large number of video images containing human actions. It can accurately extract the semantic features of human body movements, clothing texture features, and edge features between humans and backgrounds, and combine these features into a 128-dimensional feature vector.

[0120] The Euclidean distance measurement method based on feature distance is used to calculate the similarity between adjacent frame sequences. A similarity threshold is set to 0.7. When the feature distance of adjacent frame sequences is greater than the threshold, it is determined that there may be a scene change or a large action change.

[0121] The optical flow algorithm is combined with the Lucas-Kanade method to preliminarily estimate the motion vector between frame sequences. The motion trajectory and speed change of the character are determined by statistically analyzing the size and direction of the motion vector.

[0122] Based on the above calculation and analysis results, the scene change points in the video, such as stage scene switching, dance movement segment transitions, etc., the intensity of movement, such as rapid rotation, rapid jumping movements, and the key frame position are accurately determined. The start, climax and end frames of the dance movements are selected as key frames to generate a detailed video frame sequence analysis report.

[0123] The preprocessed video sequence is input into the deep learning based frame interpolation model.

[0124] In the feature extraction layer, a multi-scale convolutional neural network with three convolution kernels of different scales is used to extract frame sequence feature maps, which can simultaneously capture the macro and micro features of the video image and provide richer information for subsequent interpolation.

[0125] The motion estimation layer uses the optical flow algorithm combined with the spatiotemporal attention mechanism to calculate the motion vector between adjacent frame sequences. The spatiotemporal attention mechanism focuses on the moving areas of the characters, making the calculation of motion vectors more accurate, especially when the characters move quickly and perform complex movements.

[0126] The image generation layer generates interpolated frames based on feature maps and motion vectors through a generative adversarial network. After a lot of training, the generator in the generative adversarial network can generate high-quality interpolated frames that naturally transition with the original frames. The discriminator continuously optimizes its ability to judge the authenticity of the interpolated frames, thereby driving the generator to continuously improve the interpolation effect.

[0127] The generated interpolated frames are evenly inserted into the original video sequence frames in chronological order, so that the frame rate of the video is increased from the original 25fps to 50fps, forming a new video sequence. This effectively improves the smoothness of video playback and makes the character's dance movements more coherent and natural.

[0128] The guided filtering algorithm is used to perform edge smoothing on the new video sequence, which removes possible interpolation artifacts while keeping the edges of the characters clear, making the video more delicate.

[0129] The frame rate of the video is further optimized and adjusted according to the target playback platform and user needs of the video. For example, if it is used for playback on an online video platform, considering the differences in network bandwidth and user device performance, the frame rate is stabilized at 30fps, and adaptive bit rate control technology is used to ensure smooth playback and high-quality presentation of the video.

[0130] A combination of subjective and objective methods is used to evaluate the quality of the new post-processed video sequences. Subjectively, professional video editors and ordinary users are invited to score and evaluate the visual effects and smoothness of the videos. Objectively, indicators such as mean square error (MSE) and information entropy are used for quantitative evaluation.

[0131] According to the evaluation results, if the video is found to be blurry or jittery in certain action scenes, the convolution kernel parameters of the deep convolutional neural network in the frame interpolation process are adjusted to increase the ability to extract detail features. At the same time, the guided filter parameters in the post-processing process are optimized to enhance the edge smoothing effect. Through multiple iterative optimizations, the quality and smoothness of the video are continuously improved, ultimately meeting the user's high-quality requirements in video editing and playback.

[0132] The same or similar reference numerals correspond to the same or similar components;

[0133] The terms used in the drawings to describe positional relationships are only used for illustrative purposes and should not be construed as limiting this patent;

[0134] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not limitations on the implementation methods of the present invention. For ordinary technicians in the relevant field, other different forms of changes or modifications can be made on the basis of the above description. It is not necessary and impossible to list all the implementation methods here. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the claims of the present invention.

Claims

1. A system for intelligent frame interpolation and fluency improvement of video content, characterized in that: include: A preprocessing module is used to obtain the original video sequence and perform denoising, color correction and brightness equalization preprocessing operations on it; A frame extraction module receives the preprocessed video sequence and decomposes it into a frame sequence, extracts the semantic features, texture features and edge features of the video image using a deep convolutional neural network for each frame sequence, and combines these features into a multi-dimensional feature vector; A frame analysis module, which calculates the similarity between adjacent frame sequences using cosine similarity or a metric method based on feature distance according to the multidimensional feature vector, and uses an optical flow algorithm to preliminarily estimate the motion vector between the frame sequences and analyze the motion information between the frame sequences, and determines the scene change points, motion intensity and key frame positions in the video image based on the similarity and motion information, and generates a video frame sequence analysis report; A frame interpolation module, comprising a feature extraction layer, a motion estimation layer and an image generation layer, wherein the preprocessed video sequence is input into a frame interpolation model based on deep learning according to the video frame sequence analysis report generated by the frame analysis module to generate an interpolated frame, wherein the feature extraction layer uses a multi-scale convolutional neural network to extract a feature map of the frame sequence, the motion estimation layer uses an optical flow algorithm combined with a spatiotemporal attention mechanism to calculate a motion vector between adjacent frame sequences, and the image generation layer generates the interpolated frame based on the feature map and the motion vector through a generative adversarial network; A fusion module, used for fusing the generated interpolation frames with the original video sequence frames in time order to form a new video sequence; A post-processing module, used for performing edge smoothing and frame rate adjustment processing operations on the new video sequence to improve the fluency of the video; The fluency optimization module is used to evaluate the quality of the new video sequence and optimize the parameters of the frame interpolation module and the post-processing module according to the evaluation result.

2. The intelligent frame interpolation and fluency improvement system for video content according to claim 1, characterized in that: The multi-scale convolutional neural network includes a hole convolution unit and a grouped convolution unit; The atrous convolution unit captures the frame sequence feature map at different scales by applying an atrous convolution kernel with a atrous rate to skip a certain number of pixels and adjust the step size, and obtains a multi-scale global representation of the frame sequence feature map by combining the outputs of convolution layers with different atrous rates; The grouped convolution unit groups the input frame sequence feature maps into groups, each group containing a specific number of channels, performs a convolution operation on each group to extract local features, and recombines the processed frame sequence feature maps between different groups through cross-group connections or feature map fusion to obtain richer feature representation.

3. The intelligent frame interpolation and fluency improvement system for video content according to claim 1, characterized in that: The motion estimation layer includes an optical flow calculation unit and a spatiotemporal attention unit; The optical flow calculation unit preliminarily calculates the pixel motion vectors between adjacent frame sequences using an optical flow algorithm; The spatiotemporal attention unit is used to perform weighted processing on the pixel motion vector through a spatiotemporal attention mechanism, focus on the area of ​​the moving object in the video sequence, and improve the accuracy of the motion vector; The spatiotemporal attention unit includes a time attention subunit and a space attention subunit. The time attention subunit assigns weights according to the time sequence of the frame sequence, and the space attention subunit assigns weights according to pixel positions and feature values.

4. The intelligent frame interpolation and fluency improvement system for video content according to claim 1, characterized in that: The generative adversarial network in the image generation layer includes a generator unit and a discriminator unit; The generator unit uses a residual structured U-Net network to gradually restore detail information of the interpolation frame through upsampling and downsampling operations; The discriminator unit uses a multi-layer convolutional neural network to extract features from the input interpolated frames and real video sequence frames, and outputs the judgment results through a fully connected layer. According to the feedback information of the discriminator unit, the generator unit continuously adjusts the parameters for generating the interpolated frames to improve the quality of the generated interpolated frames.

5. The intelligent frame interpolation and fluency improvement system for video content according to claim 1, characterized in that: The fusion module fuses the interpolated frames and the original video sequence frames using a weighted average method, and adaptively adjusts the weights of the interpolated frames and the original video sequence frames according to the intensity of motion of the video content. For areas with slow motion, the weight of the original video sequence frames is increased, and for areas with intense motion, the weight of the interpolated frames is increased.

6. The intelligent frame interpolation and fluency improvement system for video content according to claim 1, characterized in that: The post-processing module includes an edge smoothing unit and a frame rate adjustment unit; The edge smoothing unit adopts a bilateral filtering algorithm to smooth the pixel values ​​of the video sequence frames while maintaining the edge information of the new video sequence, so as to reduce the edge jagged phenomenon generated during the interpolation process; The frame rate adjustment unit is used to perform dynamic adjustment according to the performance of the video playback device and the viewing needs of the user, and is achieved by copying or deleting part of the video sequence frames.

7. The intelligent frame interpolation and fluency improvement system for video content according to claim 1, characterized in that: The fluency optimization module includes a quality assessment unit and a parameter optimization unit; The quality assessment unit is used to quantitatively assess the video quality using a peak signal-to-noise ratio, a structural similarity index, and a video multi-method assessment fusion index; The parameter optimization unit is used to feed back the evaluation results to the frame interpolation model and the post-processing module, and adjust the frame interpolation model parameters and the post-processing parameters through the back propagation algorithm to improve the intelligent frame interpolation and fluency improvement effects of the video content.

8. A method for intelligent frame interpolation and fluency improvement of video content, used to implement a system for intelligent frame interpolation and fluency improvement of video content as claimed in any one of claims 1 to 7, characterized in that: The following steps are involved: Acquire an original video sequence, and preprocess the video sequence, including denoising, color correction, and brightness equalization; Perform frame decomposition on the preprocessed video sequence to split it into multiple frame sequences, perform feature extraction on each frame sequence, use a deep convolutional neural network to extract semantic features, texture features and edge features of the video image, and combine these features into a multi-dimensional feature vector; According to the multidimensional feature vector, the similarity between adjacent frame sequences is calculated by using cosine similarity or a measurement method based on feature distance, and the motion vector between the frame sequences is preliminarily estimated by using an optical flow algorithm, and the motion information between the frame sequences is analyzed, and the scene change points, motion intensity and key frame positions in the video image are determined based on the calculated similarity and the analyzed motion information, and a video frame sequence analysis report is generated; According to the generated video frame sequence analysis report, the preprocessed video sequence is input into a frame interpolation model based on deep learning, wherein the frame interpolation model includes a feature extraction layer, a motion estimation layer and an image generation layer, wherein the feature extraction layer uses a multi-scale convolutional neural network to extract the feature map of the frame sequence, the motion estimation layer uses an optical flow algorithm combined with a spatiotemporal attention mechanism to calculate the motion vector between adjacent frame sequences, and the image generation layer generates an interpolated frame based on the feature map and the motion vector through a generative adversarial network; The generated interpolated frames are merged with the original video sequence frames in time order to form a new video sequence; Post-processing the new video sequence, including edge smoothing and frame rate adjustment to improve the smoothness of the video; The quality of the new video sequence after post-processing is evaluated, and the model parameters of the frame interpolation process and the parameters of the post-processing process are optimized and adjusted according to the evaluation results.

9. An electronic device, characterized in that: The electronic device comprises: A memory storing executable program code; a processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the intelligent frame interpolation and fluency improvement method of video content as described in claim 8.

10. A computer storage medium, characterized in that: The computer storage medium stores computer instructions, and when the computer instructions are called, they are used to execute the intelligent frame interpolation and fluency improvement method of video content as described in claim 8.

Citation Information

Cited By

  • Liquid crystal display module driving control system and method supporting high refresh rate

    CN120853518A

  • A liquid crystal display module driving control system and method supporting high refresh rate

    CN120853518B

  • Infrared image ultrahigh frame frequency processing method and system

    CN120916036A