Live video image quality enhancement method based on deep learning
Through the pseudo-flow field estimation and adaptive spatiotemporal feature adjustment algorithm based on deep learning, the problems of inaccurate image processing and poor adaptability in the enhancement of live video image quality are solved, and the fine recovery of image details and timing consistency are achieved, and the video image quality is improved.
Patent Information
- Application Number
- CN202510672399.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2045-05-23
Smart Images

Figure CN120568151A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image data processing technology, and in particular to a method for enhancing the image quality of live video based on deep learning. Background Art
[0002] Live video quality enhancement technology is an important research direction in the current multimedia field. Improving video quality while maintaining real-time performance, especially in dynamic scenes, remains a constant challenge. Existing image and video enhancement methods typically rely on traditional denoising, super-resolution reconstruction, or enhancement algorithms, but these methods often face challenges with temporal consistency and dynamic scene detail recovery when processing live video streams. Traditional enhancement methods often lack precision in handling dynamic scenes with temporal changes, which can easily lead to detail loss or artifact generation. This is especially true in situations of high-speed motion or complex backgrounds, where the quality of live video often cannot be fully improved.
[0003] With the development of deep learning technology, new techniques such as convolutional neural networks (CNNs) and generative adversarial networks (GANs) have been widely applied in video processing. Deep learning methods can leverage large amounts of data to not only effectively improve image detail recovery but also develop enhanced strategies more suitable for dynamic scenes. However, existing deep learning methods often fail to effectively consider frame alignment and temporal consistency when processing the spatiotemporal relationships between video frames. This often compromises the effectiveness of video enhancement when processing high-speed motion or complex dynamic scenes.
[0004] At the same time, the above-mentioned existing technologies also have technical problems such as inaccurate image processing corresponding to live video and poor adaptability of the image quality enhancement processing process. Summary of the Invention
[0005] The present invention provides a live video quality enhancement method based on deep learning to solve the technical problems of inaccurate image processing corresponding to live video and poor adaptability of the image quality enhancement processing process.
[0006] The present invention provides a method for enhancing live video quality based on deep learning, which specifically includes the following technical solutions: A method for enhancing the quality of live video based on deep learning, comprising the following steps: S1. Acquire multiple image frames and input them into the pseudo-flow field estimation subnetwork to obtain a motion vector field. Based on the motion vector field, register the multiple image frames to obtain aligned multi-frame images. Feature extraction and fusion are then performed on the aligned multi-frame images to obtain multi-frame image feature data. S2. Enhance the multi-frame image feature data using an adaptive spatiotemporal feature adjustment and reconstruction enhancement algorithm to obtain enhanced image feature data; and generate enhanced video frames based on the enhanced image feature data.
[0007] Preferably, the S1 specifically includes: A pseudo flow field estimation subnetwork is introduced to adjust the image position by estimating the motion vector field between the previous and next frames. The pseudo flow field estimation subnetwork structure includes an optical flow estimation layer and a residual network layer.
[0008] Preferably, the S1 specifically includes: The optical flow estimation layer generates a motion vector field by calculating the pixel-level brightness and gradient information between different frame images, that is, the motion vector field from the previous frame image to the current frame image, and from the current frame image to the next frame image. The residual network layer is used to refine the motion vector field.
[0009] Preferably, the S2 specifically includes: The adaptive spatiotemporal feature adjustment and reconstruction enhancement algorithm enhances the feature data of multiple frames of images by introducing a combination of spatiotemporal feature adjustment, dynamic fuzzy kernel processing and nonlinear reconstruction.
[0010] Preferably, the S2 specifically includes: In the process of implementing the adaptive spatiotemporal feature adjustment and reconstruction enhancement algorithm, the multi-frame image feature data is normalized and preprocessed to obtain the normalized multi-frame image feature data; based on the normalized multi-frame image feature data, the spatiotemporal gradient difference is calculated to measure the changes of the image in time and space.
[0011] Preferably, the S2 specifically includes: In the implementation process of the adaptive spatiotemporal feature adjustment and reconstruction enhancement algorithm, weights are dynamically assigned to each frame of the image, and the adaptive weights are dynamically calculated based on the spatiotemporal gradient differences; based on the adaptive weights, the normalized multi-frame image feature data are weighted to obtain weighted multi-frame image feature data.
[0012] Preferably, the S2 specifically includes: In the implementation process of the adaptive spatiotemporal feature adjustment and reconstruction enhancement algorithm, an adaptive fuzzy kernel is introduced to smooth the weighted multi-frame image feature data. During the smoothing process, the L2 norm of the weighted multi-frame image feature data is calculated to obtain the dynamic intensity value of the weighted multi-frame image feature data.
[0013] Preferably, the S2 specifically includes: In the implementation process of the adaptive spatiotemporal feature adjustment and reconstruction enhancement algorithm, the blur kernel of the weighted multi-frame image feature data is calculated based on the dynamic intensity value, and a blur kernel set is formed; by performing a convolution operation on the weighted multi-frame image feature data and the blur kernel set, the smoothed multi-frame image feature data is obtained.
[0014] Preferably, the S2 specifically includes: In the process of implementing the adaptive spatiotemporal feature adjustment and reconstruction enhancement algorithm, a nonlinear reconstruction function is introduced to reconstruct the smoothed multi-frame image feature data to obtain the final enhanced image feature data.
[0015] The beneficial effects of the technical solution of the present invention are: 1. By adjusting the spatiotemporal features of multiple frames, performing dynamic blur kernel processing, and performing nonlinear adaptive reconstruction, image details are effectively restored, especially in dynamic scenes. By enhancing image details through spatiotemporal weighting and nonlinear activation functions, the image's texture, edges, and other high-frequency information can be finely enhanced, avoiding the blur and distortion found in traditional methods.
[0016] 2. During the processing of multiple frames of images, the multiple frames are input into the pseudo-flow field estimation subnetwork to obtain the motion vector field. Based on the motion vector field, the existing optical flow interpolation technology is used to align the multiple frames of images, ensuring the temporal consistency of the images and effectively capturing the motion information between the multiple frames of images. This avoids image inconsistencies caused by shooting angles, motion blur, and dynamic scenes, and ensures smooth and natural transitions during continuous video playback. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 This is a flow chart of a method for enhancing live video quality based on deep learning according to the present invention. DETAILED DESCRIPTION
[0018] In order to further illustrate the technical means and effects adopted by the present invention to achieve the predetermined purpose of the invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0019] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0020] The following describes in detail a method for enhancing the quality of live video based on deep learning provided by the present invention with reference to the accompanying drawings.
[0021] Refer to the attached Figure 1 , which shows a flow chart of a live video quality enhancement method based on deep learning provided by an embodiment of the present invention, the method includes the following steps: S1. Acquire multiple image frames and input them into the pseudo-flow field estimation subnetwork to obtain a motion vector field; perform registration processing on the multiple image frames based on the motion vector field to obtain aligned multiple image frames; perform feature extraction and fusion processing on the aligned multiple image frames to obtain feature data of the multiple image frames; Get three consecutive live video frames from the video capture end, that is, multiple frames, including the current frame , previous frame image , the next frame image ; The current frame is a frame used for reference and is the basis for registration of the remaining frames; the previous frame is a frame with an earlier timestamp; the next frame is a frame with a later timestamp. The three consecutive frames of live video will differ spatially due to shooting angles, motion blur and dynamic scenes. A pseudo flow field estimation subnetwork is introduced to estimate the motion vector field between the previous and next frames and adjust the image position so that images at different time points can be accurately aligned, thereby providing a consistency basis for subsequent feature extraction and fusion. The pseudo flow field estimation subnetwork is based on a convolutional neural network and consists of multiple convolutional layers. The specific pseudo flow field estimation subnetwork structure includes an optical flow estimation layer and a residual network layer. The optical flow estimation layer generates a motion vector field by calculating the brightness difference between adjacent frame images. The residual network layer further refines the motion vector field to reduce errors and improve the accuracy of motion estimation. The specific processing process includes: Multiple frames of image Input to the pseudo flow field estimation subnetwork, in the optical flow estimation layer, by calculating the pixel-level brightness and gradient information between different frame images, a motion vector field is generated, that is, the motion vector field is obtained from the previous frame image. To the current frame image Motion vector field , and from the current frame image To the next frame Motion vector field The two motion vector fields describe the To the current frame image And from the current frame image To the next frame In this way, the pseudo flow field estimation subnetwork can capture the motion information between multiple frames.
[0022] Furthermore, based on the estimated motion vector field and , use existing optical flow interpolation technology (such as bilinear interpolation) to align multiple frames of images. Specifically, through the motion vector field , the previous frame image can be The pixel points in the image are adjusted to the current frame according to the displacement indicated by the motion vector field. Similarly, using the motion vector field The current frame image The pixels in the image are adjusted to the next frame. Thus, through two optical flow interpolations, multiple frames of images All are aligned to images with the same spatial structure, that is, the aligned multi-frame images are obtained ,The aligned multi-frame images have the same spatial coordinate system. After registration, they can enter the subsequent feature extraction and fusion processing process.
[0023] Furthermore, after completing the registration, the aligned multi-frame images Perform feature extraction and fusion processing. Specifically, use a convolutional neural network (CNN) with shared weights to extract the features of the aligned multi-frame images and obtain multi-frame image feature data. .
[0024] S2. Enhance the feature data of multiple frames of images using an adaptive spatiotemporal feature adjustment and reconstruction enhancement algorithm to obtain enhanced image feature data; and generate enhanced video frames based on the enhanced image feature data.
[0025] The adaptive spatiotemporal feature adjustment and reconstruction enhancement algorithm is used to enhance the feature data of multiple frames of images. The adaptive spatiotemporal feature adjustment and reconstruction enhancement algorithm combines spatiotemporal feature adjustment, dynamic blur kernel processing, and nonlinear reconstruction to appropriately enhance the features of the image at different time points and spatial regions, while ensuring the consistency of the image in time sequence and the restoration of details in dynamic scenes. The specific implementation process is as follows: Multi-frame image feature data Perform normalization preprocessing to obtain normalized multi-frame image feature data , , In order to measure the changes of the image in time and space, the temporal and spatial gradient differences are calculated, so as to dynamically assign weights to each frame of the image, strengthen the temporal and spatial consistency and detail performance of the image, and dynamically calculate the adaptive weight of each frame of the image based on the temporal and spatial gradient differences. Taking time as an example, the specific formula is as follows: , , in, is the difference in spatiotemporal gradient, indicating The difference in the change of the moment image in the time dimension and space dimension; yes The gradient of the moment image in the time direction represents the difference between the current frame image and the previous and next frame images in the time dimension, which is obtained through the time difference operation; yes The gradient of the image in the spatial direction at a moment in time indicates the change in local features between the current frame image and the previous and next frames in the spatial dimension. The spatial gradient reflects the changes in texture, edge, shape and other information of the local area in the image and is calculated using convolution or discrete difference methods. is the adaptive weight, which represents the adaptive importance weight of the current frame image in the image enhancement process; It is a regulating factor that controls the spatiotemporal weighted sensitivity. It is used to control the influence of spatiotemporal gradient differences on the adaptive weight calculation. It is determined according to expert experience and the reference value is ; is the time index of multiple frames of images; Based on the adaptive weight, the normalized multi-frame image feature data is weighted to obtain the weighted multi-frame image feature data ,Right now , , .
[0026] Furthermore, in order to reduce background noise and enhance the details of dynamic areas, thereby avoiding excessive blur and detail loss and maintaining image clarity, an adaptive blur kernel is introduced to smooth the weighted multi-frame image feature data. Specifically: Taking the moment as an example, the L2 norm of the weighted multi-frame image feature data is calculated to obtain Moment Weighted image feature data Dynamic strength value ; , in, yes Moment Weighted image feature data; yes Moment Weighted image feature data; yes Moment Weighted image feature data; Based on the dynamic strength value, calculate Moment Weighted image feature data Blur kernel , and constitute Fuzzy kernel set of image feature data after time weighting ,based on The calculation formula of the blur kernel obtained by the function is: , in, It is the fuzzy kernel control factor, which is used to adjust the relationship between the fuzzy kernel and the dynamic intensity value, and control the influence of the change of the dynamic intensity of the image area on the fuzzy kernel. It is obtained through experimental methods such as Bayesian optimization, and the reference range is ; It is the dynamic intensity threshold, which is used to control the reference point of the dynamic intensity value when determining the size of the blur kernel. It is determined according to the expert experience method, and the reference value is .
[0027] Furthermore, by taking the weighted image feature data and fuzzy kernel set Perform convolution operation to obtain smoothed image feature data .
[0028] The smoothed multi-frame image feature data is obtained by the same processing as above , , .
[0029] Furthermore, a nonlinear reconstruction function is introduced to reconstruct the smoothed multi-frame image feature data, which can restore image details more finely and suppress background noise, and serve as the final enhanced image feature data. The nonlinear reconstruction function is obtained based on the feature fusion and nonlinear activation technology in deep learning. The specific formula is as follows: , in, is the enhanced image feature data, which represents the image feature data obtained through the nonlinear reconstruction process; It is a global adjustment factor used to adjust the intensity of image detail enhancement, which determines the intensity of the final enhanced image feature data, ensuring that the details of the restoration process will not be over-enhanced, thereby avoiding artifacts or distortion. It is set according to actual application requirements. The reference value is ; It is a nonlinear activation function used to enhance the details of the image; is the adaptive weight; yes Product, which means element-wise multiplication; It is an activation function used to normalize the L2 norm of the smoothed multi-frame image feature data; yes Adjustment factor, used to control The response speed of the function is determined according to the expert experience method, and the reference value is ; Through the regulating factor Dynamically adjust the total intensity of the image so that the feature recovery process takes into account the overall image intensity while avoiding over-enhancement or loss of important details; It is a nonlinear normalization of the intensity of image features; By smoothing the image features, the image details are restored while suppressing the noise; Indicates weighted summation of the smoothed image feature data of the current frame and the previous and next frames; It shows the result of weighted and nonlinear processing of the smoothed image feature data of the current frame and the previous and next frames.
[0030] Finally, based on the enhanced image feature data, enhanced video frames are generated through existing technical means such as generative adversarial networks or convolutional neural networks to achieve image quality enhancement.
[0031] In summary, a live video quality enhancement method based on deep learning has been completed.
[0032] The order in which the embodiments of the invention are presented is for illustrative purposes only and does not necessarily represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0033] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.
[0034] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. A method for enhancing live video quality based on deep learning, characterized in that: The following steps are involved: S1. Acquire multiple image frames and input them into the pseudo-flow field estimation subnetwork to obtain a motion vector field. Based on the motion vector field, register the multiple image frames to obtain aligned multi-frame images. Feature extraction and fusion are then performed on the aligned multi-frame images to obtain multi-frame image feature data. S2. Enhance the multi-frame image feature data using an adaptive spatiotemporal feature adjustment and reconstruction enhancement algorithm to obtain enhanced image feature data; and generate enhanced video frames based on the enhanced image feature data.
2. The method for enhancing live video quality based on deep learning according to claim 1, characterized in that: Said S1 specifically includes: A pseudo flow field estimation subnetwork is introduced to adjust the image position by estimating the motion vector field between the previous and next frames. The pseudo flow field estimation subnetwork structure includes an optical flow estimation layer and a residual network layer.
3. The method for enhancing live video quality based on deep learning according to claim 2, characterized in that: Said S1 specifically includes: The optical flow estimation layer generates a motion vector field by calculating the pixel-level brightness and gradient information between different frame images, i.e., the motion vector field from the previous frame image to the current frame image, and the motion vector field from the current frame image to the next frame image; the residual network layer is used to refine the motion vector field.
4. The method for enhancing live video quality based on deep learning according to claim 1, characterized in that: Said S2 specifically includes: The adaptive spatiotemporal feature adjustment and reconstruction enhancement algorithm enhances the feature data of multiple frames of images by introducing a combination of spatiotemporal feature adjustment, dynamic fuzzy kernel processing and nonlinear reconstruction.
5. The method for enhancing live video quality based on deep learning according to claim 4, characterized in that: Said S2 specifically includes: In the process of implementing the adaptive spatiotemporal feature adjustment and reconstruction enhancement algorithm, the multi-frame image feature data is normalized and preprocessed to obtain the normalized multi-frame image feature data; based on the normalized multi-frame image feature data, the spatiotemporal gradient difference is calculated to measure the changes of the image in time and space.
6. The method for enhancing live video quality based on deep learning according to claim 5, characterized in that: Said S2 specifically includes: In the implementation process of the adaptive spatiotemporal feature adjustment and reconstruction enhancement algorithm, weights are dynamically assigned to each frame of the image, and the adaptive weights are dynamically calculated based on the spatiotemporal gradient differences; based on the adaptive weights, the normalized multi-frame image feature data are weighted to obtain weighted multi-frame image feature data.
7. The method for enhancing live video quality based on deep learning according to claim 6, characterized in that: Said S2 specifically includes: In the implementation process of the adaptive spatiotemporal feature adjustment and reconstruction enhancement algorithm, an adaptive fuzzy kernel is introduced to smooth the weighted multi-frame image feature data. During the smoothing process, the L2 norm of the weighted multi-frame image feature data is calculated to obtain the dynamic intensity value of the weighted multi-frame image feature data.
8. The method for enhancing live video quality based on deep learning according to claim 7, characterized in that: Said S2 specifically includes: In the implementation process of the adaptive spatiotemporal feature adjustment and reconstruction enhancement algorithm, the blur kernel of the weighted multi-frame image feature data is calculated based on the dynamic intensity value, and a blur kernel set is formed; by performing a convolution operation on the weighted multi-frame image feature data and the blur kernel set, the smoothed multi-frame image feature data is obtained.
9. The method for enhancing live video quality based on deep learning according to claim 8, characterized in that: Said S2 specifically includes: In the process of implementing the adaptive spatiotemporal feature adjustment and reconstruction enhancement algorithm, a nonlinear reconstruction function is introduced to reconstruct the smoothed multi-frame image feature data to obtain the final enhanced image feature data.
Citation Information
Patent Citations
Video frame rate up-conversion method and system for improving motion fluency intelligently
CN106210767A
Method for estimating non-uniform blurring kernel of image based on U-Net
CN115018726A
Compressed video quality enhancement method based on reconstructed flow field
CN116012272A
Panoramic video stability enhancement method based on space-time consistency
CN119671876A
Low-light video processing method, device and storage medium
US20230196721A1