Video quality recovery method, apparatus, device, and storage medium

By performing content complexity analysis on video frames and dynamically adjusting the network depth of the target restoration model, and sharing the weight sets of the encoder and decoder, the problems of high computational resource requirements and slow processing speed in existing technologies are solved, achieving efficient and flexible video quality restoration.

CN121053039BActive Publication Date: 2026-01-27MT TITLIS BEIJING CONTROL TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511588393.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-03
Publication Date
2026-01-27
Estimated Expiration
2045-11-03

AI Technical Summary

Technical Problem

Existing neural network-based video quality restoration methods have high computational resource requirements and poor speed and flexibility when processing complex content.

Method used

By performing content complexity analysis on video frames, the network depth of the target restoration model is dynamically adjusted, and the weight sets of the encoder and decoder are shared to optimize the video quality restoration process.

Benefits of technology

It significantly reduces the computational resource requirements, improves the robustness and processing speed of the model when dealing with unseen data, and enhances the flexibility and efficiency of video quality restoration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121053039B_ABST
    Figure CN121053039B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of image optimization, in particular to a video quality recovery method and device, equipment and a storage medium, wherein the method comprises the following steps: performing frame analysis on an original video to obtain video frames, and calculating the content complexity of the video frames; performing network depth adjustment on a target recovery model based on the content complexity to obtain a depth-adjusted model; wherein an encoder and a decoder in the target recovery model share the same weight set; processing the video frames based on the depth-adjusted model to obtain reconstructed frames; and obtaining a reconstructed video based on the reconstructed frames. The application facilitates reducing the requirement for computing resources and improving the processing speed and flexibility during video quality recovery.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image optimization technology, and in particular to a video quality restoration method, apparatus, device, and storage medium. Background Technology

[0002] With the continuous development of video technology, video has been widely used in industries such as security monitoring, online education, entertainment, and social media. However, the video used in various industries often suffers from video quality degradation due to factors such as shooting conditions, equipment limitations, and compression during transmission. Common video quality degradation phenomena include: blurring, noise, compression artifacts, and lighting problems.

[0003] To ensure that videos can be used normally in the corresponding industries, it is often necessary to restore the quality of the initially completed videos. The commonly used video quality restoration method is the neural network-based video quality restoration method. However, this method has high requirements for computing resources and poor processing speed and flexibility when dealing with complex content. Summary of the Invention

[0004] To facilitate video quality restoration by reducing the computational resource requirements and improving processing speed and flexibility, this application provides a video quality restoration method, apparatus, device, and storage medium.

[0005] Firstly, this application provides a video quality restoration method, comprising:

[0006] Frame parsing is performed on the original video to obtain video frames, and the content complexity of the video frames is calculated.

[0007] Based on the aforementioned content complexity, the network depth of the target recovery model is adjusted to obtain a depth-adjusted model; wherein, the encoder and decoder in the target recovery model share the same weight set;

[0008] The video frames are processed based on the depth post-adjustment model to obtain reconstructed frames;

[0009] Based on the reconstructed frames, the reconstructed video is obtained.

[0010] Secondly, this application provides a video quality restoration apparatus, comprising:

[0011] The parsing and calculation module is used to perform frame parsing on the original video to obtain video frames and calculate the content complexity of the video frames.

[0012] The depth adjustment module is used to perform network depth adjustment on the target recovery model based on the content complexity to obtain a depth-adjusted model; wherein the encoder and decoder in the target recovery model share the same weight set;

[0013] The frame reconstruction module is used to process the video frames based on the depth post-adjustment model to obtain reconstructed frames;

[0014] The video reconstruction module is used to obtain the reconstructed video based on the reconstructed frames.

[0015] Thirdly, this application provides a computer device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps in the method described above.

[0016] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the above-described method.

[0017] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the steps in any of the above method embodiments.

[0018] The aforementioned video quality restoration method, apparatus, device, and storage medium obtain video frames by frame parsing of the original video and calculate the content complexity of the video frames; adjust the network depth of the target restoration model based on the content complexity to obtain a depth-adjusted model; wherein the encoder and decoder in the target restoration model share the same weight set; process the video frames based on the depth-adjusted model to obtain reconstructed frames; and obtain the reconstructed video based on the reconstructed frames. Through the above implementation, since the encoder and decoder in the target restoration model use the same set of weights, when processing the input video frame sequence, whether in the encoding or decoding stage, the target restoration model uses the same parameters for learning and generation. This sharing mechanism not only reduces the complexity of the model, significantly reducing the demand for computing resources, but also helps the target restoration model generalize better because all video frames are processed through the same weights, thereby improving the model's robustness when processing unseen data; furthermore, by dynamically adjusting the network depth of the target restoration model in real time according to the content complexity of the video frames, the target restoration model can achieve a higher processing speed when the content complexity is high, thus effectively improving the processing speed and flexibility of the target restoration model when processing video frames.

[0019] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description

[0020] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a flowchart of a video quality restoration method provided in the embodiments of this application;

[0022] Figure 2 This is a schematic diagram of the structure of a video quality restoration device provided in the embodiments of this application;

[0023] Figure 3 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application;

[0024] Figure 4 This is an internal structural diagram of a computer-readable storage medium provided in an embodiment of this application. Detailed Implementation

[0025] To make the objectives, technical solutions, and advantages of this disclosure clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this disclosure.

[0026] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings herein are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, apparatus, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.

[0027] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0028] Example 1

[0029] Figure 1This is a flowchart of a video quality restoration method provided in Embodiment 1 of this application, for reference. Figure 1 The method can be executed by a device that performs the method, which can be implemented in software and / or hardware, and the method includes:

[0030] S110. Perform frame parsing on the original video to obtain video frames, and calculate the content complexity of the video frames.

[0031] Among these, security monitoring, online education, online entertainment, and social media all require video as a medium to provide corresponding services. Taking one video application area as an example, the initial video footage may suffer from poor quality due to factors such as poor shooting environment (e.g., insufficient ambient light), poor shooting equipment performance, and compression affecting video quality during transmission. Poor video quality can manifest in the following ways: overall blurriness, excessive noise, and compression artifacts. Therefore, it is necessary to improve the video quality of the initial footage.

[0032] The initially captured video is denoted as the original video V. The original video consists of multiple video frames. To improve the video quality of the original video V, it is necessary to improve the image quality of each video frame one by one. To this end, it is necessary to first obtain each video frame in the original video V. In this embodiment, the original video V is parsed to obtain each video frame X(t) in the original video V, where t represents time t.

[0033] It should be noted that the original video V contains multiple video frames X(t), each containing different content. Some video frames X(t) have more complex content (a large number of objects, such as many people or items), or contain a lot of moving scenes (high-speed movement of people or objects), resulting in higher content complexity for those video frames X(t). Conversely, some video frames X(t) have simpler content, or contain static scenes, resulting in lower content complexity for those video frames X(t). Subsequent steps require reconstructing each video frame based on a pre-defined image optimization model to improve the image quality of each frame. Video frames with higher content complexity require higher computational resources from the image optimization model, while video frames with lower content complexity require only lower computational resources. If the image optimization model can determine the computational resources to be allocated to each video frame X(t) based on its content complexity, then higher computational resources can be dynamically allocated to video frame X(t) when its content complexity is high. This not only improves processing efficiency but also captures detailed dynamic information on video frame X(t), thereby enhancing the image optimization effect. Conversely, lower computational resources can be allocated to video frame X(t) when its content complexity is low, thus reducing unnecessary consumption of computational resources.

[0034] Therefore, after parsing the video frame X(t) in the original video V, this embodiment further processes the video frame X(t) using a preset complexity calculation tool to obtain the content complexity corresponding to the video frame X(t).

[0035] S120. Based on the content complexity, the network depth of the target recovery model is adjusted to obtain the depth-adjusted model; wherein, the encoder and decoder in the target recovery model share the same weight set.

[0036] The image optimization model used to optimize the image quality of video frame X(t) is denoted as the target restoration model. It should be noted that the target restoration model is a model that has already been trained. The target restoration model includes an encoder and a decoder. The encoder is responsible for extracting and encoding the features of the video frame. It usually contains multiple convolutional layers or other types of feature extraction layers. The purpose is to transform the high-dimensional input data (video frame) into a compressed and information-rich feature representation. The decoder reconstructs a high-quality image or video frame from the feature representation output by the encoder. The structure of the decoder is usually symmetrical to that of the encoder. It uses deconvolutional layers or upsampling layers to gradually restore the size and details of the image.

[0037] It should be noted that in traditional models, each functional layer (convolutional layer, pooling layer, etc.) in the encoder and decoder needs to have its own weights set. However, this results in a large number of parameters (weights) in the model, which increases the model's complexity, thereby increasing the computational resources required for image quality optimization and reducing the model's processing speed. To reduce the complexity of the target restoration model, this embodiment pre-sets a shareable weight set for the encoder and decoder. The parameters in this weight set can be shared across the functional layers in both the encoder and decoder, thus significantly reducing the number of parameters (weights) in the target restoration model.

[0038] It should also be noted that when processing sequential data such as video, weight sharing can improve model efficiency and reduce the number of parameters. By sharing the same set of parameters among multiple processing units, we can significantly reduce the required memory and computing resources while maintaining or even improving model performance.

[0039] Specifically, the weights required by the encoder are denoted as W_encoder, the weights required by the decoder are denoted as W_decoder, and the preset weight set is denoted as W_shared. In this embodiment, the weights are... Thus, when processing input video frames, the target restoration model uses the same parameters for learning and generation, whether in the encoding or decoding phase. This sharing mechanism not only reduces model complexity but also helps the model generalize better because all data is processed with the same weights, improving the robustness of the target restoration model when dealing with unseen data.

[0040] Both the encoder and decoder consist of multiple functional layers, meaning the target recovery model has a corresponding network depth. The more functional layers the encoder and decoder have, the higher the network depth of the target recovery model. If the network depth of the target recovery model is higher, the computational performance of the target recovery model is higher, but the computational resources required are more. Conversely, the computational performance of the target recovery model is lower, but the computational resources required are less.

[0041] In this embodiment, when processing video frame X(t) through the target restoration model, the network depth (depth(t) of the target restoration model is dynamically adjusted according to the content complexity of the video frame. Specifically, this embodiment pre-defines a content complexity evaluation function content_complexity(), which is used to evaluate the content complexity of video frame X(t), and this content complexity is denoted as content_complexity(X(t)). This embodiment also pre-defines a network depth mapping function f, which is used to map the content complexity content_complexity(X(t)) to the corresponding network depth (depth(t)). ; depth(t) represents the network depth that the target restoration model should have when processing the video frame corresponding to time t; after obtaining the network depth, adjust the number of functional layers in the encoder and decoder applied in the target restoration model when processing video frame X(t) according to the network depth, and denote the new model obtained after the target restoration model completes the network depth adjustment as the depth-adjusted model.

[0042] It's important to note that for video frames with high content complexity, the network depth mapping function `f` returns a larger value, indicating that a deeper network is needed to capture detailed dynamic information. Conversely, for video frames with low content complexity, `f` returns a smaller value to reduce unnecessary computation and resource consumption. The network depth of the target reconstruction model can be adjusted for each video frame. This dynamic adjustment mechanism allows the model to adapt to various processing requirements, effectively reducing the overall computational burden while maintaining high performance. For example, when processing different segments of a long video, dynamic adjustment can optimize the network structure for the specific characteristics of each segment, thereby improving processing speed while ensuring reconstruction quality.

[0043] In summary, weight sharing and dynamic adjustment of network depth are two key strategies for optimizing video image restoration tasks. Weight sharing reduces the risk of overfitting and improves computational efficiency by decreasing the number of model parameters, while dynamically adjusting network depth allows the model to flexibly handle video content of varying complexity, which is particularly important for real-time video processing.

[0044] S130. Process the video frame based on the depth post-adjustment model to obtain the reconstructed frame.

[0045] In this process, by inputting the video frame X(t) into the depth post-adjustment model for image quality optimization, a new video frame with effectively optimized image quality can be obtained, and this new video frame is denoted as the reconstructed frame.

[0046] S140. Based on the reconstructed frame, the reconstructed video is obtained.

[0047] By combining the reconstructed frames corresponding to each time point in chronological order, a new video with effectively optimized video quality can be obtained, and this new video is recorded as the reconstructed video.

[0048] It should be noted that this embodiment obtains video frames by parsing the original video and calculating the content complexity of the video frames; based on the content complexity, the network depth of the target restoration model is adjusted to obtain a depth-adjusted model; wherein, the encoder and decoder in the target restoration model share the same weight set; the video frames are processed based on the depth-adjusted model to obtain reconstructed frames; and based on the reconstructed frames, a reconstructed video is obtained. Through the above implementation, since the encoder and decoder in the target restoration model use the same set of weights, when processing the input video frame sequence, whether in the encoding or decoding stage, the target restoration model uses the same parameters for learning and generation. This sharing mechanism not only reduces the complexity of the model, significantly reducing the demand for computing resources, but also helps the target restoration model generalize better because all video frames are processed through the same weights, thereby improving the robustness of the model when processing unseen data; in addition, by dynamically adjusting the network depth of the target restoration model in real time according to the content complexity of the video frames, the target restoration model can have a higher processing speed when the content complexity is high, thereby effectively improving the processing speed and flexibility of the target restoration model when processing video frames.

[0049] Example 2

[0050] This application provides a video quality restoration method in Embodiment 2, which optimizes the "training steps of the target restoration model" in Embodiment 1. It should be noted that for parts not detailed in this embodiment, please refer to the descriptions in other embodiments. The method includes:

[0051] S210. Perform frame parsing on the original video to obtain video frames, and calculate the content complexity of the video frames.

[0052] S220. Based on the content complexity, the network depth of the target recovery model is adjusted to obtain the depth-adjusted model; wherein, the encoder and decoder in the target recovery model share the same weight set.

[0053] The training steps for the target recovery model include:

[0054] A210. Calculate the model loss based on the training frames, the original recovery model, and the dynamic shared weights corresponding to the training frames.

[0055] To ensure the efficiency and accuracy of the target restoration model in optimizing the image quality of video frames, the untrained original restoration model needs to be trained first to obtain a target restoration model with higher efficiency and accuracy in image quality optimization. To train this original restoration model, this embodiment pre-sets training data, which contains multiple video frames used for training, and these video frames are denoted as training frame Y(t).

[0056] Taking a training frame Y(t) as an example, the original recovery model processes the training frame Y(t) to obtain the prediction frame Y_hat(t) corresponding to the training frame Y(t). It should be noted that during the training process, the weights of the original recovery model are not static, but can be adjusted according to the characteristics of the training frame Y(t) at different times. This dynamic adjustment strategy allows the model to focus on different features and tasks at different stages, thereby making more efficient use of training data and time. The model weights corresponding to the training frame Y(t) at time t are denoted as the dynamic shared weights C_shared(t).

[0057] Taking a training frame Y(t) as an example, the single-frame loss L(t) corresponding to the training frame Y(t) can be constructed using the training frame Y(t), the predicted frame Y_hat(t) obtained by processing the training frame Y(t) by the original recovery model, and the dynamic shared weight C_shared(t) corresponding to the training frame Y(t). In this way, the single-frame loss L(t) corresponding to each training frame Y(t) can be calculated. Furthermore, by summing the single-frame losses L(t) corresponding to each training frame Y(t), the model loss Loss corresponding to the original recovery model can be obtained.

[0058] A220. Optimize the original recovery model based on the model loss to obtain an intermediate recovery model.

[0059] The model loss can be used to iteratively optimize the model parameters in the original model until a preset number of iterations is reached or the model converges, thereby obtaining a model that has been initially trained. This initially trained model is referred to as the intermediate recovery model.

[0060] A230. Based on the preset pruning rate, the training frames, the real frames corresponding to the training frames, and the dynamic shared weights, the intermediate recovery model is adjusted to obtain the target recovery model.

[0061] It should be noted that the intermediate restoration model already has high image quality optimization performance, but the number of model parameters of the intermediate restoration model is large. In order to minimize the amount of computation required for subsequent intermediate restoration model processing of video frames, the number of model parameters of the intermediate restoration model can be reasonably reduced. For this purpose, this embodiment presets a pruning rate based on historical experience data. The intermediate restoration model can be pruned by the pruning rate, that is, the number of model parameters of the intermediate restoration model can be reduced.

[0062] It should also be noted that this embodiment aims to enable the final trained recovery model to be applied to a specific use case, such as real-time video surveillance. Therefore, after completing the model pruning, it is necessary to further fine-tune the intermediate recovery model, that is, fine-tune the dynamic shared weight C_shared(t) in the intermediate recovery model.

[0063] Specifically, to fine-tune the dynamic shared weight C_shared(t) in the intermediate recovery model, the training frame Y(t), the corresponding real frame Y_true(t), and the dynamic shared weight C_shared(t) need to be processed according to the preset weight adjustment function fine_tune() to obtain the adjusted new weights. These new weights then replace the dynamic shared weight C_shared(t) in the intermediate recovery model, thereby adjusting the intermediate recovery model. The new model obtained after the intermediate recovery model is adjusted is denoted as the target recovery model.

[0064] Here, the real frame Y_true(t) corresponding to the training frame Y(t) is the target output in the training dataset, that is, the real frame that the model tries to match during the training process.

[0065] S230. Process the video frame based on the depth adjustment model to obtain the reconstructed frame.

[0066] S240. Based on the reconstructed frame, the reconstructed video is obtained.

[0067] Example 3

[0068] This application provides a video quality restoration method in Embodiment 3, which optimizes the "adjustment of the intermediate restoration model based on a preset pruning rate, the training frame, the real frame corresponding to the training frame, and the dynamic shared weight to obtain a target restoration model" in Embodiment 2. It should be noted that for parts not detailed in this embodiment, please refer to the descriptions in other embodiments. The method includes:

[0069] S310. Perform frame parsing on the original video to obtain video frames, and calculate the content complexity of the video frames.

[0070] S320. Based on the content complexity, the network depth of the target recovery model is adjusted to obtain the depth-adjusted model; wherein, the encoder and decoder in the target recovery model share the same weight set.

[0071] The training steps for the target recovery model include:

[0072] A310. Calculate the model loss based on the training frames, the original recovery model, and the dynamic shared weights corresponding to the training frames.

[0073] A320. Optimize the original recovery model based on the model loss to obtain an intermediate recovery model.

[0074] A331. The intermediate recovery model is pruned based on a preset pruning rate to obtain a pruned recovery model.

[0075] In this embodiment, a pruning rate (pruning_rate) is preset. This pruning rate is a value between 0 and 1, representing the proportion of parameters to be removed during the pruning process. For example, a pruning_rate of 0.5 means that approximately 50% of the model parameters will be deleted. This pruning rate is used to remove some model parameters from the intermediate recovery model, thereby pruning the intermediate recovery model to minimize the computational load required for subsequent intermediate recovery model processing of video frames, thus improving the processing efficiency of video frames. The new model obtained after pruning the intermediate recovery model is denoted as the pruned recovery model.

[0076] A332. Based on the training frames, the real frames corresponding to the training frames, and the dynamic shared weights, calculate the model adjustment weights.

[0077] In this implementation, a weight adjustment function fine_tune() is preset. This weight adjustment function fine_tune() is used to process the input training frame Y(t), the real frame Y_true(t) corresponding to the training frame, and the dynamic shared weight C_shared(t) to obtain the adjusted weight, and the adjusted weight is recorded as the model adjusted weight W_adjusted.

[0078] A333. Adjust the weights of the pruning recovery model based on the model to obtain the target recovery model.

[0079] Among them, the model adjustment weight W_adjusted is used to replace the dynamic shared weight C_shared(t) in the pruning recovery model, and the new model obtained after the pruning recovery model completes the weight adjustment is denoted as the target recovery model.

[0080] It should be noted that by sequentially pruning and fine-tuning the weights of the intermediate recovery model, secondary training of the original recovery model is achieved. This allows the resulting intermediate recovery model to not only operate effectively in resource-constrained environments but also to provide optimized results for specific task requirements. For example, in real-time video surveillance recovery tasks, pruning reduces computational load, and fine-tuning allows for precise adjustments to the model to adapt to specific scenarios and environmental changes, thereby significantly improving the model's practicality and effectiveness.

[0081] S330. Process the video frame based on the depth adjustment model to obtain the reconstructed frame.

[0082] S340. Based on the reconstructed frame, the reconstructed video is obtained.

[0083] Example 4

[0084] This application provides a video quality restoration method in Embodiment 2, which optimizes the "calculation of the content complexity of the video frame" in Embodiment 4. It should be noted that for parts not detailed in this embodiment, please refer to the descriptions in other embodiments. The method includes:

[0085] S411. Perform frame parsing on the original video to obtain video frames.

[0086] S412. Adjust the size of the video frame based on the preset frame input size to obtain a standard video frame.

[0087] The video frames are subsequently used as input to a preset restoration model for image quality optimization. The restoration model has size requirements for the input video frames, and these size requirements are denoted as the frame input size. Specifically, the frame input size includes the frame input height target_height and the frame input width target_width. The video frames can be resized using this frame input size so that the resized video frames can meet the size requirements of the restoration model. The resized video frames are then denoted as standard video frames.

[0088] It should be noted that by adjusting all video frames to the same size, it can be ensured that the recovery model applies its weights evenly throughout the entire video sequence.

[0089] S413. Normalize the standard video frames based on the statistics corresponding to the preset training frame set to obtain normalized frames.

[0090] In this embodiment, to improve the learning ability of the subsequent recovery model on video frames, after obtaining the standard video frame, the standard video frame is further normalized. To achieve this normalization, a preset training frame set needs to be statistically analyzed to obtain the corresponding statistical measure. The standard video frame is then normalized using this statistical measure. The preset training frame set is the set of training frames required to train the initial recovery model. In this embodiment, the statistical measure obtained by analyzing the training frame set is specifically the mean and standard deviation (std) of the image pixel values ​​corresponding to all training frames in the training frame set. The video frame obtained by normalizing the standard video frame using this statistical measure is denoted as the normalized frame X_normalized(t).

[0091] In this embodiment, the formula for calculating the normalized frame X_normalized(t) is:

[0092]

[0093] Where X_resized(t) is a standard video frame.

[0094] It should be noted that normalizing standard video frames, that is, scaling standard video frames to a range with zero mean and unit variance, helps the model learn more effectively because it avoids numerical instability that the model may encounter during training.

[0095] S414. Perform content parsing on the normalized frame to obtain frame content, wherein the frame content includes at least one of the following: object quantity, object movement speed, and color change information.

[0096] In this embodiment, a content parsing tool is pre-set to perform content parsing on the normalized frame. The content parsing tool is used to parse the content of the normalized frame to obtain the frame content of the normalized frame. In this embodiment, the frame content includes at least one of the following: the number of objects, the object movement speed, and color change information.

[0097] Among them, the number of objects is the total number of objects in the normalized frame, the object movement speed is the movement speed of each object in the normalized frame, and the color change information is the change of each image pixel in the normalized frame relative to the corresponding image pixel in the adjacent normalized frame.

[0098] S415. Process the frame content based on the preset complexity calculation formula to obtain the content complexity.

[0099] In this embodiment, a complexity calculation formula for processing the frame content of a normalized frame is preset. This complexity calculation formula can be used to calculate the frame content of a normalized frame in order to obtain the content complexity corresponding to the normalized frame.

[0100] S420. Based on the content complexity, the network depth of the target recovery model is adjusted to obtain the depth-adjusted model; wherein, the encoder and decoder in the target recovery model share the same weight set.

[0101] S430. Process the video frame based on the depth adjustment model to obtain the reconstructed frame.

[0102] S440. Based on the reconstructed frame, the reconstructed video is obtained.

[0103] Example 5

[0104] This application provides a video quality restoration method in Embodiment 5, which optimizes the "processing the video frame based on the depth post-adjustment model to obtain the reconstructed frame" in Embodiment 1. It should be noted that for parts not detailed in this embodiment, please refer to the descriptions in other embodiments. The method includes:

[0105] S510. Perform frame parsing on the original video to obtain video frames, and calculate the content complexity of the video frames.

[0106] S520. Based on the content complexity, the network depth of the target recovery model is adjusted to obtain the depth-adjusted model; wherein, the encoder and decoder in the target recovery model share the same weight set.

[0107] S531. Determine the adjacent frames of the video frame.

[0108] It should be noted that, in order to optimize the image quality of video frames, it is necessary to fully acquire the image features of the video frames first. These image features specifically include image spatial information and image temporal information, etc., without any specific limitation. Taking the video frame X(t) corresponding to time t as an example, if only the image features of the video frame X(t) are acquired, then the subsequent depth adjustment model will only construct the reconstructed frame corresponding to the video frame X(t) based on the image features of the video frame X(t) when processing the video frame X(t). The change information between the video frame X(t) and the video frames before and after it can also be used to assist in constructing the reconstructed frame corresponding to the video frame X(t), and can improve the image quality of the reconstructed frame. Therefore, in this embodiment, based on determining the video frame X(t), the preceding video frame X(t-1) and the following video frame X(t+1) are further acquired as the adjacent frames of the video frame X(t).

[0109] S532. Based on the depth post-adjustment model, perform 3D convolution on the video frame and the adjacent frames to obtain the feature matrix corresponding to the video frame.

[0110] Among them, the depth-adjusted model can perform 3D convolution on video frames and their corresponding neighboring frames. The purpose of 3D convolution is to obtain the dynamic changes between video frames and their corresponding neighboring frames, such as changes in the movement of objects or changes in brightness, so as to extract the frame features of the video frames and record the frame features as a feature matrix.

[0111] It should be noted that, compared with traditional 2D convolution, 3D convolution also processes the changes of video frames on the time axis. This allows the depth-adjusted model to understand the motion and other temporal features in the video sequence, thereby obtaining better frame features and facilitating the improvement of the image optimization quality of the subsequently constructed reconstructed frames.

[0112] S533. Based on the feature matrix, the video frame, and the depth post-adjustment model, determine the reconstructed frame corresponding to the video frame.

[0113] The depth post-adjustment model can optimize the image quality of video frames by processing the video frames and the feature matrices corresponding to those video frames. The new video frame obtained after the image quality optimization is called the reconstructed frame.

[0114] It should be noted that in this embodiment, the video frames in steps S531-S533 can also be normalized frames obtained from video frames as shown in Embodiment 4. In other embodiments, no specific limitations are made.

[0115] S540. Based on the reconstructed frame, the reconstructed video is obtained.

[0116] Example 6

[0117] This application provides a video quality restoration method in Embodiment Six, which optimizes the step in Embodiment Five of "determining the reconstructed frame corresponding to the video frame based on the feature matrix, the video frame, and the depth post-adjustment model." It should be noted that for parts not detailed in this embodiment, please refer to the descriptions in other embodiments. The method includes:

[0118] S610. Perform frame parsing on the original video to obtain video frames, and calculate the content complexity of the video frames.

[0119] S620. Based on the content complexity, the network depth of the target recovery model is adjusted to obtain the depth-adjusted model; wherein, the encoder and decoder in the target recovery model share the same weight set.

[0120] S631. Determine the adjacent frames of the video frame.

[0121] S632. Based on the depth post-adjustment model, perform 3D convolution on the video frame and the adjacent frames to obtain the feature matrix corresponding to the video frame.

[0122] S633A. Process the video frame based on the depth post-adjustment model to obtain the frame hiding state corresponding to the video frame.

[0123] Among them, after obtaining the feature matrix of the video frame, the depth post-tuning model can also be used to process the video frame separately to obtain the frame hidden state of the video frame.

[0124] It should be noted that the frame hidden state is a core variable used by the depth post-tuning model to dynamically store and transmit sequence information when processing video frame sequences. It can capture the dependencies in the sequence that change over time.

[0125] For example, when processing a video sequence of "a person raising their hand":

[0126] The hidden state h1 of the first frame (initial state) contains only the features of the first frame (such as "arm drooping").

[0127] The hidden state h2 in frame 2 integrates the information from frame 1 ("arm hanging down") and the new feature from frame 2 ("arm begins to rise") to form a representation of the "initial stage of the arm raising action".

[0128] The hidden state h3 in frame 3 further integrates h2 and the feature of frame 3 ("arm raised to the middle") to form a representation of "arm raising action in progress";

[0129] In this way, the subsequent hidden states will continuously accumulate the timing information of the actions until the entire process of "raising the hand" is fully captured.

[0130] Therefore, the hidden state of the video frame can be used together with the feature matrix to construct the reconstructed frame of the video frame, and further improve the image optimization quality of the reconstructed frame.

[0131] It should be noted that the frame hidden state is used to reflect the temporal continuity and dependency between adjacent frames. The depth post-tuning model can obtain the frame hidden state H(t) corresponding to the current time t by processing the frame hidden state H(t-1) corresponding to the previous time t-1 and the video frame X(t) corresponding to the current time t.

[0132] S633B. Based on the feature matrix and the frame hiding state, construct the reconstructed frame corresponding to the video frame.

[0133] The depth-adjusted model can output the reconstructed frame corresponding to the video frame by processing the feature matrix and the hidden state of the video frame.

[0134] It should be noted that constructing reconstructed frames using frame hidden states is particularly useful for recovering damage that gradually develops or changes within a video sequence, such as lighting variations, occlusion, and motion blur, because frame hidden states reflect the temporal continuity and dependencies between adjacent frames. Furthermore, it can significantly improve the quality of reconstructed frame recovery, especially in dynamic scenes where an understanding of the relationships between consecutive frames is crucial for accurate video content recovery.

[0135] S640. Based on the reconstructed frame, the reconstructed video is obtained.

[0136] Example 7

[0137] This application provides a video quality restoration method in Embodiment 7, which optimizes the "obtaining a reconstructed video based on the reconstructed frame" step in Embodiment 1. It should be noted that for parts not detailed in this embodiment, please refer to the descriptions in other embodiments. The method includes:

[0138] S710. Perform frame parsing on the original video to obtain video frames, and calculate the content complexity of the video frames.

[0139] S720. Based on the content complexity, the network depth of the target recovery model is adjusted to obtain the depth-adjusted model; wherein, the encoder and decoder in the target recovery model share the same weight set.

[0140] S730. Process the video frame based on the depth adjustment model to obtain the reconstructed frame.

[0141] S741. Perform color correction on the reconstructed frame to obtain a corrected frame.

[0142] It should be noted that the initially reconstructed frames may exhibit color distortion due to imperfections in some model parameters of the depth-adjusted model. Therefore, color correction is required for the reconstructed frames. The depth-adjusted model includes a color correction function, which is used to correct the colors of the reconstructed frames. The new video frame obtained after color correction is denoted as the corrected frame Y_corrected(t), where:

[0143]

[0144] Here, color_correction() is the color correction function, and color_params are the color correction parameters. For example, color correction parameters include: color temperature adjustment parameters, saturation enhancement parameters, contrast adjustment parameters, etc., which depend on the requirements of the original video frame and the visual effect of the target output.

[0145] It should be noted that color correction not only improves visual effects, but also helps to solve color distortion problems caused by factors such as camera settings and changes in ambient light, making the video more suitable for viewing and analysis.

[0146] S742. Encode the corrected frame to obtain the reconstructed video.

[0147] In this process, by encoding the correction frames corresponding to each video frame, the correction frames can be combined in time sequence to form a video with optimized quality corresponding to the original video, which is called video reconstruction.

[0148] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0149] The implementation of the above embodiments has the following specific effects:

[0150] 1. Significantly Reduced Resource Consumption: Traditional deep learning models typically require substantial computational resources for video image restoration, limiting their application in resource-constrained environments such as mobile devices and embedded systems. Through a dynamic weight sharing mechanism, this invention significantly reduces the number of model parameters, thereby reducing the demand for computational resources. This not only reduces memory usage and processor load during model runtime but also enables the model to run efficiently on lower-power devices, expanding its application scope.

[0151] 2. Improved Model Versatility and Flexibility: Existing models are typically designed for specific types of image degradation and lack the ability to handle multiple degradation types. The dynamic network depth adjustment of this invention can automatically adjust the model structure according to the characteristics of the input video, enabling a single model to effectively address various types of image degradation problems. This flexibility greatly enhances the model's practicality and reduces the need to develop and maintain multiple models for different degradation types.

[0152] 3. Enhanced Real-Time Processing Capabilities: Existing technologies, while ensuring high recovery quality, often fail to meet the speed requirements of real-time video processing. By introducing model pruning and efficient pre-training strategies, this invention significantly improves processing speed and reduces latency. This makes this invention highly suitable for applications such as real-time video surveillance and real-time communication, providing rapid response without sacrificing video quality.

[0153] 4. Reduced Training and Maintenance Costs: Existing deep learning models require large amounts of training data and long training cycles, resulting in high training and subsequent maintenance costs. This invention, through efficient pre-training and weight sharing strategies, not only accelerates model training but also optimizes data utilization efficiency during training, significantly reducing training and maintenance costs.

[0154] In summary, the technical solution of this invention directly solves the problems of high resource consumption, lack of flexibility, insufficient real-time processing capability, and high cost in the prior art through innovative dynamic weight sharing and network depth adjustment methods.

[0155] Example 8

[0156] Based on the same inventive concept, this embodiment also provides a video quality restoration apparatus for implementing the video quality restoration method described above. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations in one or more video quality restoration apparatus embodiments provided below can be found in the limitations of the video quality restoration method described above, and will not be repeated here.

[0157] In this embodiment, as Figure 2 As shown, a video quality restoration device is provided, comprising:

[0158] The parsing and calculation module is used to perform frame parsing on the original video to obtain video frames and calculate the content complexity of the video frames.

[0159] The depth adjustment module is used to perform network depth adjustment on the target recovery model based on the content complexity to obtain a depth-adjusted model; wherein the encoder and decoder in the target recovery model share the same weight set;

[0160] The frame reconstruction module is used to process the video frames based on the depth post-adjustment model to obtain reconstructed frames;

[0161] The video reconstruction module is used to obtain the reconstructed video based on the reconstructed frames.

[0162] Each module in the aforementioned video quality restoration device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.

[0163] It should be noted that this embodiment obtains video frames by parsing the original video and calculating the content complexity of the video frames; based on the content complexity, the network depth of the target restoration model is adjusted to obtain a depth-adjusted model; wherein, the encoder and decoder in the target restoration model share the same weight set; the video frames are processed based on the depth-adjusted model to obtain reconstructed frames; and based on the reconstructed frames, a reconstructed video is obtained. Through the above implementation, since the encoder and decoder in the target restoration model use the same set of weights, when processing the input video frame sequence, whether in the encoding or decoding stage, the target restoration model uses the same parameters for learning and generation. This sharing mechanism not only reduces the complexity of the model, significantly reducing the demand for computing resources, but also helps the target restoration model generalize better because all video frames are processed through the same weights, thereby improving the robustness of the model when processing unseen data; in addition, by dynamically adjusting the network depth of the target restoration model in real time according to the content complexity of the video frames, the target restoration model can have a higher processing speed when the content complexity is high, thereby effectively improving the processing speed and flexibility of the target restoration model when processing video frames.

[0164] In an optional embodiment, regarding the training of the target recovery model, the depth adjustment module is specifically used for:

[0165] The model loss is calculated based on the training frames, the original recovery model, and the dynamic shared weights corresponding to the training frames.

[0166] The original recovery model is optimized based on the model loss to obtain an intermediate recovery model;

[0167] Based on the preset pruning rate, the training frames, the corresponding real frames, and the dynamic shared weights, the intermediate recovery model is adjusted to obtain the target recovery model.

[0168] In an optional embodiment, regarding adjusting the intermediate recovery model based on a preset pruning rate, the training frames, the corresponding real frames, and the dynamic shared weights to obtain the target recovery model, the depth adjustment module is specifically used for:

[0169] The intermediate recovery model is pruned based on a preset pruning rate to obtain a pruning recovery model;

[0170] Based on the training frames, the corresponding real frames, and the dynamic shared weights, the model adjustment weights are calculated.

[0171] The pruning recovery model is adjusted based on the model weights to obtain the target recovery model.

[0172] In an optional embodiment, the parsing calculation module is specifically used for calculating the content complexity of the video frame as follows:

[0173] The video frames are resized based on a preset frame input size to obtain standard video frames;

[0174] The standard video frames are normalized based on the statistics corresponding to the preset training frame set to obtain normalized frames.

[0175] The normalized frame is parsed to obtain the frame content, which includes at least one of the following: the number of objects, the object movement speed, and color change information.

[0176] The frame content is processed based on a preset complexity calculation formula to obtain the content complexity.

[0177] In an optional embodiment, in processing the video frame based on the depth post-adjustment model to obtain the reconstructed frame, the frame graph reconstruction module is specifically used for:

[0178] Determine the adjacent frames of the video frame;

[0179] Based on the depth post-adjustment model, 3D convolution is performed on the video frame and the adjacent frames to obtain the feature matrix corresponding to the video frame;

[0180] Based on the feature matrix, the video frame, and the depth post-adjustment model, the reconstructed frame corresponding to the video frame is determined.

[0181] In an optional embodiment, regarding determining the reconstructed frame corresponding to the video frame based on the feature matrix, the video frame, and the depth post-adjustment model, the frame graph reconstruction module is specifically used for:

[0182] The video frame is processed based on the depth post-adjustment model to obtain the frame hidden state corresponding to the video frame;

[0183] Based on the feature matrix and the hidden state of the frame, a reconstructed frame corresponding to the video frame is constructed.

[0184] In an optional embodiment, in obtaining the reconstructed video based on the reconstructed frame, the video reconstruction module is specifically used for:

[0185] The reconstructed frame is color-corrected to obtain a corrected frame;

[0186] The corrected frame is encoded to obtain the reconstructed video.

[0187] Example 9

[0188] In this embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows. Figure 3 As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The database stores data. The network interface communicates with external terminals via a network connection. When executed by the processor, the computer program implements a video quality restoration method.

[0189] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present disclosure and does not constitute a limitation on the computer device to which the present disclosure is applied. A specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0190] Example 10

[0191] In this embodiment, a computer-readable storage medium is provided, such as... Figure 4 As shown, a computer program is stored thereon, and when the computer program is executed by the processor, it implements the steps in the above-described method embodiments.

[0192] Example 11

[0193] In this embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0194] It should be noted that the information collected is information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data all comply with the relevant laws, regulations and standards of the relevant countries and regions, necessary confidentiality measures have been taken, and it does not violate public order and good morals. Corresponding operation portals are provided for users to choose to authorize or refuse.

[0195] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this disclosure can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this disclosure may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this disclosure may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0196] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0197] The embodiments described above are merely illustrative of several implementations of this disclosure, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent disclosure. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this disclosure, and these all fall within the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be determined by the appended claims.

Claims

1. A video quality restoration method, characterized in that, include: Frame parsing is performed on the original video to obtain video frames, and the content complexity of the video frames is calculated. Based on the aforementioned content complexity, the network depth of the target recovery model is adjusted to obtain a depth-adjusted model; wherein, the encoder and decoder in the target recovery model share the same weight set; The video frames are processed based on the depth post-adjustment model to obtain reconstructed frames; Based on the reconstructed frames, the reconstructed video is obtained.

2. The method according to claim 1, characterized in that, The training steps of the target recovery model include: The model loss is calculated based on the training frames, the original recovery model, and the dynamic shared weights corresponding to the training frames. The original recovery model is optimized based on the model loss to obtain an intermediate recovery model; Based on the preset pruning rate, the training frames, the corresponding real frames, and the dynamic shared weights, the intermediate recovery model is adjusted to obtain the target recovery model.

3. The method according to claim 2, characterized in that, The process of adjusting the intermediate recovery model based on a preset pruning rate, the training frames, the corresponding real frames, and the dynamic shared weights to obtain the target recovery model includes: The intermediate recovery model is pruned based on a preset pruning rate to obtain a pruning recovery model; Based on the training frames, the corresponding real frames, and the dynamic shared weights, the model adjustment weights are calculated. The pruning recovery model is adjusted based on the model weights to obtain the target recovery model.

4. The method according to claim 1, characterized in that, The calculation of the content complexity of the video frame includes: The video frames are resized based on a preset frame input size to obtain standard video frames; The standard video frames are normalized based on the statistics corresponding to the preset training frame set to obtain normalized frames. The normalized frame is parsed to obtain the frame content, which includes at least one of the following: the number of objects, the object movement speed, and color change information. The frame content is processed based on a preset complexity calculation formula to obtain the content complexity.

5. The method according to claim 1, characterized in that, The process of processing the video frame based on the depth post-adjustment model to obtain the reconstructed frame includes: Determine the adjacent frames of the video frame; Based on the depth post-adjustment model, 3D convolution is performed on the video frame and the adjacent frames to obtain the feature matrix corresponding to the video frame; Based on the feature matrix, the video frame, and the depth post-adjustment model, the reconstructed frame corresponding to the video frame is determined.

6. The method according to claim 5, characterized in that, The step of determining the reconstructed frame corresponding to the video frame based on the feature matrix, the video frame, and the depth post-adjustment model includes: The video frame is processed based on the depth post-adjustment model to obtain the frame hidden state corresponding to the video frame; Based on the feature matrix and the hidden state of the frame, a reconstructed frame corresponding to the video frame is constructed.

7. The method according to claim 1, characterized in that, The process of obtaining the reconstructed video based on the reconstructed frame includes: The reconstructed frame is color-corrected to obtain a corrected frame; The corrected frame is encoded to obtain the reconstructed video.

8. A video quality restoration device, characterized in that, The device includes: The parsing and calculation module is used to perform frame parsing on the original video to obtain video frames and calculate the content complexity of the video frames. The depth adjustment module is used to perform network depth adjustment on the target recovery model based on the content complexity to obtain a depth-adjusted model; wherein the encoder and decoder in the target recovery model share the same weight set; The frame reconstruction module is used to process the video frames based on the depth post-adjustment model to obtain reconstructed frames; The video reconstruction module is used to obtain the reconstructed video based on the reconstructed frames.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Network depth dynamic adjustment method and device, computer equipment and storage medium

    CN114997373A

  • Video stream acquisition method based on deep learning

    CN119299703A