Model training method, video enhancement method, device, equipment and program product
By performing feature extraction and loss value training on the video enhancement model, the problems of low inter-frame consistency and inter-frame jitter are solved, and a more stable video enhancement effect is achieved.
Patent Information
- Application Number
- CN202411998236.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-06
AI Technical Summary
The inter-frame consistency of video enhancement results is low and there is a problem of inter-frame jitter in video.
By obtaining the pre-constructed initial model and sample video frame sequence, performing noise processing and feature extraction, generating a predicted enhanced frame sequence, and training the initial model according to the loss value, a video enhanced model is obtained.
It effectively avoids the uncertainty of frame enhancement content caused by the uncertainty generated by the diffusion model, reduces jitter between video frames, strengthens the inter-frame consistency of video enhancement results, and improves the robustness of video enhancement tasks.
Smart Images

Figure CN119941556A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of video processing technology, and in particular to a model training method, a video enhancement method, a model training device, a video enhancement device, an electronic device and a computer program product. Background Art
[0002] Video enhancement refers to the process of repairing missing or damaged parts of a video and restoring the complete visual content through algorithms or manual means. With the richness and diversification of video content, video enhancement technology has become increasingly important in many practical applications, especially in film and television production, surveillance systems, virtual reality, and autonomous driving. With the development of deep learning technology, the effects and application scenarios of video enhancement have also been greatly improved.
[0003] Video enhancement, as a multidisciplinary research field, originated from image enhancement. Image enhancement was originally aimed at filling in the missing parts of static images, with the goal of restoring the visual content of the image by utilizing the contextual information of the surrounding area. Traditional image enhancement methods include texture synthesis, filling algorithms, and object-based enhancement methods. However, compared with image enhancement, video enhancement faces greater challenges because videos not only involve spatial information, but also temporal information. The missing parts need to be consistent with the previous and next frames, and not only the spatial structure and texture details must be considered, but also the natural transition and temporal consistency of moving objects must be guaranteed. Therefore, the research on video enhancement is much more complicated than static image enhancement, involving more computing resources and technical methods. Summary of the invention
[0004] The present disclosure provides a model training method, a video enhancement method, a model training device, a video enhancement device, an electronic device, a computer-readable storage medium, and a computer program product, so as to at least solve the problem that the inter-frame consistency of the video enhancement result is low and there is jitter between video frames in the related art. The technical solution of the present disclosure is as follows:
[0005] According to a first aspect of the present disclosure, a model training method is provided, comprising: obtaining a pre-constructed initial model and a sample video frame sequence, wherein the sample video frame sequence comprises a degraded video frame sequence and a reference video frame sequence; performing noise processing on the reference video frame sequence to obtain a noisy reference frame sequence, and determining noisy reference frame features and noisy reference sequence timing information corresponding to the noisy reference frame sequence; determining degraded frame features and degraded sequence timing information corresponding to the degraded video frame sequence; generating a predicted enhanced frame sequence by the initial model based on the degraded frame features, the degraded sequence timing information, the noisy reference frame features and the noisy reference sequence timing information; and training the initial model according to a loss value between the predicted enhanced frame sequence and the reference video frame sequence to obtain a video enhancement model.
[0006] In an exemplary embodiment of the present disclosure, the noisy reference frame sequence includes multiple noisy reference frames, and the noisy processing of the reference video frame sequence to obtain the noisy reference frame sequence includes: encoding the reference video frame sequence to obtain an encoded reference frame sequence; and noisy processing each encoded reference frame in the encoded reference frame sequence to obtain a noisy reference frame corresponding to each encoded reference frame.
[0007] In an exemplary embodiment of the present disclosure, the noisy reference frame sequence includes multiple noisy reference frames, and determining the noisy reference frame features and noisy reference sequence timing information corresponding to the noisy reference frame sequence includes: performing image feature extraction processing on the noisy reference frame to obtain the noisy reference frame features; acquiring the current noisy reference frame features corresponding to the noisy reference frame; updating the current noisy reference frame features according to adjacent noisy reference frames of the noisy reference frame to obtain a first noisy reference frame feature; and updating the first noisy reference frame features according to the noisy reference frame sequence to obtain a second noisy reference frame feature.
[0008] In an exemplary embodiment of the present disclosure, the initial model includes a temporal information extraction structure, the temporal information extraction structure includes a sparse attention module, and the updating operation is performed on the current noisy reference frame feature according to the adjacent noisy reference frame of the noisy reference frame to obtain the first noisy reference frame feature, including: determining the first focus point vector corresponding to the noisy reference frame according to the current noisy reference frame feature; determining the first information point vector and the first detailed content vector corresponding to the noisy reference frame according to the adjacent noisy reference frame features of the adjacent noisy reference frames by the sparse attention module; and updating the current noisy reference frame feature according to the first focus point vector, the first information point vector and the first detailed content vector to obtain the first noisy reference frame feature.
[0009] In an exemplary embodiment of the present disclosure, the initial model includes a temporal information extraction structure, the temporal information extraction structure includes a temporal attention module, and the updating operation is performed on the first noisy reference frame feature according to the noisy reference frame sequence to obtain a second noisy reference frame feature, including: determining a second focus point vector corresponding to the noisy reference frame according to the first noisy reference frame feature; determining, by the temporal attention module, a second information point vector and a second detailed content vector corresponding to the noisy reference frame according to a plurality of noisy reference frame features included in the noisy reference frame sequence; and updating the first noisy reference frame feature according to the second focus point vector, the second information point vector and the second detailed content vector to obtain the second noisy reference frame feature.
[0010] In an exemplary embodiment of the present disclosure, the initial model generates a predicted enhanced frame sequence based on the degraded frame features, the degraded sequence timing information, the noisy reference frame features and the noisy reference sequence timing information, including: determining the current degraded frame features and the current degraded sequence timing information corresponding to the current degradation round; using the current degraded frame features and the current degraded sequence timing information as feature generation constraints for the next video enhancement round; generating multiple predicted video frame latent vectors based on the feature generation constraints, the noisy reference frame features and the noisy reference sequence timing information; and decoding the multiple predicted video frame latent vectors to generate the predicted enhanced frame sequence.
[0011] In an exemplary embodiment of the present disclosure, the decoding processing of the multiple predicted video frame latent vectors to generate the predicted enhanced frame sequence includes: obtaining a variational codec in the initial model, the variational codec including a temporal residual structure; and decoding processing of the multiple predicted video frame latent vectors through the temporal residual structure to generate the predicted enhanced frame sequence.
[0012] In an exemplary embodiment of the present disclosure, the initial model is trained according to the loss value between the predicted enhanced frame sequence and the reference video frame sequence to obtain the video enhancement model, including: determining the frame sequence loss value between the predicted enhanced frame sequence and the reference video frame sequence; based on the frame sequence loss value, adjusting the model parameters of the timing information extraction structure in the initial model until the model training end condition is reached to obtain the video enhancement model.
[0013] According to a second aspect of an embodiment of the present disclosure, a video enhancement method is provided, comprising: obtaining an initial degraded video; obtaining a pre-trained video enhancement model, wherein the video enhancement model is trained based on a model training method; inputting the initial degraded video into the video enhancement model, and performing video enhancement processing on the initial degraded video by the video enhancement model to obtain a target enhanced video.
[0014] According to a third aspect of the present disclosure, a model training device is provided, comprising: a sample frame sequence acquisition module, used to acquire a pre-constructed initial model and a sample video frame sequence, wherein the sample video frame sequence comprises a degraded video frame sequence and a reference video frame sequence; a noisy frame feature determination module, used to perform degradation processing on the reference video frame sequence to obtain a noisy reference frame sequence, and determine the noisy reference frame features and the noisy reference sequence timing information corresponding to the noisy reference frame sequence; a degraded frame feature determination module, used to determine the degraded frame features and the degraded sequence timing information corresponding to the degraded video frame sequence; a predicted frame sequence generation module, used to generate a predicted enhanced frame sequence based on the degraded frame features, the degraded sequence timing information, the noisy reference frame features and the noisy reference sequence timing information by the initial model; and a model training module, used to train the initial model according to the loss value between the predicted enhanced frame sequence and the reference video frame sequence to obtain a video enhancement model.
[0015] In an exemplary embodiment of the present disclosure, the noisy reference frame sequence includes multiple noisy reference frames, and the noisy frame feature determination module includes a noisy frame generation unit, which is used to: encode the reference video frame sequence to obtain a coded reference frame sequence; and perform noise processing on each coded reference frame in the coded reference frame sequence to obtain a noisy reference frame corresponding to each coded reference frame.
[0016] In an exemplary embodiment of the present disclosure, the noisy reference frame sequence includes multiple noisy reference frames, and the noisy frame feature determination module includes a noisy frame feature determination unit, which is used to: perform image feature extraction processing on the noisy reference frame to obtain the noisy reference frame feature; obtain the current noisy reference frame feature corresponding to the noisy reference frame; update the current noisy reference frame feature according to the adjacent noisy reference frames of the noisy reference frame to obtain the first noisy reference frame feature; and update the first noisy reference frame feature according to the noisy reference frame sequence to obtain the second noisy reference frame feature.
[0017] In an exemplary embodiment of the present disclosure, the initial model includes a temporal information extraction structure, the temporal information extraction structure includes a sparse attention module, and the noisy frame feature determination unit includes a first feature updating subunit, which is used to: determine a first focus point vector corresponding to the noisy reference frame based on the current noisy reference frame feature; determine, by the sparse attention module, a first information point vector and a first detailed content vector corresponding to the noisy reference frame based on the adjacent noisy reference frame features of the adjacent noisy reference frame; and update the current noisy reference frame feature based on the first focus point vector, the first information point vector and the first detailed content vector to obtain the first noisy reference frame feature.
[0018] In an exemplary embodiment of the present disclosure, the initial model includes a temporal information extraction structure, the temporal information extraction structure includes a temporal attention module, and the noisy frame feature determination unit includes a second feature updating subunit, which: determines a second focus point vector corresponding to the noisy reference frame according to the first noisy reference frame feature; determines, by the temporal attention module, a second information point vector and a second detailed content vector corresponding to the noisy reference frame according to a plurality of noisy reference frame features contained in the noisy reference frame sequence; and performs an update operation on the first noisy reference frame feature according to the second focus point vector, the second information point vector and the second detailed content vector to obtain the second noisy reference frame feature.
[0019] In an exemplary embodiment of the present disclosure, the predicted frame sequence generation module includes a predicted frame sequence generation unit, which is used to: determine the current degraded frame features and the current degraded sequence timing information corresponding to the current degradation round; use the current degraded frame features and the current degraded sequence timing information as feature generation constraints for the next video enhancement round; generate multiple predicted video frame latent vectors based on the feature generation constraints, the noisy reference frame features and the noisy reference sequence timing information; and decode the multiple predicted video frame latent vectors to generate the predicted enhanced frame sequence.
[0020] In an exemplary embodiment of the present disclosure, the predicted frame sequence generation unit includes a predicted frame sequence generation sub-unit, which is used to: obtain the variational codec in the initial model, the variational codec includes a temporal residual structure; through the temporal residual structure, decode the multiple predicted video frame latent vectors to generate the predicted enhanced frame sequence.
[0021] In an exemplary embodiment of the present disclosure, the model training module includes a model training unit, which is used to: determine a frame sequence loss value between the predicted enhanced frame sequence and the reference video frame sequence; based on the frame sequence loss value, adjust the model parameters of the timing information extraction structure in the initial model until the model training end conditions are reached, thereby obtaining the video enhancement model.
[0022] According to a fourth aspect of the present disclosure, a video enhancement device is provided, comprising: a degraded video acquisition module, used to acquire an initial degraded video; a model acquisition module, used to acquire a pre-trained video enhancement model, wherein the video enhancement model is trained based on a model training method; and a video enhancement module, used to input the initial degraded video into the video enhancement model, and the video enhancement model performs video enhancement processing on the initial degraded video to obtain a target enhanced video.
[0023] According to a fifth aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor executable instructions; wherein the processor is configured to execute instructions to implement any one of the above-described model training methods, or to implement the above-described video enhancement method.
[0024] According to a sixth aspect of the present disclosure, a computer-readable storage medium is provided. When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute any one of the above-mentioned model training methods, or implement the above-mentioned video enhancement method.
[0025] According to a seventh aspect of the present disclosure, a computer program product is provided, including a computer program, which, when executed by a processor, implements any one of the above-mentioned model training methods or the above-mentioned video enhancement method.
[0026] The technical solution provided by the embodiments of the present disclosure brings at least the following beneficial effects:
[0027] On the one hand, by controlling the generation of prediction enhancement frame sequences through the degraded frame features of the degraded video frames, the uncertainty of frame enhancement content caused by the uncertainty generated by the diffusion model can be avoided, and the jitter between video frames can be reduced. On the other hand, based on the extracted timing information between the degraded video frames, a high-quality prediction enhancement frame sequence can be generated, which can strengthen the inter-frame consistency of the video enhancement results, improve the robustness of the video enhancement task, and obtain more stable video enhancement results.
[0028] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute improper limitations on the present disclosure.
[0030] Figure 1 The figure is a flowchart of a model training method according to an exemplary embodiment.
[0031] Figure 2 is a network architecture diagram of an image enhancement model according to an exemplary embodiment.
[0032] Figure 3 is a network architecture diagram of a timing information extraction structure according to an exemplary embodiment.
[0033] Figure 4 is a network architecture diagram of a decoder introducing a temporal residual structure according to an exemplary embodiment.
[0034] Figure 5 The figure is a flowchart of a video enhancement method according to an exemplary embodiment.
[0035] Figure 6 It is a block diagram of a model training device according to an exemplary embodiment.
[0036] Figure 7 The figure is a block diagram of a video enhancement device according to an exemplary embodiment.
[0037] Figure 8 A block diagram of an electronic device according to an exemplary embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION
[0038] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings.
[0039] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0040] Video enhancement is an important research direction in the field of computer vision, which aims to enhance missing or damaged areas in videos through automated algorithms. Traditional video enhancement methods mostly rely on manually designed features and rules. In recent years, deep learning, especially the application of diffusion models and generative adversarial networks (GAN), has greatly improved the effect and quality of video enhancement. The enhancement effect based on video diffusion models mostly relies on image enhancement diffusion models. Due to the uncertainty of the content generated by the diffusion model, the effect of single-frame video enhancement is inconsistent between frames, and the enhanced video will have jitter problems, resulting in visual unnaturalness.
[0041] Among the related schemes, most representative diffusion model-based video enhancement algorithms strengthen inter-frame consistency by adding a temporal information extraction structure on the basis of the image diffusion model. For example, the temporal consistent diffusion model of real-time video super-resolution (Upscale-A-Video model) utilizes the image enhancement capability of the image super-resolution model SDx4 Upscaler, and adds a temporal attention mechanism and a 3D convolution module to the U-shaped network (UNet network) to strengthen the constraints between frames. The motion-guided latent diffusion (MGLD) model uses a pre-trained optical flow estimation network to estimate the forward and backward optical flows during the diffusion model inference sampling process, and then uses motion-guided loss to update the latent space of different frames.
[0042] The above methods are difficult to obtain high-quality generation results when facing complex scenes in practical applications. This is because, on the one hand, the design based on the image super-resolution model has its inevitable limitations. The degradation of the video is usually diverse, and it is not enough to use only the super-resolution model. On the other hand, when training the model, the original video needs to be downsampled. The downsampling process will cause the video to shake, resulting in a decrease in the consistency of the enhancement results. If the optical flow of the original video is calculated during the sampling process to update the latent space of different frames, it is limited by the accuracy of the optical flow calculation.
[0043] Based on this, according to the embodiments of the present disclosure, a model training method, a video enhancement method, a model training device, a video enhancement device, an electronic device, a computer-readable storage medium, and a computer program product are proposed.
[0044] Figure 1 is a flow chart of a model training method and a video enhancement method according to an exemplary embodiment. Figure 1 As shown, the model training method and the video enhancement method can be used in a computer device, wherein the computer device described in the present disclosure may include mobile terminal devices such as mobile phones, tablet computers, laptop computers, PDAs, and fixed terminal devices such as desktop computers. This exemplary embodiment uses the method applied to a computer device as an example. It can be understood that the method can also be applied to a server, and can also be applied to a system including a computer device and a server, and is implemented through the interaction between the computer device and the server. Specifically, the following steps are included.
[0045] Step S110, obtaining a pre-built initial model and a sample video frame sequence, where the sample video frame sequence includes a degraded video frame sequence and a reference video frame sequence;
[0046] Step S120, performing noise processing on the reference video frame sequence to obtain a noisy reference frame sequence, and determining a noisy reference frame feature corresponding to the noisy reference frame sequence and noisy reference sequence timing information;
[0047] Step S130, determining the degraded frame features and the degraded sequence timing information corresponding to the degraded video frame sequence;
[0048] Step S140, generating a predicted enhanced frame sequence by the initial model based on the degraded frame features, the degraded sequence timing information, the noisy reference frame features and the noisy reference sequence timing information;
[0049] Step S150, training the initial model according to the loss value between the predicted enhanced frame sequence and the reference video frame sequence to obtain a video enhancement model.
[0050] According to the model training method in this example embodiment, on the one hand, the generation of the prediction enhancement frame sequence can be controlled by the degraded frame features of the degraded video frame, which can avoid the uncertainty of the frame enhancement content caused by the uncertainty generated by the diffusion model and reduce the jitter between video frames. On the other hand, based on the extracted timing information between the degraded video frames, a high-quality prediction enhancement frame sequence can be generated, which can enhance the inter-frame consistency of the video enhancement result, improve the robustness of the video enhancement task, and obtain a more stable video enhancement result.
[0051] The model training method in this example embodiment will be further described below.
[0052] In step S110, a pre-built initial model and a sample video frame sequence are obtained, where the sample video frame sequence includes a degraded video frame sequence and a reference video frame sequence.
[0053] In an exemplary embodiment of the present disclosure, a sample video frame sequence may be a video frame sequence for model training, and a sample video frame sequence may be a video frame sequence containing timing information between multiple video frames. For example, a sample video frame sequence may be a video frame sequence generated by sorting multiple video frames contained in a continuous video segment according to timing information. A degraded video frame sequence may be a frame sequence composed of multiple degraded video frames, and a degraded video frame sequence may be a frame sequence corresponding to a degraded video segment generated after the image quality of a certain video segment is reduced due to factors such as image degradation. A reference video frame sequence may be a frame sequence corresponding to a video segment with a higher picture quality. The video contents of the degraded video frame sequence and the reference video frame sequence correspond, that is, the degraded video frame sequence may be generated after the reference video frame sequence is degraded or degraded.
[0054] Obtain a pre-built initial model. The initial model in the present disclosure may be an image enhancement model based on a control network (ControlNet) architecture based on a diffusion in transformer (DiT) model. The initial model may include a DiT category generation model and a ControlNet control branch to implement a video enhancement model that generates high-quality images by controlling low-quality images.
[0055] A sample video frame sequence for model training is obtained, and the sample video frame sequence may include a degraded video frame sequence and a reference video frame sequence. The reference video frame sequence may be a frame sequence corresponding to a video segment whose image quality index value is greater than or equal to a predefined quality index threshold. For example, the reference video frame sequence may be a frame sequence corresponding to a video segment with high-definition picture quality, or a frame sequence corresponding to a video segment whose video resolution is greater than a resolution threshold (such as a resolution of 1080P). In addition, the reference video frame sequence has a high inter-frame consistency, that is, the video segment usually does not have a jitter phenomenon.
[0056] The degraded video frame sequence may be a frame sequence corresponding to a video clip whose image quality index value of the video frame is less than a predefined quality index threshold. For video usage scenarios, during the video formation, recording, processing and transmission process, the video picture quality is degraded due to the imperfections of the video acquisition system, recording equipment, transmission medium and processing method. The frame sequence corresponding to the video clip whose picture quality is degraded can be used as the degraded video frame sequence of the present disclosure. After obtaining the reference video frame sequence, the reference video frame sequence can be degraded or downgraded to obtain the corresponding degraded video frame sequence.
[0057] In step S120, a noise adding process is performed on the reference video frame sequence to obtain a noisy reference frame sequence, and noisy reference frame features corresponding to the noisy reference frame sequence and noisy reference sequence timing information are determined.
[0058] In an exemplary embodiment of the present disclosure, the noisy reference frame sequence may be a frame sequence obtained by adding random noise to each reference video frame in the original reference video frame sequence. The noisy reference frame feature may be a feature contained in each noisy reference frame in the noisy reference frame sequence. The noisy reference sequence timing information may be timing information contained between multiple noisy reference frames in the noisy reference frame sequence.
[0059] For the acquired reference video frame sequence, each reference video frame in the reference video frame sequence is subjected to noise processing one by one to obtain the corresponding noisy reference frame sequence. Then, the pre-built initial model is used to perform feature extraction processing on the noisy reference frame sequence to obtain the noisy reference frame features and noisy reference sequence timing information contained in the noisy reference frame sequence.
[0060] In an exemplary embodiment of the present disclosure, for step S120, the reference video frame sequence is subjected to noise processing to obtain a noisy reference frame sequence, including: encoding the reference video frame sequence to obtain a coded reference frame sequence; and noise processing is performed on each coded reference frame in the coded reference frame sequence to obtain a noisy reference frame corresponding to each coded reference frame.
[0061] The coded reference frame sequence may be a video frame sequence obtained by encoding the reference video frames in the reference video frame sequence one by one, that is, the coded reference frame sequence may be a sequence composed of multiple coded reference frames. The coded reference frame may be a video frame obtained by encoding the reference video frame included in the reference video frame sequence. The noisy reference frame may be a video frame obtained by adding random noise to the coded reference frame.
[0062] refer to Figure 2 , Figure 2 FIG. 1 is a network architecture diagram of an image enhancement model according to an exemplary embodiment. Figure 2 As can be seen from the figure, the input of the initial model can be a video of length n, denoted by x 1 ,x 2 ,…,x n The reference video frame sequence includes n reference video frames, namely, high quality video frames (High Quality image), such as reference video frame 1 (image1), reference video frame 2 (image2), ..., reference video frame n (imagen).
[0063] After obtaining the reference video frame sequence, the variational auto-encoder (VAE) can be used to encode the reference video frame sequence, that is, the VAE encoder is used to encode each reference video frame one by one to obtain an encoded reference frame sequence. The encoded reference frame sequence obtained after encoding by the VAE encoder is z 1 ,z 2 ,…,z n .
[0064] After obtaining the coding reference frame sequence, each coding reference frame in the coding reference frame sequence is subjected to noise addition processing. For different coding reference frame sequences, the noise addition processing can perform noise addition processing on the frame sequence to different degrees, and obtain the noisy reference frames corresponding to each coding reference frame, so as to adapt to various types of video inputs in the video enhancement scenario.
[0065] In an exemplary embodiment of the present disclosure, for step S120, determining the noisy reference frame features and the noisy reference sequence timing information corresponding to the noisy reference frame sequence includes: performing image feature extraction processing on the noisy reference frame to obtain the noisy reference frame features; obtaining the current noisy reference frame features corresponding to the noisy reference frame; updating the current noisy reference frame features according to the adjacent noisy reference frames of the noisy reference frame to obtain the first noisy reference frame features; and updating the first noisy reference frame features according to the noisy reference frame sequence to obtain the second noisy reference frame features.
[0066] The current noisy reference frame feature may be a video frame feature determined by a feature extraction operation on the current round of the noisy reference frame. The adjacent noisy reference frame may be a frame adjacent to the currently processed noisy reference frame, such as the adjacent frames of video frame 2 may be video frame 1 and video frame 3. The first noisy reference frame feature may be a frame feature obtained by performing feature update processing on the current noisy reference frame feature based on the frame feature of the adjacent noisy reference frame.
[0067] Continue to refer Figure 2 After obtaining the noisy reference frame sequence, the image feature extraction process of the noisy reference frame is performed based on the DiT block (DiT-Block) in the initial model to obtain the noisy reference frame feature, which is recorded as f 1 ,f 2 ,…,f n The frame features obtained after the DiT-Block feature extraction process are used as the current noisy reference frame features corresponding to the noisy reference frame.
[0068] One of the reasons why DiT generates jitter when processing a single frame of video is due to the uncertainty generated by the diffusion model, which leads to uncertainty in the content of different frame enhancements, causing jitter in the enhanced video. In order to solve the above problem, a temporal information extraction structure, namely the Motion-Block, is added to the DiT base model and the ControlNet control branch of the initial model. The DiT base model provides the function of generating images by categories. In order to extract temporal information, that is, to enable it to have the ability to generate videos, a temporal information extraction structure is added after each DiT-Block to extract the temporal features in the video clips.
[0069] Specifically, each temporal information extraction structure is composed of a sparse attention mechanism and a temporal attention mechanism module. The sparse attention mechanism first updates the features of the current noisy reference frame according to the adjacent noisy reference frames of the noisy reference frame to obtain the first noisy reference frame features. The first noisy reference frame features are then used as the input of the temporal attention mechanism module, and the temporal attention mechanism module updates the first noisy reference frame features according to the noisy reference frame sequence to obtain the second noisy reference frame features. By updating the features of the current noisy reference frame through the temporal information extraction structure, the temporal information contained between the video frames can be extracted as the basis for subsequent video enhancement processing.
[0070] It should be noted that in some other exemplary embodiments of the present disclosure, other adjacent frames of the noisy reference frame can also be used to update its features. For example, 2 or 3 adjacent frames before and after the noisy reference frame can be selected for use in the feature update process of the noisy reference frame. The specific number of adjacent frames can be configured according to specific needs, and the present disclosure does not make any special limitation on this.
[0071] In an exemplary embodiment of the present disclosure, the initial model includes a temporal information extraction structure, the temporal information extraction structure includes a sparse attention module, and a current noisy reference frame feature is updated according to adjacent noisy reference frames of the noisy reference frame to obtain a first noisy reference frame feature, including: determining a first focus point vector corresponding to the noisy reference frame according to the current noisy reference frame feature; determining, by the sparse attention module, a first information point vector and a first detailed content vector corresponding to the noisy reference frame according to adjacent noisy reference frame features of adjacent noisy reference frames; and updating the current noisy reference frame feature according to the first focus point vector, the first information point vector and the first detailed content vector to obtain a first noisy reference frame feature.
[0072] Among them, the first focus point vector can be a vector extracted directly by the attention mechanism in the timing module based on the current noisy reference frame features, that is, the query vector in the attention mechanism, denoted by Q. The adjacent noisy reference frame features can be related features of the adjacent frames before and after the noisy reference frame. The first information point vector can be one of the vectors extracted by the attention mechanism in the timing module based on the adjacent noisy reference frame features, that is, the key vector in the attention mechanism, denoted by K. The first detailed content vector can be another vector extracted by the attention mechanism in the timing module based on the adjacent noisy reference frame features, that is, the value vector in the attention mechanism, denoted by V.
[0073] refer to Figure 3 , Figure 3: This is a network architecture diagram of a temporal information extraction structure according to an exemplary embodiment. After DiT-Block extracts features from the noisy reference frame and obtains the features of the current noisy reference frame corresponding to the current round, the features of the current noisy reference frame are input into the temporal module (Motion-Block). The sparse attention module (Sparse Attention) in Motion-Block first performs vector encoding on it to obtain the first attention point vector corresponding to the noisy reference frame, which is recorded as Q i In addition, the adjacent noisy reference frame features of the adjacent noisy reference frame are obtained, and the sparse attention module performs vector encoding on the adjacent noisy reference frame features to obtain the first information point vector K corresponding to the noisy reference frame i With the first detailed content vector V i .
[0074] For example, the feature of the currently processed noisy reference frame is f i , to update f i For example, the query, key, and value in the attention mechanism are Q, K, and V respectively. i ,f i-1 ,f i+1 Input to the sparse attention module, where Q i By i It is calculated that K i ,V i By i-1 ,f i+1 Then, the first focus point vector Q is determined based on i , the first information point vector K i With the first detailed content vector V i , the feature of the current noisy reference frame is updated to obtain the first noisy reference frame feature, and the specific updating process is shown in Formula 1. In particular, if the current frame is the first frame or the last frame, the feature vector of the missing adjacent frame can be set to 0.
[0075]
[0076] Among them, f i ' can represent the first noisy reference frame feature; Q i Can represent the first focus vector; K i Can represent the first information point vector; V i may represent a first detailed content vector; d may represent a constant; and softmax() may represent a normalized exponential function.
[0077] The sparse attention module of the timing module can update the features of all frames in the noisy reference frame sequence in sequence according to the above formula 1. The attention mechanism of the timing module in the present disclosure is mainly to extract relevant information from other frames, strengthen the constraints between different frames, and improve consistency. The sparse attention module mainly considers the relationship between the current frame and the previous and next frames, so as to apply the learned relationship between the previous and next frames to the video enhancement process.
[0078] In an exemplary embodiment of the present disclosure, the initial model includes a temporal information extraction structure, the temporal information extraction structure includes a temporal attention module, and a first noisy reference frame feature is updated according to a noisy reference frame sequence to obtain a second noisy reference frame feature, including: determining a second focus point vector corresponding to the noisy reference frame according to the first noisy reference frame feature; determining, by the temporal attention module, a second information point vector and a second detailed content vector corresponding to the noisy reference frame according to a plurality of noisy reference frame features included in the noisy reference frame sequence; and updating the first noisy reference frame feature according to the second focus point vector, the second information point vector and the second detailed content vector to obtain a second noisy reference frame feature.
[0079] Among them, the second focus point vector can be a vector extracted directly by the attention mechanism in the timing module based on the updated first noisy reference frame features. The multiple noisy reference frame features can be all frames in the noisy reference frame sequence, or a specified part of the frames. The second information point vector can be one of the vectors extracted by the attention mechanism in the timing module based on the multiple noisy reference frame features. The second detailed content vector can be another vector extracted by the attention mechanism in the timing module based on the multiple noisy reference frame features.
[0080] Continue to refer Figure 3 After obtaining the first noisy reference frame feature, the first noisy reference frame feature is sent to the next layer of temporal attention module (Temporal Attention). The temporal attention module mainly considers the relationship between the current frame and all frames. For the temporal attention module, the first noisy reference frame feature f obtained in the previous layer (sparse attention module) is sent to the next layer of temporal attention module. 1 ′,…,f n ′ as input, for the noisy reference frame currently being processed, the second focus vector Q of the frame i ′ is composed of the first noisy reference frame feature f i 'Calculated, the second information point vector K i ′ and the second detailed content vector V i ′ can be obtained by the first noise-added reference frame feature f of all frames in the noise-added reference frame sequence 1 ′,…,f n Then the second focus point vector Q is determined based on i', the second information point vector K i ′ and the second detailed content vector V i ′, update the first noisy reference frame feature to obtain the second noisy reference frame feature. The specific updating process is shown in Formula 2.
[0081]
[0082] Among them, f i ″ can represent the second noisy reference frame feature; Q i ′ Can represent the first focus vector; K i ' can represent the second information point vector; V i ′ may represent a second detailed content vector; d may represent a constant; and softmax() may represent a normalized exponential function.
[0083] The temporal attention module in the timing module can update the features of all frames in the noisy reference frame sequence in sequence based on the above formula 2. The temporal attention module mainly considers the relationship between the current frame and all frames, and takes all frame features in the noisy reference frame sequence as the basis for feature updating.
[0084] In step S130, degraded frame features and degraded sequence timing information corresponding to the degraded video frame sequence are determined.
[0085] In an exemplary embodiment of the present disclosure, the degraded frame feature may be a feature included in the degraded video frame. The degraded sequence timing information may be timing information included between multiple degraded video frames in the degraded video frame sequence.
[0086] Continue to refer Figure 2 , Figure 2 The skeleton model of ControlNet is used as the controller of the DiT-based model, and the degraded video frame sequence (LQ video) is used as the model input to realize the control of generating high-quality video through low-quality video. In order to extract the timing information of the low-quality video and make the generated high-quality video consistent with the low-quality video in timing, Motion-Block is also introduced after each DiT-Block. Its calculation process is similar to that of the DiT-based model, and this disclosure will not repeat it. Through Motion-Block, the degraded frame features and degraded sequence timing information corresponding to the degraded video frame sequence can be obtained.
[0087] In step S140, the initial model generates a predicted enhanced frame sequence based on the degraded frame features, the degraded sequence timing information, the noisy reference frame features and the noisy reference sequence timing information.
[0088] In an exemplary embodiment of the present disclosure, the predicted enhanced frame sequence may be a video enhancement result predicted after the initial model performs video enhancement processing.
[0089] After the initial model determines the degraded frame features, degraded sequence timing information, noisy reference frame features and noisy reference sequence timing information, the above information can be used as the basis for video enhancement processing, and the predicted enhanced frame sequence is generated through the network structure of the initial model.
[0090] In an exemplary embodiment of the present disclosure, for step S140, the initial model generates a predicted enhanced frame sequence based on the degraded frame features, the degraded sequence timing information, the noisy reference frame features and the noisy reference sequence timing information, including: determining the current degraded frame features and the current degraded sequence timing information corresponding to the current degradation round; using the current degraded frame features and the current degraded sequence timing information as feature generation constraints for the next video enhancement round; generating multiple predicted video frame latent vectors based on the feature generation constraints, the noisy reference frame features and the noisy reference sequence timing information; and decoding the multiple predicted video frame latent vectors to generate a predicted enhanced frame sequence.
[0091] The current degraded frame feature may be an image feature corresponding to a degraded video frame in a current feature extraction round, and the current degraded sequence timing information may be timing information included between multiple degraded video frames in a current feature extraction round.
[0092] Continue to refer Figure 2 ,from Figure 2 It can be seen that both the ControlNet network branch and the DiT base model introduce Motion-Block, and the number of Motion-Blocks can be multiple, such as 32, etc. In some other exemplary embodiments of the present disclosure, the number of Motion-Blocks can also be configured according to the expected performance requirements of the model training, and the present disclosure does not make any special limitations on this.
[0093] For the processing of each Motion-Block, the processing of the current Motion-Block in the ControlNet network branch can be considered as the current degradation round, and the current degradation frame features and the current degradation sequence timing information corresponding to the current degradation round are determined based on the ControlNet network branch. Then, the determined current degradation frame features and the current degradation sequence timing information are used as the feature generation constraint conditions for the next video enhancement round of the DiT base model, that is, as one of the inputs of the next Motion-Block in the DiT base model.
[0094] The initial model can perform video enhancement processing based on feature generation constraints, noisy reference frame features, and noisy reference sequence timing information, and finally output multiple predicted video frame latent vectors. The obtained predicted video frame latent vectors need to be decoded to generate video clips that can be used for display, that is, to obtain predicted enhanced frame sequences. The initial model can use the learned relevant frame features and timing information of degraded video frames in the video enhancement process to improve the video enhancement effect.
[0095] In an exemplary embodiment of the present disclosure, a plurality of predicted video frame latent vectors are decoded to generate a predicted enhanced frame sequence, including: obtaining a variational codec in an initial model, the variational codec including a temporal residual structure; and decoding a plurality of predicted video frame latent vectors through the temporal residual structure to generate a predicted enhanced frame sequence.
[0096] The temporal residual structure may be a network result for enhancing temporal consistency between video frames. The predicted video frame latent vector may be a video latent vector outputted by the category generation model in the initial model after performing video enhancement processing.
[0097] The obtained predicted video frame latent vector is input into the VAE decoder, which decodes it. Another reason for video jitter when the image enhancement model processes a single frame of video is that the VAE decoder based only on the image is still prone to flickering artifacts when decoding the latent sequence. In order to solve the above problem, an additional temporal residual structure (temporal 3D residual block) is introduced in the VAE decoder of the present disclosure to enhance the low-level temporal consistency.
[0098] refer to Figure 4 , Figure 4 is a network architecture diagram of a decoder introducing a temporal residual structure according to an exemplary embodiment. Figure 4 The temporal residual structure in the VAE decoder contains two groups of 2D residual blocks (ResBlock2D) and 1D residual blocks (ResBlock1D). After decoding multiple predicted video frame latent vectors through the temporal residual structure in the VAE decoder, a predicted enhanced frame sequence can be obtained.
[0099] In step S150, the initial model is trained according to the loss value between the predicted enhanced frame sequence and the reference video frame sequence to obtain a video enhancement model.
[0100] In an exemplary embodiment of the present disclosure, the video enhancement model may be used as a network model for processing video enhancement tasks.
[0101] After obtaining the predicted enhanced frame sequence, the predicted enhanced frame sequence can be compared with the reference video frame sequence in the training sample set, and the initial model can be trained according to the loss value between the two until the model training end condition is met to obtain a video enhancement model.
[0102] In an exemplary embodiment of the present disclosure, for step S150, the initial model is trained according to the loss value between the predicted enhanced frame sequence and the reference video frame sequence to obtain a video enhancement model, including: determining the frame sequence loss value between the predicted enhanced frame sequence and the reference video frame sequence; based on the frame sequence loss value, adjusting the model parameters of the timing information extraction structure in the initial model until the model training end conditions are reached to obtain the video enhancement model.
[0103] The frame sequence loss value may be a loss value between the predicted enhanced frame sequence and the reference video frame sequence. The model training end condition may be that the current model is in a convergence state, or that the current training times have reached a predefined training round threshold.
[0104] The loss value between the predicted enhanced frame sequence and the reference video frame sequence is determined as the frame sequence loss value. Then, the model parameters of the temporal information extraction structure in the initial model are adjusted according to the frame sequence loss value, that is, when fine-tuning the model parameters, the pre-trained image enhancement model is kept unchanged, and only the temporal module in the newly added temporal layer is trained. These temporal layers are trained on video data through a hybrid loss function. For example, the hybrid loss function can include an L1 norm loss function, a learned perceptual image patch similarity (LPIPS) loss function, etc.
[0105] In the model training stage, the present disclosure trains the initial model through the above loss function until the model training end condition is met, and obtains a video enhancement model that can be used to process video enhancement tasks. The video enhancement model trained through the above steps can effectively learn the frame features and inter-frame timing information of multiple video frames, and apply them to video enhancement tasks.
[0106] Furthermore, the image-based VAE decoder is still prone to flickering artifacts when decoding the latent sequence. The present disclosure introduces an additional temporal 3D residual block in the VAE decoder to enhance low-level temporal consistency. The training method of the temporal 3D residual block of the VAE decoder is similar to that of training the temporal U-Net, that is, the pre-trained spatial layer is kept unchanged, and only the newly added temporal layer is trained until the training end condition is met, thereby obtaining a VAE decoder that introduces the temporal 3D residual block.
[0107] In summary, the model training method disclosed in the present invention obtains a pre-built initial model and a sample video frame sequence, wherein the sample video frame sequence includes a degraded video frame sequence and a reference video frame sequence; performs noise processing on the reference video frame sequence to obtain a noisy reference frame sequence, determines the noisy reference frame features and the noisy reference sequence timing information corresponding to the noisy reference frame sequence; determines the degraded frame features and the degraded sequence timing information corresponding to the degraded video frame sequence; generates a predicted enhanced frame sequence based on the degraded frame features, the degraded sequence timing information, the noisy reference frame features and the noisy reference sequence timing information by the initial model; trains the initial model according to the loss value between the predicted enhanced frame sequence and the reference video frame sequence to obtain a video enhancement model. On the one hand, by controlling the generation of the predicted enhanced frame sequence through the degraded frame features of the degraded video frame, the uncertainty of the frame enhancement content caused by the uncertainty generated by the diffusion model can be avoided, and the jitter between video frames can be reduced. On the other hand, based on the extracted temporal information between degraded video frames, a high-quality predicted enhanced frame sequence is generated, which can enhance the inter-frame consistency of the video enhancement results, improve the robustness of the video enhancement task, and obtain more stable video enhancement results. On the other hand, the sparse attention module in the newly introduced temporal module mainly considers the relationship between the current frame and the previous and next frames, and the temporal attention module mainly considers the relationship between the current frame and all frames, so that the model can learn the temporal information between frames and use it in subsequent video enhancement tasks, thereby improving the robustness of video enhancement tasks.
[0108] Figure 5 is a flow chart of a video enhancement method according to an exemplary embodiment. Figure 5 As shown, the video enhancement method can be used in a computer device. This exemplary embodiment uses the method applied to a computer device as an example. It can be understood that the method can also be applied to a server, and can also be applied to a system including a computer device and a server, and is implemented through the interaction between the computer device and the server. Specifically, the following steps are included.
[0109] Step S510: obtaining an initial degraded video.
[0110] Step S520, obtaining a pre-trained video enhancement model, where the video enhancement model is trained based on a model training method.
[0111] Step S530: input the initial degraded video into a video enhancement model, and the video enhancement model performs video enhancement processing on the initial degraded video to obtain a target enhanced video.
[0112] In an exemplary embodiment of the present disclosure, the initial degraded video may be a video clip to be subjected to video enhancement processing. The video enhancement model may be a video enhancement model of the ControlNet architecture based on the DiT base model. The timing module is introduced into both the DiT base model and the ControlNet network branch to improve the consistency of the video enhancement result, reduce the inter-frame jitter of the video, and obtain a network model with a more stable video enhancement result. The target enhanced video may be an enhanced video clip obtained after the video enhancement model performs video enhancement processing.
[0113] The initial degraded video is input into a pre-trained video enhancement model. Since the video enhancement model has pre-learned the temporal features between video frames and applied them to the video enhancement process, a target enhanced video with a more stable enhancement result can be obtained.
[0114] According to the video enhancement method in this example embodiment, the robustness of the video enhancement task can be improved, a more stable video enhancement result can be obtained, and the user viewing experience can be improved.
[0115] Figure 6 is a block diagram of a model training device according to an exemplary embodiment. Figure 6 The model training device 600 includes: a sample frame sequence acquisition module 610, a noisy frame feature determination module 620, a degraded frame feature determination module 630, a predicted frame sequence generation module 640 and a model training module 650.
[0116] Specifically, the sample frame sequence acquisition module 610 is used to obtain a pre-constructed initial model and a sample video frame sequence, wherein the sample video frame sequence includes a degraded video frame sequence and a reference video frame sequence; the noisy frame feature determination module 620 is used to perform degradation processing on the reference video frame sequence to obtain a noisy reference frame sequence, and determine the noisy reference frame features and the noisy reference sequence timing information corresponding to the noisy reference frame sequence; the degraded frame feature determination module 630 is used to determine the degraded frame features and the degraded sequence timing information corresponding to the degraded video frame sequence; the predicted frame sequence generation module 640 is used to generate a predicted enhanced frame sequence based on the degraded frame features, the degraded sequence timing information, the noisy reference frame features and the noisy reference sequence timing information by the initial model; the model training module 650 is used to train the initial model according to the loss value between the predicted enhanced frame sequence and the reference video frame sequence to obtain a video enhancement model.
[0117] In an exemplary embodiment of the present disclosure, the noisy reference frame sequence includes multiple noisy reference frames, and the noisy frame feature determination module 620 includes a noisy frame generation unit, which is used to: encode the reference video frame sequence to obtain a coded reference frame sequence; and perform noise processing on each coded reference frame in the coded reference frame sequence to obtain a noisy reference frame corresponding to each coded reference frame.
[0118] In an exemplary embodiment of the present disclosure, the noisy reference frame sequence includes multiple noisy reference frames, and the noisy frame feature determination module 620 includes a noisy frame feature determination unit, which is used to: perform image feature extraction processing on the noisy reference frame to obtain noisy reference frame features; obtain current noisy reference frame features corresponding to the noisy reference frame; update the current noisy reference frame features according to adjacent noisy reference frames of the noisy reference frame to obtain first noisy reference frame features; and update the first noisy reference frame features according to the noisy reference frame sequence to obtain second noisy reference frame features.
[0119] In an exemplary embodiment of the present disclosure, the initial model includes a temporal information extraction structure, the temporal information extraction structure includes a sparse attention module, and the noisy frame feature determination unit includes a first feature updating subunit, which is used to: determine a first focus point vector corresponding to the noisy reference frame based on the current noisy reference frame feature; determine the first information point vector and the first detailed content vector corresponding to the noisy reference frame based on the adjacent noisy reference frame features of the adjacent noisy reference frame by the sparse attention module; and update the current noisy reference frame feature based on the first focus point vector, the first information point vector and the first detailed content vector to obtain the first noisy reference frame feature.
[0120] In an exemplary embodiment of the present disclosure, the initial model includes a temporal information extraction structure, the temporal information extraction structure includes a temporal attention module, and the noisy frame feature determination unit includes a second feature updating subunit, which: determines a second focus point vector corresponding to the noisy reference frame according to a first noisy reference frame feature; determines a second information point vector and a second detailed content vector corresponding to the noisy reference frame according to a plurality of noisy reference frame features included in a noisy reference frame sequence by the temporal attention module; and performs an update operation on the first noisy reference frame feature according to the second focus point vector, the second information point vector and the second detailed content vector to obtain a second noisy reference frame feature.
[0121] In an exemplary embodiment of the present disclosure, the predicted frame sequence generation module 640 includes a predicted frame sequence generation unit, which is used to: determine the current degraded frame features and the current degraded sequence timing information corresponding to the current degradation round; use the current degraded frame features and the current degraded sequence timing information as feature generation constraints for the next video enhancement round; generate multiple predicted video frame latent vectors based on the feature generation constraints, the noisy reference frame features and the noisy reference sequence timing information; and decode the multiple predicted video frame latent vectors to generate a predicted enhanced frame sequence.
[0122] In an exemplary embodiment of the present disclosure, the predicted frame sequence generation unit includes a predicted frame sequence generation sub-unit, which is used to: obtain a variational codec in an initial model, the variational codec includes a temporal residual structure; through the temporal residual structure, decode multiple predicted video frame latent vectors to generate a predicted enhanced frame sequence.
[0123] In an exemplary embodiment of the present disclosure, the model training module 650 includes a model training unit, which is used to: determine the frame sequence loss value between the predicted enhanced frame sequence and the reference video frame sequence; based on the frame sequence loss value, adjust the model parameters of the timing information extraction structure in the initial model until the model training end conditions are reached to obtain a video enhancement model.
[0124] Figure 7 is a block diagram of a video enhancement device according to an exemplary embodiment. Figure 7 The video enhancement device 700 includes: a degraded video acquisition module 710, a model acquisition module 720 and a video enhancement module 730.
[0125] Specifically, the degraded video acquisition module 710 is used to acquire the initial degraded video; the model acquisition module 720 is used to acquire a pre-trained video enhancement model, where the video enhancement model is trained based on a model training method; the video enhancement module 730 is used to input the initial degraded video into the video enhancement model, and the video enhancement model performs video enhancement processing on the initial degraded video to obtain a target enhanced video.
[0126] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0127] Reference below Figure 8 800 according to this embodiment of the present disclosure is described. Figure 8 The electronic device 800 shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0128] like Figure 8As shown, the electronic device 800 is in the form of a general computing device. The components of the electronic device 800 may include, but are not limited to: the at least one processing unit 810, the at least one storage unit 820, a bus 830 connecting different system components (including the storage unit 820 and the processing unit 810), and a display unit 840.
[0129] The storage unit stores program codes, which can be executed by the processing unit 810, so that the processing unit 810 executes the steps according to various exemplary embodiments of the present disclosure described in the above “Exemplary Method” section of this specification.
[0130] The storage unit 820 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 821 and / or a cache memory unit 822 , and may further include a read-only memory unit (ROM) 823 .
[0131] The storage unit 820 may include a program / utility 824 having a set (at least one) of program modules 825, such program modules 825 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
[0132] Bus 830 may represent one or more of several types of bus structures, including a memory unit bus or memory unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.
[0133] The electronic device 800 may also communicate with one or more external devices 870 (e.g., keyboards, pointing devices, Bluetooth devices, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 800, and / or communicate with any device that enables the electronic device 800 to communicate with one or more other computing devices (e.g., routers, modems, etc.). Such communication may be performed via an input / output (I / O) interface 850. Furthermore, the electronic device 800 may also communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) via a network adapter 860. As shown, the network adapter 860 communicates with other modules of the electronic device 800 via a bus 830. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 800, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0134] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory including instructions, and the above instructions can be executed by a processor of the device to complete the above model training method or video enhancement method. Optionally, the computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0135] In an exemplary embodiment, a computer program product is also provided, including a computer program, which implements any one of the above-mentioned model training methods or video enhancement methods when executed by a processor.
[0136] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.
[0137] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A model training method, characterized in that: include: Acquire a pre-built initial model and a sample video frame sequence, wherein the sample video frame sequence includes a degraded video frame sequence and a reference video frame sequence; Performing noise processing on the reference video frame sequence to obtain a noisy reference frame sequence, and determining noisy reference frame features and noisy reference sequence timing information corresponding to the noisy reference frame sequence; Determining degraded frame features and degraded sequence timing information corresponding to the degraded video frame sequence; The initial model generates a predicted enhanced frame sequence based on the degraded frame features, the degraded sequence timing information, the noisy reference frame features and the noisy reference sequence timing information; The initial model is trained according to the loss value between the predicted enhanced frame sequence and the reference video frame sequence to obtain a video enhancement model.
2. The method according to claim 1, characterized in that The noisy reference frame sequence includes a plurality of noisy reference frames, and the noisy processing is performed on the reference video frame sequence to obtain the noisy reference frame sequence, including: Encoding the reference video frame sequence to obtain an encoded reference frame sequence; Noise processing is performed on each coded reference frame in the coded reference frame sequence to obtain a noisy reference frame corresponding to each coded reference frame.
3. The method according to claim 1, characterized in that The noisy reference frame sequence includes a plurality of noisy reference frames, and determining the noisy reference frame features and noisy reference sequence timing information corresponding to the noisy reference frame sequence includes: Performing image feature extraction processing on the noisy reference frame to obtain features of the noisy reference frame; Acquire current noisy reference frame features corresponding to the noisy reference frame; According to the adjacent noisy reference frame of the noisy reference frame, the feature of the current noisy reference frame is updated to obtain the first noisy reference frame feature; According to the noisy reference frame sequence, the first noisy reference frame feature is updated to obtain a second noisy reference frame feature.
4. The method according to claim 3, characterized in that: The initial model includes a temporal information extraction structure, the temporal information extraction structure includes a sparse attention module, and the updating operation is performed on the current noisy reference frame feature according to the adjacent noisy reference frame of the noisy reference frame to obtain the first noisy reference frame feature, including: Determine a first focus point vector corresponding to the noisy reference frame according to the current noisy reference frame feature; The sparse attention module determines a first information point vector and a first detailed content vector corresponding to the noisy reference frame according to adjacent noisy reference frame features of the adjacent noisy reference frame; The current noisy reference frame feature is updated according to the first focus point vector, the first information point vector and the first detailed content vector to obtain the first noisy reference frame feature.
5. The method according to claim 3, characterized in that: The initial model includes a temporal information extraction structure, the temporal information extraction structure includes a temporal attention module, and the updating operation is performed on the first noisy reference frame feature according to the noisy reference frame sequence to obtain the second noisy reference frame feature, including: Determining a second focus point vector corresponding to the noisy reference frame according to the first noisy reference frame feature; The temporal attention module determines a second information point vector and a second detailed content vector corresponding to the noisy reference frame according to a plurality of noisy reference frame features included in the noisy reference frame sequence; According to the second focus point vector, the second information point vector and the second detailed content vector, the first noisy reference frame feature is updated to obtain the second noisy reference frame feature.
6. The method according to claim 1, characterized in that The generating of the predicted enhanced frame sequence by the initial model based on the degraded frame feature, the degraded sequence timing information, the noisy reference frame feature and the noisy reference sequence timing information includes: Determine the current degradation frame features and the current degradation sequence timing information corresponding to the current degradation round; Using the current degraded frame feature and the current degraded sequence timing information as feature generation constraint conditions for the next video enhancement round; Generate multiple predicted video frame latent vectors according to the feature generation constraint, the noisy reference frame feature and the noisy reference sequence timing information; The plurality of predicted video frame latent vectors are decoded to generate the predicted enhanced frame sequence.
7. The method according to claim 6, characterized in that The decoding process of the plurality of predicted video frame latent vectors to generate the predicted enhanced frame sequence includes: Obtaining a variational codec in the initial model, wherein the variational codec includes a temporal residual structure; The multiple predicted video frame latent vectors are decoded through the temporal residual structure to generate the predicted enhanced frame sequence.
8. The method according to claim 1, characterized in that: The step of training the initial model according to the loss value between the predicted enhanced frame sequence and the reference video frame sequence to obtain a video enhancement model comprises: Determining a frame sequence loss value between the prediction enhanced frame sequence and the reference video frame sequence; Based on the frame sequence loss value, the model parameters of the temporal information extraction structure in the initial model are adjusted until the model training end condition is reached to obtain the video enhancement model.
9. A video enhancement method, characterized in that: include: Get the initial degraded video; Obtain a pre-trained video enhancement model, wherein the video enhancement model is trained based on the model training method according to any one of claims 1 to 8; The initial degraded video is input into the video enhancement model, and the video enhancement model performs video enhancement processing on the initial degraded video to obtain a target enhanced video.
10. A model training device, characterized in that: include: A sample frame sequence acquisition module, used to acquire a pre-built initial model and a sample video frame sequence, wherein the sample video frame sequence includes a degraded video frame sequence and a reference video frame sequence; A noisy frame feature determination module is used to perform degradation processing on the reference video frame sequence to obtain a noisy reference frame sequence, and determine the noisy reference frame features and noisy reference sequence timing information corresponding to the noisy reference frame sequence; A degraded frame feature determination module, used to determine the degraded frame features and degraded sequence timing information corresponding to the degraded video frame sequence; A prediction frame sequence generation module, configured to generate a prediction enhanced frame sequence based on the degraded frame features, the degraded sequence timing information, the noisy reference frame features and the noisy reference sequence timing information by the initial model; The model training module is used to train the initial model according to the loss value between the predicted enhanced frame sequence and the reference video frame sequence to obtain a video enhancement model.
11. A video enhancement device, characterized in that: include: A degraded video acquisition module, used to acquire an initial degraded video; A model acquisition module, used to acquire a pre-trained video enhancement model, wherein the video enhancement model is trained based on the model training method according to any one of claims 1 to 8; The video enhancement module is used to input the initial degraded video into the video enhancement model, and the video enhancement model performs video enhancement processing on the initial degraded video to obtain a target enhanced video.
12. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the model training method as described in any one of claims 1 to 8, or to implement the video enhancement method as described in claim 9.
13. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, it implements the model training method as described in any one of claims 1 to 8, or implements the video enhancement method as described in claim 9.