A video enhancement processing method, apparatus, device and storage medium

By performing feature extraction, distortion feature removal, resolution enhancement, and noise reduction and diffusion processing on compressed videos, the interference problem caused by information loss in compressed video enhancement is solved, improving the resolution and realism of the video and achieving better enhancement results.

CN122120485APending Publication Date: 2026-05-29BEIJING ZITIAO NETWORK TECH CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2024-11-28
Publication Date
2026-05-29

Smart Images

  • Figure CN122120485A_ABST
    Figure CN122120485A_ABST
Patent Text Reader

Abstract

The embodiment of the present disclosure discloses a video enhancement processing method, device, equipment and storage medium, comprising: obtaining a to-be-enhanced video; for each video frame in the to-be-enhanced video, performing feature extraction on the video frame to obtain a feature frame of the video frame; performing distortion feature screening processing on the feature frame to obtain an enhanced feature frame of the video frame; performing resolution enhancement processing on the enhanced feature frame to obtain a first enhanced video frame of the video frame; performing denoising and diffusion enhancement on the first enhanced video of the video frame to obtain a second enhanced video frame of the video frame; and obtaining a target enhanced video of the to-be-enhanced video according to the second enhanced video of all video frames in the to-be-enhanced video. By using the method, the interference of the lost video information in the compressed video on the video enhancement processing can be effectively avoided, and the noise features and the to-be-repaired texture features in the compressed video can be more accurately distinguished, so that only the to-be-enhanced texture features are diffused and enhanced, and the sense of reality of the enhanced video is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer vision processing technology, and in particular to a video enhancement processing method, apparatus, device, and storage medium. Background Technology

[0002] In video applications, due to limitations in transmission bandwidth or storage space, videos are typically compressed to create compressed videos. Compressed videos often suffer from low resolution, low quality, and compression noise. In scenarios where playback quality and resolution are critical, it's necessary to increase the resolution of the compressed video and remove compression noise to achieve enhanced video processing.

[0003] However, existing video enhancement methods primarily perform super-resolution enhancement, which is not suitable for compressed videos. Specifically, when video frames are highly compressed at low bitrates, the quantization process inevitably leads to some information loss. Upgrading the resolution of such information-scarce video cannot avoid the interference of the lost video information on the resolution improvement; in fact, it may degrade the enhancement effect of the compressed video.

[0004] Currently, no enhancement processing method for compressed video has been found to effectively avoid the above problems. Summary of the Invention

[0005] This disclosure provides a video enhancement processing method, apparatus, device, and storage medium, which effectively enhances compressed videos and can effectively avoid interference from lost video information in compressed videos on video enhancement processing.

[0006] In a first aspect, embodiments of this disclosure provide a video enhancement processing method, the method comprising:

[0007] Obtain the video to be enhanced;

[0008] For each video frame in the video to be enhanced, feature extraction is performed on the video frame to obtain the feature frame of the video frame; distortion feature removal processing is performed on the feature frame to obtain the enhanced feature frame of the video frame; resolution enhancement processing is performed on the enhanced feature frame to obtain the first enhanced video frame of the video frame; noise reduction and diffusion enhancement are performed on the first enhanced video frame of the video frame to obtain the second enhanced video frame of the video frame.

[0009] The target enhanced video of the video to be enhanced is obtained based on the second enhanced video of all video frames in the video to be enhanced.

[0010] Secondly, embodiments of this disclosure also provide a video enhancement processing apparatus, the apparatus comprising:

[0011] The acquisition module is used to acquire the video to be enhanced;

[0012] The enhancement processing module is used to perform feature extraction on each video frame in the video to be enhanced to obtain a feature frame of the video frame; perform distortion feature removal processing on the feature frame to obtain an enhanced feature frame of the video frame; perform resolution enhancement processing on the enhanced feature frame to obtain a first enhanced video frame of the video frame; and perform noise reduction and diffusion enhancement on the first enhanced video frame of the video frame to obtain a second enhanced video frame of the video frame.

[0013] The result generation module is used to obtain the target enhanced video of the video to be enhanced based on the second enhanced video of all video frames in the video to be enhanced.

[0014] Thirdly, embodiments of this disclosure also provide a computer device, the computer device comprising:

[0015] One or more processors;

[0016] Storage device for storing one or more programs.

[0017] When the one or more programs are executed by the one or more processors, the one or more processors implement the video enhancement processing method provided in any embodiment of this disclosure.

[0018] Fourthly, embodiments of this disclosure also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the video enhancement processing method provided in any embodiment of this disclosure.

[0019] Fifthly, embodiments of this disclosure also provide a computer program product, including a computer program that, when executed by a processor, implements the video enhancement processing method provided in any embodiment of this disclosure.

[0020] The technical solution of this disclosure embodiment specifically discloses a video enhancement processing method, apparatus, device, and storage medium. The method first acquires a video to be enhanced; then, for each video frame in the video to be enhanced, features are extracted from the video frame to obtain a feature frame of the video frame; distortion feature removal processing is performed on the feature frame to obtain an enhanced feature frame of the video frame; resolution enhancement processing is performed on the enhanced feature frame to obtain a first enhanced video frame of the video frame; noise reduction and diffusion enhancement are performed on the first enhanced video frame of the video frame to obtain a second enhanced video frame of the video frame; finally, the target enhanced video of the video to be enhanced is obtained based on the second enhanced videos of all video frames in the video to be enhanced. This embodiment's technical solution can be considered as an enhancement process for compressed video. Addressing the distortion problem inherent in compressed video, this solution employs feature extraction and filtering to remove distorted features, thereby reducing interference from distorted video features on the diffusion-based video enhancement process and effectively mitigating the degradation of video enhancement caused by lost video information in the compressed video. Simultaneously, as part of the video enhancement process, this solution also performs resolution enhancement to improve the resolution of the compressed video. Furthermore, this solution considers using a diffusion mechanism to enhance video texture. Unlike existing technologies, this solution introduces noise feature removal during diffusion enhancement, enabling better identification of noise features and texture features to be enhanced during the diffusion enhancement of compressed video. This allows diffusion enhancement to be performed only on the texture features to be enhanced. Compared to existing technologies, the enhancement method provided in this embodiment better reflects spatial fidelity and temporal consistency, thus improving the realism of the enhanced video and enhancing the subjective visual quality of the video viewing experience. Attached Figure Description

[0021] To more clearly illustrate the technical solutions of the exemplary embodiments of this disclosure, the accompanying drawings used in describing the embodiments are briefly introduced below. Obviously, the accompanying drawings described are only a portion of the embodiments to be described in this disclosure, and not all of them. For those skilled in the art, other drawings can be obtained from these drawings without any creative effort.

[0022] Figure 1 A schematic flowchart illustrating a video enhancement processing method provided in an embodiment of this disclosure;

[0023] Figure 2 This is a schematic diagram of the structure of a video enhancement processing device provided in an embodiment of the present disclosure;

[0024] Figure 3 This is a schematic diagram of the structure of a computer device provided in an embodiment of this disclosure. Detailed Implementation

[0025] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0026] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0027] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0028] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules, or units, and are not used to limit the order of functions performed by these devices, modules, or units or their interdependencies. It should also be noted that the modifications of "a" and "a plurality of" mentioned in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0029] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0030] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0031] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0032] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0033] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0034] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0035] Figure 1 This is a flowchart illustrating a video enhancement processing method provided in an embodiment of the present disclosure. This embodiment is applicable to enhancing compressed videos. The method can be executed by a video enhancement processing device, which can be implemented by software and / or hardware and can be configured in a terminal and / or server to implement the video enhancement processing method in this embodiment of the present disclosure.

[0036] like Figure 1 As shown, a video enhancement processing method provided in this embodiment may include:

[0037] S101. Obtain the video to be enhanced.

[0038] In this embodiment, the video to be enhanced can be considered as a compressed video after compression. The video to be enhanced may have a certain playback duration and contain a certain number of video frames.

[0039] It is known that in video compression, some video resolution and clarity are often sacrificed to limit the size of the compressed video, leading to content distortion and noise. In existing video enhancement processes, the content distortion and noise present in compressed videos often act as interfering factors, affecting the enhancement effect. Therefore, this embodiment provides an enhancement method for compressed videos, and this step can be used to obtain the compressed video as the video to be enhanced.

[0040] S102. For each video frame in the video to be enhanced, feature extraction is performed on the video frame to obtain the feature frame of the video frame; distortion feature removal processing is performed on the feature frame to obtain the enhanced feature frame of the video frame; resolution enhancement processing is performed on the enhanced feature frame to obtain the first enhanced video frame of the video frame; noise reduction and diffusion enhancement are performed on the first enhanced video frame of the video frame to obtain the second enhanced video frame of the video frame.

[0041] In this embodiment, the enhancement processing of the video to be enhanced can be further optimized to enhancement processing of each video frame in the video to be enhanced.

[0042] In this step, for each video frame in the video to be enhanced, features are first extracted from the video frame, and then distortion features are identified and removed from the extracted feature frames to obtain the enhanced feature frames. Compared to the video frames of the video to be enhanced, these enhanced feature frames can be considered as feature frames formed by removing distortion features that affect the enhancement process. For the video to be enhanced on a frame-by-frame basis, each video frame can be considered as an image to be enhanced, and the image features formed by feature extraction from each image to be enhanced can constitute the feature frames of that video frame.

[0043] One implementation method is to use a pre-trained distortion modulation network model to perform feature extraction and feature removal of video frames. Specifically, the distortion modulation network model can first extract features from each video frame in the video to be enhanced. Then, based on the pre-determined filtering conditions for distortion feature selection, distortion features can be determined from the feature frames extracted from each video frame, and the determined distortion features can be deleted. The feature frames obtained after deleting distortion features from each video frame can be used as the enhanced feature frames of the video frame.

[0044] In this embodiment, the resolution of the compressed video is lower than that of the original video. Therefore, enhancing the video resolution is also part of the enhancement process for the compressed video, and the resolution enhancement process can also be performed through this step.

[0045] Specifically, the resolution enhancement processing can be performed on the enhanced feature frames of each video frame. Specifically, the enhanced video frames can be input into a network model with an upsampling network structure. Through this network model, the feature information in the enhanced video frames can be upsampled to increase the resolution of the video frames to the same level as the original video, or to the resolution value of the set super-resolution. Finally, the video frames obtained after resolution enhancement can be used as the first enhanced video frames corresponding to the video frames in the video to be enhanced.

[0046] Furthermore, analysis revealed that enhancing compressed video, besides increasing resolution, primarily involves improving image quality. This is equivalent to deblurring the video frame, enhancing the richness and clarity of the content. One way to improve image richness and clarity is to use existing video content as generative priors, leveraging a diffusion mechanism to enhance video texture features.

[0047] It is known that the more video association information contained in its generative prior content, the better the authenticity and coherence of the video texture generated by the diffusion mechanism can be guaranteed. In this embodiment, this step can also be used to further perform denoising diffusion enhancement processing on the first enhanced video frame of each of the above video frames. The denoising diffusion enhancement can be considered as identifying and removing noise features in the first enhanced video frame, and identifying which video features are to be enhanced by the diffusion mechanism for texture enhancement.

[0048] This step ensures that the features involved in generative texture enhancement are all effective features after noise removal. The effective features are then used to generate texture details through diffusion to enhance the realism of the video content.

[0049] One implementation method for the denoising enhancement in this step can be described as using a diffusion model. Specifically, this diffusion model creates a first encoder with a convolutional network structure and a second encoder with a deep generation mechanism. The second encoder can be a variational autoencoder. The first encoder can first perform downsampling feature extraction on the first enhanced video frame to filter out noise features and obtain a small number of denoised feature frames that can play a decisive role. The second encoder can also downsample the feature channels of the first enhanced video frame to obtain the control condition features required for texture diffusion generation.

[0050] Following the above description, the diffusion model also includes a first decoder corresponding to the first encoder and a second decoder corresponding to the second encoder. The first decoder can perform feature-aware diffusion based on denoised feature frames and control condition features, which is equivalent to performing video feature enhancement processing on the video frames in the video to be enhanced again, thereby obtaining richer effective video features. The second decoder can perform texture-aware diffusion on the effective video features after feature enhancement, which is equivalent to performing texture diffusion enhancement on the video frames in the video to be enhanced, thereby obtaining richer video texture features. Finally, after feature enhancement and texture enrichment are performed on the first enhanced video frames of each video frame, the enhanced video features of the video frame can be obtained, and the second enhanced video frame of the video frame can be formed based on these enhanced video features.

[0051] It should be noted that, under the condition that computing resources permit, this embodiment can perform the enhancement processing of this step on different video frames separately in parallel to speed up the processing time of the video to be enhanced; alternatively, under the condition that computing resources are limited, it can perform the enhancement processing of this step on each video frame in the video to be enhanced sequentially. This embodiment does not specifically limit the operation mode of video enhancement processing.

[0052] S103. Obtain the target enhanced video of the video to be enhanced based on the second enhanced video of all video frames in the video to be enhanced.

[0053] It can be seen that, relative to all video frames in the video to be enhanced, after the above steps have completed the effective enhancement processing of distortion features, resolution enhancement, noise reduction, and texture clarity enhancement, this embodiment can re-merge the second enhanced video frames obtained from each video frame according to the video playback sequence. The final video can be considered as the target enhanced video after video enhancement processing of the video to be enhanced in this embodiment.

[0054] This embodiment provides a video enhancement processing method, which can be considered as enhancement processing of compressed video. Addressing the distortion problem inherent in compressed video, this technical solution proposes feature filtering to remove distorted features, thereby reducing the interference of distorted video features on diffusion-based video enhancement processing and effectively avoiding the degradation of video enhancement processing caused by lost video information in compressed video. Simultaneously, as part of the video enhancement processing, this technical solution also performs resolution enhancement processing to improve the resolution of the compressed video. Furthermore, this technical solution also considers achieving video texture enhancement through a diffusion mechanism. Unlike existing technologies, this technical solution introduces noise feature removal processing in diffusion enhancement, enabling better identification of noise features and texture features to be enhanced when performing diffusion enhancement on compressed video. This allows diffusion enhancement to be performed only on the texture features to be enhanced. Compared to existing technologies, the enhancement processing method provided in this embodiment better reflects spatial fidelity and temporal consistency, thus better improving the realism of the video after enhancement processing and ultimately enhancing the subjective visual quality of video viewing.

[0055] As a first optional embodiment of this example, based on the above embodiment, feature extraction can be performed on each video frame in the video to be enhanced to obtain the feature frame of the video frame; distortion feature removal processing can be performed on the feature frame to obtain the enhanced feature frame of the video frame. This can be further specified as follows:

[0056] a1) The video frame is input into the feature extraction module of the feature extraction sub-model in the video enhancement processing model to extract features from the video frame and obtain the feature frame of each video frame.

[0057] In this embodiment, the enhancement processing of the video to be enhanced can be considered to be mainly achieved through a video enhancement processing model. The video enhancement processing model can include sub-models for executing different enhancement processing logics. Specifically, the feature removal sub-model can be considered as a sub-model built within the video enhancement processing model, and can be considered as a network model participating in distortion feature modulation. Specifically, a network structure with multiple codecs can be created to constitute the feature extraction module in the feature removal sub-model, and a feature filtering network structure can be created to constitute the feature filtering module in the feature removal sub-model.

[0058] The feature filtering sub-model in this embodiment can be considered as having been pre-trained and is a usable model for direct feature extraction and filtering. Different types of high-resolution, high-definition real images can be randomly selected, and then low-quality images are obtained by processing the real images with different degradation intensities. These low-quality images and corresponding real images can then be used to form sample image pairs to participate in the training of the feature filtering sub-model. Through continuous iterative optimization, the feature filtering sub-model used in this embodiment can be obtained.

[0059] In this embodiment, each video frame in the video to be enhanced can be regarded as a single-frame image and used as input data for the feature filtering sub-model. The feature filtering sub-model can first extract features from each video frame to obtain the feature frames corresponding to each video frame.

[0060] b1) Input the feature frame into the feature filtering module of the feature filtering sub-model, determine the distortion features that do not meet the set feature constraints from the feature frame and delete the distortion features to obtain the enhanced feature frame of the video frame.

[0061] In this embodiment, this step can be considered as one of the steps executed in the logical processing of the feature removal sub-model. Specifically, it can be executed through the feature filtering module in the feature removal sub-model. The feature filtering module is equivalent to having some preset thresholds for feature filtering. These thresholds can form feature constraints. Then, the extracted feature frames can be judged based on the feature constraints, thereby determining the video features that do not meet the feature constraints as distorted features. In this embodiment, the feature frames retained after deleting distorted features can be determined as enhanced feature frames.

[0062] The above-described technical solution in this embodiment provides a specific implementation for feature extraction and distortion feature screening of the video to be enhanced. Through this technical solution, distortion features can be effectively screened out, thereby reducing the interference of distortion video features on diffusion-type video enhancement processing and effectively avoiding the degradation of video enhancement processing caused by video information lost in compressed video.

[0063] In this first optional embodiment, as one implementation method, the resolution enhancement processing of the enhanced feature frame to obtain the first enhanced video frame of the video frame can be further specified as follows: through the resolution enhancement sub-model in the video enhancement processing model, the enhanced feature frame is input into the upsampling network structure contained in the resolution enhancement sub-model, and the enhanced feature frame is subjected to resolution upsampling processing to obtain the first enhanced video frame of the video frame.

[0064] In this embodiment, in addition to the feature removal sub-model, the video enhancement processing model may also include a resolution enhancement sub-model for enhancing the resolution of the enhanced feature frames obtained after distortion removal of video frames, and may also include a diffusion sub-model for denoising and texture diffusion.

[0065] As one implementation method, this embodiment can further upsample the resolution of the enhanced feature frames of each video frame using a resolution enhancement sub-model in the video enhancement processing model. Specifically, the resolution enhancement sub-model can be considered to contain an upsampling network structure, which can enhance the resolution of the enhanced feature frames, so that the enhanced resolution of each video frame can reach the target resolution.

[0066] This embodiment is equivalent to obtaining an enhanced video frame after the resolution reaches the target resolution, and this enhanced video frame can be recorded as the first enhanced video frame.

[0067] It should be noted that the target resolution can be set differently depending on the application scenario. When performing super-resolution enhancement on compressed video, the target resolution can be set to the product of the original resolution of the original video before compression and the super-resolution ratio. Alternatively, when restoring the resolution of compressed video, the target resolution can be set to the original resolution of the compressed video.

[0068] It can be seen that the first enhanced video frame determined by the above steps corresponds to a video frame in the video to be enhanced. Compared with the video frame in the video to be enhanced, the first enhanced video can be considered to have completed the removal of distortion features and the enhancement of resolution.

[0069] The above-described technical solution in this embodiment provides a specific implementation for enhancing the resolution of the video to be enhanced. Resolution enhancement processing, as a part of video enhancement processing and the main execution logic of this embodiment, can be used to improve the resolution of compressed video.

[0070] As a second optional embodiment of this embodiment, based on the above embodiment, the first enhanced video of the video frame can be denoised and diffused to obtain the second enhanced video frame of the video frame, which is specifically described as follows: through the diffusion sub-model in the video enhancement processing model, according to the encoder network structure and decoder network structure included in the diffusion sub-model, the first enhanced video frame of the video frame is subjected to video feature denoising and texture diffusion enhancement to obtain the second enhanced video frame of the video frame.

[0071] In this embodiment, the first enhanced video frame of each video frame can be subjected to diffusion-generative enhancement through a diffusion sub-model in the video enhancement processing model. This diffusion sub-model can be considered as a sub-model that incorporates the relevant network structure into the video enhancement processing model, and it can output a second enhanced video frame for each first enhanced video frame.

[0072] It is known that the logical processing mechanism of conventional diffusion models focuses on using realistic generative capabilities to enhance image texture reuse. Generally, diffusion models can generate texture details lost due to degradation features such as low resolution and compression noise based on strong generative priors. However, diffusion models themselves lack the ability to perceive these degradation features. When performing texture detail enhancement using diffusion, it is difficult to accurately distinguish which features are noise features and which are features that need to be restored. Instead, they may enhance the texture details of degraded features, affecting the enhancement quality. At the same time, when enhancing video features, conventional diffusion models often randomly diffuse the video features. This random diffusion method can cause inter-frame flickering, affecting the temporal consistency of the enhanced video.

[0073] Based on this, the video enhancement processing method proposed in this embodiment optimizes and improves the diffusion sub-model in the video enhancement processing model. Specifically, it introduces the detection and removal of noise features in the video, and introduces additional conditional information to constrain the diffusion trajectory, thereby improving the texture fidelity and realism of the diffusion sub-model in texture diffusion enhancement.

[0074] For example, the introduced noise feature detection and removal can be achieved through a compressed sensing mechanism, specifically a convolutional encoder that introduces feature downsampling to achieve compressed sensing of video features. Simultaneously, the introduced diffusion trajectory constraint can be achieved through a spatiotemporal attention mechanism, specifically a directional diffusion encoder that introduces feature channel downsampling to obtain the control condition features upon which diffusion depends.

[0075] As described above, by introducing denoising and orientation mechanisms into the diffusion sub-model, noise features and features to be enhanced can be effectively distinguished. By controlling the limitations of conditional features, the diffusion trajectory can be constrained, thereby obtaining a second enhanced video frame with texture enhancement relative to the output of each first enhanced video frame.

[0076] Based on this second optional embodiment, as one implementation method, the process of performing video feature denoising and texture diffusion enhancement on the first enhanced video frame of the video frame according to the encoder network structure and decoder network structure included in the diffusion sub-model to obtain the second enhanced video frame of the video frame can be further specified as the following steps:

[0077] a2) Input the first enhanced video frame of the video frame into the first encoder in the diffusion sub-model, and remove noise features from the video features of the first enhanced video frame through the downsampling network structure in the first encoder to obtain the denoised feature frame of the first enhanced video frame.

[0078] In this embodiment, the operations described above and below can be performed on the first enhanced video frame of each video frame. It is understood that the preferred diffusion sub-model in this embodiment includes a first encoder, a second encoder, a first decoder corresponding to the first encoder, and a second decoder corresponding to the second encoder.

[0079] The first encoder and the first decoder can be considered as the embodiment of a compressed sensing mechanism with a convolutional network structure. This step is mainly performed by the first encoder, which can first downsample the video features of the first enhanced video frame to extract and retain more critical video features through downsampling, which is equivalent to filtering out noise features. This step can also record the video features obtained by downsampling as denoised feature frames.

[0080] b2) Input the first enhanced video frame of the video frame into the second encoder in the diffusion sub-model, and filter the feature channels of the first enhanced video frame through the diffusion control network structure in the second encoder to obtain the diffusion control feature frame of the first enhanced video frame.

[0081] In this embodiment, the second encoder and the second decoder can be considered as manifestations of a spatiotemporal attention mechanism with a directional diffusion network structure. In this embodiment, the directional diffusion network constituting the second encoder also incorporates a control network structure for constraining texture diffusion. This step uses the second encoder to downsample the feature channels of the first enhanced video frame, obtaining the key feature channels that have been filtered out. Then, through the introduced control network structure, diffusion control feature frames for texture diffusion can be generated based on the video features corresponding to the filtered key feature channels.

[0082] c2) Input the denoising feature frame and the diffusion control feature frame into the first decoder in the diffusion sub-model. Through the upsampling network structure in the first decoder, the diffusion control feature frame is used as a diffusion constraint to perform feature-aware diffusion on the denoising feature frame to obtain the denoising enhancement features of the first enhanced video frame.

[0083] In this embodiment, this step specifically uses the first decoder corresponding to the first encoder to process the denoised feature frames and diffusion control feature frames as input data. During the logical operation of the first decoder, it can perform feature-aware diffusion on the denoised feature frames using the diffusion control feature frames, thereby generating refined features to fill the filtered noise features. This step can determine the denoising enhancement features of the first enhanced video frame based on the features output by the first decoder.

[0084] It can be seen that, compared with the video features initially extracted from the first enhanced video frame, this denoising enhancement feature retains key video features and derives more in-depth effective video features from these key video features. This better ensures that features such as compression noise in the compressed video do not affect texture diffusion, and improves the diffusion sub-model's ability to perceive noise features. This allows it to effectively avoid noise features in texture diffusion and reduce the interference of noise features on video enhancement.

[0085] d2) Input the denoising enhancement features into the second decoder in the diffusion sub-model, and perform texture-aware diffusion on the denoising enhancement features through the spatiotemporal attention mechanism of the second decoder to obtain the second enhanced video frame of the first enhanced video frame.

[0086] In this embodiment, the texture features can also be refined by a second decoder through diffusion generation. This texture feature refinement is based on the denoising and enhancement features output by the above steps, and the spatiotemporal attention mechanism of the second decoder is used to achieve directional diffusion of texture features. In this embodiment, the enhanced video features output by the second decoder can be determined as diffusion enhancement features.

[0087] It can be seen that, compared with the denoising enhancement feature, this diffusion enhancement feature further realizes the effective diffusion of video texture features, and through directional diffusion, it better ensures the temporal consistency of the diffused content. This processing can effectively avoid the problem of inter-frame flickering in video playback caused by the instability of random diffusion.

[0088] In this embodiment, after obtaining the extended enhancement features of the first enhanced video frame through the aforementioned diffusion sub-model, this step can process the extended enhancement features using a fully connected operation to obtain the video frame restored by the extended enhancement features. This video frame can be denoted as the second enhanced video frame of the first enhanced video frame.

[0089] This embodiment provides a specific implementation of forming a second enhanced video frame after performing texture enhancement processing on a first enhanced video frame. Specifically, this embodiment considers achieving video texture enhancement through a diffusion mechanism. Unlike existing technologies, this solution introduces noise feature removal and directional diffusion processing in the diffusion enhancement process. This allows for better identification of noise features and texture features to be enhanced when performing diffusion enhancement on compressed video, thus enabling directional diffusion enhancement only on the texture features to be enhanced. Compared to existing technologies, the enhancement processing method provided in this embodiment better reflects spatial fidelity and temporal consistency, thereby better improving the realism of the video after enhancement processing, and ultimately enhancing the subjective visual quality of video viewing.

[0090] As a third optional embodiment of this example, based on the above embodiments, the training steps of the video enhancement processing model can be specified as follows:

[0091] a3) Obtain a sample training set and an initial augmentation processing model, wherein the sample training set includes at least one sample tuple, and the sample tuple includes a real video sample and a compressed video sample of the real video sample.

[0092] In this embodiment, this step can be used to obtain the initially constructed initial enhancement processing model and acquire a sample training set for model training. It is understood that the sample training set contains sample pairs specifically determined for model training. Each sample pair includes a sample real video and its corresponding sample compressed video. The sample real video can be considered as the initially generated video that has undergone secondary processing due to limitations in transmission bandwidth or storage space; this secondary processing may include video compression. The sample compressed video can be considered as the real video that has undergone compression processing.

[0093] It is known that, compared to the real sample video, the compressed sample video has a reduced resolution and may also suffer from noise, distortion, or blurring caused by compression. In this embodiment, the compressed sample video can be considered as the input data for video enhancement training, and the real sample video can be used as the label video for video enhancement training.

[0094] In this embodiment, the initial enhancement processing model can be considered as a pre-constructed network model that includes the network structure of each sub-model, and each sub-model in the initial enhancement processing model can be considered to have initial network parameters. The included sub-models can be resolution enhancement sub-models and diffusion sub-models, etc.

[0095] b3) Input the sample compressed video into the initial enhancement processing model to obtain the actual enhanced video output by the initial enhancement processing model.

[0096] In this embodiment, the process of training the video enhancement processing model can be regarded as an iterative training process. In each iteration, one or more sample pairs can be selected from the sample training set to participate in the model training. Specifically, the sample compressed video in the selected sample pairs can be used as input data to the initial enhancement processing model.

[0097] It is understood that the initial enhancement processing model can sequentially enhance the sample compressed video according to its constituent sub-models. For example, the feature removal sub-model in the initial enhancement processing model can be used to extract features and remove distortion features from video frames in the sample compressed video, obtaining the sample enhanced feature frames output by the feature removal sub-model. Then, the sample enhanced feature frames can be used as input data to the resolution enhancement sub-model, which performs resolution enhancement processing on the sample enhanced feature frames to obtain the first enhanced video output by the resolution enhancement sub-model.

[0098] As described above, the obtained sample first enhanced video can be divided into video frames to form at least one sample first sub-video. Then, the first enhanced video frame in each sample first sub-video is used as input data for the diffusion sub-model. Through the enhancement processing of the diffusion sub-model, sample second sub-videos can be obtained relative to each sample first sub-video. Finally, the actual enhanced video output by the initial enhancement processing model can be obtained by merging each sample second sub-video.

[0099] c3) Based on the sample real video and the actual enhanced video, determine the first loss function value of the first loss function, the second loss function value of the second loss function, and the third loss function value of the third loss function.

[0100] In this embodiment, the iterative training process of the video enhancement processing model can be considered as a process of reverse adaptive adjustment of network parameters, and the loss function value can be used as the information value required for reverse adaptive adjustment. This embodiment can achieve the determination of the loss function value required for reverse adaptive adjustment of network parameters through this step and step d3) below.

[0101] Analysis reveals that the initial enhancement processing model comprises multiple sub-models, each with different network structures and processing logic. Except for the pre-trained feature filtering sub-model, all other sub-models are in their initial network parameter state. To achieve effective training of each sub-model, this embodiment considers setting different loss functions from multiple perspectives. For example, four loss functions can be designed, each focusing on the training of one or more sub-models. The weighted sum of the loss function values ​​corresponding to the different loss functions, obtained through the following steps, yields the final target loss function that adapts the network parameters.

[0102] It should be noted that the fourth loss function in this embodiment is calculated in advance in the initial enhancement processing model. This is equivalent to determining the fourth loss function based on the function corresponding to the fourth loss function during the initial enhancement processing model's one iteration. Based on this, this step preferably determines the remaining first loss function value, second loss function value, and third loss function value.

[0103] Based on this third optional embodiment, as one implementation method, the determination of the first loss function value of the first loss function, the second loss function value of the second loss function, and the third loss function value of the third loss function based on the sample real video and the actual enhanced video can be specified as follows:

[0104] c31) Using the first loss function, a first average pixel difference is determined based on the first pixel value of each video frame in the optical flow prediction video and the second pixel value of each video frame in the sample real video, and the first average pixel difference is used as the value of the first loss function. The optical flow prediction video is formed by optical flow prediction frames generated by optical flow inter-frame alignment of the actual enhanced video.

[0105] In this embodiment, a trained optical flow alignment prediction model is provided. The actual enhanced video can be input into the optical flow alignment prediction model to obtain the optical flow prediction video corresponding to the actual enhanced video. Each video frame included in the optical flow prediction video can be considered as a video frame predicted by the optical flow features of adjacent video frames in the actual enhanced video.

[0106] In this embodiment, the first loss function can be set to correlate the pixel information of each video frame in the sample real video with the pixel information of each video frame in the optical flow prediction video. Specifically, the first pixel value of each pixel in each video frame of the optical flow prediction video and the second pixel value of each pixel in each video frame of the sample real video can be used. Then, the difference between two pixel values ​​with equal pixel coordinates can be obtained to obtain the pixel difference value corresponding to each pixel coordinate. Finally, the average absolute value of the pixel difference values ​​of all pixel coordinates can be calculated to obtain the first average pixel difference value. This first average pixel difference value can be used as the first loss function value of the first loss function.

[0107] c32) Using the second loss function, a second average pixel difference is determined based on the third pixel value of each video frame in the actual enhanced video and the second pixel value of each video frame in the sample real video, and the second average pixel difference is used as the value of the second loss function.

[0108] In this embodiment, the second loss function can be set to be related to the pixel information of each video frame in the sample real video and each video frame in the actual enhanced video. Specifically, the third pixel value of each pixel in each video frame in the actual enhanced video and the fourth pixel value of each pixel in each video frame in the sample real video can be used. Then, the difference between two pixel values ​​with equal pixel coordinates can be obtained to obtain the pixel difference value corresponding to each pixel coordinate. Finally, the average absolute value of the pixel difference values ​​of all pixel coordinates can be calculated to obtain the second average pixel difference value, which can be used as the second loss function value of the second loss function.

[0109] c33) Using the third loss function, the average feature difference is determined based on the first feature value corresponding to the pixel point in each video frame of the first feature video and the second feature value corresponding to the pixel point in each video frame of the second feature video, and the average feature difference is used as the value of the third loss function.

[0110] In this embodiment, the third loss function is set as the feature video correlation after feature extraction processing of the actual enhanced video and the sample real video. Specifically, the first feature video is obtained after feature extraction processing of the actual enhanced video, which is mainly generated by feature extraction of each video frame in the actual enhanced video; the second feature video is obtained after feature extraction processing of the sample real video, which is mainly generated by feature extraction of each video frame in the sample real video. In this step, the first feature value corresponding to each pixel in each video frame of the first feature video and the second feature value corresponding to each pixel in each video frame of the second feature video can be obtained. Then, the difference between two feature values ​​with the same pixel coordinates can be obtained to obtain the feature difference value corresponding to each pixel coordinate. Finally, the average absolute value of the feature difference values ​​of all pixel coordinates can be calculated to obtain the average feature difference value, which can be used as the third loss function value of the third loss function.

[0111] d3) Determine the value of the fourth loss function from the diffusion sub-model included in the initial enhancement processing model using the fourth loss function, and determine the target loss function value based on the first loss function value, the second loss function value, the third loss function value, and the fourth loss function value.

[0112] In this embodiment, in addition to obtaining the first, second, and third loss function values ​​determined in the above steps, a fourth loss function value determined in this iteration can also be determined from the diffusion sub-model included in the initial enhancement processing model. This fourth loss function value can be considered as being determined by the diffusion sub-model through adversarial comparison based on the input and output results.

[0113] This step can be used to weight the first, second, third, and fourth loss function values. The weight value corresponding to each loss function value can be set empirically. Finally, the weighted calculation result can be used as the target loss function value in this embodiment.

[0114] e3) Update the network parameters of each sub-model included in the initial enhancement processing model according to the target loss function value, return and re-execute the step of inputting the sample compressed video into the initial enhancement processing model until the training end condition is met, and determine the initial enhancement processing model corresponding to the training end as the video enhancement processing model.

[0115] In this embodiment, the target loss function value determined above can be used as a learning metric and fed back to each sub-model of the initial enhancement processing model through gradient descent or gradient ascent, so as to achieve adaptive adjustment of the network parameters of each sub-model.

[0116] The adjusted initial enhancement processing model is equivalent to completing one iteration. Since the training termination condition has not been met, it is necessary to return to step b3 and start a new round of iteration. This process is repeated until the training termination condition is met. The initial enhancement processing model obtained after training can be considered as a video enhancement model that can participate in the enhancement processing of compressed videos involved in practical applications.

[0117] The above technical solution in this embodiment provides a training implementation of the video enhancement processing model. The video enhancement processing model formed through training can better achieve effective enhancement processing of compressed videos.

[0118] Figure 2 This is a schematic diagram of a video enhancement processing apparatus provided in an embodiment of the present disclosure. This embodiment is applicable to enhancing compressed videos. The apparatus can be implemented by software and / or hardware and can be configured in a terminal and / or server to implement the video enhancement processing method in this embodiment. Specifically, the apparatus may include: an acquisition module 21, an enhancement processing module 22, and a result generation module 23.

[0119] Among them, the acquisition module 21 is used to acquire the video to be enhanced;

[0120] The enhancement processing module 22 is used to perform feature extraction on each video frame in the video to be enhanced to obtain a feature frame of the video frame; perform distortion feature removal processing on the feature frame to obtain an enhanced feature frame of the video frame; perform resolution enhancement processing on the enhanced feature frame to obtain a first enhanced video frame of the video frame; and perform noise reduction and diffusion enhancement on the first enhanced video frame of the video frame to obtain a second enhanced video frame of the video frame.

[0121] The result generation module 23 is used to obtain the target enhanced video of the video to be enhanced based on the second enhanced video of all video frames in the video to be enhanced.

[0122] This embodiment provides a video enhancement processing device, which can be considered as an enhancement processing for compressed video. Addressing the distortion problem inherent in compressed video, this technical solution proposes feature filtering to remove distorted features, thereby reducing the interference of distorted video features on diffusion-based video enhancement processing and effectively avoiding the degradation of video enhancement processing caused by lost video information in compressed video. Simultaneously, as part of the video enhancement processing, this technical solution also performs resolution enhancement processing to improve the resolution of the compressed video. Furthermore, this technical solution also considers achieving video texture enhancement through a diffusion mechanism. Unlike existing technologies, this technical solution introduces noise feature removal processing in diffusion enhancement, enabling better identification of noise features and texture features to be enhanced when performing diffusion enhancement on compressed video. This allows diffusion enhancement to be performed only on the texture features to be enhanced. Compared to existing technologies, the enhancement processing method provided in this embodiment better reflects spatial fidelity and temporal consistency, thus better improving the realism of the video after enhancement processing and ultimately enhancing the subjective visual quality of video viewing.

[0123] Furthermore, the enhanced processing module 22 can specifically be used for:

[0124] The video frames are input into the feature extraction module of the feature filtering sub-model in the video enhancement processing model to extract features from the video frames and obtain feature frames for each video frame.

[0125] The feature frame is input into the feature filtering module of the feature filtering sub-model. Distortion features that do not meet the set feature constraints are identified from the feature frame and deleted to obtain the enhanced feature frame of the video frame.

[0126] Furthermore, the enhanced processing module 22 can also be used specifically for:

[0127] The enhanced feature frame is input into the upsampling network structure contained in the resolution enhancement sub-model of the video enhancement processing model, and the enhanced feature frame is subjected to resolution upsampling processing to obtain the first enhanced video frame of the video frame.

[0128] Furthermore, the enhanced processing module 22 may specifically include:

[0129] The denoising diffusion unit performs video feature denoising and texture diffusion enhancement on the first enhanced video frame of the video frame through the diffusion sub-model in the video enhancement processing model, based on the encoder network structure and decoder network structure included in the diffusion sub-model, to obtain the second enhanced video frame of the video frame.

[0130] Furthermore, the noise reduction and diffusion unit can specifically be used for:

[0131] The first enhanced video frame of the video frame is input into the first encoder in the diffusion sub-model. The video features of the first enhanced video frame are removed by the downsampling network structure in the first encoder to obtain the denoised feature frame of the first enhanced video frame.

[0132] The first enhanced video frame of the video frame is input into the second encoder in the diffusion sub-model. Through the diffusion control network structure in the second encoder, the feature channels of the first enhanced video frame are filtered out to obtain the diffusion control feature frame of the first enhanced video frame.

[0133] The denoising feature frame and the diffusion control feature frame are input into the first decoder in the diffusion sub-model. Through the upsampling network structure in the first decoder, the diffusion control feature frame is used as a diffusion constraint to perform feature-aware diffusion on the denoising feature frame to obtain the denoising enhancement features of the first enhanced video frame.

[0134] The denoising enhancement features are input into the second decoder in the diffusion sub-model. Through the spatiotemporal attention mechanism of the second decoder, the denoising enhancement features are subjected to texture-aware diffusion to obtain the second enhanced video frame of the first enhanced video frame.

[0135] Furthermore, the device also includes a model training module, which specifically may include:

[0136] An acquisition unit is used to acquire a sample training set and an initial augmentation processing model, wherein the sample training set includes at least one sample tuple, and the sample tuple includes a real video of the sample and a compressed video of the real video of the sample.

[0137] An input unit is used to input the sample compressed video into the initial enhancement processing model to obtain the actual enhanced video output by the initial enhancement processing model.

[0138] The first loss unit is used to determine the first loss function value of the first loss function, the second loss function value of the second loss function, and the third loss function value of the third loss function based on the sample real video and the actual enhanced video.

[0139] The second loss unit is used to determine the value of the fourth loss function from the diffusion sub-model included in the initial enhancement processing model through the fourth loss function, and to determine the target loss function value based on the first loss function value, the second loss function value, the third loss function value and the fourth loss function value;

[0140] The iterative training unit is used to update the network parameters of each sub-model included in the initial enhancement processing model according to the target loss function value, return to re-execute the step of inputting the sample compressed video into the initial enhancement processing model, until the training termination condition is met, and determine the initial enhancement processing model corresponding to the training termination as the video enhancement processing model.

[0141] Furthermore, the first loss unit can specifically be used for:

[0142] The first average pixel difference is determined by the first loss function based on the first pixel value of each video frame in the optical flow prediction video and the second pixel value of each video frame in the sample real video. The first average pixel difference is used as the value of the first loss function. The optical flow prediction video is formed by optical flow prediction frames generated by optical flow inter-frame alignment of the actual enhanced video.

[0143] The second average pixel difference is determined by the second loss function based on the third pixel value of each video frame in the actual enhanced video and the second pixel value of each video frame in the sample real video, and the second average pixel difference is used as the value of the second loss function.

[0144] The average feature difference is determined by the third loss function based on the first feature value corresponding to the pixel point in each video frame of the first feature video and the second feature value corresponding to the pixel point in each video frame of the second feature video, and the average feature difference is used as the value of the third loss function.

[0145] The first feature video is generated by extracting features from each video frame in the actual enhanced video, and the second feature video is generated by extracting features from each video frame in the sample real video.

[0146] The above-described apparatus can execute the methods provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the methods.

[0147] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.

[0148] Figure 3 This is a schematic diagram of the structure of a computer device provided in an embodiment of this disclosure. Reference is made below. Figure 3 It illustrates a computer device suitable for implementing embodiments of the present disclosure (e.g., Figure 3The diagram below shows the structure of the terminal device or server 30. The terminal device in this embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and vehicle terminals (e.g., vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 3 The computer device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0149] like Figure 3 As shown, the computer device 30 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 31, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 32 or a program loaded from a storage device 38 into a random access memory (RAM) 33. The RAM 33 also stores various programs and data required for the operation of the computer device 30. The processing unit 31, the ROM 32, and the RAM 33 are interconnected via a bus 35. An edit / output (I / O) interface 34 is also connected to the bus 35.

[0150] Typically, the following devices can be connected to I / O interface 34: input devices 36 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 37 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 38 including, for example, magnetic tapes, hard disks, etc.; and communication devices 39. Communication device 39 allows computer device 30 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 A computer device 30 with various devices is shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have instead.

[0151] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 39, or installed from a storage device 38, or installed from a ROM 32. When the computer program is executed by the processing device 31, it performs the functions defined in the methods of embodiments of this disclosure.

[0152] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0153] The computer device provided in this embodiment and the video enhancement processing method provided in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.

[0154] This disclosure provides a computer storage medium storing a computer program that, when executed by a processor, implements the video enhancement processing method provided in the above embodiments.

[0155] It should be noted that the computer-readable medium described above in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof.

[0156] In this disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0157] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0158] The aforementioned computer-readable medium may be included in the aforementioned computer device; or it may exist independently and not assembled into the computer device.

[0159] The aforementioned computer-readable medium carries one or more programs that, when executed by the computer device, cause the computer device to:

[0160] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0161] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0162] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a unit does not necessarily limit the unit itself; for example, the first acquisition unit can also be described as "a unit that acquires at least two Internet Protocol addresses".

[0163] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0164] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0165] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0166] Furthermore, although the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while some specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0167] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. A video enhancement processing method, characterized in that, include: Obtain the video to be enhanced; For each video frame in the video to be enhanced, feature extraction is performed on the video frame to obtain the feature frame of the video frame; The feature frames are subjected to distortion feature removal processing to obtain the enhanced feature frames of the video frames; The enhanced feature frame is subjected to resolution enhancement processing to obtain the first enhanced video frame of the video frame; the first enhanced video frame of the video frame is subjected to noise reduction and diffusion enhancement to obtain the second enhanced video frame of the video frame. The target enhanced video of the video to be enhanced is obtained based on the second enhanced video of all video frames in the video to be enhanced.

2. The method according to claim 1, characterized in that, For each video frame in the video to be enhanced, feature extraction is performed on the video frame to obtain the feature frame of the video frame; Distortion feature removal processing is performed on the feature frames to obtain enhanced feature frames of the video frames, including: The video frames are input into the feature extraction module of the feature filtering sub-model in the video enhancement processing model to extract features from the video frames and obtain feature frames for each video frame. The feature frame is input into the feature filtering module of the feature filtering sub-model. Distortion features that do not meet the set feature constraints are identified from the feature frame and deleted to obtain the enhanced feature frame of the video frame.

3. The method according to claim 1, characterized in that, The step of performing resolution enhancement processing on the enhanced feature frame to obtain the first enhanced video frame of the video frame includes: The enhanced feature frame is input into the upsampling network structure contained in the resolution enhancement sub-model of the video enhancement processing model, and the enhanced feature frame is subjected to resolution upsampling processing to obtain the first enhanced video frame of the video frame.

4. The method according to claim 1, characterized in that, The step of performing denoising and diffusion enhancement on the first enhanced video frame of the video frame to obtain the second enhanced video frame of the video frame includes: Using the diffusion sub-model in the video enhancement processing model, and based on the encoder network structure and decoder network structure included in the diffusion sub-model, the first enhanced video frame of the video frame is subjected to video feature denoising and texture diffusion enhancement to obtain the second enhanced video frame of the video frame.

5. The method according to claim 4, characterized in that, The step of performing video feature denoising and texture diffusion enhancement on the first enhanced video frame of the video frame according to the encoder network structure and decoder network structure included in the diffusion sub-model to obtain the second enhanced video frame of the video frame includes: The first enhanced video frame of the video frame is input into the first encoder in the diffusion sub-model. The video features of the first enhanced video frame are removed by the downsampling network structure in the first encoder to obtain the denoised feature frame of the first enhanced video frame. The first enhanced video frame of the video frame is input into the second encoder in the diffusion sub-model. Through the diffusion control network structure in the second encoder, the feature channels of the first enhanced video frame are filtered out to obtain the diffusion control feature frame of the first enhanced video frame. The denoising feature frame and the diffusion control feature frame are input into the first decoder in the diffusion sub-model. Through the upsampling network structure in the first decoder, the diffusion control feature frame is used as a diffusion constraint to perform feature-aware diffusion on the denoising feature frame to obtain the denoising enhancement features of the first enhanced video frame. The denoising enhancement features are input into the second decoder in the diffusion sub-model. Through the spatiotemporal attention mechanism of the second decoder, the denoising enhancement features are subjected to texture-aware diffusion to obtain the second enhanced video frame of the first enhanced video frame.

6. The method according to any one of claims 2-5, characterized in that, The training steps of the video enhancement processing model include: Obtain a sample training set and create an initial augmentation processing model, wherein the sample training set includes at least one sample tuple, and the sample tuple includes a real video sample and a compressed video sample of the real video sample. The sample compressed video is input into the initial enhancement processing model to obtain the actual enhanced video output by the initial enhancement processing model; Based on the sample real video and the actual enhanced video, determine the first loss function value of the first loss function, the second loss function value of the second loss function, and the third loss function value of the third loss function; The fourth loss function value is determined from the diffusion sub-model included in the initial enhancement processing model using the fourth loss function, and the target loss function value is determined based on the first loss function value, the second loss function value, the third loss function value, and the fourth loss function value. Update the network parameters of each sub-model included in the initial enhancement processing model according to the target loss function value, return to re-execute the step of inputting the sample compressed video into the initial enhancement processing model, until the training end condition is met, and determine the initial enhancement processing model corresponding to the training end as the video enhancement processing model.

7. The method according to claim 6, characterized in that, The step of determining the first loss function value of the first loss function, the second loss function value of the second loss function, and the third loss function value of the third loss function based on the sample real video and the actual enhanced video includes: The first average pixel difference is determined by the first loss function based on the first pixel value of each video frame in the optical flow prediction video and the second pixel value of each video frame in the sample real video. The first average pixel difference is used as the value of the first loss function. The optical flow prediction video is formed by optical flow prediction frames generated by optical flow inter-frame alignment of the actual enhanced video. The second average pixel difference is determined by the second loss function based on the third pixel value of each video frame in the actual enhanced video and the second pixel value of each video frame in the sample real video, and the second average pixel difference is used as the value of the second loss function. The average feature difference is determined by the third loss function based on the first feature value corresponding to the pixel point in each video frame of the first feature video and the second feature value corresponding to the pixel point in each video frame of the second feature video, and the average feature difference is used as the value of the third loss function. The first feature video is generated by extracting features from each video frame in the actual enhanced video, and the second feature video is generated by extracting features from each video frame in the sample real video.

8. A video enhancement processing apparatus, characterized in that, include: The acquisition module is used to acquire the video to be enhanced; The enhancement processing module is used to extract features from each video frame in the video to be enhanced, and obtain the feature frames of the video frames. The feature frames are subjected to distortion feature removal processing to obtain the enhanced feature frames of the video frames; The enhanced feature frame is subjected to resolution enhancement processing to obtain the first enhanced video frame of the video frame; the first enhanced video frame of the video frame is subjected to noise reduction and diffusion enhancement to obtain the second enhanced video frame of the video frame. The result generation module is used to obtain the target enhanced video of the video to be enhanced based on the second enhanced video of all video frames in the video to be enhanced.

9. A computer device, characterized in that, The computer device includes: One or more processors; a storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the video enhancement processing method as described in any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the video enhancement processing method as described in any one of claims 1-7.

11. A computer program product comprising a computer program that, when executed by a processor, implements the video enhancement processing method according to any one of claims 1-7.