A video processing method, device, storage medium and electronic device

By segmenting the video into scenes and determining the patterns, and using a diffusion model to generate reference image detail information, the problems of blurring and distortion in high-definition video production are solved, and high-resolution video production is achieved.

CN120751217BActive Publication Date: 2025-11-04HUNAN HAPPLY SUNSHINE INTERACTIVE ENTERTAINMENT MEDIA CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511182305.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-22
Publication Date
2025-11-04
Estimated Expiration
2045-08-22

AI Technical Summary

Technical Problem

In existing high-definition video production, image super-resolution technology based on reference images is prone to blurring or distortion, resulting in low video resolution.

Method used

By segmenting the video to be processed into scenes, determining the processing mode and model, generating reference images using a diffusion model, extracting their detailed information, and performing super-resolution processing on single scene video segments using a video super-resolution model, a high-resolution video sequence is generated.

Benefits of technology

It improves the resolution of video production, avoids the time-consuming and unstable nature of frame-by-frame generation by the diffusion model, and achieves high-quality video super-resolution effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120751217B_ABST
    Figure CN120751217B_ABST
Patent Text Reader

Abstract

The application discloses a video processing method and device, a storage medium and an electronic device, relates to the technical field of image processing, and generates a diffusion model generation graph as a reference graph in a processing mode corresponding to a processing requirement, and performs super-resolution processing on a single scene video segment and a reference image through a video super-resolution model. Since the super-resolution processing is a processing mode of supplementing the details in the reference image to the image after super-resolution, the richness of the details of the image is improved, time consumption and instability caused by frame-by-frame generation of the diffusion model are effectively avoided, video production is performed according to a reconstructed resolution video sequence, and the resolution of the video production is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, and more particularly to a video processing method and device, a storage medium and an electronic device. BACKGROUND

[0002] The existing video high-definition production is realized based on reference image-based image super-resolution (RefSR). The reference image-based RefSR is a technology for improving low-resolution (LR) image super-resolution reconstruction by using additional reference image information.

[0003] The core idea of the reference image-based RefSR technology is to transfer the texture information with the highest similarity in the reference image to the LR image, and fuse it with the information of the LR image, so as to reconstruct a high-resolution (HR) image with clearer texture details.

[0004] However, the existing reference image-based RefSR technology may cause blurring or distortion when sampling images, thereby resulting in low resolution of the produced video.

[0005] Therefore, how to improve the resolution of video production is a problem to be solved by the present application. SUMMARY

[0006] Therefore, the present application discloses a video processing method and device, a storage medium and an electronic device, which aims to improve the resolution of video production.

[0007] In order to achieve the above-mentioned purpose, the disclosed technical solution is as follows:

[0008] The first aspect of the present application discloses a video processing method, which comprises:

[0009] obtaining a to-be-processed video and a processing requirement;

[0010] performing scene segmentation processing on the to-be-processed video to obtain a plurality of single-scene video segments;

[0011] determining a processing mode and a processing model corresponding to the processing requirement;

[0012] performing super-resolution processing on the plurality of single-scene video segments and reference images through a video super-resolution model under the processing mode corresponding to the processing requirement to obtain a reconstructed resolution video sequence; wherein the reference images are obtained through a diffusion model; the super-resolution processing is a processing of supplementing the detail information in the reference images to the super-resolution image; and the super-resolution image is obtained by performing image super-resolution on the single-scene video segments through the video super-resolution model;

[0013] performing video production according to the reconstructed resolution video sequence.

[0014] Preferably, the scene cut processing is performed on the video to be processed to obtain a plurality of single-scene video segments, including:

[0015] The video to be processed is cut by a scene cut detection algorithm to obtain a plurality of video segments without scene switching;

[0016] The plurality of video segments without scene switching are determined as a plurality of single-scene video segments.

[0017] Preferably, the determination of the processing mode and the processing model corresponding to the processing requirement includes:

[0018] The processing requirement is analyzed to obtain a processing type corresponding to the processing requirement;

[0019] If the processing type is a fast processing type, it is determined that the processing mode corresponding to the fast processing type is a fast mode, and it is determined that the processing model corresponding to the fast processing type is a fast processing model;

[0020] If the processing type is a fine-tuning processing type, it is determined that the processing mode corresponding to the fine-tuning processing type is a fine-tuning processing mode, and it is determined that the processing model corresponding to the fine-tuning processing type is a fine-tuning processing model.

[0021] Preferably, the reconstructed resolution video sequence at least includes a reconstructed resolution video sequence and a fine-tuned reconstructed resolution video sequence, and the super-resolution processing of the plurality of single-scene video segments and the reference image is performed by a video super-resolution model under the processing mode corresponding to the processing requirement to obtain the reconstructed resolution video sequence, including:

[0022] If the processing mode corresponding to the processing requirement is the fast mode, a primary feature extraction module is used to extract primary features corresponding to the single-scene video segment; wherein the primary features at least include a convolution group or a residual group;

[0023] For a plurality of single-scene video segments, forward optical flow corresponding to the single-scene video segment is extracted;

[0024] The primary features and the forward optical flow are fused by a forward feature fusion module to obtain fused features with inter-frame information;

[0025] A first frame with a preset resolution is obtained from the reference image;

[0026] A multi-scale feature fusion module is used to process the first frame with the preset resolution and the fused features to obtain forward enhanced features with preset resolution information;

[0027] The forward enhancement feature is processed by a video super-resolution model and a feature reconstruction module to obtain a reconstructed resolution video sequence of a fast mode.

[0028] If the processing mode is the refining processing mode, reverse fusion features with inter-frame information are obtained.

[0029] The reverse fusion features and a preset resolution video tail frame obtained in advance are processed by a multi-scale feature fusion module to obtain reverse enhancement features with preset resolution information of the video tail frame.

[0030] The reverse enhancement features and the forward enhancement features are added to obtain added features.

[0031] The added features are processed by a video super-resolution model and a fine-tuning feature reconstruction module to obtain a refined reconstruction resolution video sequence of the refining processing mode.

[0032] Preferably, the training process of the fast processing model comprises:

[0033] In the process of training the fast mode, the plurality of single-scene video clips and the preset resolution first frame obtained from the reference image are sent to the fast mode model.

[0034] The fast processing model is trained by a loss function constraint, a weight and a learning rate; wherein the loss function constraint is a loss function constraint using the preset resolution first frame corresponding to the plurality of single-scene video clips.

[0035] Preferably, the video production according to the reconstruction resolution video sequence comprises:

[0036] For the fast mode, the first frame in the reconstruction resolution video sequence is generated in the fast mode to complete the process of video production.

[0037] For the refining processing mode, the first frame and the tail frame in the reconstruction resolution video sequence are generated in the refining processing mode to complete the process of video production.

[0038] The second aspect of the present application discloses a video processing device, the device comprises:

[0039] A first acquisition unit is configured to acquire a to-be-processed video and processing requirements.

[0040] A first processing unit is configured to perform scene segmentation processing on the to-be-processed video to obtain a plurality of single-scene video clips.

[0041] A determination unit is configured to determine a processing mode and a processing model corresponding to the processing requirements.

[0042] a second processing unit, configured to perform super-resolution processing on the multiple single-scene video clips and reference images by using a video super-resolution model in a processing mode corresponding to the processing requirement, to obtain a reconstructed resolution video sequence; wherein the reference images are obtained by using a diffusion model; the super-resolution processing is a processing of supplementing detailed information in the reference images to the super-resolved images; the super-resolved images are obtained by performing image super-resolution on the single-scene video clips by using the video super-resolution model;

[0043] a video production unit, configured to produce a video according to the reconstructed resolution video sequence.

[0044] Preferably, the first processing unit comprises:

[0045] a segmentation module, configured to perform scene segmentation on the video to be processed by using a scene segmentation detection algorithm, to obtain multiple video clips without scene switching;

[0046] a first determination module, configured to determine the multiple video clips without scene switching as the multiple single-scene video clips.

[0047] A third aspect of the present application discloses a storage medium, which comprises stored instructions, wherein the instructions, when executed, control a device in which the storage medium is located to perform the video processing method according to any one of the first aspect.

[0048] A fourth aspect of the present application discloses an electronic device, comprising a memory and one or more instructions, wherein the one or more instructions are stored in the memory and configured to be executed by one or more processors to perform the video processing method according to any one of the first aspect.

[0049] According to the above technical solutions, the present application discloses a video processing method, device, storage medium and electronic device, which obtains a video to be processed and a processing requirement, performs scene segmentation on the video to be processed to obtain multiple single-scene video clips, determines a processing mode and a processing model corresponding to the processing requirement, performs super-resolution processing on the multiple single-scene video clips and reference images by using a video super-resolution model in the processing mode corresponding to the processing requirement, to obtain a reconstructed resolution video sequence, wherein the reference images are obtained by using a diffusion model, the super-resolution processing is a processing of supplementing detailed information in the reference images to the super-resolved images, the super-resolved images are obtained by performing image super-resolution on the single-scene video clips by using the video super-resolution model, and a video is produced according to the reconstructed resolution video sequence.

[0050] Through the above scheme, in the processing mode corresponding to the processing requirement, the diffusion model generation image is taken as a reference image, and the single scene video segment and the reference image are subjected to super-resolution processing through the video super-resolution model. Since the super-resolution processing is a processing mode of extracting the detailed information in the reference image to supplement the image after super-resolution, the richness of the image details is improved, the time consumption and instability caused by the frame-by-frame generation of the diffusion model are effectively avoided, the video is produced according to the reconstructed resolution video sequence, and the purpose of improving the resolution of the video production is achieved. BRIEF DESCRIPTION OF DRAWINGS

[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of the provided drawings.

[0052] Figure 1 A flowchart of a video processing method disclosed in an embodiment of the present application;

[0053] Figure 2 A schematic diagram of video production of a video reconstruction model disclosed in an embodiment of the present application;

[0054] Figure 3 A schematic diagram of a multi-scale feature fusion module disclosed in an embodiment of the present application;

[0055] Figure 4 A schematic diagram of a video processing main body framework disclosed in an embodiment of the present application;

[0056] Figure 5 A structural schematic diagram of a video processing device disclosed in an embodiment of the present application;

[0057] Figure 6 A structural schematic diagram of an electronic device disclosed in an embodiment of the present application. DETAILED DESCRIPTION

[0058] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0059] In this application, the terms "comprising", "containing" or any other variant thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements, but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. Without more limitations, the element defined by the phrase "comprising a" does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.

[0060] As can be known from the background art, the core idea of the RefSR technology of the reference image is to transfer the texture information with the highest similarity in the reference image to the LR image, and fuse it with the information of the LR image, to reconstruct a high-resolution (HR) image with clearer texture details. However, the existing RefSR technology of the reference image is prone to blurring or distortion when sampling the image, thereby resulting in a low resolution of the produced video. Therefore, how to improve the resolution of video production is a problem that needs to be solved by the present application.

[0061] To solve the above problems, the present application discloses a video processing method, device, storage medium and electronic equipment, which generates a diffusion model generated image as a reference image in a processing mode corresponding to a processing requirement, and performs super-resolution processing on a single scene video segment and a reference image through a video super-resolution model. Since the super-resolution processing is a processing mode of extracting detailed information in the reference image to supplement the image after super-resolution, the richness of the image details is improved, the time-consuming and instability caused by the frame-by-frame generation of the diffusion model are effectively avoided, video production is performed according to the reconstructed resolution video sequence, and the resolution of the video production is improved. The specific implementation is described in detail in the following embodiments.

[0062] Reference Figure 1 As shown in the figure, the video processing method disclosed by the embodiment of the present application mainly includes the following steps:

[0063] S101: Obtain a to-be-processed video and a processing requirement.

[0064] It should be noted that up-sampling is to collect analog signals when the sampling frequency is higher than the highest frequency of the signal, also known as interpolation / interpolation / over-sampling. Sampling is to convert a signal that is continuous in time and amplitude into a signal that is discrete in time and amplitude under the action of a sampling pulse.

[0065] The processing requirement includes a fast mode requirement, a fine-tuning processing requirement, etc.

[0066] S102: scene segmentation processing is performed on the to-be-processed video to obtain a plurality of single-scene video clips.

[0067] In S102, the to-be-processed video is subjected to scene segmentation by a scene segmentation detection algorithm to obtain a plurality of video clips without scene switching, each of which is a single-scene video clip.

[0068] The single-scene video clip is obtained by segmenting the to-be-processed video by scene, and each scene after segmentation is a single content clip.

[0069] Since the to-be-processed video is a low-resolution video, the single-scene clip is a sequence of low-resolution video frames without transitions. A sequence of low-resolution video frames without scene switching is obtained from the single-scene video clip.

[0070] S103: determine the processing mode and processing model corresponding to the processing requirement.

[0071] The processing mode includes a fast mode and a fine processing mode. The processing model corresponding to the fast mode is a fast processing model; and the processing model corresponding to the fine processing mode is a fine processing model.

[0072] For the segmented single-scene video clip, two models can be selected for video ultra-high definition generation according to the processing requirement. When the processing requirement is a requirement for selecting the fast mode, only the first frame is subjected to high-resolution image generation; and when the processing requirement is a requirement for selecting the fine model, the first and last frames are subjected to high-resolution image generation.

[0073] For the image generation diffusion model part, the present application does not perform separate training, and an image generation model that has been open-sourced and can be directly used is selected.

[0074] The process of determining the processing mode and the pre-trained processing model corresponding to the processing requirement according to the processing requirement is shown as A1-A3.

[0075] A1: analyze the processing requirement to obtain a processing type corresponding to the processing requirement.

[0076] A2: if the processing type is a fast processing type, determine that the processing mode corresponding to the fast processing type is a fast mode, and determine that the processing model corresponding to the fast processing type is a fast processing model.

[0077] A3: if the processing type is a fine processing type, determine that the processing mode corresponding to the fine processing type is a fine processing mode, and determine that the processing model corresponding to the fine processing type is a fine processing model.

[0078] The training process of the fast processing model is shown as B1-B2.

[0079] B1: In the process of training the fast mode, a plurality of single-scene video clips (i.e. low-resolution video frame sequences) and a preset resolution first frame obtained from a reference image are sent to the fast mode model.

[0080] The preset resolution first frame is the high-resolution first frame. In the process of training the fast mode, the low-resolution video frame and the high-resolution video first frame are sent to the fast mode model together.

[0081] B2: The fast processing model is trained by loss function constraint, weight and learning rate; wherein the loss function constraint is to use the preset resolution first frame corresponding to the plurality of single-scene video clips to constrain the loss function.

[0082] In the process of training the fast mode, the low-resolution video frame and the high-resolution video first frame are sent to the fast mode model together, and the high-resolution frame corresponding to the low-resolution video frame is used for loss function constraint, such as selecting 15 video frames for each training, using 8 A6000 machines to train for a week, the initial learning rate is 0.0001, the learning rate is halved every 2 days, and the loss function uses mean square error (MSE).

[0083] The model for the fast mode is trained with fixed weights, and it is noted that the forward optical flow estimation, the forward feature fusion module and the multi-scale feature fusion module do not participate in the training.

[0084] For example, the model can be trained using a learning rate of 0.00005, using 8 A6000 cards (i.e. servers or workstations equipped with 8 NVIDIA RTX A6000 GPUs), training for 5 days, and halving the learning rate every 2 days.

[0085] S104: In the processing mode corresponding to the processing requirement, the plurality of single-scene video clips and the reference image are processed by the video super-resolution model to obtain a reconstructed resolution video sequence; wherein the reference image is obtained by the diffusion model; the super-resolution processing is a process of extracting detailed information in the reference image to supplement the image after super-resolution; the image after super-resolution is obtained by the video super-resolution model processing the single-scene video clip.

[0086] Super-resolution is a method of improving the resolution of the original image by hardware or software, and the process of obtaining a high-resolution image from a series of low-resolution images is super-resolution reconstruction.

[0087] The scheme designs a flexible video super-resolution system, which can autonomously select different modes such as fast mode, fine processing mode, etc. according to the requirements, to obtain super-resolution videos with different detail richness. This video super-resolution system provides greater flexibility, allows guiding texture creation through text prompts, and adjustable noise levels to balance repair and generation, achieving a trade-off between speed and quality.

[0088] Specifically in the processing mode corresponding to the processing requirement, the single-scene video segment and the reference image are processed by the video super-resolution model to obtain the process of the reconstructed resolution video sequence, as shown in Figure 2 Figure 2 A schematic diagram of video production of a video reconstruction model is shown. Figure 2 A video high-definition production scheme based on high-resolution images is shown.

[0089] In Figure 2 , (1) if the processing mode corresponding to the processing requirement is the fast mode, the primary features corresponding to the single-scene video segment are extracted by the primary feature extraction module; wherein the primary features at least include a convolution group or a residual group.

[0090] Among them, in the processing mode corresponding to the processing requirement, the low-resolution video frame sequence is extracted by the primary feature extraction module to the primary feature corresponding to each video frame. The primary feature extraction module can select a convolution group or a residual group.

[0091] (2) Extract the forward optical flow corresponding to the single-scene video segment.

[0092] Among them, the low-resolution video frame sequence is subjected to forward optical flow extraction, wherein the optical flow estimation module can select but is not limited to using a deep learning optical flow estimation framework (Optical Flow Estimation Using a Spatial Pyramid Network, SpyNet).

[0093] (3) The primary features and the forward optical flow are fused by the forward feature fusion module to obtain the fused features with inter-frame information.

[0094] (4) Obtain the preset resolution first frame (i.e. high-resolution first frame) from the reference image.

[0095] For example, the input video frame is 960x540 low resolution, and the resolution of the high-resolution first frame is 1920x1080 after two times super-resolution enlargement. After 4 times enlargement, the resolution of the high-resolution first frame is 384x2160. That is, several times super-resolution, the resolution of the high-resolution first frame corresponds to several times enlargement.

[0096] ​(5) The preset resolution first frame and the fusion feature are processed by the multi-scale feature fusion module to obtain forward enhancement features with preset resolution information.

[0097] In step (5), the high-resolution first frame and the fusion feature with inter-frame information are sent to the multi-scale feature fusion module to obtain forward enhancement features with high-resolution information.

[0098] The process of obtaining forward enhancement features with preset resolution information is specifically described in combination with Figure 3 . Figure 3 A schematic diagram of the multi-scale feature fusion module is shown.

[0099] Figure 3 In the multi-scale feature fusion module, the similarity between the high-resolution video frame feature and the forward fusion feature is calculated by using the cross-attention mechanism for adaptive fusion. At the same time, a multi-scale feature pyramid is constructed by two-level down-sampling to realize sufficient information interaction between different scale features.

[0100] The process of obtaining forward enhancement features with preset resolution information is as follows:

[0101] First, the high-resolution video frame feature is down-sampled to obtain high-resolution video frame down-sampling feature 1, and then down-sampled to obtain high-resolution video frame down-sampling feature 2, to construct a multi-scale feature pyramid. The multi-scale feature pyramid is used to establish a horizontal cross-scale connection and a vertical time sequence propagation path at 1 / 2 and 1 / 4 scales through two-level progressive down-sampling, i.e., stride=2 convolution combined with atrous pooling;

[0102] Then, the high-resolution video frame feature, the high-resolution video frame down-sampling feature 1, and the high-resolution video frame down-sampling feature 2 are respectively dynamically calculated with the correlation weight of the fusion feature (and the subsequent derived fusion feature 1 and fusion feature 2) by using the bidirectional cross-attention mechanism, to realize adaptive feature enhancement and redundancy suppression of the local detail features of the high-resolution video frame and the forward fusion global context features, to obtain fusion output feature 2 and fusion output feature 1; wherein the fusion output feature 2 is up-sampled and added to the feature output by the cross-attention mechanism 2 to participate in the construction of the fusion output feature 1.

[0103] Finally, the fusion output feature 1 is up-sampled and added to the feature output by the cross-attention mechanism 2 to obtain the final fusion output feature.

[0104] (6) The forward enhancement features are processed by the feature reconstruction module to obtain a reconstructed resolution video sequence in the fast mode, i.e., a reconstructed high-resolution video frame.

[0105] In step (6), the feature reconstruction module is a key link for the network to generate a high-resolution video sequence. A basic architecture of 16 residual network stacks is adopted to construct a deep feature recovery channel.

[0106] The feature reconstruction module supports flexible configuration of the architecture. A dense residual network can be selected and configured to allow dense transmission of texture details and semantic information between residual blocks in each layer through a feature reuse mechanism to strengthen the continuity of the features. Alternatively, a Transformer model module can be used to capture long-range temporal and spatial dependencies through a self-attention mechanism to accurately align cross-frame features and adapt to the high-resolution reconstruction requirements in complex motion scenarios, thereby providing diversified feature fusion and recovery capabilities for the final output of the reconstructed high-resolution video frames.

[0107] (7) If the processing mode is the fine-tuning processing mode, obtain reverse fusion features with inter-frame information;

[0108] In step (7), for the fine-tuning processing mode, a reverse optical flow estimation and a reverse feature multi-scale fusion module are added to obtain reverse fusion features with inter-frame information.

[0109] (8) Process the reverse fusion features and the pre-acquired preset resolution video tail frame through the multi-scale feature fusion module to obtain reverse enhancement features with video tail frame preset resolution information, wherein the reverse enhancement features with video tail frame preset resolution information are reverse enhancement features with video tail frame high-resolution information.

[0110] In step (8), the reverse fusion features and the high-resolution video tail frame are processed through the multi-scale feature fusion module to obtain reverse enhancement features with video tail frame high-resolution information.

[0111] The reverse fusion features and the pre-acquired preset resolution video tail frame are processed through the multi-scale feature fusion module to obtain reverse enhancement features with video tail frame preset resolution information.

[0112] The structure of the multi-scale feature fusion module can be referred to the description in step (5) above, and will not be described here again.

[0113] (9) Add the reverse enhancement features and the forward enhancement features to obtain added features.

[0114] (10) Process the added features through the fine-tuning feature reconstruction module to obtain a fine-tuning reconstruction resolution video sequence of the fine-tuning processing mode, i.e., a fine-tuning reconstruction high-resolution video sequence.

[0115] The reconstructed high-resolution video sequence includes the reconstructed resolution video sequence of the fast mode and the fine-tuning reconstruction high-resolution video sequence of the fine-tuning processing mode.

[0116] To precisely improve the reconstruction quality of high-resolution video sequences, two-stage training processes are designed for the network to fully utilize bidirectional temporal information, as follows:

[0117] Stage one: first-frame-guided forward network training (solid line module):

[0118] Focus on the first-frame-guided forward reconstruction path of high-resolution video (corresponding to Figure 2 the solid line part, covering primary feature extraction, forward optical flow estimation, forward feature fusion, multi-scale feature fusion, and feature reconstruction modules).

[0119] Data preparation:

[0120] Select a dataset containing low-resolution video frame sequences, corresponding high-resolution first frames, and complete high-resolution sequence labels, preprocess the data according to task requirements, and ensure temporal frame alignment:

[0121] Training configuration:

[0122] Set a unified training strategy, including an optimizer (such as Adam, with an initial learning rate of 1e-4), a loss function (L1 pixel loss + perceptual loss combination can be used to enhance detail and semantic consistency), and learning rate scheduling (cosine annealing, with fine-tuning in the later period).

[0123] Training process:

[0124] Input the low-resolution frame sequence and high-resolution first frame into the solid line network, obtain the basic features through primary feature extraction, model the inter-frame motion through forward optical flow estimation, align the motion and appearance information through forward feature fusion, and aggregate multi-granularity features through the multi-scale feature fusion module (based on the downsampling pyramid and cross-attention) through hierarchical interaction. Finally, output the high-resolution sequence through feature reconstruction. Take the loss between the reconstructed sequence and the real high-resolution sequence as the optimization objective, update the solid line module parameters through backpropagation, and continuously iterate until training convergence, so that the model has basic first-frame-guided reconstruction capability;

[0125] Stage two: incremental training of reverse temporal module (dashed line module + fixed solid line):

[0126] After stage one converges, if further temporal information of the video needs to be extracted and the reconstruction quality bottleneck needs to be broken, the solid line module parameters are fixed, and the reverse temporal path ( Figure 2 the dashed line part, containing reverse optical flow estimation and reverse feature multi-scale fusion module) is introduced, the training strategy of stage one is reused, and bidirectional temporal constraints are strengthened;

[0127] Data expansion:

[0128] Supplementing high-resolution video tail frame data, and co-inputting it with low-resolution sequence and first frame, constructs a bidirectional temporal correlation (low-resolution frames, after primary feature extraction, are split into separate paths for forward / backward optical flow estimation).

[0129] Training logic:

[0130] In the forward path reuse stage, the trained parameters (parameters for primary feature extraction, forward optical flow, and forward fusion modules are fixed) are used. Features from low-resolution frames are captured by reverse optical flow estimation to obtain reverse motion. Then, through the reverse feature multi-scale fusion module, they interact with the forward path features in the multi-scale fusion module, jointly providing bidirectional temporal information for feature reconstruction. Aiming at the loss between the final reconstructed sequence and the real sequence, only the parameters of the dashed-line module are updated. By supplementing with reverse temporal information, the consistency of motion between frames and the reconstruction effect of occluded regions are optimized, further improving the overall reconstruction quality.

[0131] The core training logic for Phase One and Phase Two:

[0132] By employing a two-stage strategy of "first training the first frame to guide the forward basic process → then supplementing the reverse temporal module after fixing parameters", a unified training strategy (optimizer, loss function, learning rate scheduling) is reused to gradually mine bidirectional temporal information in the video.

[0133] The first phase involves establishing a reconstruction baseline guided by the first frame;

[0134] The second stage introduces a reverse module to strengthen temporal constraints, iteratively improve the accuracy of high-resolution sequence reconstruction, adapt to the needs of tasks sensitive to temporal features such as video super-resolution and action recognition, and make the model more robust in motion consistency and detail completion.

[0135] S105: Produce video based on the reconstructed resolution video sequence.

[0136] For the fast mode, the first frame of the reconstructed resolution video sequence is used to generate an image to complete the video production process.

[0137] For the fine-tuning mode, images are generated from the first and last frames of the reconstructed resolution video sequence to complete the video production process.

[0138] To facilitate understanding of the process of reconstructing the resolution video sequence, combined with Figure 4 To explain, Figure 4 A schematic diagram of the main framework for video processing is shown.

[0139] Figure 4 In the process, obtain the video to be processed and the processing requirements;

[0140] The video to be processed is subjected to scene cut detection and scene cut processing to obtain a plurality of cut scenes; the plurality of cut scenes are single-scene video clips;

[0141] A low-resolution video frame sequence is obtained from the single-scene video clip;

[0142] According to the processing requirement, a processing mode corresponding to the processing requirement and a pre-trained processing model are determined;

[0143] If the processing mode corresponding to the processing requirement is the fast mode, the primary features corresponding to the single-scene video clip are extracted by a primary feature extraction module; wherein the primary features at least include a convolution group or a residual group;

[0144] For a plurality of single-scene video clips, forward optical flow corresponding to the single-scene video clip is extracted;

[0145] The primary features and the forward optical flow are fused by a forward feature fusion module to obtain fused features with inter-frame information;

[0146] A preset resolution first frame (high-resolution first frame) is obtained from a reference image; wherein the reference image is obtained by a diffusion model;

[0147] The preset resolution first frame and the fused features are processed by a multi-scale feature fusion module to obtain forward enhanced features with preset resolution information (high-resolution information);

[0148] The forward enhanced features are processed by a video super-resolution model and a feature reconstruction module to obtain a reconstructed resolution video sequence in the fast mode;

[0149] If the processing mode is the fine-tuning processing mode, reverse fused features with inter-frame information are obtained;

[0150] The reverse fused features and a preset resolution video tail frame obtained in advance are processed by a multi-scale feature fusion module to obtain reverse enhanced features with video tail frame preset resolution information; wherein the reverse enhanced features with video tail frame preset resolution information are reverse enhanced features with video tail frame high-resolution information;

[0151] The reverse enhanced features and the forward enhanced features are added to obtain added features;

[0152] The added features are processed by a video super-resolution model and a fine-tuning feature reconstruction module to obtain a fine-tuning reconstructed high-resolution video sequence in the fine-tuning processing mode.

[0153] Due to the problem of blur or distortion when up-sampling the image in existing image / video super-resolution technology, the image after super-resolution lacks detailed information. The technical solution generates a reference image by using a diffusion model and extracts detailed information from it to supplement the image after super-resolution, greatly improving the richness of image details, making existing videos reach 4K or even 8K resolution with rich details. And the diffusion model also has some problems in the video super-resolution task, such as inter-frame flicker and slow inference. Inter-frame flicker refers to the content change between adjacent frames during video playback, which causes visual flicker. Slow inference refers to the high computational complexity of the diffusion model when generating high-resolution images, resulting in slow inference speed and inability to meet real-time processing needs. When processing video super-resolution, the technical solution only uses the diffusion model-generated image as a reference image to extract detailed information to assist the video super-resolution task, effectively avoiding the time-consuming and instability caused by the diffusion model generating frame by frame.

[0154] The present solution improves the details after super-resolution: by using the diffusion model image generation technology to generate ultra-high definition reference images, the ultra-high definition prior and detailed information provided by the improved image can effectively solve the problem of insufficient details encountered by existing deep learning-based models when processing video super-resolution tasks. By learning the noise characteristics in the data distribution, the technical solution can generate high-quality images, thereby improving the details of the generated high-resolution images;

[0155] Reduce inter-frame flicker: by referring to the detailed information of the diffusion model-generated image instead of directly generating the image after super-resolution, the technical solution solves the inter-frame flicker problem of the diffusion model in the video super-resolution task. By improving the algorithm of the diffusion model, the content change between adjacent frames is smoother, thereby reducing the visual flicker phenomenon and improving the smoothness and viewing experience of video playback;

[0156] Improve inference speed: the technical solution only uses one or two reference images generated by the diffusion model, reducing the time-consuming problem of the diffusion model applied in the video field generating frame by frame, thereby improving the inference speed. Compared with existing diffusion models, the technical solution can significantly improve the inference speed while ensuring image quality;

[0157] Strong versatility: the technical solution is applicable to various video super-resolution tasks, whether it is super-resolution of static images or super-resolution of dynamic videos, it can be realized by the technical solution. In addition, the technical solution can also adapt to videos with different resolutions and frame rates, and has wide application prospects.

[0158] The present solution uses the powerful generation capability of the diffusion model image generation:

[0159] 1. This solution utilizes high-quality reference images and text-based image-assisted techniques to solve the problem of traditional image / video super-resolution methods failing to generate more detail. This framework enables the generation of ultra-high-definition videos rich in detail.

[0160] 2. This solution designs a flexible video super-resolution system that can autonomously select different modes according to requirements to obtain super-resolution videos with varying levels of detail. This system offers greater flexibility, allowing for text-guided texture creation and adjustable noise levels to balance restoration and generation, achieving a trade-off between speed and quality.

[0161] The proposed ultra-high-definition (UHD) production scheme based on diffusion model image generation fully utilizes the rich details of the generated images while reducing time consumption and inter-frame flicker. Furthermore, due to the controllability of the generation process, it also solves the problem of obtaining high-quality reference images in real-world scenarios for image super-resolution based on reference images. This scheme not only improves the detail and resolution of UHD video production but also significantly enhances the inter-frame stability of the resulting UHD video, providing a new direction for the development of UHD video technology.

[0162] The beneficial effects of this application's embodiments are as follows: In the processing mode corresponding to the processing requirements, the diffusion model-generated image is used as the reference image. The video super-resolution model is used to perform super-resolution processing on a single scene video clip and the reference image. Since this super-resolution processing is a method of extracting detailed information from the reference image and supplementing it to the super-resolution image, the detail richness of the image is improved. This effectively avoids the time consumption and instability caused by the diffusion model's frame-by-frame generation. Video production is carried out based on the reconstructed resolution video sequence, thereby achieving the goal of improving the resolution of video production.

[0163] Based on the above embodiments Figure 1 The present application discloses a video processing method and a corresponding video processing apparatus, such as... Figure 5 As shown, the video processing device includes:

[0164] The first acquisition unit 501 is used to acquire the video to be processed and the processing requirements;

[0165] The first processing unit 502 is used to perform scene segmentation processing on the video to be processed, and obtain multiple single scene video segments.

[0166] The determining unit 503 is used to determine the processing mode and processing model corresponding to the processing requirements;

[0167] The second processing unit 504 is configured to perform super-resolution processing on the plurality of single-scene video clips and the reference image by using a video super-resolution model in a processing mode corresponding to the processing requirement, to obtain a reconstructed resolution video sequence; the reference image is obtained by using a diffusion model; the super-resolution processing is a process of supplementing detailed information in the reference image to a super-resolution image; and the super-resolution image is obtained by performing image super-resolution on the single-scene video clips by using the video super-resolution model.

[0168] The video production unit 505 is configured to produce a video according to the reconstructed resolution video sequence.

[0169] Further, the first processing unit 502 comprises:

[0170] The segmentation module is configured to perform scene segmentation on the video to be processed by using a scene segmentation detection algorithm, to obtain a plurality of video clips without scene switching.

[0171] The first determination module is configured to determine the plurality of video clips without scene switching as the plurality of single-scene video clips.

[0172] Further, the determination unit 503 comprises:

[0173] The analysis module is configured to analyze the processing requirement, to obtain a processing type corresponding to the processing requirement.

[0174] The second determination module is configured to determine, if the processing type is a fast processing type, that a processing mode corresponding to the fast processing type is a fast mode, and that a processing model corresponding to the fast processing type is a fast processing model.

[0175] The third determination module is configured to determine, if the processing type is a fine-tuning processing type, that a processing mode corresponding to the fine-tuning processing type is a fine-tuning processing mode, and that a processing model corresponding to the fine-tuning processing type is a fine-tuning processing model.

[0176] Further, the reconstructed resolution video sequence comprises at least a reconstructed resolution video sequence and a fine-tuned reconstructed resolution video sequence, and the second processing unit 504 comprises:

[0177] The first extraction module is configured to extract, if the processing mode corresponding to the processing requirement is the fast mode, primary features corresponding to the single-scene video clips by using a primary feature extraction module; the primary features comprise at least a convolution group or a residual group.

[0178] The second extraction module is configured to extract, for the plurality of single-scene video clips, forward optical flow corresponding to the single-scene video clips.

[0179] The fusion module is configured to fuse the primary features and the forward optical flow by using a forward feature fusion module, to obtain fused features with inter-frame information.

[0180] The first obtaining module is configured to obtain a preset-resolution first frame from a reference image.

[0181] The first processing module is configured to process the preset-resolution first frame and the fused features through the multi-scale feature fusion module to obtain forward-enhanced features with preset-resolution information.

[0182] The second processing module is configured to process the forward-enhanced features through the video super-resolution model and the feature reconstruction module to obtain a reconstructed-resolution video sequence in the fast mode.

[0183] The second obtaining module is configured to obtain reverse-fused features with inter-frame information if the processing mode is the fine-tuning processing mode.

[0184] The third processing module is configured to process the reverse-fused features and the preset-resolution video tail frame obtained in advance through the multi-scale feature fusion module to obtain reverse-enhanced features with preset-resolution information of the video tail frame.

[0185] The calculating module is configured to add the reverse-enhanced features and the forward-enhanced features to obtain added features.

[0186] The fourth processing module is configured to process the added features through the video super-resolution model and the fine-tuning feature reconstruction module to obtain a fine-tuning reconstructed-resolution video sequence in the fine-tuning processing mode.

[0187] Further, the second determining module of the training process of the fast processing model is specifically configured to, in the process of training the fast mode, send a plurality of single-scene video clips and the preset-resolution first frame obtained from the reference image to the fast mode model; and train the fast processing model through loss function constraints, weights, and learning rates; wherein the loss function constraints are constraints of loss functions on the preset-resolution first frames corresponding to the plurality of single-scene video clips.

[0188] Further, the video production unit 505 is configured to, for the fast mode, perform image generation on the first frame in the reconstructed-resolution video sequence in the fast mode to complete the process of video production; and for the fine-tuning processing mode, perform image generation on the first frame and the tail frame in the reconstructed-resolution video sequence in the fine-tuning processing mode to complete the process of video production.

[0189] The beneficial effects of the embodiments of the present application are: in the processing mode corresponding to the processing requirement, the diffusion model generation image is taken as a reference image, and the single scene video segment and the reference image are processed by the video super-resolution model. Since the super-resolution processing is a processing mode of extracting the detailed information in the reference image to supplement the image after super-resolution, the richness of the image details is improved, the time consumption and instability caused by the frame-by-frame generation of the diffusion model are effectively avoided, the video is produced according to the reconstructed resolution video sequence, and the resolution of the video production is improved.

[0190] The embodiments of the present application also provide a storage medium, which comprises stored instructions, wherein the instructions control a device where the storage medium is located to perform the video processing method.

[0191] The embodiments of the present application also provide an electronic device, a structure diagram of which is shown in the figure. Figure 6 The electronic device specifically comprises a memory 601 and one or more than one instruction 602, wherein the one or more than one instruction 602 is stored in the memory 601 and is configured to be executed by one or more than one processor 603 to execute the video processing method.

[0192] For the foregoing method embodiments, in order to simply describe, they are all expressed as a series of action combinations, but those skilled in the art should know that the present application is not limited to the action sequence described, because according to the present application, certain steps can be performed in other sequences or at the same time. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions and modules involved are not necessarily necessary for the present application.

[0193] It should be noted that each embodiment in the specification is described in a progressive manner, and each embodiment focuses on the difference from other embodiments. The same and similar parts between each embodiment can be referred to each other. For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiment.

[0194] The steps in the method of each embodiment of the present application can be adjusted, combined and reduced in sequence according to actual needs.

[0195] Finally, it should be noted that in this paper, relationship terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or sequence between the entities or operations.

[0196] The above description of disclosed embodiments enables one of ordinary skill in the art to make or use the application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to other embodiments without departing from the scope of the application. Thus, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0197] The above description is merely illustrative of the preferred embodiments of the present application and is not intended to limit the scope of the application. Variations and modifications exist, which would be apparent to those in the art, which fall within the scope of the present application.

Claims

1. A method of video processing, the method comprising: The method comprises: acquiring a to-be-processed video and processing requirements; performing scene segmentation processing on the to-be-processed video to obtain a plurality of single-scene video clips; determining a processing mode and a processing model corresponding to the processing requirements; performing super-resolution processing on the plurality of single-scene video clips and reference images through a video super-resolution model in the processing mode corresponding to the processing requirements to obtain a reconstructed resolution video sequence; wherein the reference images are obtained through a diffusion model; the super-resolution processing is processing of supplementing detailed information in the reference images to the super-resolved images; the super-resolved images are obtained through the video super-resolution model to perform image super-resolution on the single-scene video clips; performing video production according to the reconstructed resolution video sequence; the super-resolution processing on the plurality of single-scene video clips and reference images through the video super-resolution model in the processing mode corresponding to the processing requirements to obtain the reconstructed resolution video sequence comprises: if the processing mode corresponding to the processing requirements is a fast mode and the processing model is a fast processing model, extracting primary features corresponding to the single-scene video clips through a primary feature extraction module; wherein the primary features at least include convolution groups or residual groups; extracting forward optical flows corresponding to the single-scene video clips; fusing the primary features and the forward optical flows through a forward feature fusion module to obtain fused features with inter-frame information; acquiring a first frame with a preset resolution from the reference images; processing the first frame with the preset resolution and the fused features through a multi-scale feature fusion module to obtain forward enhanced features with preset resolution information; processing the forward enhanced features through a video super-resolution model and a feature reconstruction module to obtain a reconstructed resolution video sequence in the fast mode; if the processing mode is a fine-tuning processing mode and the processing model is a fine-tuning processing model, acquiring reverse fused features with inter-frame information; processing the reverse fused features and a pre-acquired last frame with a preset resolution of a video through a multi-scale feature fusion module to obtain reverse enhanced features with preset resolution information of the last frame of the video; adding the reverse enhanced features and the forward enhanced features to obtain added features; processing the added features through a video super-resolution model and a fine-tuning feature reconstruction module to obtain a fine-tuning reconstructed resolution video sequence in the fine-tuning processing mode.

2. The method of claim 1, wherein, the scene segmentation processing on the to-be-processed video to obtain a plurality of single-scene video clips comprises: performing scene segmentation on the to-be-processed video through a scene segmentation detection algorithm to obtain a plurality of video clips without scene switching; determining the plurality of video clips without scene switching as the plurality of single-scene video clips.

3. The method of claim 1, wherein, the determination of the processing mode and the processing model corresponding to the processing requirements comprises: analyzing the processing requirements to obtain a processing type corresponding to the processing requirements; if the processing type is a fast processing type, determining that a processing mode corresponding to the fast processing type is a fast mode, and determining that a processing model corresponding to the fast processing type is a fast processing model; If the processing type is a refining processing type, it is determined that the processing mode corresponding to the refining processing type is a refining processing mode, and it is determined that the processing model corresponding to the refining processing type is a refining processing model.

4. The method of claim 3, wherein, The training process of the fast processing model comprises: In the process of training the fast mode, the plurality of single-scene video segments and the preset resolution first frame obtained from the reference image are sent to the fast mode model; The fast processing model is trained by using a loss function constraint, a weight, and a learning rate; wherein the loss function constraint is a constraint of a loss function using the preset resolution first frame corresponding to the plurality of single-scene video segments.

5. The method of claim 1, wherein, The video production according to the reconstructed resolution video sequence comprises: For the fast mode, the first frame in the reconstructed resolution video sequence is generated in the fast mode to complete the process of video production; For the refining processing mode, the first frame and the last frame in the reconstructed resolution video sequence are generated in the refining processing mode to complete the process of video production.

6. A video processing apparatus, comprising: The device comprises: A first acquisition unit configured to acquire a to-be-processed video and a processing requirement; A first processing unit configured to perform scene segmentation processing on the to-be-processed video to obtain a plurality of single-scene video segments; A determination unit configured to determine a processing mode and a processing model corresponding to the processing requirement; A second processing unit configured to perform super-resolution processing on the plurality of single-scene video segments and a reference image by using a video super-resolution model in the processing mode corresponding to the processing requirement to obtain a reconstructed resolution video sequence; wherein the reference image is obtained by using a diffusion model; the super-resolution processing is processing of supplementing detailed information in the reference image to a super-resolution image; and the super-resolution image is obtained by performing image super-resolution on the single-scene video segment by using the video super-resolution model; A video production unit configured to produce a video according to the reconstructed resolution video sequence; The second processing unit comprises: A first extraction module configured to, if the processing mode corresponding to the processing requirement is a fast mode and the processing model is a fast processing model, extract primary features corresponding to the single-scene video segment by using a primary feature extraction module; wherein the primary features at least include a convolution group or a residual group; A second extraction module configured to, for the plurality of single-scene video segments, extract forward optical flow corresponding to the single-scene video segment; A fusion module configured to fuse the primary features and the forward optical flow by using a forward feature fusion module to obtain fused features with inter-frame information; A first acquisition module configured to acquire a preset resolution first frame from the reference image; A first processing module configured to process the preset resolution first frame and the fused features by using a multi-scale feature fusion module to obtain forward enhanced features with preset resolution information; A second processing module configured to process the forward enhanced features by using a video super-resolution model and a feature reconstruction module to obtain a reconstructed resolution video sequence of the fast mode; The second obtaining module is configured to, if the processing mode is a fine-tuning processing mode and the processing model is a fine-tuning processing model, obtain reverse fusion features with inter-frame information. The third processing module is configured to process the reverse fusion features and a pre-obtained preset resolution video tail frame through a multi-scale feature fusion module to obtain reverse enhancement features with preset resolution information of the video tail frame. The calculation module is configured to add the reverse enhancement features and the forward enhancement features to obtain added features. The fourth processing module is configured to process the added features through a video super-resolution model and a fine-tuning feature reconstruction module to obtain a fine-tuning reconstruction resolution video sequence of the fine-tuning processing mode.

7. The apparatus of claim 6, wherein, The first processing unit comprises: The segmentation module is configured to perform scene segmentation on the to-be-processed video through a scene segmentation detection algorithm to obtain a plurality of video clips without scene switching. The first determination module is configured to determine the plurality of video clips without scene switching as a plurality of single-scene video clips.

8. A storage medium, characterized by The storage medium comprises stored instructions, wherein the instructions, when executed, control a device in which the storage medium is located to perform the video processing method according to any one of claims 1 to 5.

9. An electronic device, comprising: The device comprises a memory and one or more instructions, wherein the one or more instructions are stored in the memory and are configured to be executed by one or more processors to perform the video processing method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Space-time video stream enhancement method and device, terminal and medium

    CN116957937A

  • Embedded NPU video super-resolution reconstruction method, system and device

    CN120013764A