Video generation method and device

By acquiring the fusion of the temporal consistency characteristics and optical flow guidance characteristics of the video frame sequence, combined with sliding time window and frequency domain conversion technology, the artifact problem caused by inconsistency between video frames in dynamic scenarios is solved, and the temporal consistency and detail recovery of high-resolution video frames is achieved.

CN120343359APending Publication Date: 2025-07-18XFUSION DIGITAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510680805.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing video frame super-segment method has limited generation effects in dynamic scenarios, and inter-frame inconsistencies lead to video artifacts.

Method used

By acquiring the time consistency characteristics and optical flow guidance characteristics of the video frame sequence, the target video frame sequence is generated after fusion, and the sliding time window and attention mechanism are used to enhance time consistency, and combining frequency domain conversion and high-frequency enhancement technologies to improve detailed recovery capabilities.

Benefits of technology

Ensure the time consistency and coherence of generated video frames, reduce artifacts, and improve the resolution and detail recovery capabilities of video frames.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343359A_ABST
    Figure CN120343359A_ABST
Patent Text Reader

Abstract

The invention provides a video generation method and device. The method comprises the following steps: acquiring a time consistency feature of a video frame sequence and capturing an optical flow guide feature of the video frame sequence, wherein the time consistency feature is used for representing time-varying information of each video frame of the video frame sequence; fusing the time consistency feature and the optical flow guide feature to obtain a fusion feature; and generating a target video frame sequence with a resolution higher than that of the video frame sequence based on the fusion features. When the time consistency feature and the optical flow guide feature are fused, the time consistency feature can make up for time correlation information among partial video frames missing in the optical flow guide feature, and when a target video frame sequence is generated based on a fusion feature obtained by fusing the time consistency feature and the optical flow guide feature, the time correlation information among the partial video frames missing in the optical flow guide feature can be made up. The time consistency and coherence of video frame generation can be ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video super-resolution technology, and in particular, to a video generation method and device. Background Art

[0002] With the rapid growth of the demand for high-definition videos, how to efficiently generate high-resolution videos has become an important research direction. Most traditional video frame super-resolution methods are limited to single-frame processing or simple inter-frame stitching, especially in dynamic scenes, the generation effect is limited. Summary of the Invention

[0003] Embodiments of this application provide a video generation method and device.

[0004] According to the first aspect of the embodiments of this application, a video generation method is provided, including:

[0005] Obtain the temporal consistency feature of the video frame sequence and capture the optical flow guiding feature of the video frame sequence, where the temporal consistency feature is used to characterize the information of each video frame in the video frame sequence changing over time;

[0006] Fuse the temporal consistency feature and the optical flow guiding feature to obtain a fused feature;

[0007] Based on the fused feature, generate a target video frame sequence with a resolution higher than that of the video frame sequence.

[0008] In the embodiments of this application, the temporal consistency feature directly obtained from the video frame sequence includes the information of all video frames in the video frame sequence changing over time. Therefore, when fusing this temporal consistency feature with the optical flow guiding feature, this temporal consistency feature can make up for the missing temporal correlation information between some video frames in the optical flow guiding feature. Further, when generating a target video frame sequence based on the fused feature obtained by fusing the temporal consistency feature and the optical flow guiding feature, the temporal consistency and coherence of the generated video frames can be ensured, thereby alleviating the problem of video artifacts caused by inter-frame inconsistency when generating a high-resolution video frame sequence in the related art.

[0009] In some embodiments, according to the video generation method described in the first aspect of the embodiments of this application, obtaining the temporal consistency feature of the video frame sequence includes:

[0010] Encode the video frames in the video frame sequence to obtain the spatial features of the video frames;

[0011] Fuse the spatial features of the video frames within a set sliding time window in the video frame sequence to obtain an inter-frame fused feature;

[0012] Extract window time consistency features corresponding to the set sliding time window from the inter-frame fusion features;

[0013] Obtain the time consistency features of the video frame sequence based on the window time consistency features.

[0014] In the embodiments of the present application, using a sliding time window to obtain the time consistency features of a video frame sequence helps ensure that the obtained time consistency features can include information on the changes of all video frames in the video frame sequence over time, enhancing the consistency in the time dimension during video generation.

[0015] In some embodiments, according to the video generation method described in the first aspect of the embodiments of the present application, the set sliding time window includes at least one of the following:

[0016] The first sliding time window;

[0017] The second sliding time window;

[0018] The third sliding time window;

[0019] The inter-frame span defined by the first sliding time window is less than the inter-frame span defined by the second sliding time window, and the inter-frame span defined by the second sliding time window is less than the inter-frame span defined by the third sliding time window.

[0020] In the embodiments of the present application, the set sliding time window can be at least one of the first sliding time window, the second sliding time window, and the third sliding time window. Different sliding time windows can provide inter-frame feature alignments at different scales. In a long video frame sequence and a complex dynamic scene, these sliding time windows can ensure the time consistency and coherence of the generated video frames.

[0021] In some embodiments, according to the video generation method described in the first aspect of the embodiments of the present application, obtaining the time consistency features of the video frame sequence based on the window time consistency features includes:

[0022] When the set sliding time window includes at least two of the first sliding time window, the second sliding time window, and the third sliding time window, perform weighted fusion on the window time consistency features of the at least two sliding time windows to obtain the time consistency features.

[0023] In the embodiments of the present application, by fusing the window time consistency features provided by the set sliding time windows at multiple scales, it helps reduce the inter-frame cumulative error, avoid the accumulation of optical flow errors in a long sequence, improve the performance when processing ultra-long video frames, reduce artifacts, and enhance inter-frame smoothness.

[0024] In some embodiments, according to the video generation method described in the first aspect of the embodiments of the present application, the set sliding time window includes the first sliding time window; extracting window time consistency features corresponding to the set sliding time window from the inter-frame fusion features includes:

[0025] Using a model implemented based on a temporal attention mechanism to extract window time consistency features corresponding to the first sliding time window from the inter-frame fusion features;

[0026] The set sliding time window includes the second sliding time window; extracting window time consistency features corresponding to the set sliding time window from the inter-frame fusion features includes:

[0027] Using a temporal convolutional model to extract window time consistency features corresponding to the second sliding time window from the inter-frame fusion features;

[0028] The set sliding time window includes the third sliding time window; extracting window time consistency features corresponding to the set sliding time window from the inter-frame fusion features includes:

[0029] Using a model implemented based on a cross-frame self-attention mechanism to extract window time consistency features corresponding to the third sliding time window from the inter-frame fusion features.

[0030] In the embodiments of the present application, the attention mechanism is used to weight and fuse information at different time scales, thereby enhancing the consistency in the time dimension and avoiding flickering artifacts.

[0031] In some embodiments, according to the video generation method described in the first aspect of the embodiments of the present application, based on the fusion features, generating a target video frame sequence with a resolution higher than that of the video frame sequence includes:

[0032] Converting the fusion features into the frequency domain to obtain a frequency domain representation of the fusion features;

[0033] Performing high-frequency enhancement on the frequency domain representation to obtain an enhanced frequency domain representation;

[0034] Converting the enhanced frequency domain representation back to the spatial domain to obtain enhanced fusion features;

[0035] Decoding the enhanced fusion features to generate the target video frame sequence.

[0036] In the embodiments of the present application, when generating a target video frame sequence based on the fusion features, by performing frequency domain conversion and high-frequency enhancement operations, the deficiencies of the existing high-frequency shuttle mechanism are supplemented, thereby improving the ability to restore details.

[0037] In some embodiments, for the video generation method according to the first aspect of the embodiments of the present application, high-frequency enhancement is performed on the frequency-domain representation to obtain an enhanced frequency-domain representation, including:

[0038] Obtain the amplitude spectrum and phase spectrum in the frequency-domain representation;

[0039] Obtain the high-frequency amplitude spectrum and low-frequency amplitude spectrum in the amplitude spectrum;

[0040] Enhance the high-frequency amplitude spectrum to obtain an enhanced high-frequency amplitude spectrum;

[0041] Sum the enhanced high-frequency amplitude spectrum and the low-frequency amplitude spectrum to obtain an enhanced amplitude spectrum;

[0042] Generate the enhanced frequency-domain representation based on the enhanced amplitude spectrum and the phase spectrum.

[0043] In the embodiments of the present application, high-frequency enhancement is achieved by enhancing the high-frequency amplitude spectrum, realizing selective strengthening of high-frequency components in the frequency domain, thereby effectively improving the detail restoration ability of small-scale objects.

[0044] In some embodiments, for the video generation method according to the first aspect of the embodiments of the present application, enhancing the high-frequency amplitude spectrum in the amplitude spectrum to obtain an enhanced high-frequency amplitude spectrum includes:

[0045] Obtain the attention weight and enhancement intensity factor in the fusion feature, where the attention weight is used to indicate the local spatial region to be enhanced, and the enhancement intensity factor is used to characterize the degree to which the high-frequency amplitude spectrum needs to be enhanced;

[0046] Perform element-wise multiplication calculation on the attention weight, the enhancement intensity factor, and the high-frequency amplitude spectrum to obtain an element-wise multiplication calculation result;

[0047] Sum the element-wise multiplication calculation result and the high-frequency amplitude spectrum to obtain the enhanced high-frequency amplitude spectrum.

[0048] In the embodiments of the present application, adaptive enhancement of the high-frequency amplitude is achieved through the attention weight and enhancement intensity factor in the fusion feature, which helps to ensure the integrity of key textures and details.

[0049] In some embodiments, for the video generation method according to the first aspect of the embodiments of the present application, obtaining the high-frequency amplitude spectrum and low-frequency amplitude spectrum in the amplitude spectrum includes:

[0050] Perform element-wise multiplication on the amplitude spectrum and a preset high-frequency mask to obtain the high-frequency amplitude spectrum;

[0051] Multiply the amplitude spectrum element by element with a preset low-frequency mask to obtain the low-frequency amplitude spectrum, where the sum of the high-frequency amplitude spectrum and the low-frequency amplitude spectrum is 1.

[0052] In the embodiments of the present application, the amplitude spectrum is divided into high-frequency and low-frequency components through a high-frequency mask and a low-frequency mask, which helps to make the extracted high-frequency components accurately correspond to high-frequency information such as details, edges, and textures.

[0053] According to the second aspect of the embodiments of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the steps of the method in the first aspect.

[0054] According to the third aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which a computer program / instructions are stored, and when the computer program / instructions are executed by a processor, the steps of the method in the first aspect are implemented.

[0055] According to the fourth aspect of the embodiments of the present application, a computer program product is provided, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the method in the first aspect are implemented.

[0056] As will be described in detail below, in the video generation method of the embodiments of the present application, the temporal consistency features directly obtained from the video frame sequence include the information of all video frames in the video frame sequence changing over time. Therefore, when fusing the temporal consistency features with the optical flow-guided features, the temporal consistency features can compensate for the missing temporal correlation information between some video frames in the optical flow-guided features. Further, when generating the target video frame sequence based on the fused features obtained by fusing the temporal consistency features and the optical flow-guided features, the temporal consistency and coherence of the generated video frames can be ensured, thereby alleviating the problem of video artifacts caused by frame-to-frame inconsistency when generating a high-resolution video frame sequence in the related art.

[0057] It should be understood that both the foregoing general description and the following detailed description are exemplary and are intended to provide further explanation of the claimed technology. Brief Description of the Drawings

[0058] By describing the embodiments of the present application in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the embodiments of the present application will become more apparent. The drawings are used to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the embodiments of the present application and do not constitute a limitation to the embodiments of the present application. In the drawings, the same reference numerals generally represent the same components or steps.

[0059] Figure 1It is a flowchart showing the video generation method according to an embodiment of the present application.

[0060] Figure 2 It is a schematic diagram showing a video generation device according to an embodiment of the present application.

[0061] Figure 3 It is another schematic diagram showing a video generation device according to an embodiment of the present application.

[0062] Figure 4 It is a hardware block diagram of an electronic device according to an embodiment of the present application.

[0063] Figure 5 It is a schematic diagram of a computer program product according to an embodiment of the present application. Detailed implementation manners

[0064] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are only some of the embodiments of the present application, rather than all of the embodiments. Usually, the components of the embodiments of the present application described and illustrated in the accompanying drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed embodiments of the present application, but merely represents selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative efforts fall within the scope of protection of the embodiments of the present application.

[0065] It should be noted that like reference numerals and letters denote like items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0066] The term "and / or" in this article merely describes an association relationship and means that three relationships may exist. For example, A and / or B may mean: A exists alone, A and B exist simultaneously, or B exists alone. In addition, the term "at least one" in this article means any one of a plurality or any combination of at least two of a plurality. For example, including at least one of A, B, and C may mean including any one or more elements selected from the set composed of A, B, and C.

[0067] The embodiments of the present application provide a video generation method and device. The temporal consistency features directly obtained from the video frame sequence include the information of all video frames in the video frame sequence changing over time. Therefore, when fusing the temporal consistency features with the optical flow guidance features, the temporal consistency features can make up for the missing temporal correlation information between some video frames in the optical flow guidance features. Further, when generating the target video frame sequence based on the fused features obtained by fusing the temporal consistency features and the optical flow guidance features, the temporal consistency and coherence of the generated video frames can be ensured, thereby alleviating the problem of video artifacts caused by inter-frame inconsistency when generating a high-resolution video frame sequence in the related art.

[0068] To facilitate the understanding of this embodiment, first, a video generation method disclosed in the embodiments of the present application will be introduced in detail. The execution subject of the video generation method provided in the embodiments of the present application is generally an electronic device with a certain computing ability. The electronic device includes, for example: a terminal device, a server, or other processing devices. In some possible implementation manners, the video generation method can be implemented by a processor calling computer-readable instructions stored in a memory.

[0069] See Figure 1 As shown in the flowchart of the video generation method provided in the embodiments of the present application, the method includes the following steps:

[0070] Step 101, obtain the temporal consistency features of the video frame sequence and capture the optical flow guidance features of the video frame sequence. The temporal consistency features are used to characterize the information of each video frame in the video frame sequence changing over time.

[0071] In this embodiment, the video frame sequence can be a segment of a video stream with a lower resolution. Each video frame in the video frame sequence is displayed at a fixed time interval. The temporal consistency features reflect the temporal correlation between frames in the video frame sequence. For example, for a video frame sequence containing a moving vehicle, the temporal consistency features describe the information of each object in the video frame sequence changing over time. The object here can be a moving object (such as a vehicle) in the video frame sequence, or a relatively stationary object (such as a road, a tree, etc.) in the video frame sequence. As time changes in the video frame sequence, parameters such as the position, posture, and / or size of these objects change in the video frame sequence.

[0072] In this embodiment, the feature dimension of the temporal consistency feature is higher than that of the optical flow guidance feature. The temporal consistency feature is a high-dimensional feature describing the change of video frames over time, while the optical flow guidance feature is a low-dimensional feature including information on the change of video frames over time. The temporal consistency feature includes more abstract and higher-level features compared to the optical flow guidance feature. The temporal consistency feature is a global high-dimensional vector, with each dimension corresponding to a different feature. The temporal consistency feature includes information on the change of each video frame over time, while the optical flow guidance feature usually includes information on the change of some video frames in the video frame sequence over time. When fusing the temporal consistency feature and the optical flow guidance feature, the temporal consistency feature can compensate for the missing information on the change of video frames over time in the optical flow guidance feature.

[0073] In some embodiments, the temporal consistency feature of the video frame sequence can be obtained by presetting a sliding time window. Specifically, step 101 may include the following steps:

[0074] Encode the video frames in the video frame sequence to obtain the spatial features of the video frames;

[0075] Fuse the spatial features of the video frames within the preset sliding time window in the video frame sequence to obtain an inter-frame fusion feature;

[0076] Extract the window temporal consistency feature corresponding to the preset sliding time window from the inter-frame fusion feature;

[0077] Obtain the temporal consistency feature of the video frame sequence based on the window temporal consistency feature.

[0078] In this embodiment, an encoder can be used to encode each video frame in the video frame sequence to obtain the spatial features of each video frame.

[0079] In this embodiment, the preset sliding time window specifies the number of video frames that can be obtained at one time and the inter-frame interval of these video frames. For example, if the preset sliding time window is the first sliding time window of +1, it means that the spatial features of two adjacent video frames in the video frame sequence can be obtained each time. Another example is that if the preset sliding time window is the second sliding time window of +3, it means that 3 video frames in the video frame sequence can be obtained each time, and the sequence number difference between two adjacent video frames among these 3 video frames is 3. For example, based on the second sliding time window, the spatial features of the 1st, 4th, and 7th video frames, the spatial features of the 2nd, 5th, and 8th video frames, the spatial features of the 3rd, 6th, and 9th video frames, etc. are obtained.

[0080] In some embodiments, the preset sliding time window may include at least one of the following:

[0081] The first sliding time window;

[0082] The second sliding time window;

[0083] The third sliding time window;

[0084] The frame interval defined by the first sliding time window is less than that defined by the second sliding time window, and the frame interval defined by the second sliding time window is less than that defined by the third sliding time window.

[0085] Among them, the frame interval includes the number of video frames within the set sliding time window and the frame interval between these video frames. That the frame interval defined by the first sliding time window is less than that defined by the second sliding time window means that the number of video frames defined by the first sliding time window is less than that defined by the second sliding time window, and at the same time, the frame interval defined by the first sliding time window is less than that defined by the second sliding time window. That the frame interval defined by the second sliding time window is less than that defined by the third sliding time window means that the number of video frames defined by the second sliding time window is less than that defined by the third sliding time window, and at the same time, the frame interval defined by the second sliding time window is less than that defined by the third sliding time window.

[0086] In one example, the number of video frames defined in the first sliding time window is 2, and the frame interval is +1, that is, the spatial features of two adjacent video frames are fused based on the first sliding time window. The number of video frames defined in the second sliding time window is 3, and the frame interval is +3, that is, the spatial features of three video frames are fused based on the second sliding time window, and the frame interval between two adjacent video frames among these three video frames is 3. For example, the second sliding time window fuses the spatial features of the 1st, 4th, 7th, the 2nd, 5th, 8th, or the 3rd, 6th, 9th video frames. The number of video frames defined in the third sliding time window is 10, and the frame interval is +10, that is, the spatial features of ten video frames are fused based on the third sliding time window, and the frame interval between two adjacent video frames among these ten videos is 10. For example, the third sliding time window fuses the spatial features of the 1st, 11th, 21st, 31st, 41st, 51st, 61st, 71st, 81st, 91st video frames.

[0087] In this embodiment, the first sliding time window captures local time dynamics, the second sliding time window smooths the intermediate frame transitions, and the third sliding time window integrates the global features of long time series, reducing the frame-by-frame cumulative error. The attention mechanism is used to weight and fuse the information of different time scales, thereby enhancing the consistency in the time dimension and avoiding flicker artifacts.

[0088] In this embodiment, when the number of set sliding time windows is one, the window time consistency feature of the set sliding time window is the time consistency feature of the video frame sequence. When the set sliding time window includes at least two sliding time windows among the first sliding time window, the second sliding time window, and the third sliding time window, the window time consistency feature of a certain set sliding time window can be randomly used as the time consistency feature of the video frame sequence.

[0089] In some embodiments, in order to further enhance the consistency in the time dimension and effectively avoid flicker artifacts, when the set sliding time window includes at least two sliding time windows among the first sliding time window, the second sliding time window, and the third sliding time window, the window time consistency features of at least two sliding time windows are weighted and fused to obtain the time consistency feature.

[0090] In this embodiment, a pre-trained weighted fusion model can be used to perform weighted fusion on multiple window time consistency features. After inputting the window time consistency features of multiple sliding time windows into the pre-trained weighted fusion model, the result output by the model is the fused time consistency feature.

[0091] In application, when training the weighted fusion model, the training samples used are multiple features to be fused, and the expected fused feature is also set in the training samples. After inputting the multiple features to be fused into the weighted fusion model, the weighted fusion model outputs the predicted fused feature. Based on the expected fused feature and the predicted fused feature, the model loss can be calculated, and this model loss can be used to optimize the weighted fusion model until the weighted fusion model converges.

[0092] In some embodiments, different models are used to extract the window time consistency features from the inter-frame fusion features obtained for different sliding time windows, which helps to improve the accuracy of the extracted window time consistency features.

[0093] Specifically, when the set sliding time window includes the first sliding time window, a model based on the time attention mechanism is used to extract the window time consistency feature corresponding to the first sliding time window from the inter-frame fusion features;

[0094] When the set sliding time window includes the second sliding time window, a temporal convolutional model is used to extract the window time consistency feature corresponding to the second sliding time window from the inter-frame fusion features;

[0095] When the set sliding time window includes the third sliding time window, a model based on the cross-frame self-attention mechanism is used to extract the window time consistency feature corresponding to the third sliding time window from the inter-frame fusion features.

[0096] In this embodiment, the optical flow guided attention method can be used to capture the optical flow guided features in the video frame sequence, and the optical flow guided features are used to characterize the motion information in the video frame sequence.

[0097] Step 102: Fuse the temporal consistency features and the optical flow guided features to obtain the fused features.

[0098] Step 103: Generate a target video frame sequence with a resolution higher than that of the video frame sequence based on the fused features.

[0099] In this embodiment, a decoder can be directly used to decode the fused features to obtain the target video frame sequence.

[0100] In an alternative embodiment, when generating the target video frame sequence based on the fused features, an adaptive frequency domain enhancement algorithm can be used to enhance the reconstruction ability of small-scale objects and texture details in the low-resolution video, solve the problem of insufficient capture of high-frequency details by existing spatial domain methods, and more accurately restore the high-frequency components of the image.

[0101] Briefly, the fused features in the spatial domain are transformed to the frequency domain through Fourier transform to obtain the amplitude spectrum and the phase spectrum. The amplitude spectrum is separated into a low-frequency amplitude spectrum and a high-frequency amplitude spectrum through a high-frequency mask. The attention weight and enhancement intensity factor calculated based on the fused features are used to adaptively enhance the high-frequency amplitude spectrum. The enhanced high-frequency amplitude spectrum is recombined with the original low-frequency amplitude spectrum, and combined with the original phase spectrum, and then transformed back to the spatial domain through inverse Fourier transform for subsequent use by the decoder.

[0102] Specifically, step 103 may include the following steps:

[0103] Convert the fused features to the frequency domain to obtain the frequency domain representation of the fused features;

[0104] Perform high-frequency enhancement on the frequency domain representation to obtain the enhanced frequency domain representation;

[0105] Convert the enhanced frequency domain representation back to the spatial domain to obtain the enhanced fused features;

[0106] Decode the enhanced fused features to generate the target video frame sequence.

[0107] The fused features obtained in this embodiment are spatial domain features. The Fourier transform (FourierTransform, FFT) can be used to convert the spatial domain fused features into a frequency domain representation in complex form. This frequency domain representation can be decomposed into an amplitude spectrum and a phase spectrum At the same time, in the frequency domain, the amplitude spectrum can be divided into a low-frequency component and a high-frequency component The low-frequency components mainly contain the overall structure and basic information of the image, while the high-frequency components correspond to high-frequency information such as details, edges, and textures.

[0108] Among them,

[0109]

[0110] Among them: Mask HF ∈{0, 1} H×W is a learnable high-frequency mask, where the region where the value of Mask HF is 1 corresponds to the high-frequency components, and the region where the value is 0 corresponds to the low-frequency components. Correspondingly, the low-frequency mask Mask LF = 1 - Mask HF .

[0111] To solve the problem that the existing spatial domain methods are insufficient in capturing high-frequency details and at the same time minimize the computational amount as much as possible, this embodiment focuses on enhancing the high-frequency components. Specifically, in an optional embodiment, high-frequency enhancement is performed on the frequency domain representation to obtain an enhanced frequency domain representation, which may include the following steps:

[0112] Obtain the amplitude spectrum and phase spectrum in the frequency domain representation;

[0113] Obtain the high-frequency amplitude spectrum and low-frequency amplitude spectrum in the amplitude spectrum;

[0114] Enhance the high-frequency amplitude spectrum to obtain an enhanced high-frequency amplitude spectrum;

[0115] Sum the enhanced high-frequency amplitude spectrum and the low-frequency amplitude spectrum to obtain an enhanced amplitude spectrum;

[0116] Generate an enhanced frequency domain representation based on the enhanced amplitude spectrum and the phase spectrum.

[0117] In an optional embodiment, an adaptive mechanism is designed to selectively enhance the high-frequency components. Specifically, when enhancing the high-frequency amplitude spectrum in the amplitude spectrum to obtain an enhanced high-frequency amplitude spectrum, it may include the following steps:

[0118] Obtain the attention weight and enhancement intensity factor in the fused feature. The attention weight is used to indicate the local spatial region that needs to be enhanced, and the enhancement intensity factor is used to characterize the degree to which the high-frequency amplitude spectrum needs to be enhanced;

[0119] Perform element-wise multiplication calculation on the attention weight, enhancement intensity factor, and high-frequency amplitude spectrum to obtain the element-wise multiplication calculation result;

[0120] Sum the element-wise multiplication calculation result and the high-frequency amplitude spectrum to obtain an enhanced high-frequency amplitude spectrum.

[0121] In terms of F inWhen representing the fused features in the spatial domain, an attention weight A ∈ R can be learned C×H×W and an enhancement intensity factor M ∈ R C×H×W . These two factors are used to adjust the high-frequency amplitude spectrum. The enhanced high-frequency amplitude spectrum can be expressed as:

[0122]

[0123] where ⊙ is element-wise multiplication. The attention weight A helps the model focus on the local spatial regions that need to be enhanced, while the enhancement intensity factor M controls the degree of enhancement of these high-frequency components. The low-frequency amplitude spectrum remains unchanged in this step.

[0124] The enhanced high-frequency amplitude spectrum is recombined with the original low-frequency amplitude spectrum to form the enhanced amplitude spectrum

[0125]

[0126] Then, is combined with the original phase spectrum to form the enhanced frequency-domain representation

[0127]

[0128] Finally, the inverse Fourier transform (IFFT) is used to transform back to the spatial domain to obtain the final enhanced fused features.

[0129] In the solution provided in this embodiment, the temporal consistency features directly obtained from the video frame sequence include the information of all video frames in the video frame sequence changing over time. Therefore, when fusing the temporal consistency features with the optical flow-guided features, the temporal consistency features can compensate for the missing temporal correlation information between some video frames in the optical flow-guided features. Further, when generating the target video frame sequence based on the fused features obtained by fusing the temporal consistency features and the optical flow-guided features, the temporal consistency and coherence of the generated video frames can be ensured, thereby alleviating the problem of video artifacts caused by frame-to-frame inconsistency when generating a high-resolution video frame sequence in the related art.

[0130] This embodiment of the present application further provides a video generation device, as Figure 2 described, the device may include:

[0131] an input processing module 21, a temporal consistency module 22, an optical flow-guided feature extraction module 23, a high-frequency detail restoration module 24, and a decoding and generation module 25.

[0132] Wherein:

[0133] The input processing module 21 is used to receive a low-resolution video frame sequence and extract the spatial features of each video frame using an encoder.

[0134] The temporal consistency module 22 includes a first sliding time window module, a second sliding time window module, and a third sliding time window module. The first sliding time window module captures local temporal dynamics, the second sliding time window module smooths mid-range inter-frame transitions, and the third sliding time window module integrates global features of long time series to reduce inter-frame cumulative errors. The attention mechanism is used to weight and fuse information at different temporal scales, thereby enhancing the consistency in the temporal dimension and avoiding flickering artifacts.

[0135] Specifically, the first sliding time window module performs feature fusion on the current frame and adjacent frames and uses the temporal attention mechanism to capture local temporal dynamic relationships.

[0136] The second sliding time window module performs feature fusion on the frames within the sliding window (such as +3 frames) and uses temporal convolutional operations to extract mid-term correlations between frames.

[0137] The third sliding time window module performs global dynamic modeling on frames with a larger time span (such as +10 frames) and uses the cross-frame self-attention mechanism to weight and fuse the global temporal features of all frames.

[0138] The optical flow-guided feature extraction module 23 is used to capture optical flow-guided features that reflect dynamic information in the optical flow field. In complex dynamic scenes, the optical flow-guided feature extraction module is introduced, and the motion vector of each pixel between adjacent frames is input into the Transformer module to capture the dynamic information in the optical flow field.

[0139] The high-frequency detail restoration module 24 mainly realizes frequency domain conversion, high-frequency enhancement, and frequency domain-spatial domain fusion, maps the frequency domain enhanced features back to the spatial domain, and fuses them with the spatial domain features for decoding. This module can effectively improve the detail restoration ability of small-scale objects (such as text and small patterns).

[0140] The decoding and generation module 25 generates a high-resolution video frame sequence with good temporal consistency and rich spatial details by combining the fusion features obtained by fusing multi-scale temporal features and high-frequency enhanced features through a decoder.

[0141] As an example, a solution for generating a high-resolution digital human video based on a low-resolution digital human video is given.

[0142] In the scenario of digital human training videos, the VideoGigaGAN model, which integrates a temporal consistency module and a high-frequency detail restoration module, is used to enhance the video resolution, effectively addressing key issues regarding video quality and authenticity. These modules significantly improve temporal consistency and spatial detail restoration capabilities, providing high-quality data support for digital human model training.

[0143] Digital human training typically requires capturing subtle changes in facial expressions, body movements, and cross-frame dynamic interactions. The temporal consistency module ensures the temporal consistency of video frames by modeling short-term, mid-term, and long-term inter-frame dependencies. This effectively reduces inter-frame flickering or incoherence, enabling training videos to smoothly and realistically represent dynamic changes. For example, when training lip synchronization or eye movements, the temporal consistency module ensures accurate and smooth transitions between frames.

[0144] The high-frequency detail restoration module enhances the clarity of small-scale features such as skin texture, hair filaments, and micro-expressions, which are often lost in low-resolution videos. Through frequency domain processing, the high-frequency detail restoration module can selectively amplify high-frequency details and re-integrate them into the spatial domain, ensuring the integrity of key textures and details. This is particularly important in applications that require fine facial animation, as the accuracy of textures and details directly affects the realism of the digital human.

[0145] Low-quality or inconsistent video data can severely impact the performance of digital human training models. The improved super-resolution method can convert low-resolution video sequences into high-resolution videos with temporal dynamic consistency and enhanced spatial details. The generated high-quality videos can be more effectively used to train digital human models, helping them learn realistic and high-quality features for applications such as virtual avatars, CGI characters, and real-time interactive agents.

[0146] The embodiment of this application also provides a video generation device, which is used to execute the video generation method provided in any of the above embodiments. As Figure 3 shown, the device includes:

[0147] An acquisition module 31, configured to acquire the temporal consistency features of the video frame sequence and capture the optical flow guidance features of the video frame sequence, where the temporal consistency features are used to characterize the information of each video frame in the video frame sequence changing over time;

[0148] A fusion module 32, configured to fuse the temporal consistency features and the optical flow guidance features to obtain fusion features;

[0149] A generation module 33, configured to generate a target video frame sequence with a resolution higher than that of the video frame sequence based on the fusion features.

[0150] In some alternative embodiments, the obtaining module 31 is configured to:

[0151] Encode the video frames in the video frame sequence to obtain the spatial features of the video frames;

[0152] Fuse the spatial features of the video frames within a set sliding time window in the video frame sequence to obtain inter-frame fusion features;

[0153] Extract window time consistency features corresponding to the set sliding time window from the inter-frame fusion features;

[0154] Obtain the time consistency features of the video frame sequence based on the window time consistency features.

[0155] In some alternative embodiments, the set sliding time window includes at least one of the following:

[0156] The first sliding time window;

[0157] The second sliding time window;

[0158] The third sliding time window;

[0159] The inter-frame span defined by the first sliding time window is less than the inter-frame span defined by the second sliding time window, and the inter-frame span defined by the second sliding time window is less than the inter-frame span defined by the third sliding time window.

[0160] In some alternative embodiments, the obtaining module 31 is configured to:

[0161] When the set sliding time window includes at least two of the first sliding time window, the second sliding time window, and the third sliding time window, perform weighted fusion on the window time consistency features of the at least two sliding time windows to obtain the time consistency features.

[0162] In some alternative embodiments, when the set sliding time window includes the first sliding time window, the obtaining module 31 is configured to:

[0163] Use a model implemented based on a time attention mechanism to extract window time consistency features corresponding to the first sliding time window from the inter-frame fusion features;

[0164] In some alternative embodiments, when the set sliding time window includes the second sliding time window, the obtaining module 31 is configured to:

[0165] Use a temporal convolutional model to extract window time consistency features corresponding to the second sliding time window from the inter-frame fusion features;

[0166] When the set sliding time window includes the third sliding time window, the obtaining module 31 is configured to:

[0167] Adopt a model implemented based on a cross-frame self-attention mechanism to extract window time consistency features corresponding to the third sliding time window from the inter-frame fusion features.

[0168] In some alternative embodiments, the generating module 33 is configured to:

[0169] Convert the fusion features into the frequency domain to obtain a frequency domain representation of the fusion features;

[0170] Perform high-frequency enhancement on the frequency domain representation to obtain an enhanced frequency domain representation;

[0171] Convert the enhanced frequency domain representation back to the spatial domain to obtain enhanced fusion features;

[0172] Decode the enhanced fusion features to generate the target video frame sequence.

[0173] In some alternative embodiments, the generating module 33 is configured to:

[0174] Obtain the amplitude spectrum and phase spectrum in the frequency domain representation;

[0175] Obtain the high-frequency amplitude spectrum and low-frequency amplitude spectrum in the amplitude spectrum;

[0176] Enhance the high-frequency amplitude spectrum to obtain an enhanced high-frequency amplitude spectrum;

[0177] Sum the enhanced high-frequency amplitude spectrum and the low-frequency amplitude spectrum to obtain an enhanced amplitude spectrum;

[0178] Generate the enhanced frequency domain representation based on the enhanced amplitude spectrum and the phase spectrum.

[0179] In some alternative embodiments, the generating module 33 is configured to:

[0180] Obtain the attention weight and enhancement intensity factor in the fusion features, where the attention weight is used to indicate the local spatial region that needs to be enhanced, and the enhancement intensity factor is used to characterize the degree to which the high-frequency amplitude spectrum needs to be enhanced;

[0181] Perform element-wise multiplication calculation on the attention weight, the enhancement intensity factor, and the high-frequency amplitude spectrum to obtain an element-wise multiplication calculation result;

[0182] Sum the element-wise multiplication calculation result and the high-frequency amplitude spectrum to obtain the enhanced high-frequency amplitude spectrum.

[0183] In some alternative embodiments, the generating module 33 is configured to:

[0184] Multiply the amplitude spectrum element by element with a preset high-frequency mask to obtain the high-frequency amplitude spectrum;

[0185] Multiply the amplitude spectrum element by element with a preset low-frequency mask to obtain the low-frequency amplitude spectrum, where the sum of the high-frequency amplitude spectrum and the low-frequency amplitude spectrum is 1.

[0186] The video generation device provided by the embodiments of the present application and the video generation method provided by the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the method adopted, run or implemented by it.

[0187] The embodiments of the present application also provide an electronic device to execute the above video generation method. Please refer to Figure 4 It shows a schematic diagram of an electronic device provided by some embodiments of the present application. As Figure 4 shown, the electronic device 4 includes: a processor 400, a memory 401, a bus 402, and a communication interface 403. The processor 400, the communication interface 403, and the memory 401 are connected through the bus 402; a computer program that can run on the processor 400 is stored in the memory 401, and when the processor 400 runs the computer program, it executes the video generation method provided by any of the foregoing embodiments of the present application.

[0188] Among them, the memory 401 may include a high-speed random access memory (RAM: Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface 403 (which can be wired or wireless), a communication connection is realized between the device network element and at least one other network element, and the Internet, wide area network, local area network, metropolitan area network, etc. can be used.

[0189] The bus 402 may be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. Among them, the memory 401 is used to store a program, and after receiving an execution instruction, the processor 400 executes the program. The video generation method disclosed in any of the foregoing embodiments of the present application can be applied to the processor 400 or implemented by the processor 400.

[0190] The processor 400 may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 400 or instructions in the form of software. The above-mentioned processor 400 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 401, and the processor 400 reads the information in the memory 401 and combines its hardware to complete the steps of the above method.

[0191] The electronic device provided by the embodiments of the present application and the video generation method provided by the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the method adopted, run, or implemented by it.

[0192] The embodiments of the present application also provide a computer-readable storage medium corresponding to the video generation method provided by the foregoing embodiments. The computer-readable storage medium is an optical disc, on which a computer program (i.e., a computer program product) is stored. When the computer program is run by a processor, it will execute the video generation method provided by any of the foregoing embodiments.

[0193] It should be noted that examples of the computer-readable storage medium may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other optical and magnetic storage media, which will not be elaborated here one by one.

[0194] The computer-readable storage medium provided by the above embodiments of the present application and the video generation method provided by the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the method adopted, run, or implemented by the application program stored in it.

[0195] The embodiments of the present application also provide a computer program product. Please refer to Figure 5 . The computer program product 500 carries program code, that is, computer program 501. The instructions included in the computer program 501 can be used to execute the steps of the video generation method described in the above method embodiments. Specifically, reference can be made to the above method embodiments, which will not be elaborated here.

[0196] Among them, the above computer program product can be specifically implemented in a way of hardware, software or a combination thereof. In an alternative embodiment, the computer program product is specifically embodied as a computer storage medium. In another alternative embodiment, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.

[0197] The basic principles of the embodiments of the present application have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the embodiments of the present application are only examples and not limitations. It cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the embodiments of the present application. In addition, the above disclosed specific details are only for the purposes of illustration and easy understanding, rather than limitations. The above details do not limit the embodiments of the present application to necessarily adopt the above specific details to implement.

[0198] The block diagrams of the devices, apparatuses, equipment, and systems involved in the embodiments of the present application are only illustrative examples and do not intend to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any way. Words such as "including", "comprising", "having", etc. are open-ended words, meaning "including but not limited to", and can be used interchangeably with each other. The words "or" and "and" used herein refer to the word "and / or", and can be used interchangeably with each other unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to", and can be used interchangeably with each other.

[0199] In addition, as used herein, the "or" used in the enumeration of items starting with "at least one" indicates a separate enumeration. For example, the enumeration of "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (that is, A and B and C). In some embodiments, the term "exemplary" does not mean that the described examples are preferred or better than other examples.

[0200] It should also be noted that in the systems and methods of the embodiments of the present application, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of the embodiments of the present application.

[0201] Various changes, substitutions, and alterations to the technology described herein may be made without departing from the teachings defined by the appended claims. In some embodiments, the scope of the claims of the embodiments of the present application is not limited to the specific aspects of the processes, machines, manufactures, compositions of events, means, methods, and acts described above. Processes, machines, manufactures, compositions of events, means, methods, or acts that currently exist or will later be developed and that perform substantially the same function or achieve substantially the same result as the corresponding aspects described herein may be utilized. Accordingly, the appended claims include such processes, machines, manufactures, compositions of events, means, methods, or acts within their scope.

[0202] The foregoing description of the disclosed aspects is provided to enable any person skilled in the art to make or use the embodiments of the present application. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the embodiments of the present application. Thus, the embodiments of the present application are not intended to be limited to the aspects shown herein but are to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0203] The foregoing description has been presented for purposes of illustration and description. In some embodiments, this description is not intended to limit the embodiments of the present application to the forms disclosed herein. Although numerous example aspects and embodiments have been discussed above, those skilled in the art will recognize some of their variations, modifications, alterations, additions, and subcombinations.

Claims

1. A video generation method, characterized in that, Including: Obtaining the temporal consistency features of a video frame sequence and capturing the optical flow guidance features of the video frame sequence, where the temporal consistency features are used to characterize the information of each video frame in the video frame sequence changing over time, and the feature dimension of the temporal consistency features is higher than that of the optical flow guidance features; Fusing the temporal consistency features and the optical flow guidance features to obtain fused features; Generating a target video frame sequence with a resolution higher than that of the video frame sequence based on the fused features.

2. The method according to claim 1, wherein Obtaining the temporal consistency features of a video frame sequence includes: Encoding the video frames in the video frame sequence to obtain the spatial features of the video frames; Fusing the spatial features of the video frames within a set sliding time window in the video frame sequence to obtain inter-frame fused features; Extracting window temporal consistency features corresponding to the set sliding time window from the inter-frame fused features; Obtaining the temporal consistency features of the video frame sequence based on the window temporal consistency features.

3. The method according to claim 2, characterized in that, The set sliding time window includes at least one of the following: The first sliding time window; The second sliding time window; The third sliding time window; The inter-frame span defined by the first sliding time window is less than the inter-frame span defined by the second sliding time window, and the inter-frame span defined by the second sliding time window is less than the inter-frame span defined by the third sliding time window.

4. The method according to claim 3, characterized in that Obtaining the temporal consistency features of the video frame sequence based on the window temporal consistency features includes: When the set sliding time window includes at least two sliding time windows among the first sliding time window, the second sliding time window, and the third sliding time window, performing weighted fusion on the window temporal consistency features of the at least two sliding time windows to obtain the temporal consistency features.

5. The method according to claim 3, wherein When the set sliding time window includes the first sliding time window; Extracting window temporal consistency features corresponding to the set sliding time window from the inter-frame fused features includes: Using a model implemented based on a temporal attention mechanism to extract window temporal consistency features corresponding to the first sliding time window from the inter-frame fused features; When the set sliding time window includes the second sliding time window; Extracting window temporal consistency features corresponding to the set sliding time window from the inter-frame fused features includes: Using a temporal convolutional model to extract window temporal consistency features corresponding to the second sliding time window from the inter-frame fused features; When the set sliding time window includes the third sliding time window; Extracting window temporal consistency features corresponding to the set sliding time window from the inter-frame fused features includes: Using a model implemented based on a cross-frame self-attention mechanism to extract window temporal consistency features corresponding to the third sliding time window from the inter-frame fused features.

6. The method according to claim 1, wherein Generating a target video frame sequence with a resolution higher than that of the video frame sequence based on the fused features includes: Converting the fused features into the frequency domain to obtain the frequency domain representation of the fused features; Perform high-frequency enhancement on the frequency-domain representation to obtain an enhanced frequency-domain representation; Convert the enhanced frequency-domain representation back to the spatial domain to obtain an enhanced fusion feature; Decode the enhanced fusion feature to generate the target video frame sequence.

7. The method according to claim 6, characterized in that, Performing high-frequency enhancement on the frequency-domain representation to obtain an enhanced frequency-domain representation includes: Obtain the amplitude spectrum and phase spectrum in the frequency-domain representation; Obtain the high-frequency amplitude spectrum and low-frequency amplitude spectrum in the amplitude spectrum; Enhance the high-frequency amplitude spectrum to obtain an enhanced high-frequency amplitude spectrum; Sum the enhanced high-frequency amplitude spectrum and the low-frequency amplitude spectrum to obtain an enhanced amplitude spectrum; Generate the enhanced frequency-domain representation based on the enhanced amplitude spectrum and the phase spectrum.

8. The method according to claim 7, wherein Enhancing the high-frequency amplitude spectrum in the amplitude spectrum to obtain an enhanced high-frequency amplitude spectrum includes: Obtain the attention weight and enhancement intensity factor in the fusion feature, where the attention weight is used to indicate the local spatial region to be enhanced, and the enhancement intensity factor is used to characterize the degree to which the high-frequency amplitude spectrum needs to be enhanced; Perform element-wise multiplication calculation on the attention weight, the enhancement intensity factor, and the high-frequency amplitude spectrum to obtain an element-wise multiplication calculation result; Sum the element-wise multiplication calculation result and the high-frequency amplitude spectrum to obtain the enhanced high-frequency amplitude spectrum.

9. The method according to claim 7, characterized in that, Obtaining the high-frequency amplitude spectrum and low-frequency amplitude spectrum in the amplitude spectrum includes: Perform element-wise multiplication on the amplitude spectrum and a preset high-frequency mask to obtain the high-frequency amplitude spectrum; Perform element-wise multiplication on the amplitude spectrum and a preset low-frequency mask to obtain the low-frequency amplitude spectrum, where the sum of the high-frequency amplitude spectrum and the low-frequency amplitude spectrum is 1.

10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1-9.