Video Restoration Method and Device for Large-area Defects Based on Frequency-domain Fusion

By adopting a large-area defect video repair method based on frequency domain fusion in video repair technology, using the frequency domain fusion residual block and time Transformer module, global information modeling and time consistency optimization of video frames is solved, and the problem of insufficient large-area defect repair capability in the existing technology is achieved, and high-quality video repair effect is achieved.

CN119863405BActive Publication Date: 2025-06-24HUAQIAO UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510341442.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-24
Estimated Expiration
2045-03-21

AI Technical Summary

Technical Problem

Existing video repair technology is difficult to effectively repair large-area defective areas, resulting in the repair results that are often unreasonable and lack sufficient reference data, which affects the smoothness and nature of the video.

Method used

A large-area video repair method based on frequency domain fusion is adopted. Through the stacked frequency domain fusion residual block and time Transformer module, the video frames are modeled globally and time consistency optimized to generate visually reasonable and smooth and natural video repair effects.

Benefits of technology

It has achieved visually reasonable video repair effects in large-area defect areas, which is significantly better than the performance of other methods on objective indicators, and has obtained positive feedback from visual evaluators.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119863405B_ABST
    Figure CN119863405B_ABST
Patent Text Reader

Abstract

A method and device for video repair of large-area defects based on frequency-domain fusion, which relate to the technical field of video processing. Aiming at the problem that the current defect video repair methods are mainly limited to small-area defect scenarios, lack the ability to repair video content with large-area defects, and it is difficult to generate reasonable visual repair results, an effective solution is proposed. The method includes the following steps: First, obtain the defective video frame sequence and downsample the video frame sequence; then, use stacked frequency-domain fusion residual blocks to perform global information modeling on the downsampled defective video frames. The frequency-domain fusion residual block is composed of two adaptive frequency-domain cross-fusion modules connected in sequence; then, use stacked temporal Transformer modules to optimize the temporal consistency between multiple frames; finally, perform upsampling to reconstruct the video frames to obtain the finally repaired video. The present invention can generate a visually reasonable and content-smooth and natural video repair effect in a large-range defective area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video processing, and particularly to a method and device for video restoration of large-area defects based on frequency-domain fusion. Background Art

[0002] Video restoration technology refers to filling and complementing the defective areas in video frames through algorithms, making the restored defective areas consistent with the surrounding content and maintaining the smoothness of the video, so as to restore the integrity and naturalness of the video. This technology is widely used in fields such as post-video editing, film restoration, video transmission damage restoration, virtual reality, etc.

[0003] Currently, the research and application of video restoration technology mainly focus on deep learning-based methods. Some video restoration methods combining optical flow propagation and Transformer have achieved remarkable results in the restoration tasks of small-area defects. These methods can generate highly coordinated complemented areas with the surrounding content and maintain temporal consistency between video frames to a certain extent. However, these methods mainly focus on restoring small-area defective areas and still lack the ability to restore large-area defects. In the scenario of small-area defects, the model can make full use of the rich context information around the defective area and reasonably predict and generate the defective content through spatial and temporal correlations. But when the area of the defective area is large, the available context information is significantly reduced, and the model lacks sufficient reference data, resulting in often unreasonable restoration results. Currently, there is still a lack of effective solutions for video restoration technology for large-area defective areas. Different from general small-range defect restoration, the challenge of large-area defect restoration lies in how to effectively fill large missing areas and keep the restored content visually reasonable. Although it is not required to be as accurate as real data, the restored area needs to have a reasonable structure and texture in space and be consistent with adjacent frames in the time domain to ensure a smooth visual experience for the audience. Therefore, video restoration of large-area defects has become a difficult problem that urgently needs to be broken through in the field of video processing and requires more innovative and targeted technologies to address this issue. Summary of the Invention

[0004] The purpose of the present invention is to provide a method and device for video restoration of large-area defects based on frequency-domain fusion, which can generate a visually reasonable, content-smooth and natural video restoration effect in a large-range defective area.

[0005] The present invention adopts the following technical solutions:

[0006] On the one hand, a method for video restoration of large-area defects based on frequency-domain fusion includes:

[0007] Obtain the sequence of defective video frames, and downsample multiple consecutive video frames in the video frame sequence to obtain corresponding first features;

[0008] Input the first features into the stacked frequency-domain fusion residual blocks, and use the stacked frequency-domain fusion residual blocks to perform global information modeling on the first features; each layer of the frequency-domain fusion residual blocks consists of two frequency-domain cross-fusion modules connected in sequence; each frequency-domain cross-fusion module receives the input features, and decomposes the input features into local features and global features along the channel dimension based on the channel ratio; input the local features into the local branch, and the local branch uses a 3×3 ordinary convolution operation on the local features to extract local spatial features; input the global features into the global branch, and the global branch converts the global features from the spatial domain to the frequency domain through the adaptive frequency-domain convolution filter module, performs the enhancement and learning of the global features, fully captures the global context information, and then restores it to the spatial domain; the local branch and the global branch achieve information complementarity through the cross-fusion mechanism and are concatenated in the channel dimension to generate output features; wherein, the two frequency-domain cross-fusion modules are respectively the first frequency-domain cross-fusion module and the second frequency-domain cross-fusion module, the input features of the first frequency-domain cross-fusion module are the first features, the input features of the second frequency-domain cross-fusion module are the output features of the first frequency-domain cross-fusion module, and the output features of the second frequency-domain cross-fusion module are the second features;

[0009] Input the second features into the stacked temporal Transformer modules, and optimize the temporal consistency between frames through the stacked temporal Transformer modules to obtain third features;

[0010] Upsample the third features to reconstruct the video frames to obtain the finally repaired video.

[0011] Preferably, each frequency-domain cross-fusion module receives the input features, and decomposes the input features into local features and global features along the channel dimension based on the channel ratio, specifically including:

[0012] Each frequency-domain cross-fusion module receives the input features ; wherein, 、 and respectively represent the height, width and number of channels of the features, represents the set of real numbers;

[0013] Decompose the input features into local features and global features along the channel dimension; wherein, represents the channel ratio allocated to the global features.

[0014] Preferably, the processing process and cross-fusion process of the local branch and the global branch are as follows:

[0015] ;

[0016] ;

[0017] ;

[0018] Among them, represents the output feature of the local branch; represents the activation function; represents normalization; represents a 3×3 ordinary convolution operation; represents the output feature of the global branch; represents normalization and the activation function; represents an adaptive frequency-domain convolution filter that transforms the feature into the frequency domain to capture global context information; represents the output feature; represents concatenating and along the channel dimension.

[0019] Preferably, the processing process of the adaptive frequency-domain convolution filter is as follows:

[0020] The adaptive frequency-domain convolution filter receives the global feature , first transforms the global feature into the corresponding frequency-domain representation using the fast Fourier transform FFT, concatenates the imaginary part and the real part in the complex result to obtain the frequency-domain feature , then updates the global information of the frequency-domain feature through the adaptively learned filtering weights, expressed as:

[0021] ;

[0022] Among them, represents the frequency-domain feature after global information update; represents element-wise multiplication; represents the adaptively learned filtering weights obtained from the frequency-domain feature.

[0023] Preferably, the learning process of the adaptive filtering weights is as follows:

[0024] First, perform feature mapping using a 1×1 grouped convolution, then apply the ReLU activation function, and finally generate the filtering weights through another 1×1 grouped convolution to complete the global information update of the frequency-domain feature; the updated frequency-domain feature is split into the real part and the imaginary part along the channel dimension, reconverted into the complex representation, and finally transformed back into the spatial-domain feature through the inverse fast Fourier transform IFFT 。

[0025] Preferably, the processing process of the temporal Transformer module is as follows:

[0026] The temporal Transformer module receives the second feature processed by the frequency-domain fusion residual block, divides each feature map in the second feature into non-overlapping cubes along the height and width dimensions, and performs multi-head self-attention operations within these cube regions to obtain the third feature 。

[0027] On the other hand, a large-area defective video repair device based on frequency-domain fusion includes:

[0028] A downsampling module for obtaining a sequence of defective video frames and downsampling multiple consecutive video frames in the video frame sequence to obtain corresponding first features;

[0029] A frequency-domain fusion residual processing module for inputting the first feature into a stacked frequency-domain fusion residual block and using the stacked frequency-domain fusion residual block to perform global information modeling on the first feature; each layer of the frequency-domain fusion residual block consists of two sequentially connected frequency-domain cross-fusion modules; each frequency-domain cross-fusion module receives the input feature and decomposes the input feature into local and global features along the channel dimension based on the channel ratio; the local feature is input into the local branch, and the local branch uses a 3×3 ordinary convolution operation on the local feature to extract local spatial features; the global feature is input into the global branch, and the global branch converts the global feature from the spatial domain to the frequency domain through an adaptive frequency-domain convolution filter module, performs enhancement and learning of the global feature, fully captures the global context information, and then restores it to the spatial domain; the local branch and the global branch achieve information complementarity through a cross-fusion mechanism and are concatenated in the channel dimension to generate an output feature; wherein, the two frequency-domain cross-fusion modules are respectively the first frequency-domain cross-fusion module and the second frequency-domain cross-fusion module, the input feature of the first frequency-domain cross-fusion module is the first feature, the input feature of the second frequency-domain cross-fusion module is the output feature of the first frequency-domain cross-fusion module, and the output feature of the second frequency-domain cross-fusion module is the second feature;

[0030] A temporal consistency processing module for inputting the second feature into a stacked temporal Transformer module and optimizing the temporal consistency between frames through the stacked temporal Transformer module to obtain the third feature;

[0031] An upsampling module for upsampling the third feature to reconstruct the video frame and obtain the finally repaired video.

[0032] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0033] (1) Through the frequency-domain cross-fusion module of the stacked frequency-domain fusion residual blocks, the present invention realizes efficient global fusion modeling of spatial information.

[0034] (2) The adaptive frequency-domain convolution filter of the present invention can transform spatial features into the frequency domain for operation, with the ability to capture global information in large-area defect scenarios, while enhancing the effective information expression of features.

[0035] (3) The temporal Transformer module of the present invention calculates attention for features in the multi-frame temporal dimension, which can effectively optimize the temporal consistency of the restoration result and ensure the coherence of the video stream. Description of the Drawings

[0036] Figure 1 It is a schematic flowchart of the method for restoring large-area defective videos based on frequency-domain fusion according to an embodiment of the present invention.

[0037] Figure 2 It is a schematic structural diagram of the frequency-domain cross-fusion module according to an embodiment of the present invention.

[0038] Figure 3 It is a schematic structural diagram of the adaptive frequency-domain convolution filter according to an embodiment of the present invention.

[0039] Figure 4 It is a schematic diagram of the temporal Transformer module and cube window division according to an embodiment of the present invention.

[0040] Figure 5 It is a schematic structural diagram of the device for restoring large-area defective videos based on frequency-domain fusion according to an embodiment of the present invention.

[0041] Figure 6 It is a schematic hardware structure diagram of the electronic device provided by an embodiment of the present invention. Detailed Embodiments

[0042] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.

[0043] Such as Figure 1As shown in the figure, a method for video restoration of large-area defects based on frequency-domain fusion in this embodiment includes the following steps: obtaining a video frame with large-area defects and performing downsampling operation on it to obtain the first feature Z1; then, the frequency-domain fusion residual blocks stacked with L layers receive the feature Z1 and perform efficient global information modeling. The frequency-domain fusion residual block is sequentially connected by two frequency-domain cross-fusion modules and constructs a residual connection, and the second feature Z2 is obtained by the output of the frequency-domain fusion residual block. Then, the time Transformer module stacked with N layers optimizes the temporal consistency of the feature Z2 to obtain the third feature Z3; finally, the feature Z3 is upsampled to reconstruct the video frame, and the finally restored video is obtained.

[0044] The stacked frequency-domain fusion residual blocks with L layers are specifically L layers of sequentially connected frequency-domain fusion residual blocks.

[0045] The following elaborates in detail on the process steps of the method for video restoration of large-area defects based on frequency-domain fusion.

[0046] First, for the defective video sequence , it is downsampled by using a convolutional neural network to obtain the first feature with the number of channels c . Among them, and respectively represent the spatial resolution and the number of channels, and represents the th frame in the video sequence.

[0047] Next, the first feature enters the stacked frequency-domain fusion residual blocks.

[0048] The frequency-domain fusion residual block is sequentially connected by two frequency-domain cross-fusion modules. The output after being processed by these two frequency-domain cross-fusion modules is added to the input of the residual block through a residual connection to complete the fusion and transmission of features.

[0049] As Figure 2 shown, it is the structural schematic diagram of the frequency-domain cross-fusion module. The frequency-domain cross-fusion module adopts the construction idea of local-global two-way branch cross-fusion, combines the local convolution operation in the spatial domain with the global modeling ability in the frequency domain, so as to effectively capture and learn the residual feature information under the large-area invalid features. The frequency-domain cross-fusion module receives the input feature , where and respectively represent the spatial resolution and the number of channels. First, the input feature is decomposed along the channel dimension into the local feature and the global feature , and represents the channel ratio allocated to the global feature. The local feature enters the local branch to capture local spatial features, and the global feature Enter the global branch to capture long - distance global context information. The two branches achieve information complementarity through a cross - fusion mechanism, enhance the expression of effective features, and are concatenated in the channel dimension to finally generate output features . Specifically, the local branch uses a 3×3 ordinary convolution operation to extract local spatial features. The global branch converts from the spatial domain to the frequency domain through an adaptive frequency - domain convolution filter module to strengthen and learn global features, fully capture global context information, and then convert it back to the spatial domain. The calculation processes and cross - fusion methods of the two branches are described as follows:

[0050] ;

[0051] ;

[0052] ;

[0053] Among them, represents the output feature of the local branch; represents the activation function; represents normalization; represents a 3×3 ordinary convolution operation; represents the output feature of the global branch; represents normalization and the activation function; represents the adaptive frequency - domain convolution filter, which converts features to the frequency domain to capture global context information; represents the output feature; represents the operation of and performing a concatenation operation in the channel dimension.

[0054] It should be noted that the two frequency - domain cross - fusion modules can be respectively the first frequency - domain cross - fusion module and the second frequency - domain cross - fusion module. The input feature of the first frequency - domain cross - fusion module is the first feature Z1, the input feature of the second frequency - domain cross - fusion module is the output feature of the first frequency - domain cross - fusion module, and the output feature of the second frequency - domain cross - fusion module is the second feature Z2.

[0055] In this embodiment, the structure of the adaptive frequency - domain convolution filter is as Figure 3 shown. The adaptive frequency - domain convolution filter converts features from the spatial domain to the frequency domain for efficient global information modeling, and then converts the updated frequency - domain features back to spatial - domain features. Specifically, the second feature receives the global feature , first, it is converted into the corresponding frequency-domain representation by using the Fast Fourier Transform (FFT). Since the calculation result of the Fast Fourier Transform contains complex numbers, to ensure that both the input and output of this module are real numbers, the imaginary part and the real part in the complex number result are concatenated to obtain the frequency-domain features. , then the frequency-domain features are updated with global information through the filtering weights learned adaptively. This process is expressed as:

[0056] ;

[0057] where, is the frequency-domain feature after global information update, represents element-wise multiplication. is the adaptive filtering weight learned from the frequency-domain features. The learning of the filtering weight is achieved through the following steps: First, a 1×1 grouped convolution is used for feature mapping, then the ReLU activation function is applied, and finally, another 1×1 grouped convolution is used to generate the filtering weight. Through the above process, the global information of the frequency-domain features is updated. The updated frequency-domain features are split into the real part and the imaginary part along the channel dimension, reconverted into the complex representation, and finally converted back to the spatial-domain features through the inverse Fast Fourier Transform. .

[0058] After passing through all the frequency-domain fusion residual blocks, the second feature is obtained. Next, it enters the stacked N-layer temporal Transformer module.

[0059] The stacked N-layer temporal Transformer module is specifically N sequentially connected temporal Transformer modules.

[0060] As Figure 4 shown, it is a schematic diagram of the temporal Transformer module and the cube window division. Its self-attention operation calculates the features between different frames. Specifically, each feature map in the feature is divided into non-overlapping cubes along the height and width dimensions, and the multi-head self-attention operation is performed within these cube regions. By stacking multiple temporal Transformer modules, the feature expression ability in the temporal dimension can be gradually enhanced while maintaining the integrity of spatial information.

[0061] After passing through all the temporal Transformer modules, the third feature is obtained. Finally, a transposed convolution is constructed to perform upsampling to reconstruct the video frame, and the final restoration result is obtained.

[0062] The test platform, test data, test process, and test results adopted in this embodiment are as follows.

[0063] Test platform: The experimental tests were conducted on a computer equipped with an Intel Core i9-13900K and an NVIDIA GeForce RTX 4090 GPU, with the operating system being Ubuntu 20.04, and implemented using Python 3.10 and PyTorch 2.10.

[0064] Test data: This method was tested on the YouTube-VOS and DAVIS datasets, and an additional custom large-area defect dataset was constructed, with the mask ratio controlled between 40% and 50%. The test video resolution was 480p, and the frame rate was 30fps, mainly including natural scenes, urban street scenes, and human activities. The defect areas were constructed by random occlusion, manual annotation, etc. to simulate large-area defects in real application scenarios.

[0065] Test process: The defective videos were input into this method for restoration, and compared with current mainstream restoration methods. LPIPS and VFID were used as evaluation metrics to quantify the restoration effect. In addition, a preset number of visual evaluation personnel were invited to subjectively score the rationality, structural integrity, and fluency of the restored content.

[0066] Test results: As shown in Table 1, it is a comparison of the objective metrics for the restoration of videos with large-area defects (mask ratio of 40% - 50%) by various restoration methods. From the experimental results in Table 1, it can be seen that in terms of objective metrics, the VFID (Video Fréchet Inception Distance, which can effectively evaluate the temporal consistency and dynamic quality of videos, and the lower the VFID value, the higher the quality of the generated video) and LPIPS (Learned Perceptual Image Patch Similarity, LPIPS focuses on the high-level features of images rather than pixel-level differences, and the smaller the value, the more similar the images) metrics of this method are significantly better than other methods.

[0067] Table 1 Comparison of objective metrics for the restoration of videos with large-area defects by various restoration methods;

[0068]

[0069] As Figure 5 shown, the present invention also discloses a large-area defect video restoration device based on frequency domain fusion, including:

[0070] A downsampling module 501, which is used to obtain a sequence of defective video frames and perform downsampling on multiple consecutive video frames in the video frame sequence to obtain corresponding first features;

[0071] The frequency-domain fusion residual processing module 502 is configured to input the first feature into a stacked frequency-domain fusion residual block, and use the stacked frequency-domain fusion residual block to perform global information modeling on the first feature; each layer of the frequency-domain fusion residual block consists of two sequentially connected frequency-domain cross-fusion modules; each frequency-domain cross-fusion module receives the input feature, decomposes the input feature into local features and global features along the channel dimension based on the channel ratio; inputs the local features into the local branch, and the local branch uses a 3×3 ordinary convolution operation on the local features to extract local spatial features; inputs the global features into the global branch, and the global branch converts the global features from the spatial domain to the frequency domain through an adaptive frequency-domain convolution filter module to perform enhancement and learning of the global features, fully capture the global context information, and then restore it to the spatial domain; the local branch and the global branch achieve information complementarity through a cross-fusion mechanism and are concatenated in the channel dimension to generate an output feature; wherein, the two frequency-domain cross-fusion modules are respectively the first frequency-domain cross-fusion module and the second frequency-domain cross-fusion module, the input feature of the first frequency-domain cross-fusion module is the first feature, the input feature of the second frequency-domain cross-fusion module is the output feature of the first frequency-domain cross-fusion module, and the output feature of the second frequency-domain cross-fusion module is the second feature;

[0072] The temporal consistency processing module 503 is configured to input the second feature into a stacked temporal Transformer module, and optimize the temporal consistency between frames through the stacked temporal Transformer module to obtain a third feature;

[0073] The upsampling module 504 is configured to upsample the third feature to reconstruct the video frame, and obtain the finally repaired video.

[0074] The specific implementation of each module of a video repair device for large-area defects based on frequency-domain fusion is the same as that of a video repair method for large-area defects based on frequency-domain fusion, and this embodiment will not be repeated here.

[0075] Figure 6 The following shows the hardware structure schematic diagram of the electronic device provided by the embodiment of the present invention. As Figure 6 shown, the electronic device of this embodiment includes: a processor 601 and a memory 602; wherein the memory 602 is used to store computer execution instructions; the processor 601 is used to execute the computer execution instructions stored in the memory to implement each step executed by the electronic device in the above embodiment. Specifically, reference can be made to the relevant descriptions in the foregoing method embodiments.

[0076] Optionally, the memory 602 can be either independent or integrated with the processor 601.

[0077] When the memory 602 is independently provided, the electronic device further includes a bus 603 for connecting the memory 602 and the processor 601.

[0078] An embodiment of the present invention further provides a computer storage medium storing computer-executable instructions, which, when executed by the processor 601, implement the above method.

[0079] An embodiment of the present invention further provides a computer program product including a computer program, which, when executed by the processor 601, implements the above method.

[0080] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be indirect couplings or communication connections through some interfaces, devices or modules, and can be in electrical, mechanical or other forms.

[0081] The modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to implement the solution of this embodiment.

[0082] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing unit, or each module can exist physically alone, or two or more modules can be integrated in a unit. The units formed by the above modules can be implemented in the form of hardware or in the form of a hardware plus software functional unit.

[0083] The above integrated modules implemented in the form of software functional modules can be stored in a computer-readable storage medium. The above software functional modules are stored in a storage medium and include several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or the processor 601 to execute some steps of the methods in various embodiments of the present application.

[0084] It should be understood that the above-mentioned processor 601 can be a Central Processing Unit (CPU), or other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), etc. The general-purpose processor can be a microprocessor or the processor 601 can also be any conventional processor 601, etc. The steps of the method disclosed in combination with the invention can be directly embodied as being executed and completed by the hardware processor 601, or by a combination of the hardware and software modules in the processor 601.

[0085] The memory 602 may include high-speed RAM memory and may also include non-volatile storage NVM, such as at least one disk memory, and can also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk, or an optical disc, etc.

[0086] The bus 603 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus 603 can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, the bus 603 in the drawings of this application is not limited to only one bus 603 or one type of bus 603.

[0087] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disc. The storage medium can be any available medium that can be accessed by a general or special-purpose computer.

[0088] An exemplary storage medium is coupled to the processor 601, enabling the processor 601 to read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor 601. The processor 601 and the storage medium can be located in an Application Specific Integrated Circuits (ASIC). Of course, the processor 601 and the storage medium can also exist as discrete components in an electronic device or a main control device.

[0089] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disks, or optical discs that can store program codes.

[0090] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for restoring large-area defective video based on frequency domain fusion, characterized in that: include: Acquire a defective video frame sequence, and downsample a plurality of continuous video frames in the video frame sequence to obtain a corresponding first feature; The first feature is input into the stacked frequency domain fusion residual block, and the stacked frequency domain fusion residual block is used to perform global information modeling on the first feature; each layer of the frequency domain fusion residual block is composed of two frequency domain cross fusion modules connected in sequence; each frequency domain cross fusion module receives the input feature, and decomposes the input feature into local features and global features along the channel dimension based on the channel ratio; the local feature is input into the local branch, and the local branch uses a 3×3 ordinary convolution operation on the local feature to extract the local spatial feature; the global feature is input into the global branch, and the global branch converts the global feature from the spatial domain to the frequency domain through an adaptive frequency domain convolution filter module, strengthens and learns the global feature, fully captures the global context information, and then restores it to the spatial domain; the local branch and the global branch complement each other through the cross fusion mechanism, and splice in the channel dimension to generate the output feature; wherein the two frequency domain cross fusion modules are respectively the first frequency domain cross fusion module and the second frequency domain cross fusion module, the input feature of the first frequency domain cross fusion module is the first feature, the input feature of the second frequency domain cross fusion module is the output feature of the first frequency domain cross fusion module, and the output feature of the second frequency domain cross fusion module is the second feature; The second feature is input into the stacked temporal Transformer module, and the temporal consistency between frames is optimized through the stacked temporal Transformer module to obtain the third feature; Upsampling the third feature to reconstruct the video frame to obtain the final restored video; Each frequency domain cross-fusion module receives input features and decomposes the input features into local features and global features along the channel dimension based on the channel ratio, including: Each frequency domain cross fusion module receives input features Among them, h, w and c represent the height, width and number of channels of the feature respectively. represents the set of real numbers; The input features Decompose into local features along the channel dimension and global features Among them, α in ∈[0,1] represents the channel ratio assigned to the global feature; The processing and cross-fusion process of local branches and global branches are as follows: Y l =ReLU(BN(f Conv (X l )+f Conv (X g )))? Y g =f BN-ReLU [f Conv (X l )+f AFFC (X g )]; And=Concat(And l ,AND g ); in, represents the output feature of the local branch; ReLU represents the activation function; BN represents normalization; f Conv Represents a 3×3 normal convolution operation; represents the output feature of the global branch; f BN-ReLU represents normalization and activation function; f AFFC represents an adaptive frequency domain convolution filter that converts features into the frequency domain to capture global context information; Represents the output feature; Concat represents Y l With Y g Perform concatenation on the channel dimension.

2. The method for repairing large-area defective video based on frequency domain fusion according to claim 1, characterized in that: The processing process of the adaptive frequency domain convolution filter is as follows: The adaptive frequency domain convolution filter receives the global feature X g First, the fast Fourier transform (FFT) is used to convert the global features into the corresponding frequency domain representation, and the imaginary part and the real part of the complex result are concatenated to obtain the frequency domain features. Then, the frequency domain feature X is adjusted by adaptively learning the filter weights. gF Update the global information, expressed as: in, represents the frequency domain features after global information update; ⊙ represents the element-by-element product; W(X gF ) represents the adaptive filter weights learned from frequency domain features.

3. The method for repairing large-area defective video based on frequency domain fusion according to claim 2, characterized in that: The learning process of adaptive filter weights is as follows: First, a 1×1 group convolution is used for feature mapping, then the ReLU activation function is applied, and finally another 1×1 group convolution is used to generate the filter weights so that the frequency domain features can complete the global information update; the updated frequency domain features Split into real and imaginary parts along the channel dimension, reconvert back to integer representation, and finally convert back to spatial domain features through inverse fast Fourier transform IFFT 4. The method for repairing large-area defective video based on frequency domain fusion according to claim 1, characterized in that: The processing of the temporal Transformer module is as follows: The temporal Transformer module receives the second feature processed by the frequency domain fusion residual block, divides each feature map in the second feature into non-overlapping cubes along the height and width dimensions, performs multi-head self-attention operations in these cube areas, and obtains the third feature 5. A large area defective video restoration device based on frequency domain fusion using the large area defective video restoration method based on frequency domain fusion according to any one of claims 1 to 4, characterized in that: include: A downsampling module, used to obtain a defective video frame sequence, and downsample a plurality of continuous video frames in the video frame sequence to obtain a corresponding first feature; The frequency domain fusion residual processing module is used to input the first feature into the stacked frequency domain fusion residual blocks, and use the stacked frequency domain fusion residual blocks to perform global information modeling on the first feature; each layer of the frequency domain fusion residual block is composed of two frequency domain cross fusion modules connected in sequence; each frequency domain cross fusion module receives the input feature, and decomposes the input feature into local features and global features along the channel dimension based on the channel ratio; the local feature is input into the local branch, and the local branch uses a 3×3 ordinary convolution operation to extract local spatial features from the local feature; the global feature is input into the global branch, and the global branch uses an adaptive frequency domain convolution filter model The block converts the global features from the spatial domain to the frequency domain, strengthens and learns the global features, fully captures the global context information, and then restores it to the spatial domain; the local branch and the global branch achieve information complementarity through the cross-fusion mechanism, and splice in the channel dimension to generate output features; wherein the two frequency domain cross-fusion modules are respectively the first frequency domain cross-fusion module and the second frequency domain cross-fusion module, the input feature of the first frequency domain cross-fusion module is the first feature, the input feature of the second frequency domain cross-fusion module is the output feature of the first frequency domain cross-fusion module, and the output feature of the second frequency domain cross-fusion module is the second feature; A time consistency processing module, used for inputting the second feature into a stacked time Transformer module, optimizing the time consistency between frames through the stacked time Transformer module, and obtaining a third feature; The upsampling module is used to upsample the third feature to reconstruct the video frame to obtain the final restored video.

Citation Information

Patent Citations

  • Seismic data reconstruction method based on spatial domain and frequency domain fusion architecture

    CN118033732A

  • Defective video restoration method and system based on strong perception Transform architecture

    CN118469876A