Video restoration detection method and system based on multi-domain feature fusion
Through the video repair detection method of multi-domain feature fusion, combined with the characteristics of the airspace, time domain and frequency domain, an accurate repair area mask map is generated, which solves the problem of poor detection effects of traditional methods and achieves more efficient and accurate video repair detection.
Patent Information
- Application Number
- CN202510025129.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-08
AI Technical Summary
Traditional video repair detection methods only analyze from the perspective of the airspace within the frame or the time domain between frames, and fail to comprehensively consider multiple features and frequency domain information, resulting in unsatisfactory detection results.
A video repair detection method based on multi-domain feature fusion is adopted to generate a mask map with video repair area detection results through the extraction and fusion of airspace, time domain and frequency domain features. Specific steps include: the airspace characteristics are obtained through high-pass filtering core convolution, the time domain characteristics are extracted by optical flow estimation method, and the frequency domain characteristics are obtained through time-frequency domain conversion and adaptive frequency band separation.
By comprehensively utilizing the time, airspace and frequency domain information of video, the multimodal characteristics of repair traces are captured comprehensively, which significantly improves the accuracy and efficiency of video repair detection, and is suitable for diverse and complex video repair scenarios.
Smart Images

Figure CN119444621B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video detection technology, and in particular to a video restoration detection method and system based on multi-domain feature fusion. Background Art
[0002] Video repair is one of the video editing tasks, which is used to remove the content of a specified area in a video and fill it accordingly or repair and restore the damaged content. Video repair technology is often used for normal editing of video content, but it is also used to maliciously tamper with the content. Malicious tampering with videos will not only cause credibility issues in video use, but may also cause problems that endanger public safety. How to effectively determine whether video content has been repaired or tampered with has become a hot research topic in the field of image and video processing.
[0003] Existing deep video restoration detection methods often only analyze restoration traces from a unilateral perspective such as the spatial domain within the video frame or the temporal domain between frames, without comprehensively considering multiple features and failing to make full use of frequency domain information, resulting in unsatisfactory detection results. Summary of the invention
[0004] The purpose of the present invention is to provide a video restoration detection method and system based on multi-domain feature fusion, aiming to solve the problem that traditional technologies only perform restoration trace analysis from the spatial domain perspective within the video frame or the temporal domain perspective between frames, resulting in unsatisfactory detection results.
[0005] In a first aspect, the present invention provides a video restoration detection method based on multi-domain feature fusion, the method comprising:
[0006] Acquire a video sequence to be detected, and segment the video sequence according to a preset segmentation rule to obtain multiple video clips;
[0007] Convolving each single frame image in each of the video clips with a high-pass filter kernel to obtain spatial domain features corresponding to the single frame image;
[0008] Obtaining a forward optical flow from the t-th frame image to the t+1-th frame image and a backward optical flow from the t+1-th frame image to the t-th frame image based on an optical flow estimation method, and extracting temporal features between adjacent frame images according to the forward optical flow and the backward optical flow;
[0009] Performing a time-frequency domain conversion on the video clip to obtain a frequency component corresponding to the video clip, obtaining a first frequency band, a second frequency band, and a third frequency band according to the frequency component and a preset analysis matrix, and performing a connection operation on the first frequency band, the second frequency band, and the third frequency band to obtain a frequency domain feature corresponding to the video clip;
[0010] The fusion features of each video clip are obtained according to the frequency domain features, all the spatial domain features and all the time domain features of the same video clip, and a mask image with a video restoration area detection result corresponding to the video clip is generated according to each of the fusion features.
[0011] Furthermore, the step of obtaining a video sequence to be detected and segmenting the video sequence according to a preset segmentation rule to obtain a plurality of video segments includes:
[0012] Divide the video sequence to be detected into multiple video segments with T frames per segment. ;
[0013] For the last video clip If the video clip If the number of frames is less than T, the video clip will be repeated The last frame of the video clip is filled in The number of frames reaches T frames.
[0014] Furthermore, the step of convolving each single frame image in each of the video clips with a high-pass filter kernel to obtain a spatial domain feature corresponding to the single frame image includes:
[0015] The convolution operation is performed according to the following formula:
[0016] ;
[0017] in, Represents the spatial characteristics, Represents a single frame image, represents the high-pass filter kernel, , Represents the convolution operation;
[0018] ;
[0019] The initial values of parameters a, b, and c are all set to 0, and they are adaptively updated during the model training process, and their value range is within [-1, 1).
[0020] Furthermore, the steps of obtaining the forward optical flow from the t-th frame image to the t+1-th frame image and the backward optical flow from the t+1-th frame image to the t-th frame image based on the optical flow estimation method, and extracting the time domain features between adjacent frame images according to the forward optical flow and the backward optical flow include:
[0021] The forward optical flow from the t-th frame image to the t+1-th frame image is obtained according to the following formula:
[0022] ;
[0023] The backward optical flow from the t+1th frame image to the tth frame image is obtained according to the following formula:
[0024] ;
[0025] in, represents the forward optical flow from the t-th frame image to the t+1-th frame image, represents the t-th frame image, represents the t+1th frame image, Represents the backward optical flow from the t+1th frame image to the tth frame image;
[0026] The time domain characteristics are calculated according to the following formula:
[0027] ;
[0028] in, Represents the time domain features between the t-th frame image and the t+1-th frame image.
[0029] Furthermore, the step of performing time-frequency domain conversion on the video clip to obtain a frequency component corresponding to the video clip includes:
[0030] The time-frequency domain conversion is performed according to the following formula:
[0031] ;
[0032] ;
[0033] in, represents the corresponding frequency component in the frequency domain, , , Both represent normalization coefficients, H represents the height of any image in the video clip, and W represents the width of any image in the video clip. Represents the pixel value of the video segment at time step i, height x, width y.
[0034] Furthermore, the step of acquiring a first frequency band, a second frequency band, and a third frequency band according to the frequency components and a preset analysis matrix, and connecting the first frequency band, the second frequency band, and the third frequency band to obtain a frequency domain feature corresponding to the video clip includes:
[0035] The first frequency band, the second frequency band, and the third frequency band are obtained according to the following formula:
[0036] ;
[0037] in, Indicates the first frequency band, Indicates the second frequency band, Indicates the third frequency band, represents the first frequency component, represents the second frequency component, represents the third frequency component, represents the identity matrix, , , All of them represent analysis matrices. When initialized, all elements of the analysis matrix are set to 0. During the training process, all elements of the analysis matrix will be adaptively updated according to the frequency response of the repair traces. The value range is Inside;
[0038] The frequency domain features are obtained according to the following formula:
[0039] ;
[0040] in, Indicates a connection operation. Represents frequency domain features.
[0041] Furthermore, the step of obtaining fusion features under each of the video clips according to the frequency domain features, all the spatial domain features and all the time domain features under the same video clip, and generating a mask map with a video restoration area detection result corresponding to the video clip according to each of the fusion features includes:
[0042] The frequency domain features, all spatial domain features, and all time domain features of the same video clip are feature mapped through a ResNet50 network with unshared parameters, so that they are projected into a feature space of uniform dimension:
[0043] ;
[0044] Take advantage of additional The network integrates three types of features: time, space, and frequency domains:
[0045] ;
[0046] The fused features are converted into query vector Q, key vector K and value vector V, and the features at each position are weighted summed through the multi-head self-attention mechanism:
[0047] ;
[0048] ;
[0049] The mask map is obtained according to the following formula:
[0050] ;
[0051] in, represents the spatial features after mapping, represents the time domain features after mapping, represents the frequency domain features after mapping, represents the attention head associated with the query vector, represents the attention head associated with the key vector, represents the attention head associated with the value vector, represents the multi-head self-attention mechanism, represents the sth attention head associated with the query vector, represents the sth attention head associated with the key vector, represents the sth attention head associated with the value vector, represents the mask image, represents a multi-layer perception operation, Represents the linear transformation matrix used for the final output mapping, represents the dimension of the key vector, Represents the characteristics of deep modeling, represents the transpose of a matrix, Represent the weight matrices used to generate query vectors, key vectors, and value vectors, respectively. Both represent bias vectors.
[0052] In a second aspect, the present invention provides a video restoration detection system based on multi-domain feature fusion, the system comprising:
[0053] The video sequence acquisition module is used to acquire the video sequence to be detected and segment the video sequence according to a preset segmentation rule to obtain multiple video clips;
[0054] A spatial feature extraction module, used for convolving each single frame image in each of the video clips with a high-pass filter kernel to obtain a spatial feature corresponding to the single frame image;
[0055] A time domain feature extraction module, used to obtain the forward optical flow from the t-th frame image to the t+1-th frame image and the backward optical flow from the t+1-th frame image to the t-th frame image based on the optical flow estimation method, and extract the time domain features between adjacent frame images according to the forward optical flow and the backward optical flow;
[0056] A frequency domain feature extraction module, used to perform time-frequency domain conversion on the video clip to obtain frequency components corresponding to the video clip, obtain a first frequency band, a second frequency band and a third frequency band according to the frequency components and a preset analysis matrix, and connect the first frequency band, the second frequency band and the third frequency band to obtain frequency domain features corresponding to the video clip;
[0057] The mask image generation module is used to obtain the fusion features of each video clip based on the frequency domain features, all spatial domain features and all time domain features of the same video clip, and generate a mask image with video repair area detection results corresponding to the video clip based on each fusion feature.
[0058] In a third aspect, the present invention provides a storage medium storing one or more programs, which, when executed by a processor, implement the above-mentioned video restoration detection method based on multi-domain feature fusion.
[0059] In a fourth aspect, the present invention provides an electronic device, the electronic device comprising a memory and a processor, wherein:
[0060] The memory is used to store computer programs;
[0061] When the processor is used to execute the computer program stored in the memory, the above-mentioned video restoration detection method based on multi-domain feature fusion is implemented.
[0062] Compared with the prior art, the present invention has the following advantages:
[0063] The present invention firstly segments the video sequence to improve the video processing efficiency, and then for the single-frame image in the segmented video clip, by analyzing the local continuity and consistency between adjacent pixels, captures the abnormal texture and structural changes in the repaired image, and obtains the fine spatial domain features reflecting the repair traces; then, combined with the information of cross-frame images, by capturing the dynamic change law of pixel values corresponding to the same spatial position at different time points, the time domain features reflecting the correlation between the repaired video frames are extracted to further enhance the time dimension discrimination ability of the repaired area; then, based on the adaptive frequency band separation and frequency domain filtering technology, the frequency component analysis of the repaired video is performed, the specific frequency mode and abnormal signal introduced by the repair operation are extracted, and the frequency domain features of the repair traces that can be used are generated; on the above basis, the deep correlation modeling and fusion strategy of multi-domain features are finally used to perform feature fusion and refinement analysis on the extracted time domain, spatial domain and frequency domain features, and finally a mask map for accurately locating the repair area is generated, so as to realize the accurate detection and labeling of the video repair traces. By comprehensively utilizing the time domain, spatial domain and frequency domain information of the repaired video, the multimodal characteristics of the repair traces are fully captured, which improves the accuracy and efficiency of video repair detection tasks and is suitable for diverse and complex video repair scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 A flowchart of a video restoration detection method based on multi-domain feature fusion proposed in one embodiment of the present invention;
[0065] Figure 2A schematic diagram of the structure of a video restoration detection system based on multi-domain feature fusion proposed in one embodiment of the present invention.
[0066] The following specific implementation manner will further illustrate the present invention in conjunction with the above-mentioned drawings. DETAILED DESCRIPTION
[0067] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention. Unless otherwise defined, the technical terms or scientific terms used herein should be understood by people with general skills in the field to which the present invention belongs. "Including" and similar words used in this article mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects.
[0068] like Figure 1 As shown, an embodiment of the present invention provides a video restoration detection method based on multi-domain feature fusion, the method comprising steps S101 to S105, wherein:
[0069] Step S101: obtaining a video sequence to be detected, and segmenting the video sequence according to a preset segmentation rule to obtain a plurality of video clips;
[0070] It should be noted that, since the video sequence contains images with a huge amount of data, in order to improve the efficiency of video detection, the video sequence needs to be segmented so that the obtained video segments each contain a small number of images. Specifically, in this step, the video sequence to be detected is first segmented into multiple video segments with a frame number of T per segment. ; Then for the last video clip If the video clip If the number of frames is less than T, the video clip will be repeated. The last frame of the video clip is filled in The number of frames reaches T frames.
[0071] Step S102: convolving each single frame image in each of the video clips with a high-pass filter kernel to obtain a spatial domain feature corresponding to the single frame image;
[0072] It should be noted that the convolution operation is performed according to the following formula:
[0073] ;
[0074] in, Represents the spatial characteristics, Represents a single frame image, represents the high-pass filter kernel, , Represents the convolution operation;
[0075] ;
[0076] Among them, the initial values of parameters a, b, and c are all set to 0, and they are adaptively updated during the model training process, and the value range is within [-1,1). By optimizing these three parameters, the model can better capture the subtle differences between adjacent pixels in the spatial domain in the horizontal, vertical, and diagonal directions, thereby more accurately extracting the edge features of the repaired area in complex scenes and further improving the detection accuracy of the repaired area.
[0077] Step S103: obtaining a forward optical flow from the t-th frame image to the t+1-th frame image and a backward optical flow from the t+1-th frame image to the t-th frame image based on an optical flow estimation method, and extracting temporal features between adjacent frame images according to the forward optical flow and the backward optical flow;
[0078] It should be noted that within a video clip, the forward optical flow between each frame can be calculated using the RAFT method:
[0079] ;
[0080] The backward optical flow from the t+1th frame image to the tth frame image is obtained according to the following formula:
[0081] ;
[0082] in, represents the forward optical flow from the t-th frame image to the t+1-th frame image, represents the t-th frame image, represents the t+1th frame image, Represents the backward optical flow from the t+1th frame image to the tth frame image, where x and y are spatial indexes;
[0083] It should also be noted that the calculation of optical flow is essentially a measure of the similarity of image features. Under the premise that the optical flow estimation is ideal and accurate, the forward and backward optical flows should meet the following conditions:
[0084] ;
[0085] In practical applications, due to possible occlusion, perspective change, motion blur and other factors, the forward and backward optical flows are often not strictly opposite. In order to reduce the deviation caused by asymmetry, the time domain features are calculated according to the following formula:
[0086] ;
[0087] in, represents the time domain features between the t-th frame image and the t+1-th frame image. For example, The weighted scheme establishes a more accurate and stable time association for pixels at the same position between any two frames, thereby providing more reliable timing information for subsequent processing.
[0088] Step S104: performing time-frequency domain conversion on the video clip to obtain frequency components corresponding to the video clip, and obtaining a first frequency band, a second frequency band, and a third frequency band according to the frequency components and a preset analysis matrix, and performing a connection operation on the first frequency band, the second frequency band, and the third frequency band to obtain frequency domain features corresponding to the video clip;
[0089] Considering the spatiotemporal frequency characteristics of the video clips, the video clips are converted into time-frequency domain according to the following formula:
[0090] ;
[0091] ;
[0092] in, represents the corresponding frequency component in the frequency domain, , , Both represent normalization coefficients, which ensure energy preservation and orthogonality in the temporal and spatial dimensions. H represents the height of any image in the video clip, and W represents the width of any image in the video clip. Represents the pixel value of the video segment at time step i, height x, width y.
[0093] In addition, according to the principle of equal energy, the obtained frequency components are divided into the first frequency component, the second frequency component and the third frequency component. For example, the third frequency component occupies the first 1 / 16 of the entire spectrum, the second frequency component is generally located between 1 / 16 and 1 / 8 of the spectrum, and the first frequency component mainly covers the last 7 / 8. This division reflects the distribution characteristics of frequency components in the image, in which the low-frequency (third frequency component) component captures the macroscopic structure of the image, the medium-frequency (second frequency component) component contains more detailed information, and the high-frequency (first frequency component) component is mainly used to describe the subtle texture and noise in the image. However, this frequency band division is difficult to achieve the effect of accurately separating the frequencies corresponding to the repair traces. To this end, in some embodiments, an analysis matrix is set, and then the first frequency band, the second frequency band and the third frequency band are obtained according to the following formula:
[0094] ;
[0095] in, Indicates the first frequency band, Indicates the second frequency band, Indicates the third frequency band, represents the first frequency component, represents the second frequency component, represents the third frequency component, represents the identity matrix, , , All of them represent analysis matrices. When initialized, all elements of the analysis matrix are set to 0. During the training process, all elements of the analysis matrix will be adaptively updated according to the frequency response of the repair traces. The value range is For the frequency components related to the repair traces, the corresponding analysis matrix element values will be increased; for the frequency components unrelated to the repair traces, the element values will be updated to be smaller. This adaptive filtering strategy weights the different frequency components in the spectrum to achieve accurate positioning of the repair traces.
[0096] The frequency domain features are obtained according to the following formula:
[0097]
[0098] in, Indicates a connection operation. Represents frequency domain features.
[0099] The above method gets rid of the reliance on a priori frequency band separation by adaptively adjusting the frequency domain features, and more effectively extracts the relevant frequency information of the repair traces.
[0100] Step S105: obtaining fusion features for each video clip according to the frequency domain features, all spatial domain features and all temporal domain features of the same video clip, and generating a mask image with video restoration area detection results corresponding to the video clip according to each fusion feature.
[0101] It should be noted that the frequency domain features, all spatial domain features and all time domain features of the same video clip are feature mapped through a ResNet50 network with unshared parameters, so that they are projected into a feature space of uniform dimension:
[0102] ;
[0103] An additional ResNet50 network is used to fuse the three types of features in the time, space and frequency domains:
[0104] ;
[0105] The fused feature P is input into the Transformer structure for deep modeling. In this process, it is first converted into a query vector Q (Query), a key vector K (Key), and a value vector V (Value), and a multi-head self-attention mechanism is used to perform a weighted summation of the features at each position to process the interdependencies of multi-domain information in the repaired video:
[0106] ;
[0107] ;
[0108] Finally, the fused features are further processed using a multi-layer perceptron (MLP). Under the optimization goal of minimizing the loss function, the model parameters are continuously adjusted to finally generate a mask map of the high-precision video repair area detection result:
[0109] ;
[0110] in, represents the spatial features after mapping, represents the time domain features after mapping, represents the frequency domain features after mapping, represents the attention head associated with the query vector, represents the attention head associated with the key vector, represents the attention head associated with the value vector, represents the multi-head self-attention mechanism, represents the sth attention head associated with the query vector, represents the sth attention head associated with the key vector, represents the sth attention head associated with the value vector, represents the mask image, represents a multi-layer perception operation, Represents the linear transformation matrix used for the final output mapping, represents the dimension of the key vector, Represents the characteristics of deep modeling, represents the transpose of a matrix, Represent the weight matrices used to generate query vectors, key vectors, and value vectors, respectively. Both represent bias vectors.
[0111] It should also be noted that the mask image contains two parts, black and white. The black area is the area that has not been repaired, and the white area is the area where there are traces of repair, corresponding to the video repair area detection result.
[0112] In summary, according to the above-mentioned video restoration detection method based on multi-domain feature fusion, the video sequence is first segmented to improve the video processing efficiency, and then for the single-frame image in the segmented video clip, the local continuity and consistency between adjacent pixels are analyzed to capture the abnormal texture and structural changes in the restored image, and obtain the fine spatial domain features reflecting the restoration traces; then, combined with the information of cross-frame images, by capturing the dynamic change law of pixel values corresponding to the same spatial position at different time points, the time domain features reflecting the correlation between the restored video frames are extracted to further enhance the time dimension discrimination ability of the restoration area; then, based on the adaptive frequency band separation and frequency domain filtering technology, the frequency component analysis of the restored video is performed to extract the specific frequency patterns and abnormal signals introduced by the restoration operation, and generate the frequency domain features of the restoration traces that can be used; on the above basis, the deep correlation modeling and fusion strategy of multi-domain features is finally used to perform feature fusion and refinement analysis on the extracted time domain, spatial domain and frequency domain features, and finally a mask map for accurately locating the restoration area is generated to achieve accurate detection and labeling of video restoration traces. By comprehensively utilizing the time domain, spatial domain and frequency domain information of the repaired video, the multimodal characteristics of the repair traces are fully captured, which improves the accuracy and efficiency of video repair detection tasks and is suitable for diverse and complex video repair scenarios.
[0113] like Figure 2 As shown, an embodiment of the present invention further proposes a video restoration detection system based on multi-domain feature fusion, the system comprising:
[0114] The video sequence acquisition module 10 is used to acquire the video sequence to be detected, and segment the video sequence according to a preset segmentation rule to obtain a plurality of video segments;
[0115] A spatial feature extraction module 20, configured to convolve each single frame image in each of the video clips with a high-pass filter kernel to obtain a spatial feature corresponding to the single frame image;
[0116] A time domain feature extraction module 30 is used to obtain a forward optical flow from the t-th frame image to the t+1-th frame image and a backward optical flow from the t+1-th frame image to the t-th frame image based on an optical flow estimation method, and extract time domain features between adjacent frame images according to the forward optical flow and the backward optical flow;
[0117] The frequency domain feature extraction module 40 is used to perform time-frequency domain conversion on the video clip to obtain frequency components corresponding to the video clip, obtain a first frequency band, a second frequency band and a third frequency band according to the frequency components and a preset analysis matrix, and connect the first frequency band, the second frequency band and the third frequency band to obtain frequency domain features corresponding to the video clip;
[0118] The mask image generation module 50 is used to obtain the fusion features of each video clip based on the frequency domain features, all spatial domain features and all time domain features of the same video clip, and generate a mask image with video repair area detection results corresponding to the video clip based on each fusion feature.
[0119] On the other hand, the present invention further proposes a storage medium on which one or more programs are stored, and when the program is executed by a processor, the above-mentioned video restoration detection method based on multi-domain feature fusion is implemented.
[0120] On the other hand, the present invention also proposes an electronic device, including a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory to implement the above-mentioned video restoration detection method based on multi-domain feature fusion.
[0121] Those skilled in the art will appreciate that the logic and / or steps represented in the flowchart or otherwise described herein, for example, may be considered as an ordered list of executable instructions for implementing logical functions, and may be specifically implemented in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in conjunction with such instruction execution systems, devices or apparatuses. For purposes of this specification, "computer-readable medium" may be any device that can contain storage, communication, propagation or transmission of a program for use by an instruction execution system, device or apparatus, or in conjunction with such instruction execution systems, devices or apparatuses.
[0122] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering or, if necessary, processing in another suitable manner, and then stored in a computer memory.
[0123] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or a combination thereof: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0124] Although the embodiments of the present invention are described in detail above, it is obvious to those skilled in the art that various modifications and variations can be made to these embodiments. However, it should be understood that such modifications and variations are within the scope and spirit of the present invention as described in the claims. Moreover, the present invention described herein may have other embodiments and may be implemented or realized in a variety of ways.
Claims
1. A video restoration detection method based on multi-domain feature fusion, characterized in that: The method comprises: Acquire a video sequence to be detected, and segment the video sequence according to a preset segmentation rule to obtain multiple video clips; Convolving each single frame image in each of the video clips with a high-pass filter kernel to obtain spatial domain features corresponding to the single frame image; The convolution operation is performed according to the following formula: ; in, Represents the spatial characteristics, Represents a single frame image, represents the high-pass filter kernel, , Represents the convolution operation; ; The initial values of parameters a, b, and c are all set to 0, and are adaptively updated during model training, and their value ranges are all within [-1, 1); Obtaining a forward optical flow from the t-th frame image to the t+1-th frame image and a backward optical flow from the t+1-th frame image to the t-th frame image based on an optical flow estimation method, and extracting temporal features between adjacent frame images according to the forward optical flow and the backward optical flow; The forward optical flow from the t-th frame image to the t+1-th frame image is obtained according to the following formula: ; The backward optical flow from the t+1th frame image to the tth frame image is obtained according to the following formula: ; in, represents the forward optical flow from the t-th frame image to the t+1-th frame image, represents the t-th frame image, represents the t+1th frame image, Represents the backward optical flow from the t+1th frame image to the tth frame image; The time domain characteristics are calculated according to the following formula: ; in, Represents the time domain features between the t-th frame image and the t+1-th frame image; Performing a time-frequency domain conversion on the video clip to obtain a frequency component corresponding to the video clip, obtaining a first frequency band, a second frequency band, and a third frequency band according to the frequency component and a preset analysis matrix, and performing a connection operation on the first frequency band, the second frequency band, and the third frequency band to obtain a frequency domain feature corresponding to the video clip; The fusion features of each video clip are obtained according to the frequency domain features, all the spatial domain features and all the time domain features of the same video clip, and a mask image with a video restoration area detection result corresponding to the video clip is generated according to each of the fusion features.
2. The video restoration detection method based on multi-domain feature fusion according to claim 1 is characterized in that: The step of obtaining a video sequence to be detected and segmenting the video sequence according to a preset segmentation rule to obtain a plurality of video segments comprises: Divide the video sequence to be detected into multiple video segments with T frames per segment. ; For the last video clip If the video clip If the number of frames is less than T, the video clip will be repeated The last frame of the video clip is filled in The number of frames reaches T frames.
3. The video restoration detection method based on multi-domain feature fusion according to claim 2 is characterized in that: The step of performing time-frequency domain conversion on the video clip to obtain a frequency component corresponding to the video clip comprises: The time-frequency domain conversion is performed according to the following formula: ; ; in, represents the corresponding frequency component in the frequency domain, , , Both represent normalization coefficients, H represents the height of any image in the video clip, and W represents the width of any image in the video clip. Represents the pixel value of the video segment at time step i, height x, width y.
4. The video restoration detection method based on multi-domain feature fusion according to claim 3 is characterized in that: The step of acquiring a first frequency band, a second frequency band, and a third frequency band according to the frequency components and a preset analysis matrix, and connecting the first frequency band, the second frequency band, and the third frequency band to obtain a frequency domain feature corresponding to the video clip comprises: The first frequency band, the second frequency band, and the third frequency band are obtained according to the following formula: ; in, Indicates the first frequency band, Indicates the second frequency band, Indicates the third frequency band, represents the first frequency component, represents the second frequency component, represents the third frequency component, represents the identity matrix, , , All of them represent analysis matrices. When initialized, all elements of the analysis matrix are set to 0. During the training process, all elements of the analysis matrix will be adaptively updated according to the frequency response of the repair traces. The value range is Inside; The frequency domain features are obtained according to the following formula: ; in, Indicates a connection operation. Represents frequency domain features.
5. The video restoration detection method based on multi-domain feature fusion according to claim 4 is characterized in that: The step of obtaining fusion features under each video clip according to the frequency domain features, all spatial domain features and all time domain features under the same video clip, and generating a mask map with a video repair area detection result corresponding to the video clip according to each fusion feature comprises: The frequency domain features, all spatial domain features, and all time domain features of the same video clip are feature mapped through a ResNet50 network with unshared parameters, so that they are projected into a feature space of uniform dimension: ; Take advantage of additional The network integrates three types of features: time, space, and frequency domains: ; The fused features are converted into query vector Q, key vector K and value vector V, and the features at each position are weighted summed through the multi-head self-attention mechanism: ; ; The mask map is obtained according to the following formula: ; in, represents the spatial features after mapping, represents the time domain features after mapping, represents the frequency domain features after mapping, represents the attention head associated with the query vector, represents the attention head associated with the key vector, represents the attention head associated with the value vector, represents the multi-head self-attention mechanism, represents the sth attention head associated with the query vector, represents the sth attention head associated with the key vector, represents the sth attention head associated with the value vector, represents the mask image, represents a multi-layer perception operation, Represents the linear transformation matrix used for the final output mapping, represents the dimension of the key vector, Represents the characteristics of deep modeling, represents the transpose of a matrix, Represent the weight matrices used to generate query vectors, key vectors, and value vectors, respectively. Both represent bias vectors.
6. A video restoration detection system based on multi-domain feature fusion, characterized in that: The system comprises: The video sequence acquisition module is used to acquire the video sequence to be detected and segment the video sequence according to a preset segmentation rule to obtain multiple video clips; A spatial feature extraction module, used for convolving each single frame image in each of the video clips with a high-pass filter kernel to obtain a spatial feature corresponding to the single frame image; The convolution operation is performed according to the following formula: ; in, Represents the spatial characteristics, Represents a single frame image, represents the high-pass filter kernel, , Represents the convolution operation; ; The initial values of parameters a, b, and c are all set to 0, and are adaptively updated during model training, and their value ranges are all within [-1, 1); A time domain feature extraction module, used to obtain the forward optical flow from the t-th frame image to the t+1-th frame image and the backward optical flow from the t+1-th frame image to the t-th frame image based on the optical flow estimation method, and extract the time domain features between adjacent frame images according to the forward optical flow and the backward optical flow; The forward optical flow from the t-th frame image to the t+1-th frame image is obtained according to the following formula: ; The backward optical flow from the t+1th frame image to the tth frame image is obtained according to the following formula: ; in, represents the forward optical flow from the t-th frame image to the t+1-th frame image, represents the t-th frame image, represents the t+1th frame image, Represents the backward optical flow from the t+1th frame image to the tth frame image; The time domain characteristics are calculated according to the following formula: ; in, Represents the time domain features between the t-th frame image and the t+1-th frame image; A frequency domain feature extraction module, used to perform time-frequency domain conversion on the video clip to obtain frequency components corresponding to the video clip, obtain a first frequency band, a second frequency band and a third frequency band according to the frequency components and a preset analysis matrix, and connect the first frequency band, the second frequency band and the third frequency band to obtain frequency domain features corresponding to the video clip; The mask image generation module is used to obtain the fusion features of each video clip based on the frequency domain features, all spatial domain features and all time domain features of the same video clip, and generate a mask image with video repair area detection results corresponding to the video clip based on each fusion feature.
7. A storage medium, characterized in that: The storage medium stores one or more programs, which, when executed by the processor, implement the video restoration detection method based on multi-domain feature fusion as described in any one of claims 1 to 5.
8. An electronic device, comprising a memory and a processor, wherein: The memory is used to store computer programs; When the processor is used to execute the computer program stored in the memory, it implements the video restoration detection method based on multi-domain feature fusion as described in any one of claims 1 to 5.