AIGC video long-range consistency maintaining method and system based on space-time attention mechanism
Through the method based on the space-time attention mechanism, the time and space characteristics of the video clips are extracted and fusion processed, which solves the shortcomings of traditional methods in maintaining long-range consistency of videos, and achieves more natural and coherent video generation.
Patent Information
- Application Number
- CN202510228389.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-30
AI Technical Summary
Traditional machine learning methods are difficult to effectively capture the long-range time dependencies of videos when processing videos, resulting in insufficient accuracy when videos maintain consistency over a longer span.
AIGC video long-range consistency maintenance method based on the space-time attention mechanism is adopted. By dividing the video into multiple video segments, the features of frame images and optical flow images are extracted, combined with time and space characteristics are fused, and consistency adjustment is performed using an exception discrimination model.
Effectively maintain long-range consistency between video clips, avoid time and space inconsistencies, and the generated videos are more natural and coherent, improving the video quality.
Smart Images

Figure CN120075492A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video detection, and more particularly to a method and system for maintaining long-range consistency of AIGC videos based on spatio-temporal attention mechanism. Background Art
[0002] With the rapid development of digital video technology, videos are increasingly widely used in various fields, including film and television production, advertising, surveillance and security, online education, etc. When watching a video, viewers expect the video content to be visually and logically coherent. In a movie, if there are frequent unreasonable changes in the appearance, clothing of characters or the layout of scenes in different scenes, this will seriously affect the viewing experience of viewers and reduce the quality and attractiveness of the video.
[0003] The complexity and diversity of video content are also increasing continuously, from simple static scene shooting to complex dynamic special effect synthesis, from short video clips to long movies and TV series. Traditional machine learning methods often face difficulties in feature extraction when dealing with such high-dimensional video data. The complexity of video data makes the extracted features may not be able to fully reflect the consistency information of the video. Moreover, these methods usually lack the ability to effectively capture the long-range temporal dependence relationship of videos, resulting in insufficient accuracy in detecting long-range consistency.
[0004] Therefore, in the process of video production and processing, ensuring the consistency of videos over a long time span has become an important challenge. Summary of the Invention
[0005] In view of this, the present invention provides a method and system for maintaining long-range consistency of AIGC videos based on spatio-temporal attention mechanism, which comprehensively considers the long-range consistency of videos in terms of time and space and can achieve the consistency of videos over a long time span.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] A method for maintaining long-range consistency of AIGC videos based on spatio-temporal attention mechanism, comprising:
[0008] Dividing the AIGC video into multiple video segments according to the video frame rate and a preset segment duration, performing frame extraction on the video segments to obtain frame images, and performing optical flow extraction on the frame images to obtain optical flow images;
[0009] Input the optical flow image and the frame image into the feature extraction network to obtain the temporal feature map and the spatial feature map. Input the temporal feature map and the spatial feature map into the preset result prediction model to obtain the temporal prediction result and the spatial prediction result of each video segment. Perform pooling processing on the temporal prediction result and the spatial prediction result to obtain the temporal fusion feature and the spatial fusion feature;
[0010] Respectively input the temporal fusion feature and the spatial fusion feature into the anomaly discrimination model to respectively determine the anomaly probability values corresponding to the frame image and the optical flow image, and preset an anomaly threshold. When the anomaly probability value exceeds the anomaly threshold, perform consistency adjustment on the next frame of the corresponding video segment.
[0011] Preferably, the preset result prediction model specifically includes: at least one Block block, a Reshape layer, a LIF layer, a fully connected layer, and a Softmax layer. Perform at least one temporal feature extraction on the frame image and the optical flow image of each video segment through the at least one Block block to obtain a first feature vector. Perform matrix transformation processing on the first feature vector through the Reshape layer to obtain a second feature vector. Perform temporal full connection processing on the second feature vector through the LIF layer and the fully connected layer to obtain a third feature vector. According to the third feature vector, determine the spatial prediction result and the temporal prediction result of each video segment through the Softmax layer.
[0012] Preferably, the Block block includes: a cascaded ConvLIF layer and a pooling layer.
[0013] Preferably, the obtaining of the first feature vector specifically includes: performing temporal convolution processing on the feature maps of the frame image and the optical flow image respectively through the ConvLIF layer to obtain a first temporal convolution vector. Performing pooling processing on the first temporal convolution vector through the pooling layer to obtain a first intermediate feature vector. When performing one-time temporal feature extraction, determine the first intermediate feature vector as the first feature vector;
[0014] When performing n - time temporal feature extraction, perform temporal convolution processing on the first intermediate feature vector through the ConvLIF layer to obtain a second temporal convolution vector, perform pooling processing on the second temporal convolution vector through the pooling layer to obtain a second intermediate feature vector, and so on. Determine the n - th intermediate feature vector as the first feature vector, where n is an integer greater than 1.
[0015] Preferably, the obtaining of the time fusion feature and the spatial fusion feature specifically includes: respectively performing dimensionality reduction processing on the time prediction result and the spatial prediction result to obtain the time average feature and the spatial average feature, and respectively performing feature fusion on the time average feature and the spatial average feature to obtain the time fusion feature and the spatial fusion feature.
[0016] Preferably, the training process of the anomaly discrimination model specifically includes:
[0017] Obtain an AIGC video dataset, where the AIGC video dataset includes multiple AIGC video sample data with real labels;
[0018] According to the AIGC video sample data, construct a time and space graph structure, where the graph structure includes graph nodes and edges;
[0019] Input the video sample data with real labels into the feature extraction network respectively to obtain the time feature and the spatial feature of each graph node;
[0020] Based on the time feature and the spatial feature of each graph node, perform iterative training on the initial anomaly discrimination model until a preset convergence condition is reached, then obtain the trained anomaly discrimination model.
[0021] Preferably, the feature fusion of the time average feature and the spatial average feature specifically includes: fusing the time average feature with the spatial prediction result to obtain the time fusion feature; fusing the spatial average feature with the time prediction result to obtain the spatial fusion feature.
[0022] An AIGC video long-range consistency maintenance system based on a spatio-temporal attention mechanism, including:
[0023] A video processing module that divides an AIGC video into multiple video segments according to the video frame rate and a preset segment duration, performs frame extraction on the video segments to obtain frame images, and performs optical flow extraction on the frame images to obtain optical flow images;
[0024] A feature fusion module that inputs the optical flow images and the frame images into a feature extraction network to obtain a time feature map and a spatial feature map, inputs the time feature map and the spatial feature map into a preset result prediction model to obtain the time prediction result and the spatial prediction result of each video segment; performs pooling processing on the time prediction result and the spatial prediction result to obtain a time fusion feature and a spatial fusion feature;
[0025] The consistency adjustment module inputs the time fusion feature and the space fusion feature into the anomaly discrimination model respectively to determine the anomaly probability values corresponding to the frame image and the optical flow image respectively, and preset an anomaly threshold. When the anomaly probability value exceeds the anomaly threshold, the next frame of the corresponding video segment is adjusted for consistency.
[0026] As can be seen from the above technical solutions, compared with the prior art, the present invention discloses a method and system for maintaining long-range consistency of AIGC videos based on spatio-temporal attention mechanism. By adopting the spatio-temporal attention mechanism, this method can effectively maintain the long-range consistency between video segments, avoid spatio-temporal inconsistencies during the generation process, make the generated videos more natural and coherent, and the fusion processing of time and space features can comprehensively consider the overall information of the video, improve the effect of anomaly detection and consistency adjustment, and make the quality of the finally generated videos higher. Description of the Drawings
[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.
[0028] Figure 1 It is a flowchart of the method provided by the present invention;
[0029] Figure 2 It is a structural schematic diagram provided by the present invention. Detailed Embodiments
[0030] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0031] The embodiments of the present invention disclose a method for maintaining long-range consistency of AIGC videos based on spatio-temporal attention mechanism, as Figure 1 shown, including:
[0032] Divide the AIGC video into multiple video segments according to the video frame rate and the preset segment duration, perform frame extraction on the video segments to obtain frame images, and perform optical flow extraction on the frame images to obtain optical flow images;
[0033] Input the optical flow image and the frame image into the feature extraction network to obtain the temporal feature map and the spatial feature map. Input the temporal feature map and the spatial feature map into a preset result prediction model to obtain the temporal prediction result and the spatial prediction result for each video segment; perform pooling processing on the temporal prediction result and the spatial prediction result to obtain the temporal fusion feature and the spatial fusion feature;
[0034] Respectively input the temporal fusion feature and the spatial fusion feature into the anomaly discrimination model to respectively determine the anomaly probability values corresponding to the frame image and the optical flow image, and preset an anomaly threshold. When the anomaly probability value exceeds the anomaly threshold, perform consistency adjustment on the next frame of the corresponding video segment.
[0035] Among them, the temporal prediction result and the spatial prediction result are used to characterize the motion trend or direction of the behavior in the video at the next moment of the video segment.
[0036] In a specific embodiment, frames are extracted from each video segment at a certain interval, where the interval is the total number of frames of the video segment divided by N 1 , to obtain N 1 frame images, and N 1 is an integer greater than 1. For the extracted N 1 frame images, calculate the optical flow by extracting the optical flow between the subsequent frame and the previous frame to obtain N 1 -1 optical flows; copy the optical flows of the second frame and the first frame as the first optical flow, and merge them with the N 1 -1 optical flows to form N 1 optical flows.
[0037] In a specific embodiment, there are 2 feature extraction networks, namely the temporal branch network and the spatial branch network. Send the optical flow image to the first feature extraction model for feature extraction to obtain temporal features; send the frame image to the second feature extraction model for feature extraction to obtain spatial features.
[0038] The temporal branch network includes the first feature extraction model, and the spatial branch network includes the second feature extraction model. After obtaining the frame image and the optical flow image, the optical flow image can be input into the first feature extraction model to obtain the temporal features corresponding to each video frame of each video segment, and the frame image can be input into the second feature extraction model to obtain the spatial features corresponding to each image block; the branch networks can both be I3D models, and the I3D model can be pre-trained through the AIGC video sample data in the AIGC video dataset.
[0039] In a specific embodiment, the preset result prediction model specifically includes: at least one Block block, a Reshape layer, a LIF layer, a fully connected layer, and a Softmax layer. At least one Block block performs at least one temporal feature extraction on the frame images and optical flow images of each video segment to obtain a first feature vector; the Reshape layer performs matrix transformation processing on the first feature vector to obtain a second feature vector; the LIF layer and the fully connected layer perform temporal fully connected processing on the second feature vector to obtain a third feature vector; according to the third feature vector, the Softmax layer determines the spatial prediction result and the temporal prediction result of each video segment.
[0040] Among them, temporal feature extraction may refer to performing temporal feature extraction processing on a feature map. Matrix transformation processing refers to the process of expanding the last few dimensions of a matrix. Temporal fully connected processing refers to fully connected processing with temporal processing. In this way, multiple pictures can be processed at one time, which can not only ensure the feature extraction effect, but also relate multiple pictures and process the temporal information between pictures, thereby improving the recognition accuracy. The result prediction model of the present invention uses a network that fuses ANN and SNN, that is, the fusion of the ConvLIF layer and the LIF layer with the normalization layer and the pooling layer. Among them, the LIF layer is a fully connected layer with time series, which can process information with time series. Its function is similar to that of LSTM in ANN, but the weight is significantly lower than that of LSTM (the computational complexity of LIF in the convolutional network of the present disclosure is only one-fourth of that of LSTM and only one-third of that of GRU), which greatly reduces the computational complexity, reduces the requirements for computing devices, and correspondingly reduces the size of the network and the storage space. The ConvLIF layer is a convolutional layer with temporal information, which can process convolutional operations with time series. In the convolution of ANN, only one picture can be processed and there is no association with the pictures before and after, while the ConvLIF layer can process multiple pictures at one time, that is, it can achieve the convolution effect in ANN and can also relate multiple pictures and process the temporal information between pictures. In addition, the weight of the ConvLIF layer is also significantly lower than that of the Conv3D layer (the weight and computational complexity of the ConvLIF2D layer in the convolutional network of the present disclosure are only one-half of those of the Conv3D layer), further reducing the computational complexity, reducing the requirements for computing devices, and reducing the size of the network and the storage space.
[0041] In an alternative embodiment, when calculating the optical flow, the Brox algorithm is adopted.
[0042] In an alternative embodiment, the Block block includes: a cascaded ConvLIF layer and a pooling layer.
[0043] The Block also includes a BN (Batch Normalization) layer cascaded between the ConvLIF layer and the pooling layer. The BN layer normalizes the temporal convolution vector, and then the normalized temporal convolution vector is pooled. Since the dimension of the output data of the Block is not suitable as the input of the LIF layer, a Reshape layer is added to process the output data of the Block, and the dimension of the data is expanded as the input of the LIF layer. For example, the output shape of the Block is (10, 2, 2, 1024). After adding the reshape layer, the output data is processed, and the last three dimensions are directly expanded to obtain data with a shape of (10, 4096). The BN (Batch Normalization) layer cascaded between the ConvLIF layer and the pooling layer performs batch normalization on the data, which can accelerate the network convergence speed and improve the stability of training.
[0044] In a specific embodiment, obtaining the first feature vector specifically includes: performing temporal convolution processing on the feature maps of the frame image and the optical flow image respectively through the ConvLIF layer to obtain the first temporal convolution vector; performing pooling processing on the first temporal convolution vector through the pooling layer to obtain the first intermediate feature vector; when performing temporal feature extraction once, the first intermediate feature vector is determined as the first feature vector;
[0045] When performing temporal feature extraction n times, performing temporal convolution processing on the first intermediate feature vector through the ConvLIF layer to obtain the second temporal convolution vector, performing pooling processing on the second temporal convolution vector through the pooling layer to obtain the second intermediate feature vector, and so on. The nth intermediate feature vector is determined as the first feature vector, where n is an integer greater than 1.
[0046] In a specific embodiment, obtaining the temporal fusion feature and the spatial fusion feature specifically includes: respectively performing dimensionality reduction processing on the temporal prediction result and the spatial prediction result to obtain the temporal average feature and the spatial average feature, and respectively performing feature fusion on the temporal average feature and the spatial average feature to obtain the temporal fusion feature and the spatial fusion feature. The temporal average feature can be the mean of all temporal features, and the spatial average feature can be the mean of all spatial features. Then, the temporal average feature and the spatial-temporal feature are subjected to feature fusion to obtain the temporal fusion feature and the spatial fusion feature for subsequent abnormal probability judgment.
[0047] In a specific embodiment, the training process of the abnormal discrimination model specifically includes:
[0048] Obtain an AIGC video dataset, where the AIGC video dataset includes multiple AIGC video sample data with real labels;
[0049] Construct a time and space graph structure based on the AIGC video sample data. The graph structure includes graph nodes and edges;
[0050] Input the video sample data with real labels into the feature extraction network respectively to obtain the time features and space features of each graph node;
[0051] Based on the time features and space features of each graph node, iteratively train the initial anomaly discrimination model until the preset convergence condition is reached, then obtain the trained anomaly discrimination model.
[0052] In a specific embodiment, the feature fusion of the time average feature and the space average feature respectively includes: fusing the time average feature with the space prediction result to obtain the time fusion feature; fusing the space average feature with the time prediction result to obtain the space fusion feature. The specific formulas are as follows:
[0053]
[0054] Among them, and can respectively represent the time feature of the time branch network and the space feature in the space branch network. N and M respectively represent the number of optical flow images in the time branch network and the number of frame images in the space branch network.
[0055] Among them, the anomaly discrimination model can also be a graph attention network (GAT) model. Through this GAT model, the features of adjacent graph nodes can be better aggregated and updated, so as to better detect the anomaly probability values of frame images and optical flow images in video segments.
[0056] An AIGC video long-range consistency maintenance system based on spatio-temporal attention mechanism, as Figure 2 shown, includes:
[0057] A video processing module divides the AIGC video into multiple video segments according to the video frame rate and the preset segment duration, performs frame extraction on the video segments to obtain frame images, and performs optical flow extraction on the frame images to obtain optical flow images;
[0058] A feature fusion module inputs the optical flow images and frame images into the feature extraction network to obtain a time feature map and a space feature map, inputs the time feature map and the space feature map into a preset result prediction model to obtain the time prediction result and the space prediction result of each video segment; performs pooling processing on the time prediction result and the space prediction result to obtain the time fusion feature and the space fusion feature;
[0059] The consistency adjustment module inputs the time fusion feature and the space fusion feature into the anomaly discrimination model respectively to determine the anomaly probability values corresponding to the frame image and the optical flow image respectively, and a preset anomaly threshold is set. When the anomaly probability value exceeds the anomaly threshold, the next frame of the corresponding video segment is adjusted for consistency.
[0060] The various embodiments in this specification are described in a progressive manner. The key point of each embodiment is the difference from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, please refer to the description in the method section.
[0061] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for maintaining long-range consistency of AIGC videos based on spatiotemporal attention mechanism, characterized in that: include: Divide the AIGC video into multiple video segments according to the video frame rate and the preset segment duration, perform frame extraction processing on the video segments to obtain frame images, and perform optical flow extraction on the frame images to obtain optical flow images; Input the optical flow image and the frame image into a feature extraction network to obtain a temporal feature map and a spatial feature map, input the temporal feature map and the spatial feature map into a preset result prediction model to obtain a temporal prediction result and a spatial prediction result for each of the video clips; perform pooling processing on the temporal prediction result and the spatial prediction result to obtain a temporal fusion feature and a spatial fusion feature; The temporal fusion features and the spatial fusion features are respectively input into the abnormality discrimination model to respectively determine the abnormality probability values corresponding to the frame image and the optical flow image, and an abnormality threshold is preset. When the abnormality probability value exceeds the abnormality threshold, the next frame of the corresponding video clip is adjusted for consistency.
2. According to claim 1, a method for maintaining long-term consistency of AIGC video based on spatiotemporal attention mechanism is characterized in that: The preset result prediction model specifically includes: at least one Block block, a Reshape layer, a LIF layer, a fully connected layer and a Softmax layer. The at least one Block block is used to perform at least one temporal feature extraction on the frame image and the optical flow image of each video clip to obtain a first feature vector; the first feature vector is subjected to matrix transformation processing through the Reshape layer to obtain a second feature vector; the second feature vector is subjected to temporal fully connected processing through the LIF layer and the fully connected layer to obtain a third feature vector; according to the third feature vector, the spatial prediction result and the temporal prediction result of each video clip are determined through the Softmax layer.
3. According to claim 2, a method for maintaining long-term consistency of AIGC video based on spatiotemporal attention mechanism is characterized in that: The Block includes: a cascaded ConvLIF layer and a pooling layer.
4. According to claim 3, a method for maintaining long-term consistency of AIGC video based on spatiotemporal attention mechanism is characterized in that: The obtaining of the first feature vector specifically includes: performing temporal convolution processing on the feature graphs of the frame image and the optical flow image respectively through the ConvLIF layer to obtain a first temporal convolution vector; performing pooling processing on the first temporal convolution vector through the pooling layer to obtain a first intermediate feature vector; when performing a temporal feature extraction, determining the first intermediate feature vector as the first feature vector; When performing time series feature extraction n times, the first intermediate feature vector is subjected to time series convolution processing through the ConvLIF layer to obtain a second time series convolution vector, and the second time series convolution vector is subjected to pooling processing through the pooling layer to obtain a second intermediate feature vector, and so on, the nth intermediate feature vector is determined as the first feature vector, and n is an integer greater than 1.
5. According to claim 2, a method for maintaining long-term consistency of AIGC video based on spatiotemporal attention mechanism is characterized in that: The obtaining of the temporal fusion features and the spatial fusion features specifically includes: performing dimensionality reduction processing on the temporal prediction results and the spatial prediction results respectively to obtain the temporal average features and the spatial average features, and performing feature fusion on the temporal average features and the spatial average features respectively to obtain the temporal fusion features and the spatial fusion features.
6. According to claim 1, a method for maintaining long-range consistency of AIGC video based on spatiotemporal attention mechanism is characterized in that: The training process of the abnormality discrimination model specifically includes: Obtain an AIGC video dataset, where the AIGC video dataset includes a plurality of AIGC video sample data with real labels; Constructing a time and space graph structure according to the AIGC video sample data, wherein the graph structure includes graph nodes and edges; Input the video sample data with real labels into the feature extraction network respectively to obtain the temporal features and spatial features of each graph node; Based on the temporal features and spatial features of each of the graph nodes, the initial anomaly discrimination model is iteratively trained until a preset convergence condition is reached, thereby obtaining the trained anomaly discrimination model.
7. According to claim 5, a method for maintaining long-term consistency of AIGC video based on spatiotemporal attention mechanism is characterized in that: The feature fusion of the temporal average feature and the spatial average feature respectively specifically includes: fusing the temporal average feature with the spatial prediction result to obtain the temporal fusion feature; fusing the spatial average feature with the temporal prediction result to obtain the spatial fusion feature.
8. A system for maintaining long-term consistency of AIGC videos based on a spatiotemporal attention mechanism, applying a method for maintaining long-term consistency of AIGC videos based on a spatiotemporal attention mechanism as described in any one of claims 1 to 7, characterized in that: include: A video processing module, which divides the AIGC video into multiple video segments according to the video frame rate and the preset segment duration, performs frame extraction processing on the video segments to obtain frame images, and performs optical flow extraction on the frame images to obtain optical flow images; A feature fusion module inputs the optical flow image and the frame image into a feature extraction network to obtain a temporal feature map and a spatial feature map, inputs the temporal feature map and the spatial feature map into a preset result prediction model to obtain a temporal prediction result and a spatial prediction result of each of the video clips; performs pooling processing on the temporal prediction result and the spatial prediction result to obtain a temporal fusion feature and a spatial fusion feature; The consistency adjustment module inputs the temporal fusion features and the spatial fusion features into the abnormality discrimination model respectively to determine the abnormality probability values corresponding to the frame image and the optical flow image respectively, and presets the abnormality threshold. When the abnormality probability value exceeds the abnormality threshold, the next frame of the corresponding video clip is adjusted for consistency.