Video content intelligent analysis and statistical method under support of large model
Through the video content analysis method supported by the big model, the Swin Transformer model and the timing modeling network integrate the timing information of video frames, the problem of insufficient accuracy of behavior recognition in the prior art is solved, and more efficient video analysis is achieved.
Patent Information
- Application Number
- CN202411966416.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing video content analysis technologies are difficult to effectively integrate the timing information between video frames, resulting in insufficient accuracy and robustness of behavior recognition.
Using the method supported by the big model, the visual features of video frames are extracted through the Swin Transformer model and organized into a time series matrix. Then, the time-dependence between video frames is captured using the timing modeling network, generating a new feature matrix containing dynamic information. Finally, space-time correlation modeling technology is used for target tracking and behavior recognition.
It improves the accuracy and robustness of behavior recognition, can more effectively capture the dynamic changes and continuity of behavior, and improves the accuracy of video analysis.
Smart Images

Figure CN119992407A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the technical field of video content analysis, and more specifically, to intelligent analysis and statistical methods of video content supported by large models. Background Art
[0002] With the widespread application of video surveillance and social media, the amount of video data has shown explosive growth. As a technology for extracting useful information from video data, video content analysis has important application value in public safety, content review, sports analysis, intelligent transportation and other fields. Video content analysis usually involves multiple aspects such as target detection, behavior recognition, and sentiment analysis. The core challenge lies in how to accurately and efficiently extract and utilize information from massive video data.
[0003] Among the existing video content analysis technologies, a common approach is based on a deep learning framework, specifically using a convolutional neural network (CNN) to extract visual features from video frames. This approach uses a CNN model to extract features from each frame in the video, and then inputs the extracted features into a classifier for behavior recognition. For example, a mature CNN model such as VGG or ResNet can be used to obtain deep features of video frames, and then a support vector machine (SVM) or softmax layer can be used to classify behaviors.
[0004] Although the CNN-based video content analysis method has achieved good results in many scenarios, this method has a core technical problem: how to effectively integrate the temporal information between video frames to improve the accuracy and robustness of behavior recognition. Specifically, the CNN model mainly focuses on the feature extraction of single-frame images, while the essence of video is a continuous sequence of frames, which contains rich temporal information. Relying solely on single-frame features for behavior recognition often fails to accurately capture the dynamic changes and continuity of behavior, especially in complex and diverse video scenes. Summary of the invention
[0005] The present invention provides a video content intelligent analysis and statistical method supported by a large model, which aims to solve the technical problem that the current existing technology relies on single-frame features to perform behavior recognition, which often cannot accurately capture the dynamic changes and continuity of the behavior.
[0006] The video content intelligent analysis and statistical method supported by the big model includes the following steps:
[0007] Step 1: Input each frame of the video into the Swin Transformer model and perform feature extraction to obtain the visual feature vector of each frame;
[0008] Step 2: Arrange the visual feature vectors of each frame into a time series in chronological order to form a time series matrix. Each row of the matrix corresponds to a time point, and each column is the visual feature of the corresponding time point.
[0009] Step 3: Input the time series matrix into the time series modeling network, capture the dependency of the time series data based on the time series modeling network, and generate a new feature matrix. The feature matrix integrates the time information with the visual features of each frame and contains the dynamic information of each frame in the time dimension.
[0010] Step 4: Use the new feature matrix as the input of the temporal behavior recognition model, assign a behavior label to each video clip through the temporal behavior recognition model, and obtain the behavior recognition result;
[0011] Step 5: Track each target through space-time correlation modeling technology and associate it with the behavior recognition results to obtain the behavior sequence, behavior timestamp, and behavior time series duration and frequency of each target.
[0012] The present invention extracts rich visual features through the Swin Transformer model and dynamically adjusts the feature weights according to the application scenario. Then, the extracted features are organized into a time series matrix to facilitate subsequent time series information fusion. Through a specially designed time series modeling network, the time dependency between video frames is captured, and a new feature matrix containing dynamic information is generated, thereby improving the accuracy and robustness of behavior recognition. Finally, the space-time association modeling technology is used to achieve accurate tracking of targets in the video, and combined with the behavior recognition results, the video content is comprehensively analyzed and counted. This method effectively improves the accuracy of video analysis and solves the problem of insufficient capture of dynamic changes and continuity in behavior recognition in the prior art.
[0013] Preferably, step 2 comprises the following steps:
[0014] Temporal coding: The temporal coding mechanism is introduced to perform temporal position coding on the visual feature vector of each frame:
[0015]
[0016] Where: represents the visual feature vector after adding time series information, where d represents the dimension of the feature vector, represents the set of all vectors consisting of d real numbers; represents the visual feature vector obtained by step 1; represents the time encoding vector;
[0017] The time coding vector is obtained by alternately calculating the time coding of all dimensions in each frame using sine and cosine functions, and then concatenating the time coding of all dimensions:
[0018]
[0019] Where: t represents the time step index, which indicates the time position of the current time frame; i represents the index of the feature dimension, which is used to determine the position of the sine and cosine functions in the time encoding vector; represents the time encoding of the 2ith dimension of the tth frame; 2i+n=d;
[0020] Forming a time series matrix: Based on the results of time series coding, sort them in chronological order to form a time series matrix:
[0021]
[0022] in: Represents the time series matrix; T represents the total number of frames in the video.
[0023] Preferably, the temporal modeling network includes a temporal adaptive graph convolutional network, a multi-scale temporal attention mechanism and a fusion network;
[0024] The temporal adaptive graph convolutional network includes an input layer, a convolutional layer and an output layer;
[0025] The input layer is used to input a time series matrix;
[0026] Convolutional layer: First, the temporal dependency between each frame is calculated based on the time series matrix, and the adjacency matrix of the graph is constructed:
[0027]
[0028] Where: and They represent the feature vectors corresponding to the video frame t′ and the video frame t, respectively, and both have dimensions of d; · represents the vector dot product operation; and Represents vectors and The L2 norm of A seq (t, t′) represents the similarity between video frame t and video frame t′;
[0029] For each pair of video frame t′ and video frame t, calculate the similarity and put each pair of similarities into the adaptive adjacency matrix A seq ;
[0030] Based on the obtained adjacency matrix A seq , perform graph convolution operation:
[0031] H GCN =σ(A seq ·F seq );
[0032] Where: σ represents the nonlinear activation function; H GCN Represents the preliminary time series feature matrix obtained after the convolution operation;
[0033] Output layer: outputs the preliminary time series feature matrix;
[0034] The multi-scale temporal attention mechanism includes:
[0035] Multi-scale time window: Introduce multiple time scales to model the temporal characteristics of different events. For each time scale m, define a time window W m ;
[0036] Calculate attention weights: For each frame t, at each time scale m, the attention weights are dynamically assigned by calculating the similarity between frames:
[0037]
[0038] Where: α t,m (t′) represents the attention weights of frame t and frame t′;
[0039] Weighted feature calculation: For each time scale m, the attention weight α t,m (t′) weights the temporal features between frames to obtain the weighted feature representation:
[0040]
[0041] Where: H att,m Represents the weighted feature representation at time scale m, which contains the temporal dependency information at this time scale;
[0042] Fusion of features at multiple scales: The features H at different scales are combined att,m Perform weighted fusion to obtain the final multi-scale time series feature matrix:
[0043]
[0044] Where: α m represents the weight coefficient of each scale; M represents the number of time scales; H multi-scale Represents a multi-scale time series feature matrix;
[0045] Fusion network: Based on concatenation or weighted summation, the multi-scale time series feature matrix H multi-scale And the preliminary time series feature matrix H GCN Fusion is performed to obtain a new feature matrix Hfused .
[0046] Preferably, the temporal behavior recognition model includes an input layer, an LSTM layer, a bidirectional gated recurrent unit layer, an attention mechanism layer, a fully connected layer, and a classification layer;
[0047] The input layer is used to input the new feature matrix output in step 3;
[0048] The LSTM layer includes a first LSTM layer and a second LSTM layer. The first LSTM layer processes the output of the input layer, wherein the first LSTM layer uses return_sequences=True to retain sequence information; the second LSTM layer uses the same number of neurons as the first LSTM layer, uses return_sequences=False, and only outputs the features of the last time point;
[0049] The bidirectional gated recurrent unit layer processes the data output by the second LSTM layer through two forward and reverse GRU layers, concatenates the forward and reverse hidden states, and forms a high-dimensional representation containing forward and reverse information;
[0050] The attention mechanism layer performs weighted summation on the features output by the bidirectional gated recurrent unit layer according to the weights of each time step, and obtains a new feature representation through weighted average;
[0051] The fully connected layer maps the output data of the attention mechanism layer to a new space through the weight matrix and bias, and then outputs a high-dimensional feature representation through nonlinear activation;
[0052] The classification layer maps the high-dimensional feature representation of the fully connected output to the final classification space, calculates the probability distribution of the category through the softmax function, and takes the category with the highest probability as the result of this classification.
[0053] Preferably, step 5 comprises the following steps:
[0054] Association between behavior labels and temporal information: For each target, in each frame of the video, the target to which each frame belongs is determined based on the target tracking results, and the behavior label corresponding to each frame is obtained from the temporal behavior recognition model;
[0055] Associating the target's behavior label with the target's spatial information at that moment, thus obtaining a complete description of the target in both time and space dimensions;
[0056] Behavior timestamp statistics: record the start timestamp and end timestamp of a certain behavior of the target, and obtain the duration of the behavior based on the start timestamp and end timestamp; obtain the behavior sequence of each target based on the start timestamp and end timestamp;
[0057] Behavior frequency:
[0058]
[0059] Where: F c represents the frequency of behavior c; t represents the index of the time frame; T represents the total time frame; Represents the indicator function, which calculates whether the behavior is c at each time step t. If it is c, then Otherwise, it is 0; c represents the category of behavior;
[0060] Output results: Output the behavior sequence of each target, the timestamp, duration and frequency of each behavior.
[0061] Preferably, the target tracking steps are as follows:
[0062] Object detection: In each frame of the image, the object detection algorithm is used to identify all objects and give the bounding box of the object in the image;
[0063] Target position prediction: Based on the Kalman filter, the target state of the previous frame is used to predict the position of the target in the current frame, and the predicted position of each target in the current frame is output;
[0064] Target appearance feature extraction: Based on the bounding box of each frame, a convolutional neural network is used to extract features from each target area to obtain a high-dimensional appearance feature vector to represent the appearance of the target;
[0065] Target matching: Calculate the matching cost based on the target position prediction obtained by the Kalman filter, the target detection box of the current frame, and the target appearance feature extraction results. The matching cost includes the motion matching cost and the appearance matching cost. Based on the matching cost between each pair of targets, a cost matrix is obtained.
[0066] Data association: The Hungarian algorithm is used to find the optimal target match in the cost matrix, and the best target pairing is found by minimizing the matching cost to achieve target tracking.
[0067] Preferably, the matching cost is calculated as follows:
[0068] C ij =α·Motion_Cost(x t,i , b t,j )+β·Appearance_Cost(f t,i , f t,j );
[0069] Where: α represents the weight coefficient of motion matching cost; β represents the weight coefficient of appearance matching cost; Motion_Cost(xt,i , b t,j ) represents the motion matching cost; Appearance_Cost(f t,i , f t,j ) represents the appearance matching cost;
[0070] in:
[0071] Motion_Cost(x t,i , b t,j )=||x t,i -b t,j ||;
[0072] Where: x t,i Indicates the predicted position of the i-th target in the current frame; b t,j represents the detection box of the i-th target in the current frame; ||·|| represents the Euclidean distance;
[0073]
[0074] Where: f t,i represents the appearance feature vector of the i-th target in the current frame, obtained based on the convolutional neural network; f t,j represents the appearance feature vector of the jth target in the previous frame; f t,i ·f t,j Represents the dot product of two feature vectors, reflecting the similarity of the two vectors; ||f t,i ||||f t,j || represents the norm of two vectors, which is used to normalize the similarity.
[0075] The beneficial effects of the present invention include:
[0076] The present invention extracts rich visual features through the Swin Transformer model and dynamically adjusts the feature weights according to the application scenario. Then, the extracted features are organized into a time series matrix to facilitate subsequent time series information fusion. Through a specially designed time series modeling network, the time dependency between video frames is captured, and a new feature matrix containing dynamic information is generated, thereby improving the accuracy and robustness of behavior recognition. Finally, the space-time association modeling technology is used to achieve accurate tracking of targets in the video, and combined with the behavior recognition results, the video content is comprehensively analyzed and counted. This method effectively improves the accuracy of video analysis and solves the problem of insufficient capture of dynamic changes and continuity in behavior recognition in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0078] Figure 1 An overall step block diagram provided for an embodiment of the present invention.
[0079] Figure 2 A schematic diagram of the structure of a timing modeling network provided in an embodiment of the present invention.
[0080] Figure 3 A schematic diagram of the structure of a temporal behavior recognition model provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0081] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0082] See also Figure 1 As shown, the video content intelligent analysis and statistical method supported by the large model includes the following steps:
[0083] Step 1: Input each frame of the video into the Swin Transformer model, perform feature extraction to obtain a visual feature vector of each frame; the Swin Transformer model is an existing model, so the present invention will not be described in detail;
[0084] Step 2: Arrange the visual feature vectors of each frame into a time series in chronological order to form a time series matrix. Each row of the matrix corresponds to a time point, and each column is the visual feature of the corresponding time point.
[0085] The step 2 comprises the following steps:
[0086] Temporal coding: The temporal coding mechanism is introduced to perform temporal position coding on the visual feature vector of each frame:
[0087]
[0088] Where: represents the visual feature vector after adding time series information, where d represents the dimension of the feature vector, represents the set of all vectors consisting of d real numbers; represents the visual feature vector obtained by step 1; represents the time encoding vector;
[0089] The time coding vector is obtained by alternately calculating the time coding of all dimensions in each frame using sine and cosine functions, and then concatenating the time coding of all dimensions:
[0090]
[0091] Where: t represents the time step index, which indicates the time position of the current time frame; i represents the index of the feature dimension, which is used to determine the position of the sine and cosine functions in the time encoding vector; represents the time encoding of the 2ith dimension of the tth frame; 2i+n=d;
[0092] Forming a time series matrix: Based on the results of time series coding, sort them in chronological order to form a time series matrix:
[0093]
[0094] in: Represents the time series matrix; T represents the total number of frames in the video.
[0095] In this embodiment, time coding is advantageous for encoding the timing information of each time point into each frame by alternately using sine and cosine functions, and the coding method is periodic and can capture the relationship between different time points.
[0096] Step 3: Input the time series matrix into the time series modeling network, capture the dependency of the time series data based on the time series modeling network, and generate a new feature matrix. The feature matrix integrates the time information with the visual features of each frame and contains the dynamic information of each frame in the time dimension.
[0097] See also Figure 2 As shown, the temporal modeling network includes a temporal adaptive graph convolutional network, a multi-scale temporal attention mechanism and a fusion network;
[0098] The temporal adaptive graph convolutional network includes an input layer, a convolutional layer and an output layer;
[0099] The input layer is used to input a time series matrix;
[0100] Convolutional layer: First, the temporal dependency between each frame is calculated based on the time series matrix, and the adjacency matrix of the graph is constructed:
[0101]
[0102] Where: and They represent the feature vectors corresponding to the video frame t′ and the video frame t, respectively, and both have dimensions of d; · represents the vector dot product operation; and Represents vectors and The L2 norm of A seq (t, t′) represents the similarity between video frame t and video frame t′;
[0103] For each pair of video frame t′ and video frame t, calculate the similarity and put each pair of similarities into the adaptive adjacency matrix A seg ;
[0104] Based on the obtained adjacency matrix A seq , perform graph convolution operation:
[0105] H GCN =σ(A seq ·F seq );
[0106] Where: σ represents the nonlinear activation function; H GCN Represents the preliminary time series feature matrix obtained after the convolution operation;
[0107] In this embodiment, multiple convolutional layers may be provided, and the output of each convolutional layer is input to the next convolutional layer for further processing, and the output of the last convolutional layer is output to the output layer;
[0108] In the above technical solution, we give the processing steps of one convolution layer. If a convolution layer needs to be added, for example, there are two convolution layers, then the calculation of the second convolution layer is based on the output result of the first layer, where the adjacency matrix is dynamically updated based on the feature output of the current layer; that is, the adjacency matrix is calculated based on the input features of each layer and the adjacency matrix is updated. For example, the output of the first convolution layer is used as the input of the second convolution layer. At this time, the adjacency matrix of the second convolution layer is calculated based on the output of the first convolution layer, that is, the input of the second convolution layer, and the calculation formula is the same as the calculation formula of the above adjacency matrix; this enables the adjacency matrix to be dynamically adjusted according to the characteristics of each layer of convolution, so as to more flexibly capture the complex dependencies between frames in time series data.
[0109] Output layer: outputs the preliminary time series feature matrix;
[0110] The multi-scale temporal attention mechanism includes:
[0111] Multi-scale time window: Introduce multiple time scales to model the temporal characteristics of different events. For each time scale m, define a time window W m ;
[0112] Calculate attention weights: For each frame t, at each time scale m, the attention weights are dynamically assigned by calculating the similarity between frames:
[0113]
[0114] Where: α t,m (t′) represents the attention weight of frame t and frame t′; Attention is the attention mechanism; Indicates that the current time step t is in the time window W m The query vector under Represents time step t′ in time window W m The key vector under d k Represents the dimension of the key vector, used for normalization;
[0115] The query vector and key vector are calculated as follows:
[0116] In each time window W m Next, the query vector and key vector By input features After different linear transformations, we calculate the query vector and key vector through different weight matrices for each time step t and each scale m:
[0117]
[0118] Where: and are trainable weight matrices corresponding to the time window scale m respectively;
[0119] Weighted feature calculation: For each time scale m, the attention weight α t,m (t′) weights the temporal features between frames to obtain the weighted feature representation:
[0120]
[0121] Where: H att,m Represents the weighted feature representation at time scale m, which contains the temporal dependency information at this time scale;
[0122] Fusion of features at multiple scales: The features H at different scales are combined att,m Perform weighted fusion to obtain the final multi-scale time series feature matrix:
[0123]
[0124] Where: α m represents the weight coefficient of each scale; M represents the number of time scales; Hmulti-scale Represents a multi-scale time series feature matrix;
[0125] Fusion network: Based on concatenation or weighted summation, the multi-scale time series feature matrix H multi-scale And the preliminary time series feature matrix H GCN Fusion is performed to obtain a new feature matrix H fused .
[0126] In this embodiment, an adaptive graph structure is constructed to dynamically capture the temporal dependencies between frames, thereby avoiding the gradient vanishing problem of traditional RNN in long-term dependency modeling and effectively capturing the complex dependencies of temporal data. A multi-scale attention mechanism is introduced to adaptively learn the dependencies at different time scales in the video, thereby more accurately capturing the temporal characteristics of different events. Finally, through the deep fusion of temporal features and spatial features, it is ensured that the model can simultaneously understand the temporal dynamics and spatial content in the video.
[0127] Step 4: Use the new feature matrix as the input of the temporal behavior recognition model, assign a behavior label to each video clip through the temporal behavior recognition model, and obtain the behavior recognition result;
[0128] See also Figure 3 As shown, the temporal behavior recognition model includes an input layer, an LSTM layer, a bidirectional gated recurrent unit layer, an attention mechanism layer, a fully connected layer, and a classification layer;
[0129] The input layer is used to input the new feature matrix output in step 3;
[0130] The LSTM layer includes a first LSTM layer and a second LSTM layer, the first LSTM layer processes the output of the input layer, wherein the first LSTM layer uses return sequences = True to retain sequence information, wherein the output dimension is (T, 512), that is, the time step number T and the hidden state of 512 dimensions; the second LSTM layer uses the same number of neurons as the first LSTM layer, uses return_sequences = False, and only outputs the features of the last time point; wherein the output dimension of the second LSTM layer is 512 dimensions, and the length of the output sequence is 1;
[0131] For the LSTM layer, the update process is performed by the following formula:
[0132] Forget Gate t :
[0133] f t =σ(W f [h t-1 , x t ]+bf );
[0134] Where: h t-1 represents the hidden state at the previous moment, x t Represents the input at the current moment; W f and b f are the weight and bias of the forget gate respectively; σ represents the sigmoid activation function;
[0135] Input Gate i t :
[0136] i t =σ(W i [h t-1 , x t ]+b i );
[0137] Where: W i and b i Represent the weight and bias of the input gate respectively;
[0138] Candidate status
[0139]
[0140] Where: W c and b c Represent the weight and bias of the candidate state respectively;
[0141] Update cell status C t :
[0142]
[0143] Where: C t-1 Indicates the cell state updated at the last moment;
[0144] Output gate o t :
[0145] o t =σ(W o [h t-1 , x t ]+b o );
[0146] Where: W o and b o Represent the weight and bias of the output gate respectively;
[0147] Update hidden state h t :
[0148] h t =o t tanh(Ct );
[0149] Where: tanh represents the hyperbolic tangent activation function;
[0150] In this embodiment, the update process of the second LSTM layer is the same as that of the first LSTM layer, and the formula is the same as that of the first layer. The only difference is that the second LSTM layer is processed based on the output of the first LSTM layer, and the output sequence length of the second layer is 1 because return_sequences=False is used, which means that only the state of the last time step of the sequence is output.
[0151] The bidirectional gated recurrent unit layer processes the data output by the second LSTM layer through two forward and reverse GRU layers, concatenates the forward and reverse hidden states, and forms a high-dimensional representation containing forward and reverse information;
[0152] The calculation formula of the bidirectional gated recurrent unit layer is as follows:
[0153] Update gate z t :
[0154] z t =σ(W z [h t-1 , x t ]+b z );
[0155] Where: W z and b z Represent the weight and bias of the update gate respectively;
[0156] Reset Gate t :
[0157] r t =σ(W r [h t-1 , x t ]+b r );
[0158] Where: W r and b r Represent the weight and bias of the reset gate respectively;
[0159] Candidate Status
[0160]
[0161] Where: W h and b h Represent the weight and bias of the candidate state respectively;
[0162] Update hidden state h t :
[0163]
[0164] The output of the BiGRU layer is the hidden state of each time step. The bidirectional structure of BiGRU includes two GRU units, forward and reverse, and the output dimension is 512*2=1024.
[0165] The attention mechanism layer performs weighted summation on the features output by the bidirectional gated recurrent unit layer according to the weights of each time step, and obtains a new feature representation through weighted average;
[0166] Attention weight calculation:
[0167]
[0168] Where: α t represents the attention weight at time step t; h t represents the hidden state at time step t; q represents the query vector; score represents the function that calculates the similarity between the current time step t and the query vector q;
[0169] Weighted output:
[0170]
[0171] In this model, the attention mechanism can perform weighted summation of features at each time step and focus on frames related to action recognition.
[0172] The fully connected layer maps the output data of the attention mechanism layer to a new space through a weight matrix and a bias, and then outputs a high-dimensional feature representation through nonlinear activation; wherein the fully connected layer includes multiple fully connected layers, and the nonlinear expression ability of the model is enhanced through multiple fully connected layers;
[0173] The classification layer maps the high-dimensional feature representation of the fully connected output to the final classification space, calculates the probability distribution of the category through the softmax function, and takes the category with the highest probability as the result of this classification;
[0174]
[0175] Where: y i represents the i-th element of the classification layer output, and C represents the number of categories;
[0176] The loss function of the model adopts the cross entropy loss function.
[0177] In this embodiment, the temporal behavior recognition model (such as a classification model based on deep learning) takes the feature matrix as input and automatically recognizes the behaviors in the video clips by learning the patterns obtained from the training data. For example, for a sports video clip, the behavior labels may be recognized as "running" or "jumping"; for a security surveillance video clip, the labels may be recognized as "fighting" or "wandering".
[0178] In this embodiment, the multi-scale temporal attention mechanism used in step 3 is mainly used to temporally encode the visual feature vector of each frame in the feature extraction stage, and to model the dependencies between video frames at different time scales. Its goal is to generate a feature matrix that incorporates temporal information to provide richer temporal features for subsequent behavior recognition. In step 4, it is used to further focus on key frames to enhance the model's ability to recognize specific behavior patterns; it can significantly improve the performance of the video content analysis model, enabling it to more accurately and robustly identify and classify behaviors in the video. The synergy of these two attention mechanisms provides the model with a comprehensive temporal information processing capability, which helps to solve the complex challenges in video content analysis.
[0179] Step 5: Track each target through space-time correlation modeling technology and associate it with the behavior recognition results to obtain the behavior sequence, behavior timestamp, and behavior time series duration and frequency of each target.
[0180] The step 5 comprises the following steps:
[0181] Association between behavior labels and temporal information: For each target, in each frame of the video, the target to which each frame belongs is determined based on the target tracking results, and the behavior label corresponding to each frame is obtained from the temporal behavior recognition model;
[0182] Associating the target's behavior label with the target's spatial information at that moment, thus obtaining a complete description of the target in both time and space dimensions;
[0183] Behavior timestamp statistics: record the start timestamp and end timestamp of a certain behavior of the target, and obtain the duration of the behavior based on the start timestamp and end timestamp; obtain the behavior sequence of each target based on the start timestamp and end timestamp;
[0184] Behavior frequency:
[0185]
[0186] Where: F c represents the frequency of behavior c; t represents the index of the time frame; T represents the total time frame; Represents the indicator function, which calculates whether the behavior is c at each time step t. If it is c, then Otherwise, it is 0; c represents the category of behavior;
[0187] Output results: Output the behavior sequence of each target, the timestamp, duration and frequency of each behavior.
[0188] The steps of target tracking are as follows:
[0189] Object detection: In each frame, an object detection algorithm (such as YOLO, Faster R-CNN, etc.) is used to identify all objects and give the bounding box of the object in the image;
[0190] Output the detection box of the target through the target detection algorithm where b t,i = [x, y, w, h] represents the detection box of the i-th target, including the coordinates of the upper left corner of the target detection box (x, y), w and h represent the width and height of the target detection box, and t represents the number of frames; N t Indicates the number of targets in the current frame;
[0191] Target position prediction: Based on the Kalman filter, the target state of the previous frame is used to predict the position of the target in the current frame, and the predicted position of each target in the current frame is output. The specific steps are as follows:
[0192] State variables: Assume the state position and velocity vector of the target Where x, y are the coordinates of the upper left corner of the target detection box. The speed of the target;
[0193]
[0194] Where: represents the predicted state at the current time k; A represents the state transfer matrix; B represents the control input matrix; u k represents the control input; w k represents process noise; x k-1 Represents the state of the target at time k-1;
[0195] Measurement equation:
[0196] z k =H·x k +v k ;
[0197] Where: z k represents the measurement value at the current moment, i.e., the target position; H represents the observation matrix; v k represents the measurement noise; x k Indicates the state of the target at the current moment;
[0198] Update equation:
[0199]
[0200] Where: K represents the Kalman gain.
[0201] As a further implementation of this embodiment, due to different scenarios involved, different requirements for the parameters of the Kalman filter are required. Therefore, in this embodiment, specific steps for dynamically adjusting the parameters of the Kalman filter are introduced so that the adjusted parameters are more in line with the corresponding application scenarios, as follows:
[0202] W k =σ p ·(1+λ scene ·D scene );
[0203] v k =σ m ·(1+μ scene ·D acene );
[0204] Where: w k represents process noise; v k represents the measurement noise; σ p represents the basic process noise coefficient; σ m represents the basic measurement noise factor; λ scene Represents the scene dynamic factor, the coefficient that affects the process noise; μ scene Represents the scene dynamic factor, the coefficient that affects the measurement noise; D scene Indicators that indicate the degree of dynamics (e.g., vehicle speed changes in traffic scenes, frequency of human movement in security scenes, etc.) can be obtained by analyzing the motion of objects in the scene. The larger the value, the more dynamic the scene.
[0205] Exemplarily, the D scene Based on the following steps:
[0206] Calculate the target speed:
[0207]
[0208] Where: p i (t) represents the position of target i at time t; Δt represents the time interval between adjacent moments; v i (t) represents the speed of the target;
[0209] Calculate the speed change:
[0210] △v i (t)=|v i (t)-v i (t-1)|;
[0211] represents the velocity fluctuation of the target between consecutive time steps;
[0212] Calculate the dynamics of the scene: represented by the standard deviation of the change in target velocity, σΔv;
[0213] Then we get D by normalization scene , which is the normalized standard deviation σΔv.
[0214] Target appearance feature extraction: Based on the bounding box of each frame, a convolutional neural network is used to extract features of each target area to obtain a high-dimensional appearance feature vector f = [f1, f2, ..., f n ], used to indicate the appearance of the target;
[0215] Target matching: Calculate the matching cost based on the target position prediction obtained by the Kalman filter, the target detection box of the current frame, and the target appearance feature extraction results. The matching cost includes the motion matching cost and the appearance matching cost. Based on the matching cost between each pair of targets, a cost matrix is obtained.
[0216] The calculation method of calculating the matching cost is as follows:
[0217] C ij =α·Motion_Cost(x t,i , b t,j )+β·Appearance_Cost(f t,i , f t,j );
[0218] Where: α represents the weight coefficient of motion matching cost; β represents the weight coefficient of appearance matching cost; Motion_Cost(x t,i , b t,j ) represents the motion matching cost; Appearance_Cost(f t,i , f t,j ) represents the appearance matching cost;
[0219] in:
[0220] Motion_Cost(x t,i , b t,j )=||x t,i -b t,j ||;
[0221] Where: x t,i Indicates the predicted position of the i-th target in the current frame; b t,j represents the detection box of the i-th target in the current frame; ||·|| represents the Euclidean distance;
[0222]
[0223] Where: f t,i represents the appearance feature vector of the i-th target in the current frame, obtained based on the convolutional neural network; f t,j represents the appearance feature vector of the jth target in the previous frame; f t,i ·f t,j Represents the dot product of two feature vectors, reflecting the similarity of the two vectors; ||f t,i ||||f t,j || represents the norm of two vectors, which is used to normalize the similarity.
[0224] In target tracking, the matching cost consists of two parts: motion matching and appearance matching. Depending on the target's behavior type, the weights of motion matching cost and appearance matching cost can be dynamically adjusted to ensure that the algorithm can focus more on the most relevant target features in different scenarios:
[0225] α=α0·(1+γ behavior ·B behavior );
[0226] β=β0·(1+δ behavior ·B behavior );
[0227] Where: α0 and β0 both represent the initial matching cost weights; γ behavior Indicates the influence factor of behavior type on motion matching; δ behavior Indicates the influence factor of behavior type on appearance matching; B behavior An indicator that represents the behavior type, indicating the characteristics of the target behavior (such as fast movement, stillness, etc.). The larger the value, the more intense the target behavior is. The motion matching cost should be increased, and the appearance matching cost should be reduced.
[0228] Data association: The Hungarian algorithm is used to find the best target match in the cost matrix, and the best target pairing is found by minimizing the matching cost to achieve target tracking;
[0229] Assumptions: Target A and Target B are from the current frame and the previous frame respectively. The cost matrix is C, whose element targets are matched. The goal of the Hungarian algorithm is to minimize the total matching cost:
[0230]
[0231] Where: minimize means minimizing the total matching cost; x ij represents a binary variable, indicating whether to match target i from the current frame to the previous frame of target j; C ij Represents the matching cost.
[0232] The present invention extracts rich visual features through the Swin Transformer model and dynamically adjusts the feature weights according to the application scenario. Then, the extracted features are organized into a time series matrix to facilitate subsequent time series information fusion. Through a specially designed time series modeling network, the time dependency between video frames is captured, and a new feature matrix containing dynamic information is generated, thereby improving the accuracy and robustness of behavior recognition. Finally, the space-time association modeling technology is used to achieve accurate tracking of targets in the video, and combined with the behavior recognition results, the video content is comprehensively analyzed and counted. This method effectively improves the accuracy of video analysis and solves the problem of insufficient capture of dynamic changes and continuity in behavior recognition in the prior art.
[0233] The above are only preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application should be included in the protection scope of the present application.
Claims
1. Video content intelligent analysis and statistical method supported by a large model, characterized by: The following steps are involved: Step 1: Input each frame of the video into the Swin Transformer model and perform feature extraction to obtain the visual feature vector of each frame; Step 2: Arrange the visual feature vectors of each frame into a time series in chronological order to form a time series matrix. Each row of the matrix corresponds to a time point, and each column is the visual feature of the corresponding time point. Step 3: Input the time series matrix into the time series modeling network, capture the dependency of the time series data based on the time series modeling network, and generate a new feature matrix. The feature matrix integrates the time information with the visual features of each frame and contains the dynamic information of each frame in the time dimension. Step 4: Use the new feature matrix as the input of the temporal behavior recognition model, assign a behavior label to each video clip through the temporal behavior recognition model, and obtain the behavior recognition result; Step 5: Track each target through space-time correlation modeling technology and associate it with the behavior recognition results to obtain the behavior sequence, behavior timestamp, and behavior time series duration and frequency of each target.
2. The video content intelligent analysis and statistical method supported by a large model according to claim 1 is characterized in that: The step 2 comprises the following steps: Temporal coding: The temporal coding mechanism is introduced to perform temporal position coding on the visual feature vector of each frame: Where: represents the visual feature vector after adding time series information, where d represents the dimension of the feature vector, represents the set of all vectors consisting of d real numbers; represents the visual feature vector obtained by step 1; represents the time encoding vector; The time coding vector is obtained by alternately calculating the time coding of all dimensions in each frame using sine and cosine functions, and then concatenating the time coding of all dimensions: Where: t represents the time step index, which indicates the time position of the current time frame; i represents the index of the feature dimension, which is used to determine the position of the sine and cosine functions in the time encoding vector; represents the time encoding of the 2ith dimension of the tth frame; 2i+n=d; Forming a time series matrix: Based on the results of time series coding, sort them in chronological order to form a time series matrix: in: Represents the time series matrix; T represents the total number of frames in the video.
3. The video content intelligent analysis and statistical method supported by a large model according to claim 1 is characterized in that: The temporal modeling network includes a temporal adaptive graph convolutional network, a multi-scale temporal attention mechanism and a fusion network; The temporal adaptive graph convolutional network includes an input layer, a convolutional layer and an output layer; The input layer is used to input a time series matrix; Convolutional layer: First, the temporal dependency between each frame is calculated based on the time series matrix, and the adjacency matrix of the graph is constructed: Where: and They represent the feature vectors corresponding to the video frame t′ and the video frame t, respectively, and both have dimensions of d; · represents the vector dot product operation; and Represents vectors and The L2 norm of A seq (t, t′) represents the similarity between video frame t and video frame t′; For each pair of video frame t′ and video frame t, calculate the similarity and put each pair of similarities into the adaptive adjacency matrix A seq ; Based on the obtained adjacency matrix A seq , perform graph convolution operation: H GCN =σ(A seq ·F seq ); Where: σ represents the nonlinear activation function; H GCN Represents the preliminary time series feature matrix obtained after the convolution operation; Output layer: outputs the preliminary time series feature matrix; The multi-scale temporal attention mechanism includes: Multi-scale time window: Introduce multiple time scales to model the temporal characteristics of different events. For each time scale m, define a time window W m ; Calculate attention weights: For each frame t, at each time scale m, the attention weights are dynamically assigned by calculating the similarity between frames: Where: α t,m (t′) represents the attention weights of frame t and frame t′; Weighted feature calculation: For each time scale m, the attention weight α t,m (t′) weights the temporal features between frames to obtain the weighted feature representation: Where: H att,m Represents the weighted feature representation at time scale m, which contains the temporal dependency information at this time scale; Fusion of features at multiple scales: The features H at different scales are combined att,m Perform weighted fusion to obtain the final multi-scale time series feature matrix: Where: α m represents the weight coefficient of each scale; M represents the number of time scales; H multi-scale Represents a multi-scale time series feature matrix; Fusion network: Based on concatenation or weighted summation, the multi-scale time series feature matrix H multi-scale And the preliminary time series feature matrix H GCN Fusion is performed to obtain a new feature matrix H fused .
4. The video content intelligent analysis and statistical method supported by a large model according to claim 1 is characterized in that: The temporal behavior recognition model includes an input layer, an LSTM layer, a bidirectional gated recurrent unit layer, an attention mechanism layer, a fully connected layer, and a classification layer; The input layer is used to input the new feature matrix output in step 3; The LSTM layer includes a first LSTM layer and a second LSTM layer. The first LSTM layer processes the output of the input layer, wherein the first LSTM layer uses return_sequences=True to retain sequence information; the second LSTM layer uses the same number of neurons as the first LSTM layer, uses return_sequences=False, and only outputs the features of the last time point; The bidirectional gated recurrent unit layer processes the data output by the second LSTM layer through two forward and reverse GRU layers, concatenates the forward and reverse hidden states, and forms a high-dimensional representation containing forward and reverse information; The attention mechanism layer performs weighted summation on the features output by the bidirectional gated recurrent unit layer according to the weights of each time step, and obtains a new feature representation through weighted average; The fully connected layer maps the output data of the attention mechanism layer to a new space through the weight matrix and bias, and then outputs a high-dimensional feature representation through nonlinear activation; The classification layer maps the high-dimensional feature representation of the fully connected output to the final classification space, and calculates the probability distribution of the category through the softmax function; The category with the highest probability is taken as the result of this classification.
5. The video content intelligent analysis and statistical method supported by a large model according to claim 1 is characterized in that: The step 5 comprises the following steps: Association between behavior labels and temporal information: For each target, in each frame of the video, the target to which each frame belongs is determined based on the target tracking results, and the behavior label corresponding to each frame is obtained from the temporal behavior recognition model; Associating the target's behavior label with the target's spatial information at that moment, thus obtaining a complete description of the target in both time and space dimensions; Behavior timestamp statistics: record the start timestamp and end timestamp of a certain behavior of the target, and obtain the duration of the behavior based on the start timestamp and end timestamp; obtain the behavior sequence of each target based on the start timestamp and end timestamp; Behavior frequency: Where: F c represents the frequency of behavior c; t represents the index of the time frame; T represents the total time frame; Represents the indicator function, which calculates whether the behavior is c at each time step t. If it is c, then Otherwise, it is 0; c represents the category of behavior; Output results: Output the behavior sequence of each target, the timestamp, duration and frequency of each behavior.
6. The video content intelligent analysis and statistical method supported by a large model according to claim 5 is characterized in that: The steps of target tracking are as follows: Object detection: In each frame of the image, the object detection algorithm is used to identify all objects and give the bounding box of the object in the image; Target position prediction: Based on the Kalman filter, the target state of the previous frame is used to predict the position of the target in the current frame, and the predicted position of each target in the current frame is output; Target appearance feature extraction: Based on the bounding box of each frame, a convolutional neural network is used to extract features from each target area to obtain a high-dimensional appearance feature vector to represent the appearance of the target; Target matching: Calculate the matching cost based on the target position prediction obtained by the Kalman filter, the target detection frame of the current frame, and the target appearance feature extraction results. The matching cost includes the motion matching cost and the appearance matching cost. Based on the matching cost between each pair of targets, a cost matrix is obtained; Data association: The Hungarian algorithm is used to find the optimal target match in the cost matrix, and the best target pairing is found by minimizing the matching cost to achieve target tracking.
7. The video content intelligent analysis and statistical method supported by a large model according to claim 6 is characterized in that: The process noise parameters and measurement noise parameters of the Kalman filter are dynamically adjusted based on the application scenario: w k =s p ·(1+λ scene ·D scene ); v k =s m ·(1+μ scene ·D scene ); Where: w k represents process noise; v k represents the measurement noise; σ p represents the basic process noise coefficient; σ m represents the basic measurement noise factor; λ scene Represents the scene dynamic factor, the coefficient that affects the process noise; μ scene Represents the scene dynamic factor, the coefficient that affects the measurement noise; D scene Indicates the dynamic degree index.
8. The video content intelligent analysis and statistical method supported by a large model according to claim 6 is characterized in that: The calculation method of calculating the matching cost is as follows: C ij =α·Motion_Cost(x t,i ,b t,j )+β·Appearance_Cost(f t,i ,f t,j ); Where: α represents the weight coefficient of motion matching cost; β represents the weight coefficient of appearance matching cost; Motion_Cost(x t,i ,b t,j ) represents the motion matching cost; Appearance_Cost(f t,i ,f t,j ) represents the appearance matching cost; in: Motion_Cost(x t,i ,b t,j )=||x t,i -b t,j ||; Where: x t,i Indicates the predicted position of the i-th target in the current frame; b t,j represents the detection box of the i-th target in the current frame; ||·|| represents the Euclidean distance; Where: f t,i represents the appearance feature vector of the i-th target in the current frame, obtained based on the convolutional neural network; f t,j represents the appearance feature vector of the jth target in the previous frame; f t,i ·f t,j Represents the dot product of two feature vectors, reflecting the similarity of the two vectors; ||f t,i ||||f t,j || represents the norm of two vectors, which is used to normalize the similarity.
9. The video content intelligent analysis and statistical method supported by a large model according to claim 8, characterized in that: The weight coefficient of the motion matching cost and the weight coefficient of the appearance matching cost are dynamically adjusted based on the application scenario: α=α0·(1+γ behavior ·B behavior ); β=β0·(1+δ behavior ·B behavior ); Where: α0 and β0 both represent the initial matching cost weights; γ behavior Indicates the influence factor of behavior type on motion matching; δ behavior Indicates the influence factor of behavior type on appearance matching; B behavior An indicator of the behavior type that represents the characteristics of the target behavior.
Citation Information
Cited By
Video behavior identification method and system based on transformer substation monitoring
CN120472376A