A theft behavior recognition method based on adaptive temporal structure deep network
By adopting an adaptive time structure deep network in video processing, combining intra-segment video frame sampling LSTM and sub-behavior prototype parameter matrix, the problems of time adaptation and semantic analysis in long-term video processing are solved, and efficient identification of theft behavior is achieved.
Patent Information
- Application Number
- CN202210087197.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-25
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-01-25
AI Technical Summary
The prior art lacks time adaptability and global semantic information analysis capabilities when processing long-term videos, resulting in semantic loss problems.
Adaptive time structure deep network is adopted to realize time adaptation through the in-segment video frame sampling LSTM module, and a sub-behavior prototype parameter matrix is designed, and the semantic description related to the sub-behavior of each video segment is learned, and global semantic analysis is performed.
It realizes effective processing of long-term videos, reduces redundant calculations within segments, enhances semantic information analysis capabilities, and can effectively identify theft behavior.
Smart Images

Figure CN114463846B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of human behavior recognition, and in particular to a theft behavior recognition method based on an adaptive time structure deep network. Background Art
[0002] In order to maintain public security and protect people's property safety, it is necessary to supervise theft. With the development of deep learning and computer vision technology, intelligent behavior recognition for theft is gradually gaining attention.
[0003] Chinese patent application publication number CN111444861A "A method for identifying vehicle theft based on surveillance video" proposes a method for identifying vehicle theft based on surveillance video. It uses three-dimensional convolution to extract features, adds a spatiotemporal joint attention mechanism, focuses on the spatiotemporal position of the behavior, and classifies to achieve the identification of theft. However, the processing effect of this method on multi-stage and long-term videos has not been verified. Chinese patent application publication number CN109800717B "A method and system for sampling video frames for behavior recognition based on reinforcement learning" proposes a method for determining the importance score of the test based on LSTM, determining the key frame according to the score, predicting the behavior of the key frame, and then obtaining the video behavior prediction. However, prediction based only on the key frame can easily cause information loss or even misjudgment. Chinese patent application publication number CN109919031B "A method for identifying human behavior based on deep neural network" proposes a method for identifying human behavior based on deep neural network, which uses convolutional neural network units for learning high-level semantic features of frame images and recurrent neural network units for learning video behavior motion features. However, this method has a large amount of calculation and it is difficult to monitor abnormal behavior in real time.
[0004] Since computationally complex problems are common in video processing, keyframe extraction and processing methods are gradually used in video character behavior recognition. Xia Limin et al. used self-splitting competitive learning to extract keyframes of simple behavior clips in "Complex Human Behavior Recognition Based on Keyframes". Finally, complex human behaviors were recognized based on the similarity of keyframes in the behavior clips. Zhou Yuxin et al. expressed the product of image average edge strength and image structure clarity as image feature quantity in "Research on Lightweight Behavior Recognition Method Based on Keyframes" to find suitable keyframes in behavior video clips. Finally, the behavior classification score was obtained. Li Mingxiao et al. used the image information evaluation method in "Video Behavior Recognition Method Based on Fragment Keyframes" to find the image frame with the largest amount of information in the video clip, which is the keyframe, and used a multi-time scale two-stream network to extract features to achieve behavior recognition.
[0005] However, the above methods do not consider the temporal adaptation problem in long videos, and lack the ability to analyze global semantic information, which will cause semantic loss. Therefore, we propose a theft behavior recognition based on an adaptive temporal structure deep network. This method uses the intra-segment video frame sampling LSTM module to achieve temporal adaptation; designs a sub-behavior prototype parameter matrix, and learns the semantic description of the sub-behavior related to each video segment to achieve global semantic analysis. Summary of the invention
[0006] The purpose of the present invention is to remedy the defects of the prior art and provide a theft behavior recognition method based on an adaptive time structure deep network.
[0007] The present invention is achieved through the following technical solutions:
[0008] A theft behavior recognition method based on an adaptive time structure deep network specifically includes the following steps:
[0009] (1) Extract the video frame features of theft behavior;
[0010] (2) estimate the intra-segment temporal structure distance threshold;
[0011] (3) Calculate the hidden state features of the LSTM sampled from the video frame within the segment;
[0012] (4) Calculate the stealing behavior score based on the sub-behavior prototype time attention;
[0013] (5) Solve the parameter set of the adaptive time structure deep network;
[0014] (6) Theft behavior recognition based on adaptive temporal structure deep network.
[0015] The specific steps of extracting the theft behavior video frame features in step (1) are as follows:
[0016] Step 1-1 Take n samples of the theft video at equal time intervals. s video segments, n s The value is 8;
[0017] Step 1-2: In each video segment, continue to take n i subintervals, use random time subscripts to collect a video frame in each subinterval, and collect n video frames in total in this video segment. i Video frames, n i is the number of samples in the video frame within the segment, n i The value is 8;
[0018] Steps 1-3 use the pre-trained ResNet50 model to extract the features of each video frame and convert the features into row vectors to obtain the feature set of the video frame, X v,s ={x v,s,i}, where X v,s is a feature set, v represents the video number, s represents the video segment number, i represents the video frame number within the segment, and x v,s,i is a row vector with a feature dimension of 1x1000.
[0019] The estimation of the intra-segment temporal structure distance threshold in step (2) is specifically performed as follows:
[0020] Step 2-1 Given the features X of all video segments in the training set v,s ={x v,s,i};
[0021] Step 2-2 Calculate the distance d between the features of two video frames within a segment v,s,i,i+1 , the distance calculation formula is two features x v,s,i and x v,s,i+1 The 2-norm between them, the distance calculation formula is expressed as:
[0022] d v,s,i,i+1 =||x v,s,i -x v,s,i+1 ||2
[0023] Step 2-3: For all segments of all videos in the theft behavior training set, calculate the distance between the features of two video frames in the segment to obtain the distance set, D = {d v,s,i,i+1}, v is the number of all videos in the training set, s is the video segment number, and i is the video frame number within the segment;
[0024] Step 2-4 finds the mean u of the distance set d , the mean formula of the distance set is expressed as:
[0025]
[0026] Where n v is the number of videos, n s is the number of segments within each video;
[0027] Step 2-5 calculates the variance of the distance set. The variance formula of the distance set is expressed as:
[0028]
[0029] The time structure distance threshold estimation formula in steps 2-6 is expressed as:
[0030] τ d =u d+σ d .
[0031] The hidden state features of the LSTM sampling of the video frame in the calculation segment described in step (3) are specifically performed as follows:
[0032] Step 3-1 Given a feature X of a video segment v,s ={x v,s,i};
[0033] Step 3-2 calculates the binary sampling mark of each video frame in the segment;
[0034] Step 3-2-1 Calculate the distance d between two video frames in the segment according to step 2-2 v,s,i,i+1 ;
[0035] Step 3-2-2 Set the feature distance threshold τ d , the formula for calculating the sampling probability of the next video frame is expressed as:
[0036]
[0037] When the distance d between two video frames v,s,i,i+1 =τ d Then the sampling probability is 0.5, and as the distance increases, there is a greater probability of being sampled; as the distance decreases, there is a smaller probability of being sampled;
[0038] Step 3-2-3 performs binary 0-1 Bernoulli sampling according to the sampling probability of the next video frame calculated in step 3-2-2 to obtain the binary sampling probability of the next video frame; that is, a random number is generated using the uniform distribution z~U[0,1] in the interval [0,1]. If the random number is greater than the intra-segment sampling weight, 0 is returned, indicating that the next video frame is sampled. If the random number is less than the intra-segment sampling threshold, 1 is returned, indicating that the next video frame is not sampled. The intra-segment video frame sampling formula is expressed as:
[0039]
[0040] z is a random number, which is converted into random sampling to ensure randomness during sample sampling, allowing the model to sample samples with small feature weights; at this time, the binary sampling mark sm of each video frame in the segment is obtained v,s,i,i+1 , the binary sampling mark corresponds to the LSTM time node, where sm v,s,1,2 The binary sampling mark is used to estimate the binary sampling mark of the second video frame according to the feature distance between the first video frame and the second video frame;
[0041] Step 3-3 initializes the time node of LSTM according to the subscript of the video frame in the segment, where the i-th video frame corresponds to the i-th time node, and the binary mark of the i-th video frame is represented by sm v,s,i-1,i ;
[0042] Step 3-4 adds an initial node to the LSTM model and sets the binary sampling mark of the initial node and the first video frame to 1, that is, sm v,s,0,1 = 1, obtain the intra-segment binary sampling tag set SM of the LSTM with the initial node added = {sm v,s,0,1 ,sm v,s,1,2 ,...,sm v,s,ni-1,ni};
[0043] Step 3-5 initializes the memory state and hidden state of the initial time node;
[0044] Step 3-5-1 Calculate the average feature of the current segment The formula is expressed as:
[0045]
[0046] Step 3-5-2 uses a three-layer perceptron to learn the memory state of each LSTM initial node for the average feature of the current segment. The formula is:
[0047]
[0048] where φ init,c (·) is a three-layer perceptron for memory state.
[0049] Step 3-5-3 uses a three-layer perceptron to learn the hidden state of each LSTM initial node for the average feature of the current segment. The formula is:
[0050]
[0051] where φ init,h (·) is a three-layer perceptron with hidden states;
[0052] Step 3-6 performs intra-segment LSTM time adaptive node update on the i-th video frame of the current segment, i.e., the i-th time node in the corresponding LSTM;
[0053] Step 3-6-1 Based on the memory state c of the previous node v,s,i-1 , hidden state h v,s,i-1 and the input x of the current node v,s,i , use the LSTM module to estimate the candidate state of the i-th node, the estimation method is:
[0054]
[0055] Among them, the first row is used to estimate the forget gate, f v,s,i is the output of the forget gate, W fx , W fh and b f represents the LSTM parameters of the forget gate; the second line is used to estimate the input gate, in v,s,i is the output of the input gate, W inx , W inh and b in represents the LSTM parameters of the input gate; the third row is used to estimate the selection gate, g v,s,i is the output of the select gate, W gx , W gh and b g represents the LSTM parameters for selecting the gate; the fourth line is used to estimate the output gate, o v,s,i is the output of the output gate, W ox , W oh and b o Represents the LSTM parameters of the output gate; the fifth line is used to estimate the candidate memory state of the next node The sixth line is used to estimate the candidate hidden state of the next node σ(·) and tanh(·) denote sigmoid and tanh activation functions, ⊙ denotes the Hadamard product;
[0056] Step 3-6-2 According to step 3-3, obtain the i-th video frame, that is, the binary sampling mark sm of the i-th time node v,s,i-1,i ;
[0057] Step 3-6-3 Mark sm according to the binary sampling v,s,i-1,i , estimate the memory state c of the i-th node v,s,i and hidden state h v,s,i , the specific method is:
[0058]
[0059] If the binary sample marker sm v,s,i-1,i If sm is 1, it means sampling the node, then the candidate memory state of the i-th node updated by LSTM is used as the memory state of the i-th node, and the candidate hidden state of the i-th node updated by LSTM is used as the hidden state of the i-th node; if the binary sampling marker sm v,s,i-1,i If it is 0, it means that the node is not sampled, then the memory state of the LSTM i-1th node is used as the memory state of the i-th node, and the hidden state of the LSTM i-1th node is used as the hidden state of the i-th node;
[0060] Steps 3-7 update all nodes in the segment and calculate the hidden state features of the last time node of the LSTM in the segment;
[0061] Step 3-7-1 If the current node is not n i , it means there is another time node in the segment, execute step 3-3-5, process the next time node in the segment, and find the memory state c of the next time node v,s,i and hidden state h v,s,i ;
[0062] Step 3-7-2 repeats step 3-6 until the current node is the last time node in the segment;
[0063] Step 3-7-3 If the current node is n i , it means that there is no next time node in this segment, and the hidden state feature h of the last time node in LSTM is returned. v,s,ni .
[0064] The calculation described in step (4) is based on the stealing behavior score of the sub-behavior prototype time attention, and the specific steps are as follows:
[0065] Step 4-1 Execute step 3 to obtain the output hidden state h of the last time node of each segment of the vth video v,s,ni , where s = 1, 2, .., n s ,nth i The time node is the last time node in this segment;
[0066] Step 4-2 Set the sub-behavior prototype parameter matrix, Z = {z k}, where k = 1, 2, .., K, a total of K prototypes, a prototype is a description of a sub-behavior of a specific category, and the vector feature dimension of each prototype is 1x1000;
[0067] Step 4-3 calculates the temporal attention of each video segment;
[0068] Step 4-3-1 calculates the hidden state of the output of the last time node of the sth video segment and the prototype distance of the kth sub-behavior. The specific form is:
[0069] dist v,s,k =||h v,s,ni -z k ||2
[0070] Step 4-3-2 repeats step 4-3-1 to calculate the distance vector between the hidden state of the sth video segment and all sub-behavior prototypes. The specific form is:
[0071] Dist v,s ={dist v,s,k}
[0072] Step 4-3-3 Input the distance vector between the hidden state of the sth video segment and all sub-behavior prototypes, use a three-layer perceptron and a sigmoid function to learn the temporal attention of the sth video segment. The specific form is:
[0073] α v,s =sigmoid(MLP(Dist v,s ))
[0074] Where MLP is a three-layer perceptron, sigmoid(.) is the sigmoid activation function;
[0075] Step 4-4 calculates the hidden state features of each video segment after the temporal attention enhancement, that is, according to the temporal attention of the sth video segment, the hidden state h v,s,ni After feature enhancement, the hidden state features after enhancement are:
[0076]
[0077] Step 4-5 concatenates the enhanced hidden state features of all video segments to obtain the behavior expression of the entire video, and uses a three-layer perceptron to predict the score of the theft behavior. The specific form is:
[0078]
[0079] The specific steps of solving the parameter aggregation of the adaptive time structure deep network described in step (5) are as follows:
[0080] Step 5-1: Execute step 1 to segment all videos in the theft behavior training set, perform video frame sampling and feature extraction, and obtain the video feature set X. v,s ={x v,s,i};
[0081] Step 5-2 executes step 3 to obtain the output hidden state h of the last time node of each segment of each video v,s,ni , where s = 1, 2, .., n s ,nth i The time node is the last time node in this segment;
[0082] Step 5-3: Execute step 4 to obtain the enhanced hidden state features And get the score of predicted stealing behavior py v ;
[0083] Step 5-4 calculates the loss function of the adaptive time structure deep network;
[0084] Step 5-4-1 Based on the actual label of the theft behavior training set and the predicted theft behavior marker PY = {py v}, calculate the prediction score loss function, the specific form is:
[0085]
[0086] Among them, the true label of the theft behavior training set It means stealing. It means it is not stealing;
[0087] Step 5-4-2 calculates the loss function of the sub-behavior prototype according to the sub-behavior prototype parameter matrix Z. The specific form is:
[0088] Loss z =||Z·Z T -I||2
[0089] Where Z is the sub-behavior prototype parameter matrix, the matrix size of Z is K×1000, I is the unit matrix, the matrix size of I is KxK, K is the number of prototypes, the diagonal elements in I are 1, and the remaining elements are 0. The loss function estimates that the vector product between each sub-behavior prototype is close to 0, that is, the features of each sub-behavior prototype are close to orthogonal to each other, and at this time, there are large differences between the features of each sub-behavior;
[0090] The total loss function of the model training in step 5-4-3 is the sum of the prediction score loss function and the sub-behavior prototype loss function, that is:
[0091] Loss=Loss y +Loss z
[0092] Step 5-5 uses the total loss function calculated in step 5-4-3 as the gradient, and uses the back propagation algorithm to solve the parameters of the model in turn. During the back propagation process, the parameters that need to be solved include: the parameters in the three-layer perceptron in step 4-5, the sub-behavior prototype parameter matrix in step 4-2, the LSTM model parameters in step 3-6, the three-layer perceptron parameters of the hidden state in step 3-5-3, and the three-layer perceptron parameters of the memory state in step 3-5-2; when the training is completed, the optimal parameter set of the model is obtained.
[0093] The specific steps of the theft behavior recognition based on the adaptive time structure deep network in step (6) are as follows:
[0094] Step 6-1: For the test video, execute step 1 to perform video segmentation, video frame sampling and feature extraction on all videos in the theft behavior training set to obtain the video feature set X. test,s ={x test,s,i};
[0095] Step 6-2 For the test video, use the optimal parameter set of the model learned in step 5.5 and execute step 3 to obtain the output hidden state h of the last time node of each segment of each video. test,s,ni , where s = 1, 2, .., n s ,nth i The time node is the last time node in this segment;
[0096] Step 6-3 executes step 4, and uses the best parameter set of the model learned in step 5-5 to obtain the enhanced hidden state features And predict the theft behavior marker py test , if the predicted tag py test >0.5 indicates that the video is stolen. If the predicted label is py test <0.5 means it is not a stolen video.
[0097] The advantages of the present invention are as follows: the present invention realizes the identification of theft behavior by acquiring the surveillance video of the parking space; the present invention adopts a segmentation method to realize the processing of long-term video; in view of the information redundancy problem caused by the high similarity of features within the segment, by estimating the distance threshold of the temporal structure within the segment, the binary sampling mark of each video frame in the segment is calculated, the redundant video frames in the segment are removed, and the calculation amount of the model within the segment is reduced; in view of the problem of estimating the importance of semantic information in the video segment, a sub-behavior prototype parameter matrix is designed, the semantic description related to the sub-behavior of each video segment is learned, and the temporal attention of the video segment is estimated to enhance the features of the video segment; finally, the features of multiple video segments are connected in series, and a three-layer perceptron is used to realize the identification of theft behavior; the present invention has strong temporal self-adaptation ability, and has good robust processing ability for redundant video frames within the segment of long-term video and semantic information analysis between segments, and can effectively realize the identification of theft behavior. BRIEF DESCRIPTION OF THE DRAWINGS
[0098] Figure 1 Flowchart of theft behavior recognition using adaptive temporal structure deep network.
[0099] Figure 2 Schematic diagram of the theft behavior recognition process based on adaptive temporal structure deep network.
[0100] Figure 3 Compute the hidden state feature module of the LSTM sampling of the video frame within the segment.
[0101] Figure 4 Module for calculating stealing behavior scores based on temporal attention of sub-behavior prototypes.
[0102] Figure 5 A module for solving parameter sets of deep networks with adaptive structures.
[0103] Figure 6 Theft behavior recognition module based on adaptive temporal structure deep network. DETAILED DESCRIPTION
[0104] like Figure 1-6 As shown in FIG. 1 , a theft behavior recognition method based on an adaptive time structure deep network is shown in FIG. 1 . The specific steps are as follows:
[0105] Step (1) Extract the video frame features of the theft behavior:
[0106] Step 1-1: Take 8 video segments of the theft behavior video at equal time intervals. s is the subscript of the video segment, and the value of s is 1, 2, ..., n. s , n s is the number of samples in the middle of the video.
[0107] Step 1-2: In each video segment, continue to take n i sub-intervals, use random time subscripts to collect a video frame in each interval, and collect n video frames in total in this video segment i Video frames, n i is the number of samples in the video frame within the segment, n i The value is 8.
[0108] Steps 1-3 use the pre-trained ResNet50 model to extract the features of each video frame and convert the features into row vectors to obtain the feature set of the video frame, X v,s ={x v,s,i}, where X v,s is a feature set, v represents the video number, s represents the video segment number, i represents the video frame number within the segment, and x v,s,i is a row vector with a feature dimension of 1x1000.
[0109] Step (2) estimates the intra-segment temporal structure distance threshold:
[0110] Step 2-1 Given the features X of all video segments in the training set v,s ={x v,s,i};
[0111] Step 2-2 Calculate the distance d between the features of two video frames within a segment v,s,i,i+1 , the distance calculation formula is two features x v,s,i and x v,s,i+1 The 2-norm between them, the distance calculation formula is expressed as:
[0112] d v,s,i,i+1 =||x v,s,i -x v,s,i+1 ||2
[0113] Step 2-3: For all segments of all videos in the theft behavior training set, calculate the distance between the features of two video frames in the segment to obtain the distance set, D = {d v,s,i,i+1}, v is the number of all videos in the training set, s is the video segment number, and i is the video frame number within the segment.
[0114] Step 2-4 finds the mean u of the distance set d , the mean formula of the distance set is expressed as:
[0115]
[0116] where n v is the number of videos, n s is the number of segments in each video.
[0117] Step 2-5 calculates the variance of the distance set. The variance formula of the distance set is expressed as:
[0118]
[0119] The time structure distance threshold estimation formula in steps 2-6 is expressed as:
[0120] τ d =u d +σ d
[0121] Step (3) calculates the hidden state features of the LSTM sampled within the video frame:
[0122] Step 3-1 Given a feature X of a video segment v,s ={x v,s,i};
[0123] Step 3-2 calculates the binary sampling mark of each video frame in the segment;
[0124] Step 3-2-1 Calculate the distance d between two video frames in the segment according to step 2-2 v,s,i,i+1 ;
[0125] Step 3-2-2 Set the feature distance threshold τ d , the formula for calculating the sampling probability of the next video frame is expressed as:
[0126]
[0127] This formula can be interpreted as when the distance d between two video frames v,s,i,i+1 =τ d The sampling probability is 0.5, and as the distance increases, there is a greater probability of being sampled; as the distance decreases, there is a smaller probability of being sampled.
[0128] Step 3-2-3 performs binary 0-1 Bernoulli sampling based on the sampling probability of the next video frame calculated in step 3-2-2, sm v,s,1,2 Get the binary sampling probability of the next video frame; that is, use the uniform distribution z~U[0,1] in the interval [0,1] to generate a random number. If the random number is greater than the intra-segment sampling weight, it returns 0, indicating that the next video frame is sampled. If the random number is less than the intra-segment sampling threshold, it returns 1, indicating that the next video frame is not sampled. The intra-segment video frame sampling formula is expressed as:
[0129]
[0130] The purpose of introducing the random number z is to avoid deterministic sampling based on the sampling weight, which will lead to over-trust in the feature weight. By converting the random number z into random sampling, the randomness of the sample sampling can be guaranteed, allowing the model to sample samples with small feature weights. At this time, the binary sampling mark sm of each video frame in the segment is obtained. v,s,i,i+1 , the binary sampling mark corresponds to the LSTM time node, where sm v,s,1,2 The binary sampling label can be interpreted as estimating the binary sampling label of the second video frame according to the feature distance between the first video frame and the second video frame.
[0131] Step 3-3 initializes the time node of LSTM according to the subscript of the video frame in the segment, where the i-th video frame corresponds to the i-th time node, and the binary label of the i-th video frame is represented by sm v,s,i-1,i ;
[0132] Step 3-4 adds an initial node to the LSTM model and sets the binary sampling mark of the initial node and the first video frame to 1, that is, sm v,s,0,1 = 1, obtain the intra-segment binary sampling tag set SM of the LSTM with the initial node added = {sm v,s,0,1 ,sm v,s,1,2 ,...,sm v,s,ni-1,ni};
[0133] Step 3-5 initializes the memory state and hidden state of the initial time node;
[0134] Step 3-5-1 Calculate the average feature of the current segment The formula is expressed as:
[0135]
[0136] Step 3-5-2 uses a three-layer perceptron to learn the memory state of each LSTM initial node for the average feature of the current segment. The formula is:
[0137]
[0138] where φ init,c (·) is a three-layer perceptron for memory state.
[0139] Step 3-5-3 uses a three-layer perceptron to learn the hidden state of each LSTM initial node for the average feature of the current segment. The formula is:
[0140]
[0141] where φ init,h (·) is a three-layer perceptron with hidden states.
[0142] Step 3-6 performs intra-segment LSTM time adaptive node update on the i-th video frame of the current segment, i.e., the i-th time node in the corresponding LSTM;
[0143] Step 3-6-1 Based on the memory state c of the previous node v,s,i-1 and hidden state h v,s,i-1 , and the input x of the current node v,s,i , use the LSTM module to estimate the candidate state of the i-th node, the estimation method is:
[0144]
[0145] Among them, the first row is used to estimate the forget gate, f v,s,i is the output of the forget gate, W fx , W fh and b f represents the LSTM parameters of the forget gate; the second line is used to estimate the input gate, in v,s,i is the output of the input gate, W inx , W inh and b in represents the LSTM parameters of the input gate; the third row is used to estimate the selection gate, g v,s,i is the output of the select gate, W gx , W gh and b g represents the LSTM parameters for selecting the gate; the fourth line is used to estimate the output gate, o v,s,i is the output of the output gate, W ox , W oh and b o Represents the LSTM parameters of the output gate; the fifth line is used to estimate the candidate memory state of the next node The sixth line is used to estimate the candidate hidden state of the next node σ(·) and tanh(·) denote sigmoid and tanh activation functions, and ⊙ denotes the Hadamard product.
[0146] Step 3-6-2 According to step 3-3, obtain the i-th video frame, that is, the binary sampling mark sm of the i-th time node v,s,i-1,i ;
[0147] Step 3-6-3 Mark sm according to the binary sampling v,s,i-1,i , estimate the memory state c of the i-th node v,s,i and hidden state h v,s,i , the specific method is:
[0148]
[0149] The updating process can be explained as follows: if the binary sample mark sm v,s,i-1,i If sm is 1, it means sampling the node, then the candidate memory state of the i-th node updated by LSTM is used as the memory state of the i-th node, and the candidate hidden state of the i-th node updated by LSTM is used as the hidden state of the i-th node; if the binary sampling marker sm v,s,i-1,i If s is 0, it means that the node is not sampled. Then the memory state of the LSTM i-1th node is used as the memory state of the i-th node, and the hidden state of the LSTM i-1th node is used as the hidden state of the i-th node. Therefore, the binary sampling mark sm v,s,i-1,i It realizes time sampling and time adaptation.
[0150] Steps 3-7 update all nodes in the segment and calculate the hidden state features of the last time node of the LSTM in the segment;
[0151] Step 3-7-1 If the current node is not n i , it means there is another time node in the segment, execute step 3-3-5, process the next time node in the segment, and find the memory state c of the next time node v,s,i and hidden state h v,s,i ;
[0152] Step 3-7-2 repeats step 3-6 until the current node is the last time node in the segment;
[0153] Step 3-7-3 If the current node is n i , it means that there is no next time node in this segment, and the hidden state feature h of the last time node in LSTM is returned. v,s,ni ;
[0154] Step (4) calculates the stealing behavior score based on the sub-behavior prototype time attention:
[0155] Step 4-1 Execute step 3 to obtain the output hidden state h of the last time node of each segment of the vth video v,s,ni, where s = 1, 2, .., n s ,nth i The time node is the last time node in this segment.
[0156] Step 4-2 Set the sub-behavior prototype parameter matrix, Z = {z k}, where k = 1, 2, .., K, a total of K prototypes, the prototype is a description of the sub-behavior of a specific category, and the vector feature dimension of each prototype is 1x1000.
[0157] Step 4-3 calculates the temporal attention of each video segment;
[0158] Step 4-3-1 calculates the hidden state of the output of the last time node of the sth video segment and the prototype distance of the kth sub-behavior. The specific form is:
[0159] dist v,s,k =||h v,s,ni -z k ||2
[0160] Step 4-3-2 repeats step 4-3-1 to calculate the distance vector between the hidden state of the sth video segment and all sub-behavior prototypes. The specific form is:
[0161] Dist v,s ={dist v,s,k}
[0162] Step 4-3-3 Input the distance vector between the hidden state of the sth video segment and all sub-behavior prototypes, use a three-layer perceptron and a sigmoid function to learn the temporal attention of the sth video segment. The specific form is:
[0163] α v,s =sigmoid(MLP(Dist v,s ))
[0164] Where MLP is a three-layer perceptron and sigmoid(.) is the sigmoid activation function.
[0165] Step 4-4 calculates the hidden state features of each video segment after the temporal attention enhancement, that is, according to the temporal attention of the sth video segment, the hidden state h v,s,ni After feature enhancement, the hidden state features after enhancement are:
[0166]
[0167] Step 4-5 concatenates the enhanced hidden state features of all video segments to obtain the behavior expression of the entire video, and uses a three-layer perceptron to predict the score of the theft behavior. The specific form is:
[0168]
[0169] Step (5) solves the parameter set of the adaptive time structure deep network:
[0170] Step 5-1: Execute step 1 to segment all videos in the theft behavior training set, perform video frame sampling, and feature extraction to obtain the video feature set X. v,s ={x v,s,i};
[0171] Step 5-2 executes step 3 to obtain the output hidden state h of the last time node of each segment of each video v,s,ni , where s = 1, 2, .., n s ,nth i The time node is the last time node in this segment;
[0172] Step 5-3: Execute step 4 to obtain the enhanced hidden state features And get the predicted theft behavior score py v .
[0173] Step 5-4 calculates the loss function of the adaptive time structure deep network;
[0174] Step 5-4-1 Based on the actual label of the theft behavior training set and the predicted theft behavior marker PY = {py v}, calculate the prediction score loss function, the specific form is:
[0175]
[0176] Among them, the true label of the theft behavior training set It means stealing. It means it is not stealing.
[0177] Step 5-4-2 calculates the loss function of the sub-behavior prototype according to the sub-behavior prototype parameter matrix Z. The specific form is:
[0178] Loss z =||Z·Z T -I||2
[0179] Where Z is the sub-behavior prototype parameter matrix, the matrix size of Z is K×1000, I is the unit matrix, the matrix size of I is KxK, K is the number of prototypes, the diagonal elements in I are 1, and the remaining elements are 0. The loss function estimates that the vector product between each sub-behavior prototype is close to 0, that is, the features of each sub-behavior prototype are close to orthogonal to each other, and at this time there are large differences between the features of each sub-behavior.
[0180] The total loss function of the model training in step 5-4-3 is the sum of the prediction score loss function and the sub-behavior prototype loss function, that is:
[0181] Loss=Loss y +Loss z
[0182] Step 5-5 uses the total loss function calculated in step 5-4-3 as the gradient, and uses the back propagation algorithm to solve the parameters of the model in turn. During the back propagation process, the parameters that need to be solved include: the parameters in the three-layer perceptron in step 4-5, the sub-behavior prototype parameter matrix in step 4-2, the LSTM model parameters in step 3-6, the three-layer perceptron parameters of the hidden state in step 3-5-3, and the three-layer perceptron parameters of the memory state in step 3-5-2. When the training is completed, the optimal parameter set of the model is obtained.
[0183] Step (6) Prediction of theft behavior based on adaptive time structure deep network:
[0184] Step 6-1: For the test video, execute step 1 to segment all the videos in the theft behavior training set, perform video frame sampling, and feature extraction to obtain the video feature set X. test,s ={x test,s,i};
[0185] Step 6-2 For the test video, use the optimal parameter set of the model learned in step 5.5 and execute step 3 to obtain the output hidden state h of the last time node of each segment of each video. test,s,ni , where s = 1, 2, .., n s ,nth i The time node is the last time node in this segment;
[0186] Step 6-3 executes step 4, and uses the best parameter set of the model learned in step 5-5 to obtain the enhanced hidden state features And predict the theft behavior marker py test , if the predicted tag py test >0.5 indicates that the video is stolen. If the predicted label is py test <0.5 means it is not a stolen video.
Claims
1. A theft behavior recognition method based on adaptive time structure deep network, characterized by: The specific steps include: (1) Extract the video frame features of theft behavior; (2) estimate the intra-segment temporal structure distance threshold; (3) Calculate the hidden state features of the LSTM sampled from the video frame within the segment; (4) Calculate the stealing behavior score based on the sub-behavior prototype time attention; (5) Solve the parameter set of the adaptive time structure deep network; (6) Identify theft behavior based on adaptive time structure deep network; The calculation described in step (4) is based on the stealing behavior score of the sub-behavior prototype time attention, and the specific steps are as follows: Step 4-1 Execute step 3 to obtain the output hidden state h of the last time node of each segment of the vth video v,s,ni , where s = 1, 2, .., n s ,nth i The time node is the last time node in this segment; Step 4-2 Set the sub-behavior prototype parameter matrix, Z = {z k }, where k = 1, 2, .., K, a total of K prototypes, a prototype is a description of a sub-behavior of a specific category, and the vector feature dimension of each prototype is 1x1000; Step 4-3 calculates the temporal attention of each video segment; Step 4-3-1 calculates the hidden state of the output of the last time node of the sth video segment and the prototype distance of the kth sub-behavior. The specific form is: dist v,s,k =||h v,s,ni -z k ||2 Step 4-3-2 repeats step 4-3-1 to calculate the distance vector between the hidden state of the sth video segment and all sub-behavior prototypes. The specific form is: Dist v,s ={dist v,s,k } Step 4-3-3 Input the distance vector between the hidden state of the sth video segment and all sub-behavior prototypes, use a three-layer perceptron and a sigmoid function to learn the temporal attention of the sth video segment. The specific form is: α v,s =sigmoid(MLP(Dist v,s )) Where MLP is a three-layer perceptron, sigmoid(.) is the sigmoid activation function; Step 4-4 calculates the hidden state features of each video segment after the temporal attention enhancement, that is, according to the temporal attention of the sth video segment, the hidden state h v,s,ni After feature enhancement, the hidden state features after enhancement are: Step 4-5 concatenates the enhanced hidden state features of all video segments to obtain the behavior expression of the entire video, and uses a three-layer perceptron to predict the score of the theft behavior. The specific form is:
2. The theft behavior recognition method based on adaptive time structure deep network according to claim 1 is characterized by: The specific steps of extracting the theft behavior video frame features in step (1) are as follows: Step 1-1 Take n samples of the theft video at equal time intervals. s video segments; Step 1-2: In each video segment, continue to take n i subintervals, use random time subscripts to collect a video frame in each subinterval, and collect n video frames in total in this video segment. i Video frames, n i is the number of samples of the video frame in the segment; Steps 1-3 use the pre-trained ResNet50 model to extract the features of each video frame and convert the features into row vectors to obtain the feature set of the video frame, X v,s ={x v,s,i }, where X v,s is a feature set, v represents the video number, s represents the video segment number, i represents the video frame number within the segment, and x v,s,i is a row vector with a feature dimension of 1x1000.
3. The theft behavior recognition method based on adaptive time structure deep network according to claim 2 is characterized by: The estimation of the intra-segment temporal structure distance threshold in step (2) is specifically performed as follows: Step 2-1 Given the features X of all video segments in the training set v,s ={x v,s,i }; Step 2-2 Calculate the distance d between the features of two video frames within a segment v,s,i,i+1 , the distance calculation formula is two features x v,s,i and x v,s,i+1 The 2-norm between them, the distance calculation formula is expressed as: d v,s,i,i+1 =||x v,s,i -x v,s,i+1 ||2 Step 2-3: For all segments of all videos in the theft behavior training set, calculate the distance between the features of two video frames in the segment to obtain the distance set, D = {d v,s,i,i+1 }, v is the number of all videos in the training set, s is the video segment number, and i is the video frame number within the segment; Step 2-4 finds the mean u of the distance set d , the mean formula of the distance set is expressed as: Where n v is the number of videos, n s is the number of segments within each video; Step 2-5 calculates the variance of the distance set. The variance formula of the distance set is expressed as: The time structure distance threshold estimation formula in steps 2-6 is expressed as: t d =u d +s d 。 4. The theft behavior recognition method based on adaptive time structure deep network according to claim 3 is characterized by: The hidden state features of the LSTM sampling of the video frame in the calculation segment described in step (3) are specifically performed as follows: Step 3-1 Given a feature X of a video segment v,s ={x v,s,i }; Step 3-2 calculates the binary sampling mark of each video frame in the segment; Step 3-2-1 Calculate the distance d between two video frames in the segment according to step 2-2 v,s,i,i+1 ; Step 3-2-2 Set the feature distance threshold τ d , the formula for calculating the sampling probability of the next video frame is expressed as: When the distance d between two video frames v,s,i,i+1 =τ d Then the sampling probability is 0.5, and as the distance increases, there is a greater probability of being sampled; as the distance decreases, there is a smaller probability of being sampled; Step 3-2-3 performs binary 0-1 Bernoulli sampling according to the sampling probability of the next video frame calculated in step 3-2-2 to obtain the binary sampling probability of the next video frame; that is, a random number is generated using the uniform distribution z~U[0,1] in the interval [0,1]. If the random number is greater than the intra-segment sampling weight, 0 is returned, indicating that the next video frame is sampled. If the random number is less than the intra-segment sampling threshold, 1 is returned, indicating that the next video frame is not sampled. The intra-segment video frame sampling formula is expressed as: z is a random number, which is converted into random sampling to ensure randomness during sample sampling, allowing the model to sample samples with small feature weights; at this time, the binary sampling mark sm of each video frame in the segment is obtained v,s,i,i+1 , the binary sampling mark corresponds to the LSTM time node, where sm v,s,1,2 The binary sampling mark is used to estimate the binary sampling mark of the second video frame according to the feature distance between the first video frame and the second video frame; Step 3-3 initializes the time node of LSTM according to the subscript of the video frame in the segment, where the i-th video frame corresponds to the i-th time node, and the binary mark of the i-th video frame is represented by sm v,s,i-1,i ; Step 3-4 adds an initial node to the LSTM model and sets the binary sampling mark of the initial node and the first video frame to 1, that is, sm v,s,0,1 = 1, obtain the intra-segment binary sampling tag set SM of the LSTM with the initial node added = {sm v,s,0,1 ,sm v,s,1,2 ,...,sm v,s,ni-1,ni }; Step 3-5 initializes the memory state and hidden state of the initial time node; Step 3-5-1 Calculate the average feature of the current segment The formula is expressed as: Step 3-5-2 uses a three-layer perceptron to learn the memory state of each LSTM initial node for the average feature of the current segment. The formula is: where φ init,c (·) is a three-layer perceptron for memory state; Step 3-5-3 uses a three-layer perceptron to learn the hidden state of each LSTM initial node for the average feature of the current segment. The formula is: where φ init,h (·) is a three-layer perceptron with hidden states; Step 3-6 performs intra-segment LSTM time adaptive node update on the i-th video frame of the current segment, i.e., the i-th time node in the corresponding LSTM; Step 3-6-1 Based on the memory state c of the previous node v,s,i-1 , hidden state h v,s,i-1 and the input x of the current node v,s,i , use the LSTM module to estimate the candidate state of the i-th node, the estimation method is: Among them, the first row is used to estimate the forget gate, f v,s,i is the output of the forget gate, W fx , W fh and b f represents the LSTM parameters of the forget gate; the second line is used to estimate the input gate, in v,s,i is the output of the input gate, W inx , W inh and b in represents the LSTM parameters of the input gate; the third row is used to estimate the selection gate, g v,s,i is the output of the select gate, W gx , W gh and b g represents the LSTM parameters for selecting the gate; the fourth line is used to estimate the output gate, o v,s,i is the output of the output gate, W ox , W oh and b o Represents the LSTM parameters of the output gate; the fifth line is used to estimate the candidate memory state of the next node The sixth line is used to estimate the candidate hidden state of the next node σ(·) and tanh(·) represent sigmoid and tanh activation functions, and e represents the Hadamard product; Step 3-6-2 According to step 3-3, obtain the i-th video frame, that is, the binary sampling mark sm of the i-th time node v,s,i-1,i ; Step 3-6-3 Mark sm according to the binary sampling v,s,i-1,i , estimate the memory state c of the i-th node v,s,i and hidden state h v,s,i , the specific method is: If the binary sample marker sm v,s,i-1,i If sm is 1, it means sampling the node, then the candidate memory state of the i-th node updated by LSTM is used as the memory state of the i-th node, and the candidate hidden state of the i-th node updated by LSTM is used as the hidden state of the i-th node; if the binary sampling marker sm v,s,i-1,i If it is 0, it means that the node is not sampled, then the memory state of the LSTM i-1th node is used as the memory state of the i-th node, and the hidden state of the LSTM i-1th node is used as the hidden state of the i-th node; Steps 3-7 update all nodes in the segment and calculate the hidden state features of the last time node of the LSTM in the segment; Step 3-7-1 If the current node is not n i , it means there is another time node in the segment, execute step 3-3-5, process the next time node in the segment, and find the memory state c of the next time node v,s,i and hidden state h v,s,i ; Step 3-7-2 repeats step 3-6 until the current node is the last time node in the segment; Step 3-7-3 If the current node is n i , it means that there is no next time node in this segment, and the hidden state feature h of the last time node in LSTM is returned. v,s,ni .
5. The theft behavior recognition method based on adaptive time structure deep network according to claim 1 is characterized by: The specific steps of solving the parameter aggregation of the adaptive time structure deep network described in step (5) are as follows: Step 5-1: Execute step 1 to segment all videos in the theft behavior training set, perform video frame sampling and feature extraction, and obtain the video feature set X. v,s ={x v,s,i }; Step 5-2 executes step 3 to obtain the output hidden state h of the last time node of each segment of each video v,s,ni , where s = 1, 2, .., n s ,nth i The time node is the last time node in this segment; Step 5-3: Execute step 4 to obtain the enhanced hidden state features And get the score of predicted stealing behavior py v ; Step 5-4 calculates the loss function of the adaptive time structure deep network; Step 5-4-1 Based on the actual label of the theft behavior training set and the predicted theft behavior marker PY = {py v }, calculate the prediction score loss function, the specific form is: Among them, the true label of the theft behavior training set It means stealing. It means it is not stealing; Step 5-4-2 calculates the loss function of the sub-behavior prototype according to the sub-behavior prototype parameter matrix Z. The specific form is: Loss z =||Z·Z T -I||2 Where Z is the sub-behavior prototype parameter matrix, the matrix size of Z is K×1000, I is the unit matrix, the matrix size of I is KxK, K is the number of prototypes, the diagonal elements in I are 1, and the remaining elements are 0. The loss function estimates that the vector product between each sub-behavior prototype is close to 0, that is, the features of each sub-behavior prototype are close to orthogonal to each other, and at this time, there are large differences between the features of each sub-behavior; The total loss function of the model training in step 5-4-3 is the sum of the prediction score loss function and the sub-behavior prototype loss function, that is: Loss=Loss y +Loss z Step 5-5 uses the total loss function calculated in step 5-4-3 as the gradient, and uses the back propagation algorithm to solve the parameters of the model in turn. During the back propagation process, the parameters that need to be solved include: the parameters in the three-layer perceptron in step 4-5, the sub-behavior prototype parameter matrix in step 4-2, the LSTM model parameters in step 3-6, the three-layer perceptron parameters of the hidden state in step 3-5-3, and the three-layer perceptron parameters of the memory state in step 3-5-2; when the training is completed, the optimal parameter set of the model is obtained.
6. The theft behavior recognition method based on adaptive time structure deep network according to claim 5 is characterized by: The specific steps of the theft behavior recognition based on the adaptive time structure deep network in step (6) are as follows: Step 6-1: For the test video, execute step 1 to perform video segmentation, video frame sampling and feature extraction on all videos in the theft behavior training set to obtain the video feature set X. test,s ={x test,s,i }; Step 6-2 For the test video, use the optimal parameter set of the model learned in step 5.5 and execute step 3 to obtain the output hidden state h of the last time node of each segment of each video. test,s,ni , where s = 1, 2, .., n s ,nth i The time node is the last time node in this segment; Step 6-3 executes step 4, and uses the best parameter set of the model learned in step 5-5 to obtain the enhanced hidden state features And predict the theft behavior marker py test , if the predicted tag py test >0.5 indicates that the video is stolen. If the predicted label is py test <0.5 means it is not a stolen video.
Citation Information
Patent Citations
A Reinforcement Learning-Based Behavior Recognition Video Frame Sampling Method and System
CN109800717B
A Human Behavior Recognition Method Based on Deep Neural Networks
CN109919031B
Video classification based on hybrid convolution and attention mechanism
CN109389055A
Vehicle stealing behavior recognition method based on monitoring video
CN111444861A