Timing Boundary Detection Methods and Timing Perceptrons
By building a detection network containing backbone network and detection model, using attention mechanism and video semantic structure, the dispersion and non-generalization problems of classless timing boundary detection in the existing technology are solved, efficient and unified timing boundary detection is achieved, and prediction accuracy and inference speed are improved.
Patent Information
- Application Number
- CN202111615241.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-27
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-12-27
AI Technical Summary
The existing classless timing boundary detection paradigm has the difference in boundary semantics and granularity, and the research is scattered in different tasks and is difficult to generalize to different types of classless boundary detection, which lacks generalization.
A general classless timing boundary detection method is proposed. By building a detection network containing backbone network and detection model, using attention mechanism and video semantic structure, compressing redundant video input into stable and reliable feature representations, achieving efficient timing boundary detection.
It realizes efficient and unified detection of different types of classless timing boundaries, reduces model complexity, improves prediction accuracy and inference speed, and has good generalization performance.
Smart Images

Figure CN114494314B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of computer software, relates to video timing boundary detection, and is a timing boundary detection method and a timing sensor. Background Art
[0002] Due to the explosive growth of video data on the Internet, video content understanding has become an important issue in the field of computer vision. In previous literature, the exploration of long video understanding is still insufficient. Category-free temporal boundary detection is an effective technology to bridge the gap between long and short video understanding. Its purpose is to segment long videos into a series of video clips. Category-free temporal boundaries are temporal boundaries that arise naturally due to semantic discontinuity. They are not constrained by any predefined semantic categories. Existing datasets include category-free temporal boundaries of different granularities, such as sub-action level, event level, and scene level. For the detection of category-free temporal boundaries of different granularities, different levels of information are needed to obtain temporal structures and contextual relationships at different scales.
[0003] Currently, the research on classless temporal boundary detection is divided into several different tasks due to the differences in temporal boundary semantics and granularity. The goal of the temporal action segmentation task is to detect sub-action-level classless temporal boundaries that segment an action instance into multiple different sub-action segments. General temporal boundary detection aims to locate event-level classless temporal boundaries, i.e., the moments when the action / theme / environment changes. Movie scene segmentation detects scene-level classless temporal boundaries, i.e., the transitions between movie scenes, marking the turning points of the high-level plot. The target videos of these tasks have the same semantic structure, and their boundary detection paradigms show similar characteristics. Previous works on these tasks mainly focus on the feature encoding carefully designed for specific boundaries and reduce the boundary detection problem to a dense prediction problem. During the prediction process, these works use complex post-processing techniques to eliminate the large number of false positives in the results that repeatedly predict the same truth value. Such complex designs and post-processing modules are highly dependent on specific boundary types, so they cannot be well generalized to different types of classless boundary detection and lack generalization. Summary of the invention
[0004] The problem to be solved by the present invention is that the existing paradigms of unbounded temporal boundary detection have similar properties, but are scattered in different tasks due to differences in boundary semantics and granularity. Existing related work mainly focuses on feature encoding carefully designed for specific boundaries, and because the dense prediction paradigm uses complex post-processing techniques to eliminate false positives, it cannot be well generalized to different types of unbounded boundary detection.
[0005] The technical solution of the present invention is: a temporal boundary detection method, which constructs a classless temporal boundary detection network to perform temporal boundary detection on a video. The detection network includes a backbone network and a detection model. The implementation method is as follows:
[0006] 1) Generate detection samples by backbone network: sample the video interval to obtain video image sequence A video segment is generated for each frame, and the i-th video segment is composed of the i-th frame image f i The backbone network generates video features for the input video segment. and continuity scoring F i and S i Score the RGB features and continuity of video segment i respectively;
[0007] 2) The detection model performs non-classified temporal action detection based on the video feature F and the continuity score S, and the detection model includes the following configuration:
[0008] 2.1) Encoder: Encoder: Encoder E includes N e The transform decoding layer is a series of layers, each of which contains a multi-head self-attention layer, a multi-head cross-attention layer and a linear mapping layer. The self-attention layer, the cross-attention layer and the linear mapping layer each have a residual structure, and the encoder introduces M hidden feature query quantities Q e , based on the continuity score S, the video features F are sorted in descending order and then input into the encoder. The encoder compresses the sorted video features into compressed features H of M frames. The initial compressed features H 0 is 0, at the jth transform decoding layer, the implicit feature query quantity Q e And the compression feature H of the current layer j After addition, the self-attention layer and its residual structure interact with the reordered video features in the cross-attention layer, and then the residual structure-linear mapping layer-residual structure transformation is used to obtain the compressed feature H j+1 , j∈[1,(N e -1)], through the stacked N e After the encoding layer, the input features are compressed and encoded to obtain compressed features.
[0009] Among them, the generation of latent feature query volume is: latent feature query volume Q e Divided into M b Boundary query volume and M c context queries, randomly initialized, are generated with training sample learning during the training detection model training process; boundary queries correspond to processing boundary region features in video features, context queries correspond to processing context region features in video features, and the top M in video features are reordered.b Features are boundary area features, and the others are context features;
[0010] 2.2) Decoder: Decoder D includes N d The decoder layer is a series of layers, each of which contains a multi-head self-attention layer, a multi-head cross-attention layer and a linear mapping layer. The self-attention layer, the cross-attention layer and the linear mapping layer each have a residual structure. For the compressed feature H obtained by the encoder, the decoder performs temporal boundary point analysis by transforming the decoder structure. The decoder defines N p Nomination query volume Q d , nomination query volume Q d As with the latent feature query volume, it is randomly initialized and then learned and generated during training, and the boundary nomination B is initialized 0 is 0, at the jth level, the nomination query volume Q d With border nomination B j After addition, after the self-attention layer and the primary residual structure, the cross-attention layer interacts with the compressed feature H, and after the residual structure-linear mapping layer-residual structure transformation, the updated boundary nomination B is obtained. j+1 ; Through the stacked N d After decoding layers, the compressed features are parsed to obtain the temporal boundary nomination representation
[0011] 2.3) Generation and scoring of temporal unclassified boundaries: The obtained temporal boundary nomination representation B is sent to two different fully connected layer branches: the positioning branch and the classification branch. The two branches are used to output the time and confidence score of the temporal unclassified boundary respectively;
[0012] 2.4) Assign training labels: A strict one-to-one training label matching strategy is adopted: According to the defined matching cost C, a set of optimal one-to-one matches is obtained using the Hungarian algorithm. Each prediction assigned to a classless boundary truth value obtains a positive sample label, and its corresponding boundary truth value is the training target; the matching cost C consists of two parts: the position cost and the classification cost. The position cost is defined based on the absolute value of the distance between the prediction time and the boundary truth time, and the classification cost is defined based on the prediction confidence;
[0013] 2.5) Submission of time series unclassified boundaries: After generating a series of time series unclassified boundaries, the most credible time series unclassified boundary moments are screened out through the confidence score threshold γ and submitted for subsequent performance measurement;
[0014] 3) Training phase: The configured model is trained with training samples, using cross entropy, L1 distance and log function as loss functions, using AdamW optimizer, and updating network parameters through back propagation algorithm, and repeating steps 1) and 2) until the number of iterations is reached;
[0015] 4) Detection: Input the video feature sequence and continuity score of the test data into the trained detection model to generate temporal category-free boundary moments and scores, and then use the method in 2.3) to obtain the temporal category-free boundary moment sequence for performance measurement.
[0016] The present invention also provides a timing sensor, which has a computer storage medium, wherein the computer storage medium is configured with a computer program, wherein the computer program is used to implement the above-mentioned categoryless timing boundary detection network, and when the computer program is executed, the above-mentioned timing boundary detection method is implemented.
[0017] The present invention proposes a universal and unified architecture to handle different types of unclassified temporal boundary detection. It can compress redundant video input into a stable and reliable feature representation based on the attention mechanism and video semantic structure, thereby reducing the complexity of the model. From a global context perspective, it can sparsely, efficiently and accurately give the position and confidence score of any unclassified temporal boundary.
[0018] Compared with the prior art, the present invention has the following advantages
[0019] This paper proposes a universal category-free temporal boundary detection paradigm, which provides an efficient and unified method for arbitrary category-free temporal boundary detection based on transformation structure and attention mechanism.
[0020] The present invention introduces a small set of learnable latent feature queries as anchor frames to compress redundant video inputs. The latent feature queries use a cross-attention mechanism to compress the input into a fixed-size latent feature space while preserving important boundary information, reducing the model's temporal and spatial complexity from the square of the input length to linear complexity.
[0021] The present invention constructs an effective latent feature query quantity for temporal boundary detection, including boundary query quantity and context query quantity. In order to better utilize the semantic structure composed of boundary and context in the video, the video features are also divided into boundary area features and context area features. The boundary query quantity extracts the boundary area features in a targeted manner; the context query quantity clusters the context area and compresses the redundant context content into multiple context clustering centers.
[0022] The present invention utilizes the alignment loss function to perform one-to-one alignment between the boundary query volume and the boundary area features, effectively reducing the training difficulty, shortening the convergence time, forming a stable compression feature, and improving the positioning prediction performance.
[0023] The present invention utilizes a transform decoder and a one-to-one training label matching strategy to effectively use global context information for boundary prediction, sparsely and efficiently generating more accurately positioned unclassified temporal boundary positions without the need for complex post-processing technology.
[0024] The present invention has the characteristics of universality, high efficiency and accuracy in the task of temporal boundary detection. Compared with the existing methods, the present invention achieves better prediction accuracy and faster reasoning speed on the sub-action level, event level and scene level unclassified temporal boundary data sets, reflecting the generalization performance of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 It is a system framework diagram used in the present invention.
[0026] Figure 2 It is a schematic diagram of frame extraction processing of the video of the present invention.
[0027] Figure 3 Schematic diagram of the encoder of the present invention.
[0028] Figure 4 Schematic diagram of a decoder of the present invention.
[0029] Figure 5 It is a schematic diagram of boundary nomination generation and submission of the present invention.
[0030] Figure 6 The comparison results of the model efficiency of the present invention on the MovieNet dataset samples are shown.
[0031] Figure 7 The results of the present invention compared with previous work on the MovieNet dataset samples are shown.
[0032] Figure 8 The results of the present invention compared with previous work on Kinetics-GEBD and TAPOS dataset samples are shown.
[0033] Fig. 9 The results of the present invention on the MovieNet dataset sample are visualized.
[0034] Fig.10 It is a schematic diagram of the overall process of the present invention. DETAILED DESCRIPTION
[0035] The present invention constructs a classless temporal boundary detection network to perform temporal boundary detection on a video, and the detection network includes a backbone network and a detection model. The detection model of the present invention is a temporal perception model (Temporal Perceiver), which is a general classless temporal boundary detection framework. It introduces a small group of latent feature query quantities as anchor frames, compresses redundant inputs to a fixed dimension through a cross-attention mechanism, and implements the classless temporal boundary detection task. The method of the present invention includes a sample generation stage, a network configuration stage, a training stage, and a testing stage. Fig.10 The specific instructions are as follows.
[0036] 1) Sample generation stage: Use the backbone network based on ResNet50 and temporal convolutional layers to generate samples for training and test videos. The samples include video features F and continuity scores S. For each video, all corresponding picture frames are sampled at intervals of τ frames to obtain a video image sequence In the video image sequence L f N f The length of the video segment is 2k frames, where the i-th video segment is an image sequence consisting of the consecutive k frames before and after the i-th frame image. N f is the length of the video image sequence, and also the number of video segments. Each video segment image sequence L s,i It is sent to the backbone network, and after the convolution layer, pooling layer and full connection layer with pre-training and fine-tuning parameters, the D-dimensional RGB feature F corresponding to the i-th frame is output. i and continuity score S i The features and scores of different video segments are spliced together in chronological order to obtain the features of the entire video and continuity scoring The sampling interval τ represents the granularity of global time division; the length of the video segment 2k represents the local receptive field range of the feature. In order to reduce the time complexity while retaining more local information, the embodiment of the present invention preferably takes τ as 3, k as 5, and D as 2048. The specific implementation is as follows:
[0037] Use denseflow to extract picture frames from the original video, sample all the picture frames of the video at intervals of 3 frames, and obtain a length of N f The video image sequence L f . Call the transforms package of the torchvision library to scale each image frame to obtain a corresponding image with a scale of 224*224. f N f The length of the video segment is 2k = 10 frames, where the i-th video segment L s,iIt consists of 5 consecutive frames before and after the i-th frame of the video, that is, The video segment sequence composed of all video segments is recorded as The video segment L s,i ∈R 10×224×224 Send it to the backbone network, the ResNet50 network that has been pre-trained and fine-tuned, and obtain the intermediate feature F mid,i ∈R 10×2048 , F mid,i After the temporal convolution layer and pooling layer, the RGB feature F of the video segment is obtained. i ∈R 2048 Finally, F i Input the fully connected layer to get the continuity score S i ∈R. The features F of different video segments i and score S i Splice them together in chronological order to get the characteristics of the entire video and continuity scoring The video features and continuity scores are processed by a sliding window method to form a series of fragments, where the window length is N. ws When the value is 100, there is no overlapping frame when sliding the window. The details are as follows:
[0038] 1. The overall video segment sequence obtained after frame extraction and sampling is as follows:
[0039]
[0040] L s,i ={f i-5 ,f i-4 ,f i-3 ,f i-2 ,f i-1 ,f i+1 ,f i+2 ,f i+3 ,f i+4 ,f i+5 ,}
[0041] Where V f Represents a sequence of video segments, which consists of N f Image sequence segment L s,i Each image sequence segment contains 2k=10 images.
[0042] 2. The process of the backbone network processing the input video image sequence is as follows:
[0043] F mid,i =Resnet50(L s,i )
[0044] F i=MaxPooling(Tconv(F mid,i ))
[0045] S i =FC(F i )
[0046]
[0047]
[0048] Among them, F mid,i represents the intermediate features obtained by processing the input video segment through the Resnet50 network, F i F mid,i The video segment features obtained after temporal convolution, S i is the continuity scoring result, F is the feature sequence of different video segments spliced together in chronological order, and S is the continuity scoring sequence of different video segments spliced together in chronological order.
[0049] 2) In the network configuration stage, based on the transform decoder structure and attention mechanism, a general category-free temporal action detection model Temporal Perceiver is established. The model includes the following configurations:
[0050] 2.1) Encoder: Based on the continuity score S generated in 1), the video features F generated in 1) are projected and sorted in descending order to obtain F rerank , introduce M learnable latent feature query quantities Q e and the compressed features H initialized to 0 0 , using the transform decoder structure to compress the reordered features into the compressed features H of M frames. The encoder E includes N e The transform decoding layer is connected in series, each layer is denoted as Encoder j , j represents the coding layer number, j∈[0,(N e -1)]. Encoder j The input is H j , the output is H j+1 Each Encoder j Contains a multi-head self-attention layer MSA j , a multi-head cross attention layer MCA j , a linear mapping layer FFN j and three residual structures formed by addition and layer normalization operations.
[0051] Multi-head self-attention layer MSA jThe inputs are key parameter K, query parameter Q and value parameter V. The key and query of the self-attention mechanism are the same inputs. The key parameter and query parameter are multiplied by themselves and normalized by the Softmax function to obtain the weight matrix A s , the value parameter is multiplied by the weight matrix to get the output. The multi-head structure divides the input into multiple parts along the parameters, inputs them into the self-attention layer respectively, and finally splices the results along the channel. Multi-head cross attention layer MCA j The input includes key, query and value parameters. The key parameter and query parameter are multiplied and Softmax normalized to get the cross attention weight A c , value parameter and cross attention weight A c The multi-head structure divides the input into multiple parts along the parameters, inputs them into the cross attention layer respectively, and finally splices the results along the channel.
[0052] The present invention preferably e =6, M = 60, the number of branches in the multi-head self-attention layer is 8, and the number of branches in the multi-head cross-attention layer is 8. MSA in the encoder j The key and query parameter are the implicit feature query quantity Q e With compression feature H j The result of addition, the value parameter is the compression feature H j , weight matrix A s ∈R M×M , remember MSA j The output after the residual structure is H′ j .MCA j The key parameter is F rerank The sum of video features and corresponding position encoding F rerank +P rerank , the value parameter is F rerank , the query parameter is MSA j The output H′ j and latent feature query quantity Q e The sum of the weight matrix A c ∈R N×M , remember MCA j The output after the residual structure is H″ j The encoder E is constructed by stacking N e = 6 encoding layers to compress and encode input features. The last encoding layer is Encoder 5 The output of the encoder E is the compressed feature H = H 6 The specific calculation is as follows:
[0053] 1. Reordering transformation of video features and position encoding:
[0054] F rerank = sort(fcproj (F),S)
[0055] P rerank =sort(P,S)
[0056] Among them, P is the position code formed by taking the sin function of the relative position of the time series corresponding to F, which corresponds to F one by one; fc proj The video feature is transformed from input dimension D = 2048 to model dimension D model =512 is the linear projection layer used, where the subscript proj is the abbreviation of projection.
[0057] 2. Encoder in a certain encoding layer j The encoding process:
[0058] H′ j =LayerNorm(H j +MSA j (H j ,Q e ))
[0059] H″ j =LayerNorm(MCA j (H′ j ,Q e ,F rerank ,P rerank )+H′ j )
[0060] H j+1 =LayerNorm(FFN j (H″ j )+H″ j )
[0061] 3. Encoding process of multi-head self-attention layer MSA:
[0062]
[0063] S h (x h ,q h )=x h W v,h A s,h
[0064] A s,h =Softmax((x h +q h )Wk,h·(x h +q h )W q,h )
[0065] Among them, N h represents the number of head branches in the multi-head self-attention layer, W k,h , W q,h , W v,h and W o is the projection matrix corresponding to the key, query, value, and output. The h subscript refers to different branches. The projection matrix parameters of each layer and each branch are not shared. h represents a single-head self-attention layer, x h ,q h The slices of feature x and its position code q along the channel dimension are N in total h , A s,h Represents the self-attention matrix of the h-th branch.
[0066] 4. The encoding process of the multi-head cross attention layer MCA:
[0067]
[0068] CA h (x h ,q x,h ,y h ,q y,h )=y h W v,h A c,h
[0069] A c,h =Softmax((y h +q y,h )W k,h ·(x h +q x,h )W q,h )
[0070] Among them, W k,h , W q,h , W v,h and W o The projection matrices corresponding to the key, query, value, and output are not shared by each layer or branch, and the projection matrix of the cross-attention layer is not shared with the self-attention layer. h represents a single-head cross attention layer, x h ,y h ,q x,h ,q y,h Encode x, y and the corresponding position q x ,q y Slices along the channel dimension, A c,h represents the cross attention matrix of the h-th branch.
[0071] 5. Encoding process of encoder:
[0072] H j+1 =Encoder j (H j ;Q e ,F rerank ,P rerank )
[0073] H=H 6 =Encoder 5 (Encoder 4 (Encoder 3 (Encoder 2 (Encoder 1 (Encoder 0 (H 0 ))))))
[0074] 2.2) The present invention introduces a hidden feature query quantity Q to the encoder e , which is generated specifically as follows: In order to better utilize the semantic structure of the video, the latent feature query quantity Q e Divided into M b Boundary query volume and M c The number of context queries, M b =48,M c = 12; the reordering features are divided into boundary region features and context region features, the first M b The features are boundary region features. The definition of boundary region features and context region features is based on the continuity score S. After the video features are reordered in descending order according to S, the top M with higher scores are selected. b Features form boundary region features, and the rest are context region features. e is a learnable parameter. According to the above description, the implicit boundary query quantity Q e Random initialization, generated by learning during model training. b The query volume is defined as the boundary query volume, and the boundary region features are combined one-to-one during the compression process; the remaining M c The query volume is the context query volume. In the compression process, the context region features are clustered to obtain M b The cluster centers represent the context information.
[0075] 2.3) Further, in order to increase the implicit boundary query quantity Q e The encoding effect of video features, aligning the latent feature query volume with the video features during training: In order to accelerate the convergence of the model and obtain more stable compression features, the present invention introduces additional supervision constraints and calculates the alignment loss function based on the last layer of cross attention matrix. The loss function is calculated by b ×Mb The sum of the diagonal weights of the region is taken as the negative logarithm, and the diagonal weights are maximized by minimizing the loss function to achieve one-to-one alignment of the boundary query volume and features. The specific calculation is as follows:
[0076] Alignment loss function L align The calculation process:
[0077]
[0078] Among them, α align = 1 is the parameter of the alignment loss function, and the alignment loss function is only calculated for the last cross-attention layer of the encoder.
[0079] 2.4) Decoder: For the compressed feature H obtained in 2.1), use N p The number of nomination queries that can be learned is Q d and the boundary nomination initialized to 0 represents B 0 , by transforming the decoder structure to perform temporal boundary point parsing, nominating the query quantity Q d As with the latent feature query, it is randomly initialized and then learned and generated during training; the decoder D includes N d The jth decoding layer is recorded as Decoder j , the input is the boundary nomination representation B j , the output is B j+1 Each decoding layer contains a multi-head self-attention layer, a multi-head cross-attention layer, a linear mapping layer, and three residual structures formed by addition and layer normalization operations. At the jth layer, the nomination query volume Q d With border nomination B j After addition, after the self-attention layer and the primary residual structure, the cross-attention layer interacts with the compressed feature H, and after the residual structure-linear mapping layer-residual structure transformation, the updated boundary nomination B is obtained. j+1 By stacking N d After a decoding layer, the compressed features are parsed and the temporal boundary nomination representation B is obtained.
[0080] The present invention preferably d =6, N p = 10, the number of branches in the multi-head self-attention layer is 8, and the number of branches in the multi-head cross-attention layer is 8. MSA j The key and query parameters are the nomination query quantity Q d With border nomination indicates B j The result of addition, the value parameter is the boundary nomination representation B j , the weight matrix Remember MSA j The output after the residual structure is B′ j .MCAj The key parameters are the compressed feature H and the latent feature query Q as the compressed position encoding. e The sum H+Q e , the value parameter is H, and the query parameter is MSA j The output B′ j and nomination query volume Q d The sum of the weight matrix MCA j The output after the residual structure is B″ j The decoder D is symmetrical to the encoder E and is constructed by stacking N d = 6 encoding layers, to achieve the decoding of boundary nomination representation, the last layer is the decoding layer Decoder 5 Output B 6 =B, which will be put into the fully connected branch as the final output of the decoder to predict the boundary position and confidence. The specific calculation is as follows:
[0081] 1. A decoding layer in the decoder j The decoding process:
[0082] B′ j =LayerNorm(B j +MSA j (B j ,Q d ))
[0083] B″ j =LayerNorm(MCA j (B′ j ,Q d ,H,Q e )+B′ j )
[0084] B j+1 =LayerNorm(FFN j (B″ j )+B″ j )
[0085] 2. The decoding process of the decoder:
[0086] B j+1 =Decoder j (B j ;Q d ,H,Q e )
[0087] B=B 6 =Decoder 5 (Decoder 4 (Decoder 3(Decoder 2 (Decoder 1 (Decoder 0 (B 0 ))))))
[0088] 2.5) Generation and scoring of temporal non-classified boundaries: The temporal boundary nomination representation B obtained in 2.4) is sent to two different fully connected layer branches: the positioning branch Head loc and classification branch Head cls , the two branches are used to output the time series without category boundary and confidence score respectively. The predicted boundary moment is a decimal between 0 and 1, which indicates the relative position in the current segment; the confidence score is two scores, including positive confidence and negative confidence, and the higher score represents a greater probability of the corresponding category. The classification branch consists of a fully connected layer, with input and output feature dimensions of 512 and 2 respectively; the positioning branch consists of a multi-layer perceptron composed of three fully connected layers and a Sigmoid activation function, with input and output dimensions of 512, 512, 512 and 512, 512, 1 respectively. The specific calculation is as follows:
[0089] 1. Prediction of the location time t of the unclassified action:
[0090] t=sigmoid(fc 2 (fc 1 (fc 0 (B))))
[0091] Among them, the three fully connected layers of the positioning branch are fc 2 ,fc 1 ,fc 0 The input and output dimensions are 512, 512, 512 and 512, 512, 1 respectively.
[0092] 2. Binary classification confidence score p pos Generation of:
[0093] p pos ,p neg =fc(B)
[0094] Among them, the fully connected layer of the confidence branch is denoted as fc, and the positive example fraction p is generally taken pos As confidence scores, the input dimension is 512 and the output dimension is 2.
[0095] 2.6) Assign training labels: Adopt a strict one-to-one training label matching strategy: According to the defined matching cost C, use the Hungarian algorithm to obtain a set of optimal one-to-one matches. Each prediction assigned to a classless boundary truth value obtains a positive sample label, and its corresponding boundary truth value is the training target; the matching cost C consists of two parts: position cost and classification cost. The position cost is defined based on the absolute value of the distance between the prediction time and the boundary truth value time, and the classification cost is defined based on the prediction confidence. The specific calculation is as follows:
[0096] 1. Optimization indicators of the Hungarian algorithm:
[0097]
[0098] The optimization index is denoted as C. The optimization index has two components: position cost and classification cost. Each component has a corresponding weight, denoted as α loc and α cls The present invention preferably loc =5,α cls =1.
[0099] 2. Definition of optimization components:
[0100]
[0101] L cls,n =-p pos,n
[0102] In the optimization component, the nth predicted positioning component L loc,n The predicted boundary time t n The position of the corresponding boundary truth value The absolute value of the distance is used to measure; the nth predicted classification component L cls,n The confidence p of the predicted moment as the boundary pos,n To measure, since C is the minimization target, the confidence is negative; σ(·) is a mapping from prediction to boundary truth value, and σ(n) is the action truth value corresponding to the nth prediction.
[0103] 2.7) Submission of time series unclassified boundaries: After generating a series of time series unclassified boundaries, the most credible time series unclassified boundary moments are screened out through the confidence score threshold γ and submitted for subsequent performance measurement. pos,n ≥γ, then the predicted position is submitted as the result. pos,n <γ, the predicted position is discarded. The preferred embodiment of the present invention is γ = 0.9.
[0104] 3) Training phase: The configured model is trained using training samples, using cross entropy and L1 distance as loss functions based on the final results, and log function as loss function based on the intermediate results. The AdamW optimizer is used to update the network parameters through the back propagation algorithm, and steps 1) and 2) are repeated until the number of iterations is reached;
[0105] 4) Testing phase: The video feature sequence and continuity score of the test data are input into the trained Temporal Perceiver model to generate temporal category-free boundary moments and scores. Then, the temporal category-free boundary moment sequence for performance measurement is obtained through the method in 2.5).
[0106] The present invention proposes a temporal perceptron model, a general classless temporal boundary detection framework. The following is further described by specific embodiments. After training and testing on the TAPOS dataset, Kinetics-GEBD dataset and MovieNet / MovieScenes dataset, high inference speed and high accuracy are achieved, preferably using Python 3.8.8 programming language and PyTorch 1.7.0 deep learning framework.
[0107] Figure 1 The system framework diagram used in the present invention is shown, and the specific implementation steps are as follows:
[0108] 1) The preparation stage of generating samples, such as Figure 2 As shown in the figure, both training data and test data are processed in the same way. Denseflow is used to extract picture frames from the video, and the picture frame sequence is sampled at intervals of τ=3. The picture frames are scaled to a scale of 224*224 using the transforms package of torchvision, and finally converted into tensor form and normalized. Taking each frame in the picture frame sequence as the center, the k=5 frames before and after it are input into the backbone network to obtain the features and continuity scores corresponding to the frame, and the video features and video continuity scores are obtained by splicing along the time dimension. The video-level features and continuity scores are divided into a series of video window segments of the same length and fed into the model.
[0109] 2) In the configuration phase of the model, we first score based on continuity, such as Figure 3As shown in the figure, for the extracted video features, the program first projects the video features to a lower feature dimension, and sorts them in descending order with the corresponding feature codes according to the continuity score, and inputs them into the encoder. The encoder includes alternating multi-head self-attention modules, multi-head cross-attention modules, linear mapping layers, and residual structures with addition and layer normalization. A set of learnable latent feature query quantities including boundary query quantities and context query quantities are input into the encoder. The compressed features are initialized to 0, participate in the calculation of the encoding layer, and accumulate the features extracted from the input features by each encoding layer. The latent feature query quantity can be regarded as the position encoding of the compressed feature. The compressed features and the position encoding are added to obtain the position information, which participates in the attention calculation.
[0110] First, the compressed features and latent feature query are added as the key and query input to the self-attention layer, and are input separately as value parameters. In the self-attention layer, the relationship between compressed features is modeled by global self-attention, and mutual information is extracted for feature update. The compressed features are added before and after update and normalized to obtain the intermediate representation. Subsequently, the intermediate representation and the latent feature query as the position code are added again as the query input to the cross-attention module, and the input video features are used as the value and added to the corresponding position code as the key input to the cross-attention layer. The cross-attention layer extracts boundary features and clusters context features from the sorted video features. Useful video features are extracted by calculating the dot product activation of the compressed features and the latent feature query on the input video sequence, and the compressed features are updated to achieve the effect of feature refinement and complexity reduction. Finally, the results of the cross-attention layer are passed through the residual structure, the projection of the linear mapping layer and another residual structure to obtain the compressed features after cumulative update.
[0111] The step of decoding the encoded features to obtain the final result representation is the aforementioned step 2.4), such as Figure 4 As shown. The input nomination query volume and boundary nomination are input into the multi-head self-attention layer to strengthen the nomination representation, and then input into the cross attention layer. The learned latent feature query volume is used as the position encoding of the compressed feature and participates in the cross attention calculation. The cross attention layer extracts the boundary position information of the compressed feature, finds the weight of each time dimension position of the compressed feature through the cross attention matrix multiplied by the compressed feature and the boundary nomination, and extracts the corresponding features. The boundary nomination accumulates the extracted boundary representation at each layer, and finally obtains the decoding result. There is an additive normalized residual transformation between the self-attention layer and the cross attention layer, and there is a second residual transformation, a linear projection layer projection and a third residual transformation after the cross attention layer.
[0112] Decoding and submission of timing boundary nominations, such as Figure 5As shown. The boundary nomination representation is input into the localization and classification branches represented by the fully connected layer to obtain the position and confidence score. The localization branch includes three fully connected layers (fc) and a sigmoid activation function, and finally obtains the position. The classification branch is a binary classification, including one fc layer, and obtains the confidence score corresponding to the current predicted position. The confidence score is screened, and if it is higher than the threshold γ=0.9, then the prediction is submitted as the final prediction.
[0113] 3) During the training phase, this example uses cross entropy, L1 distance, and negative logarithmic function as loss functions, uses the AdamW optimizer, and sets the batch size to 64, that is, 64 window samples are taken from the training set for each training. The total number of training rounds is set to 100, the initial learning rate is 2e-4, and there is no learning rate decay strategy. The model is trained on an NVIDIA RTX 2080ti GPU. The process from the original video to the temporal boundary result is divided into two stages. In the first stage, the backbone network is fine-tuned on the dataset based on the pre-trained parameters to obtain video features and continuity scores; in the second stage, TemporalPerceiver is trained and tested. In the stage of allocating positive and negative samples, the model adopts a strict one-to-one matching strategy to reduce the occurrence of false positives, so that the model can sparsely and efficiently predict unclassified temporal boundaries.
[0114] 4) Testing phase
[0115] The preprocessing of the test set is the same as that of the training data. After extracting frames, the data is compressed to a size of 224*224. The backbone network based on ResNet50 is used for RGB feature extraction. The test indicators used are different for different task datasets: the f1 score at different relative distances (Rel.Dis.) of the sub-action unclassified temporal boundary and the event unclassified temporal boundary is used as the indicator, and the scene-level unclassified temporal boundary uses AP and M. iou As an indicator. The calculation basis of F1 score is recall and precision. The recall rate refers to the ratio of the number of samples predicted correctly to the total number of true values, and the precision rate refers to the ratio of the number of samples predicted correctly to the number of samples predicted as true values. Samples with prediction errors within the relative distance can be considered as correctly predicted. The calculation basis of AP (Average Precision) is the average precision value corresponding to the recall value from 0 to 1. iou It is the weighted sum of the intersection of the distance between the predicted boundary and the true boundary and the length of the true scene. AP and M iou All indicators require that the prediction results must be consistent with the true value position to be considered a correct prediction.
[0116] On three randomly selected videos from the MovieNet dataset, such as Figure 6 As shown in the figure, TemporalPerceiver has a 7-times faster scene reasoning speed per second and nearly 200-times fewer floating-point calculations than the classic LGSS, reflecting the advantages of the model's sparse prediction and no need for post-processing modules; compared with the Transformer variant, TemporalPerceiver also has a faster scene reasoning speed per second and fewer floating-point operations, reflecting the role of feature compression in model lightweighting, further proving the efficiency of Temporal Perceiver. In terms of prediction accuracy, TemporalPerceiver has achieved a huge improvement in all indicators of all data sets compared to previous work, reflecting the versatility and generalization of the model: Figure 7 , on the MovieScenes dataset, Temporal Perceiver has the best performance in AP and M iou The indicators all exceed the classic work LGSS by nearly 3%; Figure 8 In the Kinetics-GEBD and TAPOS datasets, based on all indicators of f1@Rel.Dis., the performance of Temporal Perceiver is higher than that of the previous state-of-the-art PC. In the Kinetics-GEBD dataset, based on the f1@0.05 indicator, the performance of Temporal Perceiver is 12.6% higher than that of the previous state-of-the-art PC; on the TAPOS dataset, based on the average f1 score, the performance of Temporal Perceiver is 9% higher than that of the previous state-of-the-art PC, showing that Temporal Perceiver's prediction is more flexible and accurate. More specific prediction visualization on the MovieNet dataset is shown in the following figure. Fig. 9 As shown in Figure 2, Temporal Perceiver’s prediction avoids false positives near the true value and accurately predicts the true value of each scene.
Claims
1. Timing boundary detection method, characterized by Construct a classless temporal boundary detection network to detect temporal boundaries in videos. The detection network includes a backbone network and a detection model. The implementation method is as follows: 1) Generate detection samples by backbone network: sample the video interval to obtain video image sequence Generate a video segment for each frame, and the i-th video segment is composed of the i-th frame image f i The backbone network generates video features for the input video segment. and continuity scoring F i and S i Score the RGB features and continuity of video segment i respectively; 2) The detection model performs non-classified temporal action detection based on the video feature F and the continuity score S, and the detection model includes the following configuration: 2.1) Encoder: Encoder E includes N e The transform decoding layer is a series of layers, each of which contains a multi-head self-attention layer, a multi-head cross-attention layer and a linear mapping layer. The self-attention layer, the cross-attention layer and the linear mapping layer each have a residual structure, and the encoder introduces M hidden feature query quantities Q e , based on the continuity score S, the video features F are sorted in descending order and then input into the encoder. The encoder compresses the sorted video features into compressed features H of M frames. The initial compressed feature H0 is 0. At the jth transform decoding layer, the hidden feature query quantity Q e And the compression feature H of the current layer j After addition, the self-attention layer and its residual structure interact with the reordered video features in the cross-attention layer, and then the residual structure-linear mapping layer-residual structure transformation is used to obtain the compressed feature H j+1 ,j∈[0,(N e -1)], through the stacked N e After the encoding layer, the input features are compressed and encoded to obtain compressed features. Among them, the generation of latent feature query quantity is: latent feature query quantity Q e Divided into M b Boundary query volume and M c context queries, randomly initialized, are generated with training sample learning during the training detection model training process; boundary queries correspond to processing boundary region features in video features, context queries correspond to processing context region features in video features, and the top M in video features are reordered. b Features are boundary area features, and the others are context features; 2.2) Decoder: Decoder D includes N d The decoder layer is a series of layers, each of which contains a multi-head self-attention layer, a multi-head cross-attention layer and a linear mapping layer. The self-attention layer, the cross-attention layer and the linear mapping layer each have a residual structure. For the compressed feature H obtained by the encoder, the decoder performs temporal boundary point analysis by transforming the decoder structure. The decoder defines N p Nomination query volume Q d , nomination query volume Q d As with the latent feature query volume, it is randomly initialized and then learned and generated during training, and the boundary nomination B0 is initialized to 0. At the jth layer, the nomination query volume Q d With border nomination B j After addition, after the self-attention layer and the primary residual structure, the cross-attention layer interacts with the compressed feature H, and after the residual structure-linear mapping layer-residual structure transformation, the updated boundary nomination B is obtained. j+1 ; Through the stacked N d After decoding layers, the compressed features are parsed to obtain the temporal boundary nomination representation 2.3) Generation and scoring of temporal unclassified boundaries: The obtained temporal boundary nomination representation B is sent to two different fully connected layer branches: the positioning branch and the classification branch. The two branches are used to output the time and confidence score of the temporal unclassified boundary respectively; 2.4) Assign training labels: A strict one-to-one training label matching strategy is adopted: According to the defined matching cost C, a set of optimal one-to-one matches is obtained using the Hungarian algorithm. Each prediction assigned to a classless boundary truth value obtains a positive sample label, and its corresponding boundary truth value is the training target; the matching cost C consists of two parts: the position cost and the classification cost. The position cost is defined based on the absolute value of the distance between the prediction time and the boundary truth time, and the classification cost is defined based on the prediction confidence; 2.5) Submission of time series unclassified boundaries: After generating a series of time series unclassified boundaries, the most credible time series unclassified boundary moments are screened out through the confidence score threshold γ and submitted for subsequent performance measurement; 3) Training phase: The configured model is trained with training samples, using cross entropy, L1 distance and log function as loss functions, using AdamW optimizer, and updating network parameters through back propagation algorithm, and repeating steps 1) and 2) until the number of iterations is reached; 4) Detection: Input the video feature sequence and continuity score of the test data into the trained detection model to generate temporal class-free boundary moments and scores, and then use the method in 2.3) to obtain the temporal class-free boundary moment sequence for performance measurement.
2. The timing boundary detection method according to claim 1, characterized in that When training the detection model, align the latent feature query volume with the video feature: the boundary query volume is aligned with the boundary area feature of the video feature through the alignment loss function. The alignment loss function is calculated based on the last layer of cross-attention map. Using the condition that the number of boundary query volume and boundary area features is consistent, the diagonal attention weights formed by the corresponding relationship between the two are combined and negatively logarithmic to obtain the value of the alignment loss function. By minimizing the alignment loss function, the purpose of maximizing the diagonal weight is achieved, and the boundary query volume is ensured to correspond to the extracted boundary area features during cross-attention.
3. The timing boundary detection method according to claim 1 or 2, characterized in that The backbone network is based on ResNet50 and temporal convolutional layers to generate video samples. For each video, all picture frames corresponding to the video are sampled at intervals of τ frames to form a video image sequence. Divided into N f video segments of length 2k frames, N f is the length of the video image sequence, and also the number of video segments. The i-th video segment is an image sequence consisting of k consecutive frames before and after the i-th frame image. The image sequence L s,i It is sent to the backbone network, and after pre-training and fine-tuning of the parameters of the convolution layer, pooling layer, and fully connected layer, the RGB feature F is output. i and continuity score S i , the features and scores of different video segments are spliced together in chronological order to obtain the features of the entire video and continuity scoring Among them, the sampling interval τ represents the granularity of global time division, and the size of the video segment length 2k represents the local receptive field range of the feature.
4. The timing boundary detection method according to claim 3, characterized in that First, use the denseflow library to extract frames from all videos to obtain image frames, sample at intervals of τ frames, τ is 3, and perform the following processing: call the transforms package of the torchvision library, scale each image frame to 224*224 size, convert it into tensor form, and finally normalize it to obtain the video image sequence L f ; Traverse the image sequence of the entire video, take the k frames before and after each frame, k is 5, and get N f The stacked spatial convolution layer, maximum pooling layer and temporal convolution layer are used to extract the features of each video segment. The features are then passed through a fully connected layer to obtain a continuity score. The features and scores are concatenated in chronological order to obtain video-level features F and continuity score S.
5. The timing boundary detection method according to claim 1 or 2, characterized in that In the configuration of step 2), the convolution layer of the backbone network is composed of a convolution operation, a batch normalization operation, and a ReLU activation function, the encoder is a transform decoder structure used as an encoder, and the decoder is also a transform decoder structure.
6. The timing boundary detection method according to claim 1 or 2, characterized in that the encoder Number of layers N e Take 6, at the jth layer, compress feature H j As the value parameter, the latent feature query quantity Q e and compression feature H j After addition, they are used as key and query parameters to input into the 8-branch self-attention module. The attention matrix is formed by self-multiplication and Softmax normalization of the key and query parameters and multiplied with the value parameter to get the output. In the residual structure, it is combined with H j Add and normalize to get H j ';H j ' and latent feature query quantity Q e The video feature F is reordered in descending order based on the continuity score S as the value parameter, and is added to the position code that has also been reordered as the key parameter to enter the cross attention module. In the module, the attention matrix is formed by cross-multiplying the key and query parameters and normalizing them with Softmax, and then multiplied with the value parameter. After the residual structure and H' j After adding normalization, linear mapping and residual structure again, the updated compressed feature H is obtained j+1 .
7. The timing boundary detection method according to claim 1 or 2, characterized in that Number of decoder layers N d Take 6, at level j, the boundary nomination B j As the value parameter, the nomination query quantity Q d and Boundary Nomination B j After addition, they are used as key and query parameters to input into the self-attention module of 8 branches. The attention matrix is formed by self-multiplication and Softmax normalization of the key and query parameters and multiplied with the value parameter to get the output. j Add and normalize to get B j ';B j ' and the nomination query volume Q d Add as query parameters, the compressed feature H obtained in 2.1) as the value parameter, and the latent feature query Q as the compressed position encoding e The addition is input as the key parameter into the cross attention module, in which the attention matrix is formed by cross-multiplying the key and query parameters and normalizing them with Softmax and multiplying them with the value parameters. j 'After adding and normalizing, we get the updated boundary nomination B j+1 .
8. The timing sensor is characterized by A computer storage medium is provided, wherein the computer storage medium is configured with a computer program, wherein the computer program is used to implement the classless timing boundary detection network described in any one of claims 1 to 7, and when the computer program is executed, the timing boundary detection method described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Video content description method based on text auto-encoder
CN111079532A
Video crowd counting system and method
CN111860162A