Open set weak supervision time sequence action detection method based on sub-action unit perception

By adopting the open set weak supervision method based on sub-action unit perception in video action detection, the relationship between sub-action unit and action category is constructed, and the problem of limited performance of existing methods when detecting changes in action category is solved, and more accurate detection and identification of action categories is achieved.

CN119920009APending Publication Date: 2025-05-02CHANGZHOU INST OF MECHATRONIC TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510001893.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-05-02

AI Technical Summary

Technical Problem

The existing video action detection methods cannot effectively detect actions whose action categories are non-fixed and changing, and ignore the relationship between the overall action characteristics and the fine-grained characteristics of the sub-action units, resulting in limited detection performance.

Method used

Using the open set weakly supervised timing action detection method based on sub-action unit perception, by serializing and feature extraction of video frames, a classifier is used to predict the fragment-level class activation sequence, and a cross-entropy loss function is calculated. At the same time, the relationship between the sub-action unit and the known action category is constructed, the sub-action unit features are learned using graph convolutional networks, and the consistency loss function is calculated to constrain the classification results.

Benefits of technology

By modeling fine-grained sub-action units in the action, explaining and expressing the inherent nature of the action, the accuracy of action detection is improved, and the known and unknown action categories can be effectively identified, and the sub-action pattern expression of unknown actions can be provided to provide a basis for further classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119920009A_ABST
    Figure CN119920009A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of video motion detection, in particular to an open set weak supervision time sequence motion detection method based on sub-motion unit perception, which comprises the following steps of: extracting embedded features; predicting a fragment level class activation sequence of the embedded features; calculating a cross entropy loss function; calculating a total loss function in the baseline branch; constructing a relationship between the sub-action units and known action categories and a sequential relationship between the sub-action units and different moment segments of the video; performing action semantic relation expression and sequential relation expression modeling based on the sub-action units by taking the activation value on each sub-action unit as relation expression, and calculating action classification loss and DTW loss based on the sub-action units; calculating a consistency loss function of the two branches; calculating characterization loss and balance loss functions; and constructing a total loss function of the baseline branch and the sub-action unit branch. The invention solves the problem that the detection performance needs to be further improved when the detection action category is non-fixed and different in the existing method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video action detection, and in particular to an open set weakly supervised temporal action detection method based on sub-action unit perception. Background Art

[0002] Video motion detection, as an important technology in the field of computer vision, can be widely used in security monitoring, human-computer interaction, sports analysis, medical rehabilitation and other fields.

[0003] The patent publication number is CN 114821772A. It constructs a three-branch classification network, uses the baseline branch to predict the class activation sequence of action and background in the video, and uses the pooling branch to predict the action class activation sequence and background class activation sequence in the video respectively, thereby obtaining three video-level class activation scores, and finally uses the cross entropy classification loss to train the three-branch classification network separately.

[0004] The patent publication number is CN 116310988A. It uses threshold erasing to determine whether other clips in the video contain less significant action clips; multiplies the cascaded action attention weights with the embedded features along the time dimension to generate action attention pooled features; generates class activation sequences and fuses them to form cascaded class activation sequences; and calculates the cascade branch entropy classification loss function value through the cascaded action attention branch network.

[0005] However, the above two methods are designed for closed-set weakly supervised temporal action detection methods, that is, the action categories of the training set and the test set are the same and fixed; and with the continuous development of society, in a dynamically changing open environment, action categories are constantly changing, and action categories that humans have never seen before are constantly emerging; the above two methods cannot correctly classify and detect them.

[0006] In addition, the above two methods either study the action instance as a whole feature, or divide the entire action into only salient segments and sub-salient segments, without considering the finer-grained sub-action units, and thus ignore the relationship between the overall action features and the fine-grained features of the sub-action units, resulting in limited improvement in action detection performance. Summary of the invention

[0007] In view of the shortcomings of the existing methods, the present invention solves the problem that the detection performance of the existing methods needs to be further improved when the detection action categories are non-fixed and different.

[0008] The technical solution adopted by the present invention is: an open set weakly supervised temporal action detection method based on sub-action unit perception comprises the following steps:

[0009] Step 1: Serialize the video frames, extract video clip features from the serialized video frames, and then extract embedded features;

[0010] Step 2: Use the classifier to predict the segment-level class activation sequence of the embedded features; calculate the probability that the video belongs to each category, and calculate the cross entropy loss function between the probability of the predicted action category and the true value;

[0011] Step 3: Use the evidence deep learning loss function to calculate the probability of actions belonging to known categories, and combine it with the cross entropy loss function of the video category to obtain the loss function in the baseline branch;

[0012] As a preferred embodiment of the present invention, the formula of the loss function in the baseline branch is:

[0013]

[0014] in, is the cross entropy loss function of the video category, L edl is the loss function for evidence deep learning.

[0015] Step 4: construct the relationship between the sub-action unit and the known action category, and the relationship module between the sub-action unit and the video segments at different times;

[0016] As a preferred embodiment of the present invention, the construction of the relationship between the sub-action unit and the known action category includes:

[0017] Arbitrarily initialize a set of learnable sub-action units, characterized by Among them, M is the number of sub-action units, and E is the feature dimension of the sub-action unit;

[0018] Calculate the adjacency matrix between sub-action units Use cosine similarity to calculate the similarity between two different sub-action units;

[0019] Based on the adjacency matrix E sub Construct the edges of the sub-action unit graph convolutional network; use different weights Q for the neighbor connections and self-connections of the sub-action unit;

[0020] Implement the softmax activation operation along the category dimension, expressed as:

[0021]

[0022] in, Represents the activation value of different actions on each action unit, j=1,…,M; c=1,…,C,C+1.

[0023] As a preferred embodiment of the present invention, the construction of the relationship between the sub-action unit and the video segments at different times includes:

[0024] Get the activation value S of each sub-action unit in the video clips at different times sub ;

[0025] Softmax function activation is performed along the sub-action category dimension,

[0026] Step 5: Use the activation value on each sub-action unit as the input of the relational expression modeling branch to perform action semantic relation expression and temporal relation expression modeling based on the sub-action unit respectively;

[0027] As a preferred implementation of the present invention, step five specifically includes:

[0028] Calculate the activation scores of the video on M sub-action units, use action attention that is independent of the category to perform weighted summation along the temporal dimension, and then use softmax for normalization;

[0029] Using the Relationship Matrix The activation score of the input video on the sub-action unit is converted into a predicted value on the action category, and the corresponding cross entropy classification loss function is calculated. The formula is:

[0030]

[0031] in, is the video-level label after regularization, The activation scores of the input video on the sub-action unit are converted into prediction values ​​on the action category.

[0032] Step 6: Calculate the consistency loss function based on the class activation scores in the baseline branch and the class activation scores in the sub-action unit relationship;

[0033] As a preferred embodiment of the present invention, the formula of the consistency loss function is:

[0034]

[0035] in, is the class activation score in the baseline branch, is the class activation score in the sub-action unit relation module.

[0036] Step 7: Construct a temporal relationship expression model based on sub-action units;

[0037] As a preferred embodiment of the present invention, step seven specifically includes:

[0038] Calculate video embedding feature X emb and sub-action unit feature A sub The correlation value between The expression is:

[0039]

[0040] calculate The DTW distance is obtained and the DTW loss function is obtained.

[0041] Step 8: Introduce the representation loss and balance loss training model;

[0042] Step 9: Construct a total loss function based on the baseline branch, sub-action unit branch, representation loss and balance loss;

[0043] As a preferred embodiment of the present invention, the formula of the total loss function is:

[0044]

[0045] Among them, α1, α2 and α3 are hyperparameters; L rep is the loss function; L ban is the balance loss function; L dtw is the DTW loss function; L cons is the consistency loss function; L base is the loss function in the baseline branch; is the cross entropy classification loss function.

[0046] As a preferred embodiment of the present invention, an open-set weakly supervised temporal action detection system based on sub-action unit perception includes: a memory for storing instructions executable by a processor; a processor for executing instructions to implement an open-set weakly supervised temporal action detection method based on sub-action unit perception.

[0047] As a preferred embodiment of the present invention, a computer-readable medium stores a computer program code, and when the computer program code is executed by a processor, an open-set weakly supervised temporal action detection method based on sub-action unit perception is implemented.

[0048] Beneficial effects of the present invention:

[0049] 1. By modeling fine-grained sub-action units in actions, the present invention can more effectively explain the intrinsic nature of the expressed actions, solve various known and unknown action problems that are difficult to detect due to environmental changes, and promote the accuracy of action detection;

[0050] 2. The present invention models the differences in the semantic relationships and temporal associations of sub-actions, learns fine-grained sub-action unit feature expressions through graph convolutional networks, establishes semantic concepts and temporal associations between known actions and sub-action units, learns fine-grained sub-action unit feature expressions and more discriminative video embedding features, thereby more effectively identifying known action categories and unknown action categories, and at the same time can also express sub-action patterns for unknown actions, providing a basis for further classification of unknown actions. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 It is a logic block diagram of the open set weakly supervised sequential action detection method based on sub-action unit perception of the present invention;

[0052] Figure 2 It is a branch network block diagram of the sub-action unit of the present invention. DETAILED DESCRIPTION

[0053] The present invention is further described below in conjunction with the accompanying drawings and embodiments. This figure is a simplified schematic diagram, which only illustrates the basic structure of the present invention in a schematic manner, and therefore it only shows the components related to the present invention.

[0054] like Figure 1 As shown, the open set weakly supervised temporal action detection method based on sub-action unit perception includes the following steps:

[0055] Open-set weakly supervised temporal action detection requires the model to not only accurately identify known action categories that have appeared in the training set, but also identify unknown action categories that have never been seen in the training set, and detect the start and end position information of each action instance in the video.

[0056] Step 1: Serialize the video frames, extract video clip features from the serialized video frames, and then extract embedded features;

[0057] Video frame sequence Among them, l is the video frame number, L is the total number of video frames, v l The first frame of the video; the video segment is divided into video segments according to the interval of 16 consecutive frames;

[0058] Use the I3D network to extract features from each video clip and generate and are RGB features and optical flow features respectively; t represents the tth segment of the video, and D represents the feature dimension; Stitching to get I3D video features T is the sampling length of the video;

[0059] The feature embedding module consists of a one-dimensional temporal convolution layer and an activation layer. The I3D video features are projected into the vector space. The embedding feature expression is:

[0060] X emb =Φ emb (X,W emb ) (1)

[0061] Among them, Φ emb (·) is the convolution operation; W emb is the learning parameter; T is the video length, E is the embedding feature dimension;

[0062] Use convolution operation to aggregate the information of adjacent segments and capture the local temporal features of the video;

[0063] Step 2: Use the classifier to predict the segment-level class activation sequence of the embedded features;

[0064] Embed feature X emb Input to a classifier consisting of a 1D temporal convolutional layer to predict the sequence of class activations at the segment level The expression is as follows:

[0065]

[0066] Among them, Φ cls (·) is the convolution operation, φ cls To learn the parameters, C is the number of action categories;

[0067] The top-k means method is used to aggregate the class activation scores at the segment level along the temporal dimension. The formula is as follows:

[0068]

[0069] in, a is the class activation sequence Elements in

[0070] Then, the softmax activation function is used along the category dimension to calculate the probability that the video belongs to each category. The expression is as follows:

[0071]

[0072] in,

[0073] The predicted probability is compared with the true value using the cross entropy loss function, the formula is:

[0074]

[0075] Among them, N is the number of videos, C+1 is the unknown category, is the video level label after regularization, that is Since the training data does not contain unknown action videos, y(C+1)=0.

[0076] Embed feature X emb Input to the attention module to predict the attention score sequence that is independent of the action category. The attention module consists of a one-dimensional temporal convolutional layer; the expression is as follows:

[0077] λ=Φ att (X emb ,W att ) (6)

[0078] where Φ att (·) is the convolution operation, W att is the learning parameter; Represents the most discriminative action attention weight value in the video;

[0079] Step 3: Use the Evidential Deep Learning (EDL) theory to optimize the classification and uncertainty of video-level actions in the model. The probability of unknown category actions in the model prediction results is calculated using the Evidential Deep Learning loss function. The expression is:

[0080] L edl =EDL(X emb ) (7)

[0081] Among them, X emb is the embedding feature, and EDL(·) is the evidential deep learning operation used to consider action classification and uncertainty modeling.

[0082] Here, the θ generated in the model must ensure that 95% of the videos in the training dataset are identified as known class actions. Therefore, the loss function in the baseline branch is:

[0083]

[0084] Compared with CN 114821772 A and CN 116310988 A, the baseline branch of the present invention not only considers the classification loss of the action, but also considers the uncertain modeling in an open set environment.

[0085] like Figure 2 ,Step 4, construct the relationship module between the sub-action unit and the known action category, ,and the relationship module between the sub-action unit and the video segments at different ,times;

[0086] Design a sub-action unit module, that is, arbitrarily initialize a set of learnable sub-action unit features Among them, M is the number of sub-action units, and E is the feature dimension of the sub-action unit. The gradient descent method is used for updating during the training process. It is mainly divided into the following two parts: the first part is to model the relationship between the sub-action unit and the known action category, and the second part is to model the relationship between the sub-action unit and the video segments at different times.

[0087] Step 41: Model the relationship between the sub-action units and the known action categories; first, calculate the adjacency matrix between the sub-action units Use cosine similarity to calculate the similarity between two different sub-action units. The expression is as follows:

[0088]

[0089] in, is the adjacency matrix E sub The element in the i-th row and j-th column of i , a j They are different sub-action unit characteristics;

[0090] Based on the adjacency matrix E sub Construct the edges of the sub-action unit graph convolutional network; in order to show the differences of the sub-action units, different weights are used for the neighbor connections and self-connections of the sub-action units. The expressions are as follows:

[0091]

[0092] in, yes The Laplace regularization operation of yes The degree matrix of . To learn the parameters,

[0093] therefore, is the relationship matrix between C+1 action categories and M sub-action units; then, the softmax activation operation is implemented along the category dimension, and the expression is as follows:

[0094]

[0095] in, Represents the activation value of different actions on each action unit, j=1,…,M; c=1,…,C,C+1.

[0096] Step 42, matching modeling of sub-action units and video temporal features;

[0097] Assume that all actions in the dataset are composed of a set of sub-action units A subThe activation value of each sub-action unit in the video clip at different moments is obtained by calculating the correlation between the video clip features and the sub-action unit features at different moments in the video. The expression is as follows:

[0098] S sub =A sub X emb (12)

[0099] in, represents the activation score of the jth sub-action unit in the tth video clip, t∈{1,2,…,T}, j∈{1,2,…,M}.

[0100] Softmax function activation is performed along the sub-action category dimension, that is,

[0101] The sub-action units in the video action are modeled through step four, that is, the semantic association relationship between the sub-action unit and the known action is learned through the graph convolutional network, and a relationship model of the sub-action unit in the action is formed. The known class and unknown class actions are identified by the difference in the activation degree of the action to be detected on the sub-action unit; secondly, considering that the timing information of the sub-action unit has the characteristics of distinguishing known class actions from unknown class actions, a temporal relationship model of the sub-action unit in the action is constructed to distinguish and detect actions with different timing relationships due to the same semantic relationship pattern, which can more effectively detect known and unknown class actions and more effectively solve the action detection problem in open scenes.

[0102] Step 5: Activate the sub-action sequence generated in the sub-action unit modeling branch As the input of the relational expression modeling branch, the action semantic relation expression and temporal relation expression modeling based on sub-action units are performed respectively;

[0103] First, the activation scores of the video on M sub-action units are calculated, and the weighted summation along the temporal dimension is performed using action attention that is independent of the category. The formula is as follows:

[0104]

[0105] in,

[0106] Similarly, the softmax function is used to activate and normalize it, and we get Use the relationship matrix between known actions and sub-actions The activation score of the input video on the sub-action unit is converted into a predicted value on the action category. The expression is as follows:

[0107]

[0108] in, Finally, the cross entropy classification loss function is used, and the expression is as follows:

[0109]

[0110] in, is the video-level label after regularization.

[0111] Step 6: In order to make the classification scores in the baseline branch and the classification scores in the sub-action-based semantic relationship expression branch mutually constrained, a consistency loss function L is introduced. cons , requiring video-level classification loss through the action relation module Classification loss with baseline branch The activation scores on the learned action categories remain consistent and are expressed as follows:

[0112]

[0113] in, is the class activation score in the baseline branch, is the class activation score in the sub-action unit relation module; the consistency loss function L cons The difference between them should be as small as possible;

[0114] Step 7: In order to allow the sub-action unit to tolerate the changes in the external environment and the speed of the sub-action unit itself as much as possible, and to effectively distinguish actions with similar motion patterns but different timing relationships, a temporal relationship expression model based on the sub-action unit is constructed;

[0115] For the same type of action, the associations of its sub-action units at different times are relatively similar, while the sub-action units of different types of actions are quite different; and in each batch, for each video sample, its video embedding feature X emb and sub-action unit feature A sub The correlation value between The expression is:

[0116]

[0117] In the same batch, the samples with the same category are positive samples, and the samples with different categories are negative samples. Under the constraints of the dynamic time warping (DTW) algorithm, the DTW distances between action samples of the same category are closer, while the DTW distances between action samples of different categories are farther, resulting in a more discriminative feature S. sub , the DTW loss function expression is as follows:

[0118]

[0119] Step 8: In order to make the model learn more robust sub-action unit features, the present invention adds representation loss and balance loss to jointly train the model. The representation loss requires the model to learn a variety of sub-action units. It is expected that each clip in the video corresponds to only one sub-action unit. The representation loss function is as follows:

[0120]

[0121] Where ⊙ represents the Hadamard product, is the relationship matrix The cth column of For the matrix The element of the tth row of For the matrix The element in row t and column c. By subtracting the maximum value in the Hadamard product, the purpose is to encourage the activation score of the t-th action clip in the c-th action category to be obtained by the activation of the t-th action clip on a single action unit.

[0122] The balance loss is used as a supplement to the representation loss function to avoid the model from generating redundant sub-action unit fragments and to emphasize that sub-action units of the same action should appear in the corresponding video at the same time; the balance loss function is as follows:

[0123]

[0124] in, represents the activation score of the video on action category c, Represents the video-level activation score for each sub-action unit in the video. represents the activation score of the c-th action on the j-th sub-action unit. Similarly, The goal of using the balanced loss function is to make the activation score on a known action category c the average of the contributions of all corresponding sub-action units.

[0125] Step 9. Combining the total loss function of the baseline branch and the sub-action unit branch, as well as the representation loss and balance loss for the sub-action unit, the total loss function is:

[0126]

[0127] Among them, α1, α2 and α3 are hyperparameters used to balance various loss functions.

[0128] Using the trained model, the test video is used to classify and locate the action, and the class activation sequence generated by the recognition of sub-action semantic association expression is used to classify and locate the action, that is, First, the threshold θ calculated based on the known training data distinguishes the test video as a known class action or an unknown class action (θ ensures that 95% of the training videos are identified as known class actions); when the video-level action class activation score If θ is greater than θ, the video is considered to contain known actions, otherwise it is an unknown action, thereby identifying the unknown actions in the test set; then, in the known action category, the classification score is greater than the preset threshold θ cls The action categories are retained for action localization; for the retained categories, the localization threshold θ is set act , the class activation score exceeds the threshold θ act The continuous segments of the proposed action are combined to form candidate action proposals. For unknown actions, the action attention λ that is independent of the category is directly used to decide. Finally, the non-maximum suppression method is used to delete the action predictions with high overlap in the candidate proposals to obtain the final action detection result.

[0129] Specific experiments:

[0130] Dataset and evaluation indicators:

[0131] THUMOS14 is the most popular dataset for action detection tasks; therefore, the present invention conducts experiments on the dataset THUMOS14; THUMOS14 contains 200 verification videos and 213 test videos, with a total of 20 categories of actions, and each video contains about 15.5 action instances on average; wherein, the verification video is used for model training, and the test video is used for model reasoning test; in the experiment, 1 / 4 of the action categories in the THUMOS14 verification set are randomly erased, and the THUMOS14 test set is retained for open set action detection evaluation, and the random erasure is repeated 3 times; three evaluation indicators are used to measure the performance of the method of the present invention.

[0132] 1. Ordinary mAP, that is, the mean average precision (mAP) under different temporal intersection over union (t-IoU) thresholds is evaluated. For the THUMOS14 dataset, the t-IoU threshold ranges from 0.1 to 0.7 with a step size of 0.1.

[0133] 2. Top-KmAP: In order for the model to generate a limited number of relatively accurate action proposals rather than a large number of redundant proposals, the maximum number of generated proposals is limited by selecting proposals with top-K confidence scores; when K is not restricted, top-KmAP will be downgraded to ordinary mAP; in the THUMOS14 dataset, K is set to 10, 20, 50, and 100 respectively.

[0134] 3. Classification index: The present invention adopts the following four classification indexes:

[0135] a) AU-ROC: The area under the ROC (Receiver Operating Characteristic Curve) curve. The ROC curve is also called the receiver operating characteristic curve. It uses the false positive rate FPR (False Positive Rate) as the horizontal axis and the true positive rate TPR (True Positive Rate) as the vertical axis, and is used to predict the accuracy rate.

[0136] b) AU-PR: The area under the PR (Precision-Recall Curve) curve, with recall as the horizontal axis and precision as the vertical axis. These two indicators are used to evaluate the performance of detecting unknown action categories from known action categories in correct action localization.

[0137] c) FAR@95: In order to achieve practical operational significance, the FAR@95 value is evaluated, that is, the false alarm rate when the true positive rate reaches 95%, that is, the number of unknown actions that are misclassified as known classes. The smaller the value, the better the detection performance.

[0138] d) OSDR: The OSDR (Open Set Detection Rate) curve is the area under the curve with FPR (False Positive Rate) as the horizontal coordinate and CDR (Correct Detection Rate) as the vertical coordinate. CDR represents the part of the known action category that is correctly located and classified; FPR represents the part of the unknown action category that is incorrectly classified into the known action category but is correctly located.

[0139] Experimental setup:

[0140] The I3D network pre-trained on the Kinetics400 dataset is used to extract the RGB features and optical flow features of the video, and the feature dimensions are 1024 respectively. After splicing, the video features with dimension D of 2048 are obtained; in the embedding feature module, it is composed of a two-layer temporal convolutional network, and the output dimension is 1024; in the temporal attention module, a two-layer temporal convolution layer is also used, and the output dimensions are 256 and 1 respectively; a Dropout layer is added before the activation layer; in the training stage, the video length uses stratified random sampling, and in the test stage, uniform sampling is used, and the length sampling of each video is T; in this embodiment, T is 750; in the baseline branch of the model training stage, in the top-k averaging strategy, k is selected as 100; the number of sub-action units in the training set THUMOS14 dataset is set to 80; in the inference process, the classification threshold θ of the known action category is set cls=0.2; positioning threshold θ act =[0:0.25], and the step size is 0.025; hyperparameters α1=α2=0.1, α3=1; the Adam optimizer with a learning rate of 1e-4 is used for model optimization, and the number of training times is 200; this experimental environment is implemented using PyTorch, in the Python3.9.12 environment, running on the Ubuntu 20.04 system, equipped with a 24GB GeForce RTX 3090Ti GPU.

[0141] Performance comparison:

[0142] Table 1 shows the performance comparison between the proposed method and the currently only open set weakly supervised sequential action detection method CELL and closed set advanced methods;

[0143] Table 1 Performance comparison with related advanced closed-set and open-set weakly supervised action detection methods on the THUMOS14 dataset (%)

[0144]

[0145] From the results, it can be seen that the average mAP of the method of the present invention is 0.3% higher than that of CELL; since CELL does not give the detection performance of the model at each t-IoU threshold in detail, it is impossible to directly compare whether the method of the present invention performs better when the t-IoU threshold is lower or higher; therefore, the performance of the present invention is significantly different from that of the closed-set advanced method StochasticFormer; but this is reasonable, because under the open-set setting, the test video contains unknown action categories, and there is no label supervised learning, which is more difficult; then, compared with CMCS in the closed-set method, the detection performance is comparable; specifically, when the t-IoU threshold is lower, the performance of the proposed method is higher, which means that most action instances can be detected; when the t-IoU threshold is higher, the performance of the method is reduced, which means that the detected action instances are not complete enough; since a complete action is composed of one or more sub-actions, this shows that the sub-action unit perception modeling used in the proposed method can detect all action instances as much as possible.

[0146] Table 2 is a comparison of the OWTAL positioning performance of the method of the present invention on the test set THUMOS14 dataset, Top-K mAP and traditional mAP (%) on known action categories and unknown action categories;

[0147] Table 2 Performance comparison of Top-KmAP@Avg and traditional mAP@Avg on THUMOS14 dataset (%)

[0148]

[0149] Among them, the t-IoU threshold is 0.1 to 0.7, and the step size is 0.1; when the t-IoU threshold is the average value from 0.1 to 0.7, the detection performance mAP (%) on known action categories and unknown action categories; From the results, it can be seen that the detection performance of the proposed method in the Top-10 is slightly lower than that of CELL for known action categories, which may be due to the fact that the proposed method mainly focuses on how to identify unknown action categories, while ignoring the improvement of the detection performance of known action categories on the closed set; but the detection performance on unknown action categories is 0.2% higher, indicating that the proposed method has enhanced the ability to identify unknown action categories; As the K value in top-K gradually increases, when K = 100, the average detection performance of the proposed method on both known and unknown categories is higher than that of CELL; Therefore, on the traditional mAP, the detection performance of this method in known action categories is improved by 0.5%, and the detection performance in unknown categories is improved by 0.3%. This further verifies that the relationship modeling of sub-action units is conducive to improving the model's ability to identify unknown actions and the ability to locate action instances.

[0150] Table 3 shows the classification performance comparison results of the method of the present invention and the latest open set detection method;

[0151] Table 3 Classification performance comparison with the latest open-set temporal action detection methods on the THUMOS14 dataset

[0152]

[0153] The data in the table are the detection performance when t-IoU is 0.5; from the results, it can be seen that the FAR@95 performance of the method of the present invention is 3.55% lower than that of CELL, which shows that the proposed method can better identify unknown action categories, and also verifies that the relationship modeling based on sub-action units and temporal relationship modeling are helpful in identifying unknown actions; similarly, the two columns of AU-ROC and AU-PR also show the ability of the model to detect unknown categories from known action categories; according to the results, it can be seen that AU-ROC and AU-PR are improved by 1.1% and 0.78% respectively; this shows that the finer-grained sub-action units in this method are more diverse than the coarse-grained known action categories, and can better identify known and unknown actions; and OSDR is 0.69% lower than that of CELL, indicating that the classification accuracy of this method on known action categories has decreased, which may be because this method pays more attention to the improvement of open set detection performance, while ignoring the optimization of closed set action recognition performance.

[0154] Based on the above ideal embodiments of the present invention, the relevant staff can make various changes and modifications without departing from the technical concept of the present invention through the above description. The technical scope of the present invention is not limited to the contents of the specification, and its technical scope must be determined according to the scope of the claims.

Claims

1. An open set weakly supervised temporal action detection method based on sub-action unit perception, characterized in that: The following steps are involved: Step 1: Serialize the video frames, extract video clip features from the serialized video frames, and then extract embedded features; Step 2: Use the classifier to predict the segment-level class activation sequence of the embedded features; calculate the probability that the video belongs to each category, and calculate the cross entropy loss function between the probability of the predicted action category and the true value; Step 3: Use the evidence deep learning loss function to calculate the probability of actions belonging to known categories, and combine it with the cross entropy loss function of the video category to obtain the loss function in the baseline branch; Step 4: construct the relationship between the sub-action unit and the known action category, and the relationship module between the sub-action unit and the video segments at different times; Step 5: Use the activation value on each sub-action unit as the input of the relational expression modeling branch to perform action semantic relation expression and temporal relation expression modeling based on the sub-action unit respectively; Step 6: Calculate the consistency classification loss function based on the class activation scores in the baseline branch and the class activation scores in the sub-action unit relationship; Step 7: Construct a temporal relationship expression model based on sub-action units; Step 8: Introduce the representation loss function and balance loss function training model; Step 9. Construct a total loss function based on the baseline branch, sub-action unit branch, representation loss and balance loss.

2. The open set weakly supervised temporal action detection method based on sub-action unit perception according to claim 1 is characterized in that: The formula of the loss function in the baseline branch is: in, is the cross entropy loss function of the video category, L edl is the loss function for evidence deep learning.

3. The open set weakly supervised temporal action detection method based on sub-action unit perception according to claim 1 is characterized in that: The construction of the relationship between sub-action units and known action categories includes: Calculate the adjacency matrix between sub-action units Use cosine similarity to calculate the similarity between two different sub-action units; Based on the adjacency matrix E sub Construct the edges of the sub-action unit graph convolutional network; use different weights Q for the neighbor connections and self-connections of the sub-action unit; Implement the softmax activation operation along the category dimension, expressed as: in, Represents the activation value of different actions on each action unit, j=1,…,M; c=1,…,C,C+1.

4. The open set weakly supervised temporal action detection method based on sub-action unit perception according to claim 3 is characterized in that: The construction of the relationship between the sub-action unit and the video clips at different moments includes: Get the activation value S of each sub-action unit in the video clips at different times sub ; Softmax function activation is performed along the sub-action category dimension, 5. The open set weakly supervised temporal action detection method based on sub-action unit perception according to claim 1 is characterized in that: Step 5 specifically includes: Calculate the activation scores of the video on M sub-action units, use action attention that is independent of the category to perform weighted summation along the temporal dimension, and then use softmax for normalization; Using the Relationship Matrix The activation score of the input video on the sub-action unit is converted into a predicted value on the action category, and the corresponding cross entropy classification loss function is calculated. The formula is: in, is the video-level label after regularization, The activation scores of the input video on the sub-action unit are converted into prediction values ​​on the action category.

6. The open set weakly supervised temporal action detection method based on sub-action unit perception according to claim 1 is characterized in that: The formula of the consistency loss function is: in, is the class activation score in the baseline branch, is the class activation score in the sub-action unit relation module.

7. The open set weakly supervised temporal action detection method based on sub-action unit perception according to claim 1 is characterized in that: Step 7 specifically includes: Calculate video embedding feature X emb and sub-action unit feature A sub The correlation value between The expression is: calculate The DTW distance is obtained and the DTW loss function is obtained.

8. The open set weakly supervised temporal action detection method based on sub-action unit perception according to claim 1 is characterized in that: The formula for the total loss function is: Among them, α1, α2 and α3 are hyperparameters; L rep is the loss function; L ban is the balance loss function; L dtw is the DTW loss function; L cons is the consistency loss function; L base is the loss function in the baseline branch; is the cross entropy classification loss function.

9. An open set weakly supervised temporal action detection system based on sub-action unit perception, characterized in that: include: a memory for storing instructions executable by a processor; A processor, configured to execute instructions to implement the open set weakly supervised temporal action detection method based on sub-action unit perception as described in any one of claims 1-8.

10. A computer readable medium storing computer program code, characterized in that: When the computer program code is executed by a processor, the computer program code implements the open set weakly supervised temporal action detection method based on sub-action unit perception as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Weak supervision time sequence action detection method based on space-time correlation learning

    CN114821772A

  • Weak supervision time sequence action detection method based on cascade attention mechanism

    CN116310988A