A Single-Frame Supervised Method and System for Temporal Action Detection and Classification in Videos
By decomposing the video timing action detection and classification tasks into a two-stage framework, using action seed frame detection and action center length parameters, the problem of large solution space and insufficient supervision in single-frame supervised video timing action detection is solved, and more accurate action detection and classification is achieved.
Patent Information
- Application Number
- CN202111190861.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-13
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2041-10-13
AI Technical Summary
In the existing single-frame supervised video timing action detection methods, there are problems such as too large solution space and insufficient supervision, which leads to inaccurate action detection and easy detection of action into scattered segments, making it difficult to meet the actual application needs.
Using the strategy of division and conquer, the video timing action detection and classification tasks are broken down into a two-stage framework. First, the video is divided into a single instance video clip through action seed frame detection, and then the action example is characterized by the action center and length parameters, and action detection and classification are performed.
It significantly improves the accuracy and completeness of action detection, reduces the difficulty of model processing, and makes it easier to obtain accurate detection and classification results.
Smart Images

Figure CN113936174B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer vision and image processing. Specifically, it relates to a method and system for single-frame supervised video temporal action detection and classification, and more specifically, to a method and system for two-stage single-frame supervised video temporal action detection and class prediction based on a divide-and-conquer strategy. Background Art
[0002] With the rapid popularization of 5G communication technology, multimedia videos have seen explosive growth, and video content understanding and retrieval technologies have thus become new research hotspots. Action is the core of video content understanding. Temporal action detection in the video time dimension provides strong assistance for downstream short video recommendation, highlight retrieval, and long video editing, and has therefore received extensive attention from researchers.
[0003] The video temporal action detection task is as follows: Given a large-scale unedited long video, it is divided into two groups. One group is labeled with the action occurrence time and action category to form a training dataset, and the other group forms a test dataset. Based on the training dataset, a generalizable model is learned, which can detect the start and end time positions of actions in the test dataset and predict the action category at the same time. Each test video may contain multiple action categories and an indefinite number of action instances, and the model needs to detect all action instances and classify them correctly.
[0004] There are three existing settings for this task: strong supervision, weak supervision, and single-frame supervision. The strong supervision setting requires labeling the action categories contained in each long video in the training dataset and the start and end time positions of each action instance. Although the performance of the model under this setting is already very good, the required precise action boundary labeling is time-consuming, laborious, and extremely expensive, making it difficult to collect training data on a large scale and limiting the actual application scenarios. On the other hand, the weak supervision setting only requires labeling the action categories contained in each long video in the training dataset to perform action position detection and class prediction. However, due to the lack of explicit position annotation information, the performance of this setting lags far behind the full supervision setting and it is difficult to meet the requirements of actual applications. To balance the annotation cost and model performance, single-frame supervision emerged. Its setting is between strong supervision and weak supervision. For each video in the training dataset, a single-frame timestamp needs to be randomly labeled for each action instance, and the action category is given at the same time. Compared with the strong supervision setting, its annotation cost is very low, so it is easy to obtain a large amount of labeled data; compared with the weak supervision setting, it provides a stronger prior for the number of actions and a rough prior for the action position, so the detection performance is better and it is convenient for actual applications.
[0005] With the continuous development of deep neural networks and multi-instance learning in the field of weak supervision, the vast majority of weakly supervised temporal action detection adopts a single-stage detection framework, that is, directly predicting the action probability for each frame of the video, and generating the final detection result by post-processing the frame-level action probability. Existing single-frame supervised temporal action detection basically also follows this single-stage detection framework.
[0006] After retrieval, the Chinese invention patent application with the publication number CN112926492A and the application number CN202110291231.2 discloses a temporal behavior detection method and system based on single-frame supervision. It uses a video foreground and background mining module to expand the single-frame position supervision, and then uses the expanded local action position labels to train the single-stage detection framework to generate temporal detection results. However, the above patent does not consider that the solution space corresponding to the single-stage detection framework is too large, and ignores that the expanded local action position labels are still sparse and noisy. Using such insufficient and inaccurate position labels to handle complex detection tasks with an uncertain number of action instances will inevitably introduce incorrect background misactivations, and it is very easy to detect actions as scattered fragments, which limits the performance under the single-frame supervision setting. Summary of the Invention
[0007] Aiming at the defects in the prior art, the purpose of the present invention is to provide a single-frame supervised video temporal action detection and classification method and system.
[0008] In the first aspect of the present invention, a single-frame supervised video temporal action detection method is provided, including:
[0009] For the input long video, extract the video feature map;
[0010] Regarding the given action single frame as the seed frame, map the video feature map to the action seed frame probability map, and obtain several action seed frame positions;
[0011] Based on the action seed frame positions, divide the long video into several single-instance video segments, and calculate the corresponding single-instance video segment feature maps according to the time positions of the single-instance video segments;
[0012] Map any of the single-instance video segment feature maps to an action position proposal, which is composed of an action center and an action length;
[0013] Generate an action detection result based on the action position proposal.
[0014] Optionally, the step of extracting the video feature map for the input long video includes:
[0015] Extract optical flow motion information from the RGB data of long videos, and use a feature encoding network to map the RGB data and optical flow data into visual feature maps of dimension T*D respectively; where T represents the time length of the video, and D represents the feature dimension of the video;
[0016] Then concatenate the RGB and optical flow features in the feature dimension to generate a fused video feature map F, whose dimension is T*2D.
[0017] Optionally, map the video feature map to an action seed frame probability map, and obtain several action seed frame positions, including:
[0018] Use an action seed frame detection network composed of fully convolutional layers to map the video feature map F to an action seed frame probability map S of dimension T*1, where S represents the probability that each video frame belongs to the seed frame, and T represents the time length of the video;
[0019] Use action seed frame labels to supervise the action seed frame probability map S, calculate the loss function to train the action seed frame detection network until the loss function converges;
[0020] For unannotated test set videos, obtain several action seed frame positions by post-processing the action seed frame probability map S.
[0021] Optionally, calculate the corresponding single-instance video segment feature map according to the time position of the single-instance video segment, and align the time length of the single-instance video segment feature to T by linear interpolation W , that is, the single-instance video segment feature is denoted as F W , whose dimension is T W *2D, where D represents the feature dimension of the video.
[0022] Optionally, mapping any of the single-instance video segment feature maps to action location proposals includes:
[0023] Use an action detection network composed of convolutional layers and fully connected layers to map any single-instance video segment feature map F W to action location proposals P=(C,L), where C represents the predicted action center and L represents the predicted action length.
[0024] In the second aspect of the present invention, a single-frame supervised video temporal action classification method is provided, including:
[0025] Determine action location proposals; specifically, the techniques in the above single-frame supervised video temporal action detection method can be adopted;
[0026] Map the action location proposals to a temporal action location mask, indicating the time positions of the video frames containing actions in the video;
[0027] Calculate all action features and background features in a single-instance video clip based on the temporal action location mask, and aggregate them into video-level action features and video-level background features;
[0028] Use a classification network to map the video-level action features to action class probabilities, map the video-level background features to background class probabilities, and calculate a loss function with the action class label and background class label to train the classification network until the loss function converges;
[0029] Input a test video into the trained classification network, output the action class probabilities, and use a threshold method to generate action class predictions for the action class probabilities.
[0030] Optionally, calculating all action features and background features in a single-instance video clip based on the temporal action location mask and aggregating them into video-level action features and video-level background features includes:
[0031] Let the temporal action location mask be M and the single-instance video clip be F W All action features in and background features The calculation formulas are as follows:
[0032]
[0033] Then, perform averaging on separately in the time dimension and aggregate them into video-level action features and video-level background features Both of them have a dimension of 1*2D.
[0034] Optionally, when using the classification network to map the video-level action features to action class probabilities and map the video-level background features to background class probabilities, where:
[0035] The classification network is a classification network composed of fully connected layers, which maps the video-level action features to action class probability p A and maps the video-level background features to background class probability p B and calculates a loss function with the action class label and background class label to train the classification network until the loss function converges.
[0036] In the third aspect of the present invention, a single-frame supervised video temporal action detection system is provided, including:
[0037] Video feature map extraction module: Extract video feature maps for the input long video;
[0038] Action seed frame detection module: Regarding a given single frame of an action as a seed frame, mapping the video feature map to an action seed frame probability map, and obtaining several action seed frame positions;
[0039] Action instance division module: Based on the action seed frame positions, dividing the long video into several single-instance video segments, and calculating the corresponding single-instance video segment feature maps according to the time positions of the single-instance video segments;
[0040] Action position detection module: Mapping any of the single-instance video segment feature maps to an action position proposal, which consists of an action center and an action length;
[0041] Detection result generation module: Generating an action detection result based on the action position proposal.
[0042] In a fourth aspect of the present invention, a single-frame supervised video temporal action classification system is provided, including:
[0043] Action position proposal determination module: For an input long video, determining an action position proposal; specifically, the techniques in the above-mentioned single-frame supervised video temporal action detection method can be adopted;
[0044] Temporal mask generation module: Mapping the action position proposal to a temporal action position mask, indicating the time positions of the video frames containing actions in the video;
[0045] Action-background separation module: Based on the temporal action position mask, calculating all action features and background features in the single-instance video segment, and aggregating them into video-level action features and video-level background features;
[0046] Action-background classification module: Using a classification network to map the video-level action features to action category probabilities, mapping the video-level background features to background category probabilities, and calculating a loss function with the action category label and the background category label, and training the classification network until the loss function converges;
[0047] Classification result generation module: For the action category probabilities, using a threshold method to generate action category predictions.
[0048] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0049] 1. In the above detection method and classification method of the present invention, specifically for single-frame supervision, a divide-and-conquer strategy is carefully designed to simplify the time-series action detection and classification tasks. Specifically, the overall task is the action detection and classification of long videos. The present invention designs a two-stage framework. In the first stage, it is decomposed into several single-instance video segments, and in the second stage, the action detection and classification subtasks of single-instance video segments are performed, significantly reducing the processing difficulty of the model and making it easier to obtain accurate and complete detection and classification results.
[0050] 2. In order to decompose a long video into several single-instance video segments, the present invention makes full use of single-frame action position labels. Regarding them as seed-frame supervision, an action seed-frame detection network is trained. Through the number and position of seed frames, the number and rough position of action instances in the long video can be determined more reliably and effectively.
[0051] 3. In order to overcome the problems of large solution space and insufficient supervision faced by the existing single-stage frame-level action probability prediction paradigm, the present invention adopts a proposal-level action prediction paradigm. By using two parameters, the action center and length, to characterize an action instance, the degree of freedom of the action prediction paradigm is significantly reduced, and at the same time, the probability smoothing prior inside and outside the action is increased. Under the condition of insufficient supervision, it is easier to find complete and accurate detection and classification results. BRIEF DESCRIPTION OF THE DRAWINGS [[ID=lo]]
[0052] Other features, objectives, and advantages of the present invention will become more apparent by reading the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0053] Figure 1 It is a flowchart of the method for single-frame supervised video time-series action detection and classification according to an embodiment of the present invention;
[0054] Figure 2 It is a schematic diagram of the system for single-frame supervised video time-series action detection and classification according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0055] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that those of ordinary skill in the art can make several changes and improvements without departing from the concept of the present invention. These all belong to the protection scope of the present invention.
[0056] The present invention makes specific settings for single-frame supervision and carefully designs a divide-and-conquer strategy to simplify the time-series action detection and classification tasks.
[0057] Figure 1 It is a flowchart of the method for single-frame supervised video time-series action detection and classification according to an embodiment of the present invention.
[0058] Referring to Figure 1 as shown, in an embodiment of the present invention, the single-frame supervised video temporal action detection method provided includes the following steps:
[0059] S11, video feature map extraction step: For the input long video, use a 3D depth convolutional feature encoding network to extract a video feature map of a preset dimension;
[0060] S12, action seed frame detection step: Regard the given single action frame as the seed frame, use a seed frame detection network composed of fully convolutional layers to map the video feature map to an action seed frame probability map of a preset dimension, and calculate the loss function with the single-frame action position label. Then post-process the action seed frame probability map to obtain several action seed frame positions;
[0061] S13, action instance division step: Based on the action seed frame positions, divide the long video into several single-instance video segments, and calculate the corresponding single-instance video segment feature maps according to the time positions of the single-instance video segments;
[0062] S14, action position detection step: For any single-instance video segment feature map, use an action detection network composed of convolutional layers and fully connected layers to map it to an action position proposal, which consists of an action center and an action length;
[0063] S15, detection result generation step: Generate an action detection result based on the action position proposal. The action proposal consists of an action center and an action length. Based on these two points, the start time and end time of the action can be calculated. The start and end times are the targets of action detection.
[0064] In some embodiments, when performing the video feature map extraction step of S11, the following operations can be carried out: Extract the optical flow motion information from the RGB data, and use a feature encoding network composed of convolutional layers and fully connected layers to map the RGB data and the optical flow data to visual feature maps of T*D dimensions respectively. Wherein, T represents the time length of the video, and D represents the feature dimension of the video. Then concatenate the RGB and optical flow features in the feature dimension to generate a fused video feature map F, whose dimension is T*2D. The video feature map of the set dimension is obtained in this step.
[0065] In some embodiments, when performing the action seed frame detection step of S12, the following operations can be carried out: regarding a given single action frame as a seed frame, using an action seed frame detection network composed of fully convolutional layers to map the video feature map F into an action seed frame probability map S with a dimension of T*1, which represents the probability that each video frame belongs to the seed frame. Supervising the action seed frame probability map S with action seed frame labels, the loss function can be calculated to train the action seed frame detection network until the loss function converges; specifically, in this embodiment, the calculation formula of the loss function is as follows:
[0066]
[0067] where θ E are the parameters of the feature encoding network, θ SD are the parameters of the action seed frame detection network, (X S , Y S ) represents the distribution of the input video and the single-frame action position label, x i represents the video instance, y i is its single-frame action position label, E represents the feature encoding network, SD represents the action seed frame detection network, and H represents the cross-entropy function. The action seed frame labels can be generated by manual annotation.
[0068] For the unlabeled test set videos, several action seed frame positions can be obtained by post-processing the action seed frame probability map S. Since there is sufficient supervision for the action seed frames in the video, by detecting the seed frame positions and quantities, the number and approximate positions of action instances in the long video can be determined. Using the seed frames to divide the video into several single-instance video segments is essentially to decompose the task of video temporal action detection with an indefinite number into single-temporal action detection tasks, simplifying the problem.
[0069] In some embodiments, when performing the action instance division step of S13, the following operations can be included: based on the action seed frame positions, dividing the long video into several single-instance video segments. For a given action seed frame s t , its corresponding single-instance video segment position is [s t-1 , s t+1 , where s t-1 and s t+1 respectively represent the positions of the adjacent action seed frames before and after s t . Then, according to the positions of the single-instance video segments, the corresponding single-instance video segment feature maps are obtained, and the time lengths of the single-instance video segment features are aligned to T W by linear interpolation. That is, the single-instance video segment feature is denoted as F W , and its dimension is T W*2D. In this embodiment, in order to decompose a long video into several single-instance video segments, the single-frame action position tags are fully utilized.
[0070] In some embodiments, when performing the action position detection step of S14, the following operations may be included: For any single-instance video segment feature map F W , use an action detection network composed of a convolutional layer and a fully connected layer to map it to an action position proposal P = (C, L), where C represents the predicted action center and L represents the predicted action length. Compared with predicting the frame-level action probability, directly predicting the action proposal has a smaller solution space, and only two variables, the action center and the length, are used to characterize the action instance. In addition, the action position proposal constrains that the action probability within the action interval is 1 and the action probability outside the action interval is 0, so it naturally has strong action probability smoothness. In this embodiment, by using two parameters, the action center and the length, to characterize an action instance, the degree of freedom of the action prediction paradigm is significantly reduced, and it is easier to find complete and accurate detection results under the condition of insufficient supervision.
[0071] Refer to Figure 1 As shown, in another embodiment of the present invention, a single-frame supervised video temporal action classification method is further provided, including:
[0072] S21, action position proposal determination step: For the input long video, determine the action position proposal, which is composed of the action center and the action length;
[0073] S22, temporal mask generation step: Use a neural network composed of a fully connected layer to map the action position proposal to a temporal action position mask, indicating the time positions of the video frames containing actions in the video;
[0074] S23, action-background separation step: Based on the temporal action position mask, calculate all the action features and background features in the single-instance video segment, and aggregate them into video-level action features and video-level background features;
[0075] S24, action-background classification step: Use a classification network composed of a fully connected layer to map the video-level action features to action category probabilities, map the video-level background features to background category probabilities, and calculate the loss function with the action category label and the background category label;
[0076] S25, classification result generation step: For the action category probability, use the threshold method to generate the action category prediction. Specifically, given a test video, input it into the trained action-background classification network, and output the action category probability. Then input the action category probability into S25 (constituted by the threshold method) to output the action category prediction of the test video.
[0077] In the above embodiments of the present invention, by using two parameters, i.e., the action center and the length, to characterize an action instance, the degree of freedom of the action prediction paradigm is significantly reduced, and at the same time, the prior probability of action internal and external probability smoothing is increased. Under the condition of insufficient supervision, it is easier to find complete and accurate classification results.
[0078] In some embodiments, when performing the above S21 action position proposal determination step, it may be performed with reference to S11 - S14 in the above detection method embodiments.
[0079] In some embodiments, when performing the above S22 timing mask generation step, it may include: using a neural network composed of fully connected layers to map the action position proposal P to a timing action position mask M, representing the positions of video frames containing actions in the video, with a dimension of T W *1;
[0080] In some embodiments, when performing the above S23 action background separation step, it may include: based on the timing action position mask M, calculating all action features W in the single-instance video segment F and background features The calculation formula is as follows:
[0081]
[0082] Then, respectively average in the time dimension, and aggregate them into video-level action features and video-level background features Both of their dimensions are 1 * 2D.
[0083] In some embodiments, when performing the above S24 action background classification step, it may include: using a classification network composed of fully connected layers to map the video-level action features to an action category probability p A , map the video-level background features to a background category probability p B , and calculate a loss function with the action category label and the background category label, and train the classification network until the loss function converges; in this embodiment, the loss function, the calculation formula is as follows:
[0084]
[0085]
[0086] where θ K is the parameter of the action classification network, (X W , Y W ) represents the distribution of the input video-level features and the action category label, Representing an instance of video-level action features, representing an instance of video-level background features, y j is its action category label, K represents the action classification network, and H represents the cross-entropy function. The output of the action classification network is the probability of predicting that a single-instance video clip belongs to each category of action or background, that is, By calculating the loss function between the predicted category probability and the category label of the video clip, the prediction model updates its parameters according to the loss function, thereby forcing the predicted category probability to gradually approach the category label. In this embodiment, by separating actions and backgrounds, the model can clearly learn to distinguish between actions and backgrounds, enabling more accurate action classification.
[0087] Refer to Figure 2 As shown, based on the same technical concept as above, another embodiment of the present invention further provides a single-frame supervised video temporal action detection system, including:
[0088] Video feature map extraction module: For the input long video, use a 3D depth convolutional feature encoding network to extract video feature maps of a preset dimension;
[0089] Action seed frame detection module: Treat the given single action frame as a seed frame, and use a seed frame detection network composed of fully convolutional layers to map the video feature map to an action seed frame probability map of a preset dimension, and calculate the loss function with the single-frame action position label. Then post-process the action seed frame probability map to obtain several action seed frame positions;
[0090] Action instance division module: Based on the action seed frame positions, divide the long video into several single-instance video clips, and calculate the corresponding single-instance video clip feature maps according to the time positions of the single-instance video clips;
[0091] Action position detection module: For any single-instance video clip feature map, use an action detection network composed of convolutional layers and fully connected layers to map it to an action position proposal, which consists of an action center and an action length;
[0092] Detection result generation module: Generate action detection results based on the action position proposals.
[0093] As a preferred embodiment, the video feature map extraction module includes: Extract optical flow motion information from RGB data, and use a feature encoding network composed of convolutional layers and fully connected layers to map the RGB data and the optical flow data to visual feature maps of T*D dimensions respectively. Among them, T represents the time length of the video, and D represents the feature dimension of the video. Then concatenate the RGB and optical flow features in the feature dimension to generate a fused video feature map F, whose dimension is T*2D.
[0094] As a preferred embodiment, the action seed frame detection module includes: regarding a given single action frame as a seed frame, and using an action seed frame detection network composed of fully convolutional layers to map the video feature map F into an action seed frame probability map S with a dimension of T*1, which represents the probability that each video frame belongs to the seed frame. Using the action seed frame label to supervise the action seed frame probability map S, the loss function can be calculated to train the action seed frame detection network until the loss function converges;
[0095] The loss function has the following calculation formula:
[0096]
[0097] where θ E are the parameters of the feature encoding network, θ SD are the parameters of the action seed frame detection network, (X S , Y S ) represents the distribution of the input video and the single-frame action position label, x i represents the video instance, y i is its single-frame action position label, E represents the feature encoding network, SD represents the action seed frame detection network, and H represents the cross-entropy function. For the unlabeled test set video, several action seed frame positions can be obtained by post-processing the action seed frame probability map S. Since there is sufficient supervision for the action seed frames in the video, by detecting the seed frame positions and quantities, the number and approximate positions of action instances in the long video can be determined. Using the seed frames to divide the video into several single-instance video segments is essentially to decompose the task of video temporal action detection with an indefinite number into single temporal action detection tasks, simplifying the problem.
[0098] As a preferred embodiment, the action instance division module includes: based on the action seed frame positions, dividing the long video into several single-instance video segments. For a given action seed frame s t , its corresponding single-instance video segment position is [s t-1 , s t+1 , where s t-1 and s t+1 respectively represent the positions of the action seed frames adjacent to s t before and after. Then, according to the positions of the single-instance video segments, the corresponding single-instance video segment feature maps are obtained, and the time lengths of the single-instance video segment features are aligned to T W through linear interpolation. That is, the single-instance video segment feature is denoted as F W , and its dimension is T W *2D.
[0099] As a preferred embodiment, the action position detection module includes: for any single-instance video segment feature map F W, the action detection network composed of convolutional layers and fully connected layers maps it to an action location proposal P = (C, L), where C represents the predicted action center and L represents the predicted action length. Compared with predicting frame-level action probabilities, directly predicting action proposals has a smaller solution space, and only two variables, the action center and length, are needed to characterize action instances. In addition, the action location proposal constrains the action probability within the action interval to be 1 and the action probability outside the action interval to be 0, so it naturally has strong action probability smoothness.
[0100] Referring to Figure 2 As shown, based on the above embodiments, another embodiment of the present invention further provides a single-frame supervised video temporal action classification system, including:
[0101] Action location proposal determination module: Determine an action location proposal, which is composed of an action center and an action length;
[0102] Temporal mask generation module: Use a neural network composed of fully connected layers to map the action location proposal to a temporal action location mask, indicating the temporal positions of video frames containing actions in the video;
[0103] Action-background separation module: Based on the temporal action location mask, calculate all action features and background features in a single-instance video segment, and aggregate them into video-level action features and video-level background features;
[0104] Action-background classification module: Use a classification network composed of fully connected layers to map the video-level action features to action category probabilities, map the video-level background features to background category probabilities, and calculate a loss function with the action category label and the background category label; both the action category label and the background category label can be generated by manual annotation;
[0105] Classification result generation module: For the action category probability, use the threshold method to generate an action category prediction. Specifically, given a test video, input it into the trained classification network of the action-background classification module, and output the action category probability. Then, use the threshold method for the action category probability to output the action category prediction of the test video.
[0106] The classification result generation module is independent of the action-background classification module. In other words, the action-background classification module consists of a classification network, which inputs video action / background features and outputs an action category probability (a decimal between 0 and 1, such as 0.45), and then inputs the action category probability into the action result generation module, which consists of the threshold method and outputs an action category prediction (an integer such as 0, 1).
[0107] As a preferred embodiment, the temporal mask generation module includes: a neural network composed of fully connected layers that maps the action location proposal P to a temporal action location mask M, representing the locations of video frames containing actions in the video, with a dimension of T W *1;
[0108] As a preferred embodiment, the action-background separation module includes: based on the temporal action location mask M, calculating all the action features W and background features in the single-instance video segment F The calculation formula is as follows:
[0109]
[0110] Then, perform averaging on separately in the time dimension, and aggregate them into video-level action features and video-level background features both with a dimension of 1*2D.
[0111] As a preferred embodiment, the action-background classification module includes: a classification network composed of fully connected layers that maps the video-level action features to the action class probability p A , maps the video-level background features to the background class probability p B , and calculates the loss function with the action class label and the background class label, and trains the classification network until the loss function converges;
[0112] For the said loss function, the calculation formula is as follows:
[0113]
[0114]
[0115] Among them, θ K is the parameter of the action classification network, (X W , Y W ) represents the distribution of the input video-level features and the action class label, represents the video-level action feature instance, represents the video-level background feature instance, y j is its action class label, K represents the action classification network, and H represents the cross-entropy function. The output of the action classification network is the probability of predicting that the single-instance video segment is an action or background of each category, that is, in the formula. By calculating the loss function between the predicted class probability and the video segment class label, the prediction model updates its parameters according to the loss function, thereby forcing the predicted class probability to gradually approach the class label.
[0116] In another embodiment of the present invention, a single-frame supervised video temporal action detection and classification method is further provided, including the following steps:
[0117] Video feature map extraction step, wherein: for the input long video, a 3D depth convolutional feature encoding network is used to extract video feature maps of a preset dimension. Each action instance in the video to be detected has only single-frame position annotation and action category annotation, without accurate action boundary position annotation.
[0118] Action seed frame detection step, wherein: the given single action frame is regarded as the seed frame, and a seed frame detection network composed of fully convolutional layers is used to map the video feature map into an action seed frame probability map of a preset dimension, and calculate the loss function with the single-frame action position label. Then, the action seed frame probability map is post-processed to obtain several action seed frame positions;
[0119] Action instance division step, wherein: based on the action seed frame positions, the long video is divided into several single-instance video segments, and the corresponding single-instance video segment feature maps are calculated according to the time positions of the single-instance video segments;
[0120] Action position detection step, wherein: for any single-instance video segment feature map, an action detection network composed of convolutional layers and fully connected layers is used to map it into an action position proposal, which consists of an action center and an action length;
[0121] Temporal mask generation step, wherein: a neural network composed of fully connected layers is used to map the action position proposal into a temporal action position mask, indicating the time positions of the video frames containing actions in the video;
[0122] Action-background separation step, wherein: based on the temporal action position mask, all action features and background features in the single-instance video segment are calculated and aggregated into video-level action features and video-level background features;
[0123] Action-background classification step, wherein: a classification network composed of fully connected layers is used to map the video-level action features into action category probabilities, map the video-level background features into background category probabilities, and calculate the loss function with the action category label and the background category label;
[0124] Detection and classification result generation step, wherein: an action detection result is generated based on the action position proposal. For the action category probability, a threshold method is used to generate an action category prediction.
[0125] Specifically, the single-frame supervised video temporal action detection network framework composed of a video feature map extraction module, an action seed frame detection module, an action instance division module, an action position detection module, a temporal mask generation module, an action-background separation module, an action-background classification module, and a detection and classification result generation module is as followsFigure 2 as shown
[0126] In the system framework of the embodiment as Figure 2 shown, the video to be detected is input into the video feature map extraction module, and the visual feature F of the video to be detected is output. Its feature dimension is T*2D, where T represents the time length of the video and 2D represents the feature dimension of the video. The video feature map extraction module is composed of a series of downsampling modules composed of 3D convolutional layers (+batchnorm layer + relu layer), and existing network structures can be used, such as two-stream I3D, TSN, C3D, etc. The visual feature F of the video to be detected will be input into the action seed frame detection module and mapped to an action seed frame probability map S with a dimension of T*1, which represents the probability that each video frame belongs to the seed frame.
[0127] Using the action seed frame label to supervise the action seed frame probability map S, the loss function can be calculated to train the action seed frame detection network until the loss function converges. The loss function is calculated as follows:
[0128]
[0129] where θ E is the parameter of the feature encoding network, θ SD is the parameter of the action seed frame detection network, (X S ,Y S ) represents the distribution of the input video and the single-frame action position label, x i represents the video instance, y i is its single-frame action position label, E represents the feature encoding network, SD represents the action seed frame detection network, and H represents the cross-entropy function. For the unlabeled test set video, several action seed frame positions can be obtained by post-processing the action seed frame probability map S. Since there is sufficient supervision for the action seed frames in the video, by detecting the seed frame positions and quantities, the number and approximate positions of action instances in the long video can be determined. Using the seed frames to divide the video into several single-instance video segments is essentially to decompose the task of detecting video temporal actions with an indefinite number into a single temporal action detection task, simplifying the problem.
[0130] To divide the input long video according to the detected action seed frame positions and divide it into several single-instance video segments. In the action instance division module, for a given action seed frame s t , its corresponding single-instance video segment position is [s t-1 ,s t+1 , where s t-1 and s t+1 respectively represent s tThe positions of the action seed frames adjacent before and after. Then, according to the position of the single-instance video segment, the corresponding single-instance video segment feature map is obtained, and the temporal length of the single-instance video segment feature is aligned to T by linear interpolation W That is, the single-instance video segment feature is denoted as F W , and its dimension is T W *2D
[0131] To detect the action positions in each single-instance video segment, an arbitrary single-instance video segment feature map F W is input into the action position detection module, and the action position proposal P=(C, L) is output, where C represents the predicted action center and L represents the predicted action length. The action position detection module is composed of a series of convolutional layers and fully connected layers. Compared with predicting the frame-level action probability, directly predicting the action proposal has a smaller solution space, and only two variables, the action center and length, are used to characterize the action instance. In addition, the action position proposal constrains that the action probability within the action interval is 1 and the action probability outside the action interval is 0, so it naturally has strong action probability smoothness
[0132] To convert the predicted action position proposal into an index in the temporal dimension, the action position proposal P is input into the temporal mask generation module, and the temporal action position mask M is output, representing the positions of the video frames containing actions in the video, and its dimension is T W *1. The temporal mask generation module is a neural network composed of fully connected layers
[0133] Based on the temporal action position mask M, calculate all the action features W and background features in the single-instance video segment F The calculation formula is as follows
[0134]
[0135] Then, we perform averaging on in the time dimension respectively, and aggregate them into the video-level action feature and the video-level background feature both with the dimension of 1*2D
[0136] To adjust the action position proposal using the action category label, as Figure 2 shown, the video-level action feature and the video-level background feature are input into the action-background classification module composed of fully connected layers, and the video-level action category probability p A and the video-level background category probability p B are output, and the loss function is calculated with the action category label and the background category label, and the classification network is trained until the loss function converges
[0137] The loss function has the following calculation formula:
[0138]
[0139]
[0140] where θ K are the parameters of the action classification network, and (X W , Y W ) represents the distribution of the input video-level features and action category labels, represents the video-level action feature instance, represents the video-level background feature instance, y j is its action category label, K represents the action classification network, and H represents the cross-entropy function. The output of the action classification network is the probability of predicting a single-instance video segment as each category of action or background, that is, in the formula. By calculating the loss function between the predicted category probability and the video segment category label, the prediction model updates its parameters according to the loss function, thereby forcing the predicted category probability to gradually approach the category label.
[0141] After the overall model training is completed, action detection results are generated based on the action location proposals. In addition, the predicted action category probabilities are input into the detection classification result generation module, which uses the threshold method on the category probabilities, and the action categories higher than the threshold constitute the final classification result.
[0142] In summary, the present invention carefully designs a divide-and-conquer strategy to simplify the temporal action detection task. Specifically, the overall task is the action detection and classification of long videos. The present invention designs a two-stage framework. In the first stage, it is decomposed into several single-instance video segments, and in the second stage, the action detection and classification subtasks of the single-instance video segments are performed, significantly reducing the processing difficulty of the model; in order to decompose the long video into several single-instance video segments, the present invention makes full use of the single-frame action location labels and regards them as seed frame supervision to train the action seed frame detection network. The number and rough location of action instances in the long video are determined more reliably and effectively through the number and location of the seed frames; in order to overcome the problems of large solution space and insufficient supervision faced by the existing single-stage frame-level action probability prediction paradigm, the present invention adopts the proposal-level action prediction paradigm. By using two parameters, the action center and the length, to characterize an action instance, the degree of freedom of the action prediction paradigm is significantly reduced, and at the same time, the action inside and outside probability smoothing prior is increased. Under the condition of insufficient supervision, it is ensured that the model can easily find more complete and accurate detection and classification results.
[0143] Those skilled in the art know that, in addition to implementing the systems, devices and their respective modules provided by the present invention in the form of pure computer-readable program codes, the method steps can be logically programmed to enable the systems, devices and their respective modules provided by the present invention to be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers and embedded microcontrollers, etc. to achieve the same program. Therefore, the systems, devices and their respective modules provided by the present invention can be considered as a kind of hardware components, and the modules included therein for implementing various programs can also be regarded as the structures within the hardware components; the modules for implementing various functions can also be regarded as either software programs for implementing the methods or the structures within the hardware components.
[0144] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.
Claims
1. A single-frame supervised video temporal action detection method, characterized in that Including: For the input long video, extract video feature maps; Regarding the given single action frame as a seed frame, map the video feature maps to an action seed frame probability map, and obtain several action seed frame positions; Based on the action seed frame positions, divide the long video into several single-instance video segments, and calculate the corresponding single-instance video segment feature maps according to the time positions of the single-instance video segments; Map any of the single-instance video segment feature maps to an action position proposal, which consists of an action center and an action length; Generate an action detection result based on the action position proposal; Map the video feature maps to an action seed frame probability map, and obtain several action seed frame positions, including: Use an action seed frame detection network composed of fully convolutional layers to map the video feature map F to an action seed frame probability map S with a dimension of T*1. S represents the probability that each video frame belongs to the seed frame, and T represents the time length of the video; Supervise the action seed frame probability map S using action seed frame labels, calculate the loss function to train the action seed frame detection network until the loss function converges; The calculation formula of the loss function is as follows: Among them, θ E is the parameter of the feature encoding network, θ SD is the parameter of the action seed frame detection network, (X S , Y S ) represents the distribution of the input video and the single-frame action position label, x i represents the video instance, y i is its single-frame action position label, E represents the feature encoding network, SD represents the action seed frame detection network, and H represents the cross-entropy function; For the unlabeled test set video, obtain several action seed frame positions by post-processing the action seed frame probability map S.
2. The single-frame supervised video temporal action detection method according to claim 1, wherein Regarding the extraction of video feature maps for the input long video, including: Extract optical flow motion information from the RGB data of the long video, and use a feature encoding network to map the RGB data and the optical flow data to visual feature maps with a dimension of T*D respectively; where T represents the time length of the video, and D represents the feature dimension of the video; Then concatenate the RGB and optical flow features in the feature dimension to generate a fused video feature map F with a dimension of T*2D.
3. The single-frame supervised video temporal action detection method according to claim 1, characterized in that Calculate the corresponding single-instance video clip feature map according to the time position of the single-instance video clip, and align the time length of the single-instance video clip features to T through linear interpolation W , that is, the single-instance video clip features are denoted as F W , and its dimension is T W *2D, where D represents the feature dimension of the video 4. The single-frame supervised video temporal action detection method according to claim 1, wherein Regarding the mapping of any of the single-instance video segment feature maps to an action position proposal, including: An action detection network composed of convolutional layers and fully connected layers maps an arbitrary single-instance video clip feature map F W to an action location proposal P = (C, L), where C represents the predicted action center and L represents the predicted action length.
5. A single-frame supervised video temporal action classification method, characterized in that, Including: For the input long video, determine an action position proposal, which consists of an action center and an action length; Map the action position proposal to a temporal action position mask, which represents the time positions of the video frames containing actions in the video; Based on the temporal action position mask, calculate all the action features and background features in the single-instance video segment, and aggregate them into video-level action features and video-level background features; Use a classification network to map the video-level action features to action category probabilities, map the video-level background features to background category probabilities, and calculate the loss function with the action category labels and background category labels to train the classification network until the loss function converges; Input the test video into the trained classification network, output the action category probabilities, and use the threshold method to generate action category predictions for the action category probabilities; Regarding the determination of an action position proposal for the input long video, including: For the input long video, extract video feature maps; Regarding the given single action frame as a seed frame, map the video feature maps to an action seed frame probability map, and obtain several action seed frame positions; Based on the action seed frame positions, divide the long video into several single-instance video segments, and calculate the corresponding single-instance video segment feature maps according to the time positions of the single-instance video segments; Map any of the single-instance video segment feature maps to an action location proposal, which consists of an action center and an action length; Map the video feature map to an action seed frame probability map, and obtain several action seed frame positions, including: Use an action seed frame detection network composed of fully convolutional layers to map the video feature map F to an action seed frame probability map S of dimension T*1, where S represents the probability that each video frame belongs to a seed frame, and T represents the time length of the video; Supervise the action seed frame probability map S using action seed frame labels, calculate the loss function to train the action seed frame detection network until the loss function converges; The calculation formula of the loss function is as follows: Among them, θ E is the parameter of the feature encoding network, θ SD is the parameter of the action seed frame detection network, (X S , Y S ) represents the distribution of the input video and the single-frame action position label, x i represents the video instance, y i is its single-frame action position label, E represents the feature encoding network, SD represents the action seed frame detection network, and H represents the cross-entropy function; For unannotated test set videos, obtain several action seed frame positions by post-processing the action seed frame probability map S.
6. The single-frame supervised video temporal action classification method according to claim 5, wherein Based on the temporal action location mask, calculate all action features and background features in the single-instance video segment, and aggregate them into video-level action features and video-level background features, including: Set the timing action position mask M and the single-instance video clip F W All the action features And the background features The calculation formula is as follows: Then, perform averaging on separately in the time dimension, and aggregate them into video-level action features and video-level background features Both of them have a dimension of 1*2D, where D represents the feature dimension of the video.
7. The single-frame supervised video temporal action classification method according to claim 6, characterized in that Use a classification network to map the video-level action features to action category probabilities, and map the video-level background features to background category probabilities, where: The classification network is a classification network composed of fully connected layers, which maps the video-level action features to the action class probability p A , maps the video-level background features to the background class probability p B , and calculates the loss function with the action class label and the background class label, and trains the classification network until the loss function converges.
8. A single-frame supervised video temporal action detection system, characterized in that, including: Video feature map extraction module: For the input long video, extract the video feature map; Action seed frame detection module: Regard a given single action frame as a seed frame, map the video feature map to an action seed frame probability map, and obtain several action seed frame positions; Action instance division module: Based on the action seed frame positions, divide the long video into several single-instance video segments, and calculate the corresponding single-instance video segment feature maps according to the time positions of the single-instance video segments; Action location detection module: Map any of the single-instance video segment feature maps to an action location proposal, which consists of an action center and an action length; Detection result generation module: Generate action detection results based on the action location proposals; The action seed frame detection module maps the video feature map to an action seed frame probability map, and obtains several action seed frame positions, including: Use an action seed frame detection network composed of fully convolutional layers to map the video feature map F to an action seed frame probability map S of dimension T*1, where S represents the probability that each video frame belongs to a seed frame, and T represents the time length of the video; Supervise the action seed frame probability map S using action seed frame labels, calculate the loss function to train the action seed frame detection network until the loss function converges; The calculation formula of the loss function is as follows: Among them, θ E is the parameter of the feature encoding network, θ SD is the parameter of the action seed frame detection network, (X S , Y S ) represents the distribution of the input video and the single-frame action position label, x i represents the video instance, y i is its single-frame action position label, E represents the feature encoding network, SD represents the action seed frame detection network, and H represents the cross-entropy function; For unannotated test set videos, obtain several action seed frame positions by post-processing the action seed frame probability map S.
9. A single-frame supervised video temporal action classification system, characterized in that, including: Action location proposal determination module: Determine an action location proposal, which consists of an action center and an action length; Temporal mask generation module: Map the action location proposal to a temporal action location mask, indicating the time positions of the video frames containing actions in the video; Action-background separation module: Based on the temporal action location mask, calculate all action features and background features in the single-instance video segment, and aggregate them into video-level action features and video-level background features; Action-background classification module: Use a classification network to map the video-level action features to action class probabilities, map the video-level background features to background class probabilities, and calculate the loss function with the action class label and background class label to train the classification network until the loss function converges; Classification result generation module: Input the test video into the trained classification network, output the action class probabilities, and use the threshold method to generate action class predictions for the action class probabilities; The action location proposal determination module includes: For the input long video, extract the video feature map; Regard the given single action frame as the seed frame, map the video feature map to the action seed frame probability map, and obtain several action seed frame positions; Based on the action seed frame positions, divide the long video into several single-instance video segments, and calculate the corresponding single-instance video segment feature maps according to the time positions of the single-instance video segments; Map any of the single-instance video segment feature maps to an action location proposal, which is composed of an action center and an action length; Mapping the video feature map to the action seed frame probability map and obtaining several action seed frame positions includes: Use an action seed frame detection network composed of fully convolutional layers to map the video feature map F to an action seed frame probability map S with a dimension of T*1. S represents the probability that each video frame belongs to the seed frame, and T represents the time length of the video; Supervise the action seed frame probability map S with the action seed frame label, calculate the loss function to train the action seed frame detection network until the loss function converges; The calculation formula of the loss function is as follows: Among them, θ E is the parameter of the feature encoding network, θ SD is the parameter of the action seed frame detection network, (X S , Y S ) represents the distribution of the input video and the single-frame action position label, x i represents the video instance, y i is its single-frame action position label, E represents the feature encoding network, SD represents the action seed frame detection network, and H represents the cross-entropy function; For the unlabeled test set video, obtain several action seed frame positions by post-processing the action seed frame probability map S.
Citation Information
Patent Citations
Time sequence behavior detection method and system based on single frame supervision
CN112926492A