A Temporal Action Detection Method and Detector Based on Anchor-Free Technology

By building an anchor box-free detection network, directly regressing the left and right boundary distances of the action and using the instance-aware alignment module to explicitly extract features, the problem of poor flexibility in existing anchor box technology is solved, and more flexible and accurate action detection effects are achieved.

CN114821774BActive Publication Date: 2025-07-25NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210404413.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-18
Publication Date
2025-07-25
Estimated Expiration
2042-04-18

AI Technical Summary

Technical Problem

The existing timing action detection algorithm based on anchor box technology requires manual design of anchor box, and then return to the action boundary based on predefined anchor box. However, in fact, the time span of different actions is very huge, and the manual design of anchor box is poor in flexibility and cannot cover various actions.

Method used

A detection network based on anchor-free frame technology is constructed, including feature extraction network, timing feature pyramid, boundary offset regressor, instance-aware alignment module and refined classification regressor, directly regressing the left and right boundary distances of the action, and explicitly extracting features within the predicted action duration through the instance-aware alignment module.

Benefits of technology

It realizes more flexible and accurate action classification and boundary regression, improves robustness and efficiency, simplifies the complexity of action modeling, and has strong scalability and portability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114821774B_ABST
    Figure CN114821774B_ABST
Patent Text Reader

Abstract

A temporal action detection method and detector based on the anchor-free technology, which constructs a network to detect temporal actions in videos, including a feature extraction network, a temporal feature pyramid, a boundary offset regressor, an instance-aware alignment module, and a refined classification regressor. The feature extraction network extracts spatio-temporal features of the video, the temporal feature pyramid obtains features with different temporal resolutions, the boundary offset regressor predicts the distances of each temporal position relative to the left and right boundaries of the action at that moment, and then the start and end times of the action are obtained through transformation. The instance-aware alignment module obtains action features for fine prediction according to the start and end times of the action, and the refined classification regressor is used to predict the action category and fine-tune the action boundary to obtain the temporal action detection result. The present invention directly regresses the distances from the left and right boundaries of the action, completes the temporal localization and classification tasks of actions in videos, and is more simple and efficient than the existing detectors with anchors without the need to pre-set anchors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer software, relates to the technology of temporal action detection, and a temporal action detection method based on the anchor-free technology. Background Art

[0002] The goal of the video temporal action detection algorithm is to identify the category of the actions of the actors in the video and the specific start and end times of the actions. Temporal action detection is widely used in fields such as intelligent security, autonomous driving, and intelligent editing. The existing action detection algorithms are mainly divided into two types, the action detection algorithm based on action scores and the action detection algorithm based on anchor boxes.

[0003] The first algorithm first predicts frame by frame in the video sequence whether the frame belongs to the action category, and then through a post-processing strategy, connects the consecutive video frames predicted to be of the action category to determine the specific interval where the action occurs. The problem with this algorithm is that it often only considers a single frame of video and cannot well model the overall action interval, which leads to the instability of single-frame prediction and is prone to misjudged single-frame predictions. In this case, it is often necessary to carefully adjust the post-processing strategy to obtain reliable action detection results. This limits the robustness and universality of the algorithm, and the post-processing strategy needs to be continuously adjusted according to the business scenario.

[0004] The second algorithm relies on the anchor box strategy in two-dimensional picture object detection. This algorithm first presets several anchor boxes with different time lengths, and then assumes that the action lengths in real applications can all be divided into these several different types of anchor boxes. However, the action time lengths in real applications are infinite, while the design of the anchor boxes is limited, and it is inevitable to have errors when using the anchor box-based strategy. The design of the anchor boxes generally considers common action time lengths, and when there are particularly long or particularly short actions in real applications, it is difficult to capture these actions with "strange" lengths using the anchor box strategy. Summary of the Invention

[0005] The problem to be solved by the present invention is that the existing temporal action detection algorithms based on the anchor box technology need to manually design the anchor boxes and then regress the action boundaries based on the predefined anchor boxes. However, in fact, the duration spans of different actions are very large, and the manually designed anchor boxes have poor flexibility and cannot cover all actions.

[0006] The technical solution of the present invention is: a temporal action detection method based on the anchor-free technology, which constructs a detection network to detect the temporal actions in the video. The network structure includes a feature extraction network, a temporal feature pyramid, a boundary offset regressor, an instance-aware alignment module, and a refined classification regressor:

[0007] Feature extraction network: Use C3D as the basic network structure to extract features from the input video sequence I. For the image sequence I of continuous T frames, a video feature sequence f is extracted;

[0008] Temporal feature pyramid: Use pooling layers with different kernel sizes on the obtained video feature sequence f to construct multi-level feature maps with different time scales;

[0009] Boundary offset regressor: Feed the multi-level feature maps into a neural network sequence composed of three one-dimensional convolutional layers and a deformable convolutional layer to generate predictions of the distances of each temporal position in the temporal feature sequence from the left and right boundaries of the action. Then, apply the generated action boundary offsets to each temporal position in the feature sequence to obtain the predicted action boundaries;

[0010] Instance-aware alignment module: Map the predicted action boundaries back to the video feature sequence f obtained by the feature extraction network. Then, obtain the action feature segments belonging to the action indicated by the action boundaries on the video feature sequence. Take half of the length of the feature segment of the action as the length of the context feature, and obtain context feature segments before and after the action boundary respectively. Concatenate the two context feature segments and the action feature segment along the temporal dimension, and then pass through an adaptive max pooling layer to obtain the action features after the alignment operation;

[0011] Refined classification regressor: Input the action features obtained by the instance-aware alignment module into two branches for classification and regression respectively. In the classification branch, output class scores of (C + 1) dimensions, where C represents the number of action categories; the regression branch adopts the regression branch proposed by RCNN and is responsible for predicting the action boundary offsets corresponding to the feature sequence, that is, the normalized temporal length and the center offset in the logarithmic space; After classification and regression, obtain the action prediction results, that is, the category of the action and its boundaries in the video sequence;

[0012] Through the above network structure, use the non-maximum suppression algorithm to remove duplicates from the prediction results obtained by the refined classification regressor, and then concatenate the action detection results of each video segment belonging to the same video to obtain the final action detection results.

[0013] Furthermore, the implementation of the detection network includes a training example generation stage, a network configuration stage, a training stage, and a testing stage.

[0014] 1) Generate training examples: Extract frames from the video at the set sampling frame rate, then divide the video into several overlapping video segments. Each video segment contains the sampled consecutive frame RGB images and the corresponding action instance annotations. Finally, use the video segments as the network input;

[0015] 2) In the network configuration stage, configure the feature extraction network, the temporal feature pyramid, the boundary offset regressor, the instance-aware alignment module, and the refinement classification regressor;

[0016] 3) Training stage: In the training stage of the boundary offset regressor, use the IoU Loss to supervise the predicted action boundaries. In the training stage of the refinement classification regression, use the Cross-entropy Loss to supervise the class prediction branch and the L1 Loss to supervise the regression prediction branch. During training, use the ground truth labels to supervise the three branches to complete training independently, and then superimpose the three loss functions. Use the SGD optimizer to optimize the overall loss, and update the network parameters through the backpropagation algorithm until the number of iterations is reached;

[0017] 4) Testing stage: Collect the image sequence of the video clip in the test set, input it into the network, and obtain the temporal action detection results in the entire video to verify the detection effect.

[0018] The present invention also provides a temporal action detector based on the anchor-free technology. The detector has a computer-readable storage medium, in which a computer program is configured. The computer program is programmed according to the above detection network. When the computer program is executed, it implements the above temporal action detection method based on the anchor-free technology.

[0019] The present invention proposes a temporal action detection method based on the anchor-free technology, which directly regresses the distances from the left and right boundaries of the action. And aiming at the problem of the receptive field center offset in the anchor-free regression, an instance-aware alignment module is proposed. The features located within the predicted action duration are explicitly extracted using the action boundaries predicted by the anchor-free boundary offset regressor to achieve better action classification and action boundary regression effects.

[0020] The present invention has the following advantages compared with the prior art

[0021] The present invention proposes an anchor-free video temporal action detector to complete the temporal localization and classification tasks of actions in the video. Compared with the previous detectors with anchors, there is no need to pre-set anchors, which is more simple and efficient.

[0022] Aiming at the problem of the receptive field center offset in the anchor-free regression, the present invention proposes an instance-aware alignment module. The features located within the predicted action duration are explicitly extracted using the action boundaries predicted by the anchor-free boundary offset regressor to achieve better action classification and action boundary regression effects.

[0023] The present invention demonstrates good robustness and efficiency in the video action temporal detection task. It is more concise and efficient compared with the previous video action detectors with anchors, and has strong scalability and portability. Description of the Drawings

[0024] Figure 1 It is the detection framework diagram of the anchor-free action detection of the present invention.

[0025] Figure 2 It is the schematic diagram of the temporal feature pyramid of the present invention.

[0026] Figure 3 It is the schematic diagram of the anchor-free boundary offset regressor of the present invention.

[0027] Figure 4 It is the schematic diagram of the instance-aware alignment module proposed by the present invention. Detailed implementation manners

[0028] The present invention proposes a video action detection method based on the anchor-free technology. In the framework of the method of the present invention, the preset anchor boxes are cancelled, and instead, the distances from the left and right boundaries of the action are directly regressed. This strategy is more flexible and can cope with various different action time lengths. Compared with the algorithms based on single-frame prediction, the method of the present invention includes multi-scale prediction, and actions of different lengths are assigned to different scales for prediction. This ensures that the method of the present invention can model the entire action interval for any length of action, rather than relying only on single-frame prediction. Such a strategy improves the robustness of the prediction results.

[0029] The detection method of the present invention is implemented based on a computer program. For this, a temporal action detector based on the anchor-free technology is also provided. The detector has a computer-readable storage medium, in which a computer program is configured. The computer program is programmed according to the above detection network, and when the computer program is executed, the above temporal action detection method based on the anchor-free technology is implemented. As an embodiment, the detection network of the present invention has achieved high accuracy after training and testing on the THUMOS14 temporal action detection data set, and is specifically implemented using the Python3 programming language and the Pytorch 1.3.0 deep learning framework to obtain an anchor-free temporal action detector.

[0030] Figure 1 It is the network system framework diagram used in the present invention. The network implementation of the present invention includes a training example generation stage, a network configuration stage, a training stage, and a testing stage. The specific implementation steps are as follows:

[0031] 1) Training example generation stage: The video frames of the THUMOS14 dataset are pre-extracted and stored on the hard disk at a sampling frame rate of 25fps. The sliding window technique is used to segment the long video, with a window size of 768 frames (about 30 seconds) and a sliding step of 192 frames. If there is an action instance in the video segment and the IoA of the existing action instance and the video segment is greater than 0.7, then the video segment is selected as a training sample, and the resolution of each frame of video image is adjusted to 171×128 pixels. In the training stage, each input frame is randomly cropped to a size of 112×112 pixels. In the testing stage, the input frame is centrally cropped to a size of 112×112 pixels. To increase the training data, we not only extract data from the beginning to the end of the video using the sliding window, but also extract it once from the end to the beginning of the video, and adopt a data augmentation strategy of random horizontal flipping. Finally, after reading in the pictures, the obtained picture sequence is normalized by subtracting the mean of the three channels of the THUMOS dataset and dividing by the standard deviation of the three channels, and finally converted into the form of Tensor, batched and shuffled in the data loading order.

[0032] Taking RGB pictures as input, the frame sequence I of the training sample video segment is as follows:

[0033]

[0034] where Img i represents the i-th frame corresponding to the training sample video segment, with 3 channels, T represents the length of the frame sequence, W is the width of the input picture resolution, and H is the height of the input picture resolution.

[0035] 2) Network configuration stage:

[0036] 2.1) Feature extraction network: Using C3D as the basic network structure, the parameters of the pre-trained model in the ActivityNet action recognition dataset are loaded into the network, and the output result of the conv5 layer of the C3D network is taken as the basic video feature. Specifically, the input of the feature extraction network is the video segment after data preprocessing in step 1), with a size of 3x768x112x112, and the output feature map is 512x96x7x7.

[0037] Let the C3D network be B, and perform spatio-temporal feature extraction on the input sequence I to obtain the feature sequence f as follows:

[0038]

[0039] where 8 is the downsampling rate of the C3D network in the temporal dimension, and R is the downsampling rate of the C3D network in the spatial dimension.

[0040] 2.2) Temporal Feature Pyramid: Use pooling layers with different kernel sizes on the video feature sequence obtained from the feature extraction network to construct multi-level feature maps with different temporal scales. As Figure 2 shown, taking the configuration on the THUMOS14 dataset as an example, the video features obtained from the feature extraction network pass through 3D max pooling layers with kernel sizes of 2x7x7 and 4x7x7 respectively to generate two feature maps with different scales and The spatial information of the temporal features is compressed and summarized into 1×1 pixels, and the temporal scales are reduced to and

[0041] The above multi-scale temporal feature pyramid is calculated as follows:

[0042] Denote the 3D pooling layer of the k-th layer as P k , and the output feature of the k-th layer feature pyramid as f k :

[0043]

[0044] Among them, the convolution kernel size of the pooling layer P k is and the stride is 2 k , and the pooling layer will completely compress the spatial information of the features.

[0045] The feature pyramid technology of the present invention constructs multi-level feature maps with different temporal scales. Since the time lengths of different actions vary greatly, predicting actions based on features at a single scale will cause the receptive field to be unable to cover the entire range of all actions, resulting in a decrease in the accuracy of action detection. The present invention uses the feature pyramid technology to construct multi-level feature maps with different temporal scales. The receptive field range of the low-level feature map is small, and it can better capture the fine changes of actions, which is suitable for detecting actions with short durations. While the receptive field range of the high-level feature map is large, and it can better model the integrity of actions, which is suitable for detecting actions with long durations. Therefore, using the temporal feature pyramid can detect actions at an appropriate temporal scale and improve the accuracy of action detection.

[0046] 2.3) Boundary Offset Regressor: Send the feature maps obtained in 2.2) into a neural network sequence composed of three one-dimensional convolutional layers and a deformable convolutional layer for processing. The convolution kernel size of each layer is 3, and the stride is set to 1 to ensure that the temporal dimension size of the input features remains unchanged. Specifically, input features with dimensions of 512x48 and 512x24, and output boundary offset prediction results of 2x48 and 2x24. As Figure 3 shown, for each temporal feature position t, a regression layer generates the offset amounts (l t , r t), representing the distance from the temporal feature position t to the boundaries of the predicted action. At this time, the predicted left and right action boundaries (s t , e t ) can be calculated.

[0047] During training, given the true action labels (c*, s*, e*), where c* represents the action category, and s* and e* represent the start frame and end frame of the action duration respectively. A temporal feature position t is considered a positive sample only if it falls within the duration of the true action, and is otherwise considered a negative sample and does not participate in the boundary offset regression task. The regression loss uses IoU Loss, denoted as

[0048] The boundary offset regressor is implemented as follows:

[0049] 1. Denote the predicted boundary regression offset results of the j-th layer as (l j , r j ):

[0050]

[0051] CB = Relu(Conv1D)

[0052] Denote the boundary regression offset convolutional block as CB, which contains a 1D convolutional layer with a kernel size of 3 and a Relu activation function layer. The boundary regression offsetter contains a total of 4 boundary regression offset convolutional blocks CB, and the final output channel number is 2, representing the predicted results of the distances to the left and right boundaries of the action respectively.

[0053] 2. The calculation method for the predicted start and end times of the action at position t is as follows:

[0054] s t = t - l t

[0055] e t = t + r t

[0056] Among them, s t represents the predicted start time of the action at position t, and e t represents the predicted end time of the action at position t.

[0057] 3. During training, the loss function term generated by this branch of the boundary offset regressor is denoted as as follows:

[0058]

[0059] Among them, N p o s is the number of positive samples, and N is the number of layers of the temporal pyramid. is an indicator function, I t represents the intersection of the predicted action interval (s t , e t ) and the ground-truth action interval (s * , e * ), and U t represents the union of the predicted action interval (s t , e t ) and the ground-truth action interval (s * , e * ). Specifically:

[0060] I t = min(e t , e * ) - max(s t , s * )

[0061] U t = (et - s t ) + (e * - s * ) - I t

[0062] The present invention uses an anchor-free bounding box offset regressor to regress the multi-level temporal feature maps output by the feature pyramid to generate action boundary offsets. Different from the conventional anchor-based temporal action detection algorithms, the present invention adopts a simpler and more effective anchor-free representation, predicting the distances of each temporal position in the temporal feature sequence from the left and right boundaries of the action. Using the anchor-free method is more flexible and can take into account actions of specific lengths that cannot be covered by the anchor-based methods, not only simplifying the complexity of action modeling but also improving the processing speed and more effectively realizing action instance modeling.

[0063] 2.4) Instance-aware alignment module: As Figure 4 shown, map the action boundaries predicted in 2.3) back to the video feature sequence f obtained in the feature extraction network. The mapped start frame and end frame on f are respectively obtained by multiplying the predicted left and right boundaries (s t , e t ) by a constant λ k , where λ k represents the sampling rate from the video feature sequence f to f k . represents the duration of the action, and its value is and The difference between. The action feature mapped back to f is Then, further expand the context of the predicted action. The lengths of the front and back contexts are respectively Half of this, and these two features are respectively represented as and represent the start and end of the action. Finally, the three features pass through the adaptive max-pooling layer to obtain three fixed-length action features of 512x1x2x2, 512x2x2x2, and 512x1x2x2, and then are concatenated along the time sequence dimension to obtain the final action feature representation

[0064] The specific calculation method of the instance-aware alignment module is as follows:

[0065] 1. Map the predicted action start and end times back to the C3D features:

[0066]

[0067]

[0068] where λ k represents the upsampling coefficient of the k-th layer feature pyramid in step 3).

[0069] 2. Extract the action features and context features:

[0070]

[0071]

[0072]

[0073] Using the action start and end times predicted by the boundary regression offsetter, we can intercept the corresponding action feature F act and context features F start 、F end . In the formula represents the action duration.

[0074] 3. Obtain the final action feature representation at position t:

[0075]

[0076]

[0077] Pass the action feature F act and context features F start 、F end through the adaptive pooling layer respectively to obtain fixed-length features, and then concatenate them in the time dimension to obtain the final action feature representation F uni .

[0078] The instance-aware alignment module of the present invention aligns the action features for classification and regression with the action boundaries. When the predicted position is not at the center of the action, in order to cover all action regions during the prediction process, the receptive field must be relatively large, which leads to the introduction of too much background noise, and the action boundaries are usually blurred. To avoid this problem, an instance-aware alignment module is used to correct the features according to the action instances predicted by the anchor-free bounding box offset regressor. Specifically, first, the predicted action is mapped onto a feature map with finer temporal and spatial dimensions, then we can clearly locate the regions belonging to the action and its context, and then use the adaptive maximum pooling layer to extract the salient features. Finally, the aligned features can be used to classify the action and refine the boundaries.

[0079] 2.5) Refine the classification and regression model: The joint action feature representation F obtained by the instance-aware alignment module uni is respectively input into two branches for classification and regression. In the classification branch, the class scores of (C + 1) dimensions are output, where C represents the number of action categories, and (C + 1) represents the number of action categories and the background category; the regression branch adopts the regression branch proposed by RCNN and is responsible for predicting the parameterized two-dimensional offset vector, that is, the normalized temporal length and the central offset in the logarithmic space.

[0080] During the training process, the predicted action instances are considered positive samples only when the IoU value with the ground truth action (st, et) is greater than 0.5, otherwise they are considered background samples. The classification loss generated in the classification branch is denoted as Cross-entropy loss is adopted, and the regression loss generated in the regression branch is denoted as Smooth L1 loss is adopted.

[0081] The specific calculation process in the refinement classification and regression stage is as follows:

[0082] 1. Classification branch

[0083] c = Softmax(Linear(Relu(Linear(F uni )))) ∈ [0, 1] (C+1)×t

[0084] 2. Regression branch

[0085]

[0086] 3. During the training process, the classification loss function term generated by this branch is denoted as The regression loss function is denoted as

[0087]

[0088] lre g = SmoothL1Loss(r, r * )

[0089] 2.6) Post-processing: First, arrange the action detection results of the video clips in chronological order, and use the non-maximum suppression algorithm (NMS) to remove duplicate action nominations. The NMS threshold is set to 0.6, and the top 200 action nominations with the highest predicted action scores are retained as the action detection results for this video clip. Then, merge the action detections of all clips of the same video, arrange them in chronological order, and use the non-maximum suppression algorithm again. The purpose of this step is to remove duplicate action nominations in the overlapping parts of the clips. The NMS threshold is set to 0.3, and the top 200 action nominations with the highest predicted action scores are retained as the action detection results for this video.

[0090] 3) During the training phase, use IoU Loss as the loss function for the bounding box regression branch, use Cross-Entropy Loss to supervise the classification branch in the refined classification and regression phase, and use Smooth L1 Loss to supervise the regression branch in the refined classification and regression phase. During training, use the ground truth labels to supervise the independent training of the three branches. The losses of the three branches are weighted and added together in a ratio of 1:1:1. Use the SGD optimizer to optimize the overall loss. The initial learning rate is 5e-5, and the learning rate is reduced by 10 times after the 4th epoch. The training is completed on 8 NVIDIA Tesla P40 GPUs. The BatchSize for a single card is set to 8, and the total number of training epochs is 6.

[0091] The specific calculation process of the training loss function is as follows:

[0092]

[0093] a = 1

[0094] b = 1

[0095] including that of the bounding box regression branch in the refined classification and regression phase and

[0096] 4) During the testing phase, the input data of the test set was not data-augmented. It was directly deformed into 171x128 using bilinear interpolation, and then center-cropped to obtain frame images of 112x112. Each frame image was normalized by subtracting the respective means of the three channels of the THUMOS14 dataset and dividing by the standard deviations of the three channels. During testing, the test effect was improved by horizontal flipping. On the THUMOS test set, mAP@0.3 reached 63.7, mAP@0.4 reached 58.2, mAP@0.5 reached 49.2, mAP@0.6 reached 36.4, and mAP@0.7 reached 24.2.

Claims

1. A temporal action detection method based on the anchor-free technology, characterized in that Construct a detection network to detect temporal actions in videos. The network structure includes a feature extraction network, a temporal feature pyramid, a boundary offset regressor, an instance-aware alignment module, and a refinement classification regressor: Feature extraction network: Use C3D as the basic network structure to extract features from the input video sequence I. For the image sequence I of consecutive T frames, a video feature sequence f is extracted; Temporal feature pyramid: Use pooling layers with different kernel sizes on the obtained video feature sequence f to construct multi-level feature maps with different temporal scales; Boundary offset regressor: Send the multi-level feature maps into a neural network sequence composed of three one-dimensional convolutional layers and a deformable convolutional layer for processing, generate predictions of the distances from each temporal position in the temporal feature sequence to the left and right boundaries of the action, and then apply the generated action boundary offsets to each temporal position in the feature sequence to obtain the predicted action boundaries; Instance-aware alignment module: Map the predicted action boundaries back to the video feature sequence f obtained by the feature extraction network, then obtain the action feature segments belonging to the action indicated by the action boundaries on the video feature sequence. Take half of the length of the feature segments of this action as the length of the context features, obtain context feature segments before and after the action boundaries respectively, concatenate the two context feature segments and the action feature segments along the temporal dimension, and then obtain the action features after the alignment operation through an adaptive max pooling layer; Refinement classification regressor: Input the action features obtained by the instance-aware alignment module into two branches for classification and regression respectively. In the classification branch, output class scores of (C + 1) dimensions, where C represents the number of action categories; the regression branch adopts the regression branch proposed by RCNN and is responsible for predicting the action boundary offsets corresponding to the feature sequence, that is, the normalized temporal length and the center offset in the logarithmic space; through classification and regression, obtain the action prediction results, that is, the category of the action and its boundaries in the video sequence; Through the above network structure, use the non-maximum suppression algorithm to remove duplicates from the prediction results obtained by the refinement classification regressor, and then concatenate the action detection results of each video segment belonging to the same video to obtain the final action detection results.

2. The temporal action detection method based on the anchor-free technology according to claim 1, characterized in that The implementation of the detection network includes a training example generation stage, a network configuration stage, a training stage, and a testing stage. 1) Generate training examples: Extract frames from the video at the set sampling frame rate, then divide the video into several overlapping video segments. Each video segment contains the sampled consecutive frame RGB images and the corresponding action instance annotations. Finally, use the video segments as the network input; 2) Network configuration stage, configure the feature extraction network, the temporal feature pyramid, the boundary offset regressor, the instance-aware alignment module, and the refinement classification regressor; 3) Training phase: During the training phase of the boundary offset regressor, the predicted action boundaries are supervised using IoU Loss. During the training phase of the refined classification regression, the category prediction branch is supervised using Cross-entropy Loss, and the regression prediction branch is supervised using L1 Loss. During training, the three branches are independently trained using the ground truth labels, and then the three loss functions are superimposed. The SGD optimizer is used to optimize the overall loss, and the network parameters are updated through the backpropagation algorithm until the number of iterations is reached; 4) Testing phase: The image sequence of the video segment in the test set is collected and input into the network to obtain the temporal action detection results in the entire video, and the detection effect is verified.

3. A temporal action detection method based on an anchor-free technique according to claim 1 or 2, characterized in that When generating training examples, the sampling frame rate is 25fps, and the continuous T frames of images are used as the input sequence I of the network. T is set to 768. If there is an action instance in the video segment and the IoA of the existing action instance and the video segment is greater than 0.7, then the video segment is selected as a training sample. Also, to increase the training data, in addition to extracting data from the beginning to the end of the video using a sliding window, it is extracted once again from the end to the beginning of the video, and a data augmentation strategy of random horizontal flipping is adopted for data augmentation.

4. The temporal action detection method based on the anchor-free technology according to claim 1, characterized in that for The refined classification regressor, during the training process, a predicted action instance is considered a positive sample only when the IoU value with the ground-truth action is greater than 0.5, otherwise it is considered a background sample. The classification loss generated in the classification branch is denoted as Cross-entropy loss is adopted, and the regression loss generated in the regression branch is denoted as Smooth L1 loss is adopted; finally, non-maximum suppression is performed on the obtained prediction results to remove duplicates.

5. A temporal action detector based on the anchor-free technology, characterized in that The detector has a computer-readable storage medium, which is configured with a computer program. When the computer program is executed, it implements the temporal action detection method based on the anchor-free technology described in any one of claims 1-4.

Citation Information

Patent Citations

  • Single-stage video behavior detection method

    CN108805083A

  • Video action detection method based on central point trajectory prediction

    CN111259779A