A spatio-temporal detection method for combined actions based on enhancing the motion of interactive objects
By combining the action space-time detection model, using the interactive object motion trajectory information, the problem of insufficient generalization ability in the existing technology is solved, and efficient detection of combined actions is achieved.
Patent Information
- Application Number
- CN202210630316.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-06
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-06-06
AI Technical Summary
The prior art ignores the motion trajectory of the action object in combined action detection, resulting in insufficient generalization capabilities of the model and inability to effectively utilize interactive object information.
By mining the interactive information between objects, modeling motion trajectories, iteratively training the model, and using the combined action space-time detection model to perform combined action detection, including position generation, trajectory generation, video processing, feature extraction, local feature extraction, trajectory feature generation and feature fusion.
The generalization ability of the model is improved, the ability to classify action information is enhanced, and the space-time location and category of combined actions can be effectively detected.
Smart Images

Figure CN115100737B_ABST
Abstract
Description
Technical Field
[0001] The present invention designs a combined action spatio-temporal detection method based on the enhancement of the motion of interactive objects, and particularly relates to a method for detecting combined actions. Background Art
[0002] With the popularity of short videos, videos have replaced pictures as the current most mainstream information medium. Simply relying on manual discrimination of various actions in videos not only consumes a large amount of manpower and material resources but also takes a lot of time. Based on the above factors, using computer vision to solve action detection in videos has become a research focus in the industrial and academic fields.
[0003] Spatio-temporal action detection (also known as spatio-temporal action localization) refers to, for a video containing action segments, not only locating the spatio-temporal position of the action subject, that is, the position where the participant is at any moment, but also classifying the action performed by the subject to determine the category to which the action belongs. Previous research on spatio-temporal action detection has taken humans as the subject and ignored the influence of the object on action recognition. However, the generation of most actions in real-world scenarios is often restricted by things or the environment. For example, pouring and tearing can easily make people associate the objects involved. To make full use of object information, more and more scholars have started using methods such as relational modeling, attention mechanisms, and graph networks to assist spatio-temporal action detection. Although the above methods have achieved good results in the field of spatio-temporal action detection, they have ignored the influence of object appearance information on action classification. In fact, the occurrence of known actions often involves new objects, thus generating new action combinations.
[0004] A combined action refers to associating different nouns (objects) with verbs to form various action combinations. To improve the generalization ability of the model, in this paper, verbs and nouns are combined without overlap, that is, the objects involved in the same verb in the training set will not appear in the test set. Different from traditional spatio-temporal action detection, combined action detection (also known as combined spatio-temporal action detection) requires the model to simultaneously detect the subject and object involved in the action and classify the action category to which they belong.
[0005] Current spatio-temporal action detection tasks pay too much attention to the static features of the action subject and ignore the motion trajectories of the action objects. In fact, the motion trajectories of action objects are an indirect manifestation of temporal information. In the combined action detection task, due to the involvement of the interaction information between the subject and object, temporal information is particularly important. Therefore, enhancing the motion trajectories of interactive objects has broad application prospects and practical value for combined action detection. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide a spatio-temporal detection method for combined actions enhanced by the motion of interaction objects in view of the above deficiencies of the prior art. By mining the interaction information between objects, modeling the motion trajectories, and iteratively training the model.
[0007] To achieve the above object, the present invention provides the following solutions:
[0008] A spatio-temporal detection method for combined actions enhanced by the motion of interaction objects, which uses a spatio-temporal detection model for combined actions to perform spatio-temporal detection on a video segment to be detected; wherein, the spatio-temporal detection model for combined actions includes a position generation module, a trajectory generation module, a video processing module, a feature extraction module, a local feature extraction module, a trajectory feature generation module, a feature fusion module, and an action classification module;
[0009] In the position generation module, the pre-trained object detection model Faster R-CNN is used to perform object detection on the video segment to be detected, and the coordinate information of each object in each frame is obtained;
[0010] In the trajectory generation module, the coordinate information of each object in each frame is sent to the Transformer-based tracking model Stark to align the positions of each object in each frame, and the trajectory information of each object is obtained;
[0011] In the video processing module, a mask operation is used to occlude specific objects in the video segment to be detected, and a processed video segment is obtained;
[0012] In the feature extraction module, the pre-trained network model SlowFast is used to extract the spatio-temporal features of the processed video segment;
[0013] In the local feature extraction module, after processing the spatio-temporal features of the processed video segment using the Non-local module, local features are obtained by combining the coordinate information of each object in each frame through RoiAlign;
[0014] In the trajectory feature generation module, the trajectory information of each object is linearly transformed twice to obtain trajectory features;
[0015] In the feature fusion module, the local features and the trajectory features are fused to obtain combined features;
[0016] In the action classification module, the combined features are sent to a fully connected layer to obtain the probability of predicting each action category, and the spatio-temporal detection of the combined actions is completed.
[0017] Further, if the number of objects in the video segment to be detected is less than the maximum number of object detections of Faster R-CNN, zero vectors are filled to represent the object.
[0018] Further, the spatio-temporal features output by the feature extraction module include slow-channel spatio-temporal features F slow and fast-channel spatio-temporal features F fast , in the local feature extraction module: First, the AdaptiveAvgPool3d function is used to perform pooling operations in the temporal dimension respectively, and a concatenate operation is performed on the pooled features to obtain pooled features F pool ; Then, F pool is sent to the Non-Local Network module to obtain pooled features F nlc with an expanded receptive field; Subsequently, the coordinate information of each object in each frame is sent to the RoiAlign module, and combined with F nlc to obtain object roi features F roi ; Finally, local features F local are obtained through the maxpool2d operation.
[0019] Further, in the trajectory feature generation module: First, the coordinate information of each object in each frame is converted into a coordinate tensor T coord ; Then, after two linear transformations on T coord the final trajectory features F tracks are obtained; where each linear transformation is followed by BatchNorm and ReLU.
[0020] Further, in the feature fusion module: First, the trajectory features F tracks are converted into the same dimension as the local features F local to obtain E′ tracks ; Then, F local and E′ tracks are concatenated in the channel dimension to obtain the final combined features F com .
[0021] Further, the loss function of the combined action spatio-temporal detection model is the cross-entropy function.
[0022] The beneficial effects of the present invention are as follows:
[0023] 1) The present invention proposes a new spatio-temporal detection task, filling the gap in combined action spatio-temporal detection;
[0024] 2) The present invention suppresses the influence of appearance bias on model classification through mask occlusion operations, and at the same time uses trajectory modeling to simulate motion information, enhancing the model's ability to classify relying on action information, thereby improving the model's generalization ability. Brief Description of the Drawings
[0025] Figure 1 is the local feature extraction module of the present invention;
[0026] Figure 2 is the trajectory generation module of the present invention;
[0027] Figure 3 is the spatio-temporal detection model for combined actions of the present invention;
[0028] Figure 4 is the structural schematic diagram of the spatio-temporal detection model for combined actions of the present invention. Detailed Description of the Preferred Embodiments
[0029] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0030] The object of the present invention is to provide a spatio-temporal detection method for combined actions based on the enhancement of the motion of interactive objects, and use the spatio-temporal detection model for combined actions to perform spatio-temporal detection on the video segment to be detected; wherein, as Figure 4 shown, the spatio-temporal detection model for combined actions includes a position generation module, a trajectory generation module, a video processing module, a feature extraction module, a local feature extraction module, a trajectory feature generation module, a feature fusion module, and an action classification module.
[0031] In the position generation module, the pre-trained object detection model Faster R-CNN is used to perform object detection on the video segment to be detected, and the coordinate information of each object in each frame is obtained.
[0032] In the trajectory generation module, the coordinate information of each object in each frame is sent to the tracking model Stark based on Transformer to align the positions of each object in each frame, and the trajectory information of each object is obtained.
[0033] In the video processing module, as Figure 1 shown, the mask operation is used to occlude specific objects in the video segment to be detected, and the processed video segment is obtained.
[0034] In the feature extraction module, as Figure 1 shown, the pre-trained network model SlowFast is used to extract the spatio-temporal features of the processed video segment.
[0035] In the local feature extraction module, as Figure 1 shown, after using the Non-local module to process the spatio-temporal features of the processed video clip, the local features are obtained through RoiAlign in combination with the coordinate information of each object in each frame.
[0036] In the trajectory feature generation module, as Figure 2 shown, the trajectory information of each object is linearly transformed twice to obtain the trajectory features.
[0037] In the feature fusion module, the local features and the trajectory features are fused to obtain the combined features.
[0038] In the action classification module, the combined features are sent to a fully connected layer to obtain the probability of predicting each action category, completing the spatio-temporal detection of the combined action.
[0039] In one embodiment, a spatio-temporal detection method for combined actions based on enhancing the motion of interacting objects, as Figure 3 shown, the specific steps are as follows:
[0040] Step 1: In the training stage, the object position information is obtained from the Something-Something V2 dataset. In the testing stage, the original video clip V with length L clip is input into the Faster R-CNN model (detection tracker) to obtain the object position information. The Faster R-CNN model uses ResNet-101 as the backbone, includes a Feature Pyramid Network (FPN) module, and the model is pre-trained on the COCO dataset and fine-tuned on Something-Something V2. During detection, the maximum number of objects detected is set to 4. If there are fewer objects presented, a zero vector is filled to represent the object. Therefore, for V with length L clip , it contains a total of 4L detection results.
[0041] Step 2: In the training and testing stage, for a video clip V with length L train , V trainThe object ground-truth boxes corresponding to the dataset Something-Something V2 are fed into the Transformer-based tracker Stark to obtain the coherent motion trajectory information of each object. In the test phase, the detection boxes obtained by the Faster R-CNN model are sent to Stark to obtain the coherent motion trajectory information of each object in the test phase. Since the maximum number of objects is set to 4, after each video is input into Stark, 4 trajectory segments of length L will be obtained.
[0042] Step 3: For the input original video segment V original , in the training phase, according to the object coordinates provided by the dataset itself, a mask operation is performed on the corresponding objects with an occlusion ratio of 50%; in the test phase, according to the object coordinates obtained by the Faster R-CNN model, a mask operation is performed on the corresponding objects with an occlusion ratio of 50%. This operation obtains the video segment V mask with the objects occluded.
[0043] Step 4: First, extract the image frames of the occluded video segment V mask , and use Slowfast to extract the slow-channel video spatio-temporal features and fast-channel video spatio-temporal features where B is the sample size, C1 = 2048 is the F slow channel dimension, T1 = 8 is the F slow temporal dimension, C2 = 256 is the F fast channel dimension, T2 = 32 is the F fast temporal dimension, and H and W are the height and width respectively.
[0044] Step 5: According to the two sets of features F slow and F fast obtained in step (4), first use the AdaptiveAvgPool3d function to perform pooling operations on the two sets of features in the temporal dimension respectively, and perform a concatenate operation on the pooled features to obtain the pooled feature where C = 2304; then send F pool to the Non-Local Network module to obtain the pooled feature with an expanded receptive field Subsequently, send the bounding boxes information of the objects to the RoiAlign module (spatial size of 7×7) to obtain the object roi features Finally, obtain N local features through the maxpool2d operation.
[0045] Step 6: According to the 4 sets of trajectory segments with length L obtained in step (2), first convert the numerical coordinates into coordinate tensors That is, for any set of trajectory segments S i , its corresponding tensor can be expressed as: where 4 represents 4 coordinates; then, input T coord into the linear transformation layer to obtain a high-dimensional tensor At the same time, BatchNorm and ReLU follow each linear transformation. After two linear transformations, the final trajectory features are obtained
[0046] Step 7: First, convert the trajectory features F tracks obtained in step (6) into the same dimension as the local features F local , that is Then, concatenate the local features F local and E′ tracks in the channel dimension to obtain the final combined features
[0047] Step 8: First, convert the combined features F com obtained in step (7) into classification features through the Linear layer where C action is equal to the total number of action categories; then, use the cross-entropy function to calculate the loss between the true label and the predicted label. The formula is as follows:
[0048]
[0049] where N represents the total number of test samples; y i represents the label of sample i, with the positive class being 1 and the negative class being 0; p i represents the probability that sample i is predicted as the positive class. The loss obtained in each iteration will be passed to the next round of model iteration. Finally, calculate the accuracy mean average precision (mAP) based on the predicted labels obtained from F class to obtain the classification results for each action category.
[0050] In this article, specific examples are used to elaborate on the principles and implementation methods of the present invention. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A spatio-temporal detection method for combined actions based on enhancing the motion of interaction objects, characterized in that, This method uses a combined action spatio-temporal detection model to perform combined action spatio-temporal detection on the video segment to be detected; among them, the combined action spatio-temporal detection model includes a position generation module, a trajectory generation module, a video processing module, a feature extraction module, a local feature extraction module, a trajectory feature generation module, a feature fusion module, and an action classification module; In the position generation module, the pre-trained object detection model Faster R-CNN is used to perform object detection on the video segment to be detected, and the coordinate information of each object in each frame is obtained; in the trajectory generation module, the coordinate information of each object in each frame is sent to the tracking model Stark based on Transformer to align the positions of each object in each frame, and the trajectory information of each object is obtained; In the video processing module, the mask operation is used to occlude specific objects in the video segment to be detected, and the processed video segment is obtained; In the feature extraction module, the pre-trained network model SlowFast is used to extract the spatio-temporal features of the processed video segment; In the local feature extraction module, after processing the spatio-temporal features of the processed video segment using the Non-local module, the local features are obtained by combining the coordinate information of each object in each frame through RoiAlign; In the trajectory feature generation module, the trajectory information of each object is linearly transformed twice to obtain the trajectory features; In the feature fusion module, the local features and the trajectory features are fused to obtain the combined features; In the action classification module, the combined features are sent to a fully connected layer to obtain the probability of predicting each action category, and the combined action spatio-temporal detection is completed; The spatio-temporal features output by the feature extraction module include slow-channel spatio-temporal features F slow and fast-channel spatio-temporal features F fast , in the local feature extraction module: First, use the AdaptiveAvgPool3d function to perform pooling operations in the temporal dimension respectively, and perform a concatenate operation on the pooled features to obtain pooled feature F pool ; Then send F pool to the Non-LocalNetwork module to obtain the pooled feature F with an expanded receptive field nlc ; Subsequently, send the coordinate information of each object in each frame to the RoiAlign module, and combine it with F nlc to obtain the object roi feature F roi ; Finally, obtain the local feature F through the maxpool2d operation local .
2. The spatio-temporal detection method for combined actions based on the enhancement of the movement of interaction objects according to claim 1, wherein, If the number of objects in the video segment to be detected is less than the maximum number of object detections of Faster R-CNN, a zero vector is filled to represent the object.
3. A spatio-temporal detection method for combined actions based on the enhancement of the movement of an interaction object according to claim 1, characterized in that, In the trajectory feature generation module: First, convert the coordinate information of each object in each frame into a coordinate tensor T coord ; Next, after T coord undergoes two linear transformations, the final trajectory feature F is obtained tracks ; where BatchNorm and ReLU are followed after each linear transformation.
4. A spatio-temporal detection method for combined actions based on enhancing the movement of interaction objects according to claim 1, characterized in that In the feature fusion module: First, convert the trajectory feature F tracks to the same dimension as the local feature F local to obtain F' tracks ; then concatenate F local and F' tracks along the channel dimension to obtain the final combined feature F com .
5. The spatio-temporal detection method for combined actions based on enhancing the motion of interaction objects according to claim 1, characterized in that, The loss function of the combined action spatio-temporal detection model is the cross-entropy function.
Citation Information
Patent Citations
Video action detection method based on central point trajectory prediction
CN111259779A
Method and device for detecting human-object interaction relationship in video
CN112464875A