Action positioning method and system based on multi-instance small sample

Through the action positioning method of multi-instance small samples, feature extraction and alignment processing is used using C3D model and twin networks, the problems of high cost and low efficiency of action positioning in the prior art are solved, and efficient and accurate action positioning under small samples are achieved.

CN120298644AActive Publication Date: 2025-07-11CAPITAL UNIV OF PHYSICAL EDUCATION & SPORTS
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510430093.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-11
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

The prior art relies on a full supervision mode in action positioning to require a large amount of labeled data, resulting in high costs and long-term cycles, and lack of action positioning models under small sample conditions.

Method used

The action positioning method with multiple instances is adopted, and the feature extraction and alignment process is performed using the C3D model and the twin network. The twin network is trained through a small amount of sample data, candidate regions are generated and action prediction is performed, and the action positioning results are optimized in combination with the non-maximum suppression method.

Benefits of technology

It reduces the cost and time period of model training, improves the efficiency and accuracy of action positioning, and can achieve efficient and accurate action positioning under small sample conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298644A_ABST
    Figure CN120298644A_ABST
Patent Text Reader

Abstract

The invention discloses an action positioning method and system based on a multi-instance small sample, and relates to the field of action positioning, and the method comprises the steps: obtaining a to-be-detected video and a standard action video; the to-be-tested video comprises a plurality of to-be-tested examples; inputting the divided video clip and the standard action video into a C3D model to obtain a C3D feature map and a standard action sample feature corresponding to the video clip; inputting the C3D feature map into a center judgment model, screening out a center video clip of each to-be-detected instance, and expanding the center video clip to two sides in different frames to generate a plurality of candidate regions; aligning the C3D feature map and the candidate region by adopting a self-adaptive pooling layer to obtain an aligned feature map; and inputting the aligned feature map and the standard action sample features into the action prediction model to obtain an action prediction result, and performing optimization by adopting a non-maximum suppression method to complete action localization, thereby reducing the input cost and time period of model training, and improving the action localization efficiency and accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of action localization, and in particular, to an action localization method and system based on multi-instance small samples. Background Art

[0002] Video-based action localization technology can segment action segments in a large number of videos, providing a key basis for the evaluation of each subsequent action. Therefore, as an indispensable pre-step for action evaluation in actual scenarios, action localization plays a crucial role.

[0003] The relevant mainstream action localization methods mainly rely on the fully supervised mode, and use large-scale data sets and complex models to improve the prediction accuracy. However, the premise of this method is that a large amount of labeled data must be available to train the model. In practical applications, obtaining and labeling such a large amount of data is not only time-consuming and laborious, but also has a high input cost and a long time cycle, which greatly limits its wide application. At the same time, there is currently a lack of a model that can achieve action localization under small sample conditions. Summary of the Invention

[0004] The purpose of this application is to provide an action localization method and system based on multi-instance small samples, which can complete the training of the model through a small amount of sample data, reduce the input cost and time cycle of model training, and improve the efficiency and accuracy of action localization.

[0005] To achieve the above purpose, this application provides the following solutions:

[0006] In the first aspect, this application provides an action localization method based on multi-instance small samples, including:

[0007] Obtain a video to be measured and a standard action video; the video to be measured includes multiple instances to be measured; one instance is an action;

[0008] Divide the video to be measured to obtain multiple video segments;

[0009] Input each video segment into the C3D model to obtain a C3D feature map corresponding to each video segment;

[0010] Input the standard action video into the C3D model to obtain standard action sample features;

[0011] Input the C3D feature map into the center judgment model to screen out the central video segment of each instance to be measured; each central video segment of the instance to be measured is the video segment corresponding to the central moment of each instance to be measured; the center judgment model is obtained by training a regression model with a sample C3D feature map;

[0012] Taking the center point of the center video segment of each instance to be measured as the center, expand it to both sides of the center video segment with different frames to generate multiple candidate regions for each instance to be measured;

[0013] An adaptive pooling layer is used to align the C3D feature map and the multiple candidate regions to obtain the aligned feature map of each candidate region;

[0014] The aligned feature map of each candidate region and the standard action sample feature are input into an action prediction model to obtain the action prediction result of each candidate region; the action prediction model is obtained by training a siamese network with a sample training set; the sample training set includes: sample aligned feature maps, standard action sample features, and corresponding sample prediction results;

[0015] The non-maximum suppression method is used to optimize the action prediction results of each candidate region to obtain the action type, start time, and end time of each instance to be measured, and complete action localization.

[0016] In a second aspect, the present application provides an action localization system based on multi-instance small samples, including:

[0017] An acquisition module for acquiring a video to be measured and a standard action video; the video to be measured includes multiple instances to be measured; one instance is an action;

[0018] A division module for dividing the video to be measured to obtain multiple video segments;

[0019] A C3D module for inputting each video segment into a C3D model to obtain the C3D feature map corresponding to each video segment; the C3D module is further used to input the standard action video into the C3D model to obtain the standard action sample feature;

[0020] A center judgment module for inputting the C3D feature map into a center judgment model to screen out the center video segment of each instance to be measured; the center video segment of each instance to be measured is the video segment corresponding to the center moment of each instance to be measured; the center judgment model is obtained by training a regression model with sample C3D feature maps;

[0021] A candidate region generation module for taking the center point of the center video segment of each instance to be measured as the center and expanding it to both sides of the center video segment with different frames to generate multiple candidate regions for each instance to be measured;

[0022] An alignment module for using an adaptive pooling layer to align the C3D feature map and the multiple candidate regions to obtain the aligned feature map of each candidate region;

[0023] A prediction module, configured to input the alignment feature map of each candidate region and the standard action sample features into an action prediction model to obtain the action prediction result of each candidate region; the action prediction model is obtained by training a siamese network with a sample training set; the sample training set includes: sample alignment feature maps, standard action sample features, and corresponding sample prediction results;

[0024] An action determination module, configured to optimize the action prediction result of each candidate region by using a non-maximum suppression method to obtain the action type, start time, and end time of each instance to be measured, and complete action localization.

[0025] According to the specific embodiments provided in the present application, the present application has the following technical effects:

[0026] The present application provides an action localization method and system based on multi-instance small samples. The action is localized by a trained siamese network (i.e., an action prediction model). Since the siamese network can map the input sample pairs to a feature space through a shared neural network and judge whether the sample pairs belong to the same category by calculating the distance between the sample pairs, this metric-based learning method enables the model to learn the relationships between samples from a small number of samples without the need for a large amount of labeled data to learn the specific features of each category. Therefore, the present application only needs small sample data to complete the training of the siamese model, reducing the training cost and training duration and improving the efficiency of action localization. At the same time, multiple candidate regions generated from the central video segment can retain the integrity of the action within the video frame segment, thereby ensuring the accuracy of the action prediction result obtained by the action prediction model and further improving the accuracy of action localization. Description of the Drawings

[0027] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0028] Figure 1 It is a schematic flowchart of an action localization method based on multi-instance small samples provided by an embodiment of the present application;

[0029] Figure 2 It is a detailed flowchart of an action localization method based on multi-instance small samples provided by an embodiment of the present application;

[0030] Figure 3 It is a schematic flowchart of action localization using a siamese network;

[0031] Figure 4 Schematic diagram generated for the candidate region. Specific implementation manners

[0032] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0033] The following are the explanations of terms:

[0034] Action localization refers to predicting the start time, end time, and action type of a specific action in a long video (or segment).

[0035] Few-shot action localization means training a model only relying on a few samples to achieve the above functions, rather than relying on a large number of models.

[0036] The advantage of few-shot is that it does not require too many samples and too much manual annotation of samples, but the training difficulty of the model is great. And the present application uses a siamese network for few-shot training, thus solving the problem of great training difficulty of the model. Moreover, the present application can perform action localization on multiple instances in a video segment without determining the type of the video and the time localization of the actions in the video. In addition, the present application can also determine the central video segment, determine the center point of the central video segment as the anchor point, generate multiple candidate regions, so that both the start time and the end time of the action can be retained, and then use the model trained by few-shot to generate the action prediction result, and then use the non-maximum suppression method to optimize the start time and the end time of the action, finally ensuring the accuracy of action localization. Therefore, the present application not only reduces the input cost and time cycle of model training, but also improves the action localization efficiency and accuracy.

[0037] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific implementation manners.

[0038] In an exemplary embodiment, as Figures 1 to 3 shown, a method for few-shot action localization based on multi-instance is provided. This method is executed by a computer device, and specifically can be executed alone by a computer device such as a terminal or a server, or can be jointly executed by a terminal and a server. In the embodiments of the present application, taking this method applied to a server as an example for illustration, it includes the following steps 1 to step 8. Among them:

[0039] Step 1: Obtain the video to be measured and the standard action video; the video to be measured includes multiple instances to be measured; one instance is an action.

[0040] Step 2: Divide the video to be tested to obtain multiple video segments.

[0041] In an exemplary embodiment, the entire video is divided into video segments with 16 frames each, thereby obtaining multiple video segments.

[0042] Step 3: Input each video segment into the C3D model to obtain the C3D feature map corresponding to each video segment.

[0043] Step 4: Input the standard action video into the C3D model to obtain the standard action sample features.

[0044] Step 5: Input the C3D feature map into the center judgment model to screen out the central video segment of each instance to be tested; each central video segment of an instance to be tested is the video segment corresponding to the central moment of each instance to be tested; the center judgment model is obtained by training a regression model with a sample C3D feature map.

[0045] Specifically, Step 5 includes:

[0046] Step 51: Input the C3D feature map into the center judgment model to obtain the probability value corresponding to the C3D feature map.

[0047] Step 52: Determine whether the probability value is greater than the probability threshold. Wherein, the probability threshold is 0.5 to 1.

[0048] Step 53: If so, screen out the video segment corresponding to the C3D feature map as the central video segment.

[0049] Step 54: If not, delete the video segment corresponding to the C3D feature map.

[0050] Among them, the training process of the center judgment model specifically includes:

[0051] Input the sample C3D feature map into the regression model to obtain the detection probability value.

[0052] Construct a loss function according to the detection probability value and the sample probability value corresponding to the sample C3D feature map, and iteratively optimize the parameters of the regression model according to the loss function until the loss function reaches the minimum value or the number of iterative optimization rounds reaches the maximum value, and stop the iterative optimization to obtain the center judgment model. Wherein, the loss function is a cross-entropy loss function.

[0053] Step 6: Expand from the center point of each central video segment of an instance to be tested to both sides of the central video segment with different frames to generate multiple candidate regions for each instance to be tested.

[0054] In an exemplary embodiment, the candidate region is expanded to both sides according to a plurality of pre-arranged templates of a fixed size, as Figure 4 shown. Taking the center of a 16-frame video segment as the center point, a plurality of fixed-size lengths are selected as the radii, and the expansion is performed forward and backward according to the number of frames to generate candidate regions.

[0055] Step 7: Use an adaptive pooling layer to align the C3D feature map and multiple candidate regions to obtain the aligned feature map of each candidate region.

[0056] Step 8: Input the aligned feature map of each candidate region and the standard action sample feature into the action prediction model to obtain the action prediction result of each candidate region; the action prediction model is obtained by training a siamese network with a sample training set; the sample training set includes: sample aligned feature maps, standard action sample features, and corresponding sample prediction results. Among them, the number of sample aligned feature maps, standard action sample features, and corresponding sample prediction results in the sample training set is less than 5 pairs.

[0057] Specifically, as Figure 2 shown, the C3D feature map extracted by the C3D model and the candidate regions are used for feature extraction to obtain the aligned feature map of each candidate region. The aligned feature map of each candidate region and the standard action sample feature are simultaneously input into the siamese network for contrast learning prediction to obtain the action similarity, the offset of the center point of the candidate region, and the offset of the length of the candidate region. The standard action sample is a small sample of a certain type of action referred to in this application, that is, whether the candidate region is of this type of action is evaluated by the similarity between the aligned feature map of the candidate region and the sample of this type of action. The specific prediction of the actual action region (start time, end time) is given by the predicted offset of the center point of the candidate region and the offset of the length of the candidate region.

[0058] As Figure 3As shown in the figure, the Siamese network includes two identical sub-networks. Each sub-network contains a convolutional layer and a fully connected layer, and parameter sharing is carried out between the two sub-networks. During the contrastive learning prediction process of the Siamese network, each sub-network processes the input data and extracts features, which are used to describe the key information of the input data. Since parameter sharing exists between the sub-networks, their processing methods for the input data are consistent, ensuring that the extracted feature representations are comparable. The two sub-networks of the Siamese network share weights, which means that during the training process, the parameter updates of the two sub-networks are synchronized. This parameter sharing mechanism reduces the number of model parameters, reduces the risk of overfitting, and improves the training efficiency of the model. The input of the Siamese network is the aligned feature map of each candidate region and the standard action sample feature. The distance (such as Euclidean distance) between the aligned feature map of each candidate region and the standard action sample feature is calculated through the two sub-networks in the Siamese network. The output result of the Siamese network is finally used for similarity prediction through a fully connected layer and a Sigmoid activation function, so as to obtain the action prediction result and judge whether the sample pair belongs to the same category. This learning method of calculating the distance between the aligned feature map of each candidate region and the standard action sample feature does not require a large amount of labeled data, enabling this application to complete the training of the Siamese network with only small sample data, thus reducing the training cost and training duration and improving the efficiency of action localization.

[0059] Specifically, the training process of the action prediction model specifically includes:

[0060] Input the sample aligned feature map and the standard action sample feature into the Siamese network to obtain a prediction result.

[0061] Construct a loss function according to the prediction result and the sample action prediction result corresponding to the sample aligned feature map and the standard action sample feature, and iteratively optimize the parameters of the Siamese network according to the loss function until the loss function reaches the minimum value or the number of iterative optimization rounds reaches the maximum value, then stop the iterative optimization to obtain the action prediction model.

[0062] Step 9: Use the non-maximum suppression method to optimize the action prediction results of each candidate region to obtain the action type, start time, and end time of each instance to be measured, and complete action localization.

[0063] Specifically, the action prediction results include: action similarity, candidate region center point offset, and candidate region length offset; the action similarity is the similarity between the action corresponding to each aligned feature map and the action corresponding to the standard action sample feature. Step 9 specifically includes:

[0064] Step 91: Determine the action type of the action corresponding to the candidate region according to the action similarity and the similarity threshold.

[0065] Specifically, step 91 includes: determining whether the action similarity is greater than the similarity threshold; if so, determining that the action type of the action corresponding to the candidate region is the same as the action type corresponding to the standard action sample feature and retaining it; if not, determining that the action type of the action corresponding to the candidate region is inconsistent with the action type corresponding to the standard action sample feature and deleting it.

[0066] Step 92: Determine the start time and end time of the action corresponding to the candidate region according to the center point offset of the candidate region and the length offset of the candidate region. Specifically, the start time of the action corresponding to the candidate region is equal to the center point offset of the candidate region minus half of the length offset of the candidate region; the end time of the action corresponding to the candidate region is equal to the center point offset of the candidate region plus half of the length offset of the candidate region.

[0067] Step 93: Determine whether the overlap rate between the start time and the end time of the action corresponding to the candidate region reaches the overlap threshold. If so, execute step 94; if not, execute step 95.

[0068] Step 94: Determine that there are actions corresponding to multiple candidate regions for the instance to be tested, and use the action corresponding to the candidate region with the highest action similarity as the action type, start time, and end time corresponding to the instance to be tested.

[0069] Step 95: Determine that there is only one action corresponding to the candidate region for the instance to be tested, and use the action corresponding to the candidate region as the action type, start time, and end time corresponding to the instance to be tested, and complete action localization.

[0070] Specifically, since each candidate region predicts an action, and there are many candidate regions in a video, if there are multiple actions, they can all be localized. However, the predicted actions may have overlapping regions. Considering that this application only targets the case where there is no time overlap between actions, the method of non-maximum suppression (NMS) is used to retain the region where the highly credible (i.e., the action with the highest action similarity) action is located and delete the prediction results of other actions that have a high time overlap (e.g., 30%) with the region where the highly credible action is located.

[0071] The beneficial effects of the action localization method based on multi-instance small samples proposed in this application are mainly manifested in:

[0072] (1) This application uses an action prediction model (a trained Siamese network) to complete action localization. Since the Siamese network determines whether a sample pair belongs to the same category by calculating the distance (such as Euclidean distance) between sample pairs, this learning method of calculating the distance between sample pairs requires no large amount of labeled data, enabling this application to complete the training of the Siamese network with only small sample data, thereby reducing the training cost and training duration and improving the efficiency of action localization.

[0073] (2) Through multiple candidate regions generated from the central video segment with different frame lengths in this application, the integrity of each instance to be measured within the video frame segment can be retained, thereby ensuring the accuracy of the action prediction result obtained by the action prediction model and improving the accuracy of action localization.

[0074] Based on the same inventive concept, the embodiment of this application also provides an action localization system based on multi-instance small samples. The implementation solutions provided by this system to solve problems are similar to those described in the above method. Therefore, the specific limitations in one or more embodiments of the action localization system based on multi-instance small samples provided below can refer to the limitations on the method of the action localization system based on multi-instance small samples in the above text, and will not be elaborated here.

[0075] In an exemplary embodiment, an action localization system based on multi-instance small samples is provided, including:

[0076] An acquisition module, configured to acquire a video to be measured and a standard action video; the video to be measured includes multiple instances to be measured; one instance is one action.

[0077] A division module, configured to divide the video to be measured to obtain multiple video segments.

[0078] A C3D module, configured to input each video segment into a C3D model to obtain a C3D feature map corresponding to each video segment; the C3D module is further configured to input the standard action video into the C3D model to obtain standard action sample features.

[0079] A center judgment module, configured to input the C3D feature map into a center judgment model to screen out the central video segment of each instance to be measured; the central video segment of each instance to be measured is the video segment corresponding to the central moment of each instance to be measured; the center judgment model is obtained by training a regression model with a sample C3D feature map.

[0080] A candidate region generation module, configured to expand from the center point of the central video segment of each instance to be measured to both sides of the central video segment with different frames to generate multiple candidate regions for each instance to be measured.

[0081] An alignment module for aligning the C3D feature map and the multiple candidate regions by using an adaptive pooling layer to obtain an aligned feature map for each candidate region.

[0082] A prediction module for inputting the aligned feature map of each candidate region and the standard action sample feature into an action prediction model to obtain an action prediction result for each candidate region; the action prediction model is obtained by training a siamese network with a sample training set; the sample training set includes: sample aligned feature maps, standard action sample features, and corresponding sample prediction results.

[0083] An action determination module for optimizing the action prediction result of each instance to be measured by using a non-maximum suppression method to obtain the action type, start time, and end time of each instance to be measured, and completing action localization.

[0084] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0085] Specific examples are used in this article to elaborate on the principles and implementation manners of the present application. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. An action localization method based on multi-instance small samples, characterized in that, The action localization method based on multi-instance small samples includes: Obtain the video to be tested and the standard action video; the video to be tested includes multiple instances to be tested; one instance is an action. Divide the video to be tested to obtain multiple video segments. Input each video segment into the C3D model to obtain the C3D feature map corresponding to each video segment. Input the standard action video into the C3D model to obtain the standard action sample features. Input the C3D feature map into the center judgment model to screen out the central video segment of each instance to be tested; each central video segment of the instance to be tested is the video segment corresponding to the central moment of each instance to be tested; the center judgment model is obtained by training the regression model with the sample C3D feature map. Expand from the center point of each central video segment of the instance to be tested to both sides of the central video segment with different frames to generate multiple candidate regions for each instance to be tested. Use the adaptive pooling layer to align the C3D feature map and the multiple candidate regions to obtain the aligned feature map of each candidate region. Input the aligned feature map of each candidate region and the standard action sample features into the action prediction model to obtain the action prediction result of each candidate region; the action prediction model is obtained by training the siamese network with the sample training set; the sample training set includes: sample aligned feature map, standard action sample features and corresponding sample prediction results. Use the non-maximum suppression method to optimize the action prediction results of each candidate region to obtain the action type, start time and end time of each instance to be tested, and complete the action localization.

2. The action localization method based on multi-instance small samples according to claim 1, wherein Input the C3D feature map into the center judgment model to screen out the central video segment of each instance to be tested, specifically including: Input the C3D feature map into the center judgment model to obtain the probability value corresponding to the C3D feature map. Judge whether the probability value is greater than the probability threshold. If so, screen out the video segment corresponding to the C3D feature map as the central video segment. If not, delete the video segment corresponding to the C3D feature map.

3. The action localization method based on multi-instance small samples according to claim 2, characterized in that, The training process of the center judgment model specifically includes: Input the sample C3D feature map into the regression model to obtain the detection probability value. Construct a loss function according to the detection probability value and the sample probability value corresponding to the sample C3D feature map, and iteratively optimize the parameters of the regression model according to the loss function until the loss function reaches the minimum value or the number of iterative optimization rounds reaches the maximum value, and stop the iterative optimization to obtain the center judgment model.

4. The action localization method based on multi-instance small samples according to claim 1, characterized in that The action prediction result includes: action similarity, candidate region center point offset and candidate region length offset; the action similarity is the similarity between the action corresponding to each aligned feature map and the action corresponding to the standard action sample features. Use the non-maximum suppression method to optimize the action prediction results of each candidate region to obtain the action type, start time and end time of each instance to be tested, and complete the action localization, specifically including: Determine the action type of the action corresponding to the candidate region according to the action similarity and the similarity threshold. Determine the start time and end time of the action corresponding to the candidate region according to the center point offset of the candidate region and the length offset of the candidate region; Judge whether the overlap rate between the start time and the end time of the action corresponding to the candidate region reaches the overlap threshold; If so, determine that there are actions corresponding to multiple candidate regions for the instance to be measured, and use the action corresponding to the candidate region with the highest action similarity as the action type, start time, and end time of the corresponding instance to be measured; If not, determine that there is only one action corresponding to the candidate region for the instance to be measured, and use the action corresponding to the candidate region as the action type, start time, and end time of the corresponding instance to be measured, and complete action localization.

5. The action localization method based on multi-instance small samples according to claim 4, wherein Determine the action type of the action corresponding to the candidate region according to the action similarity and the similarity threshold, specifically including: Judge whether the action similarity is greater than the similarity threshold; If so, determine that the action type of the action corresponding to the candidate region is the action type corresponding to the standard action sample feature and retain it; If not, determine that the action type of the action corresponding to the candidate region is inconsistent with the action type corresponding to the standard action sample feature and delete it.

6. The method for action localization based on multi-instance small samples according to claim 4, characterized in that The start time of the action corresponding to the candidate region is equal to the center point offset of the candidate region minus half of the length offset of the candidate region; The end time of the action corresponding to the candidate region is equal to the center point offset of the candidate region plus half of the length offset of the candidate region.

7. The action localization method based on multi-instance small samples according to claim 1, characterized in that The training process of the action prediction model specifically includes: Input the sample alignment feature map and the standard action sample feature into the siamese network to obtain a prediction result; Construct a loss function according to the prediction result and the sample action prediction result corresponding to the sample alignment feature map and the standard action sample feature, and iteratively optimize the parameters of the siamese network according to the loss function until the loss function reaches the minimum value or the number of iterative optimization rounds reaches the maximum value, stop iterative optimization, and obtain the action prediction model.

8. An action localization system based on multi-instance small samples, characterized in that, The action localization system based on multi-instance small samples includes: An acquisition module for acquiring a video to be measured and a standard action video; the video to be measured includes multiple instances to be measured; one instance is one action; A division module for dividing the video to be measured to obtain multiple video segments; A C3D module for inputting each video segment into the C3D model to obtain a C3D feature map corresponding to each video segment; the C3D module is also used for inputting the standard action video into the C3D model to obtain standard action sample features; A center judgment module for inputting the C3D feature map into the center judgment model to screen out the central video segment of each instance to be measured; each central video segment of each instance to be measured is the video segment corresponding to the central moment of each instance to be measured; the center judgment model is obtained by training a regression model with a sample C3D feature map; A candidate region generation module for expanding from the center point of each central video segment of each instance to be measured to both sides of the central video segment by different frames to generate multiple candidate regions for each instance to be measured; An alignment module, configured to perform alignment processing on the C3D feature map and the multiple candidate regions by using an adaptive pooling layer to obtain an aligned feature map for each candidate region; A prediction module, configured to input the aligned feature map of each candidate region and the standard action sample feature into an action prediction model to obtain an action prediction result for each candidate region; the action prediction model is obtained by training a siamese network with a sample training set; the sample training set includes: sample aligned feature maps, standard action sample features, and corresponding sample prediction results; An action determination module, configured to optimize the action prediction result of each candidate region by using a non-maximum suppression method to obtain the action type, start time, and end time of each instance to be measured, thereby completing action localization.

Citation Information

Patent Citations

  • Action segment detection method and device, model training method and device

    CN113033500A

  • Video behavior identification method and system

    CN114581819A

  • Online end-to-end space-time action detection method and detector

    CN115439923A

  • Action recognition using limited data

    US20220318555A1

  • Prior-driven supervision for weakly-supervised temporal action localization

    US20240404279A1