Point supervision time sequence motion positioning method based on cooperation of motion discrimination and boundary sensitivity
Through a hierarchical learning strategy that coordinates action discriminative and boundary sensitivity, combined with discriminative guidance propagation and boundary sensitivity timing modeling, the shortcomings of existing algorithms in action boundary positioning and integrity are solved, and more efficient action proposals and positioning effects are achieved.
Patent Information
- Application Number
- CN202510235466.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-07-11
AI Technical Summary
The existing point-supervised timing action positioning algorithms have shortcomings in action boundary positioning and integrity, especially in complex and dense action examples, which are difficult to capture global information and timing structures, and fuzzy fragments may introduce interference information, resulting in incomplete action proposals.
A hierarchical learning strategy that coordinates action discrimination and boundary sensitivity is adopted, and a high-quality pseudo-label and proposal are generated through the discriminant guide propagation mechanism and boundary sensitivity, combining differential convolution and timing perturbation operations.
It improves the accuracy of action boundary positioning and the completeness of proposed expressions, enhances the model's precise positioning ability of action boundaries, and improves the overall performance of action positioning.
Smart Images

Figure CN120299080A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of computer vision and action localization, and specifically relates to a point-supervised temporal action localization method that combines a discriminative-guided knowledge propagation mechanism and a boundary-sensitive temporal modeling strategy based on deep learning technology, jointly forming action discriminability and boundary sensitivity in combination with visual features, and is applicable to fields such as intelligent monitoring, anomaly detection, and visual question answering. Background Art
[0002] With the rapid development of the Internet and the popularization of electronic shooting devices, the number of videos has shown an explosive growth. Thanks to the powerful feature extraction ability of deep learning, temporal action localization algorithms have made significant progress in the field of video understanding, and the detection performance has also been greatly improved. However, the temporal action localization task relies on accurate boundary annotations for training. In practice, it is extremely challenging to obtain large-scale and high-quality boundary annotations, not only due to the high cost of the annotation process itself, but also because manual annotation of long temporal videos is easily affected by subjective factors, resulting in poor consistency. To balance the labeling cost and detection performance, researchers have begun to turn to point-supervised temporal action localization methods (PTAL), which only require providing a single timestamp label for each action instance. Research shows that the point-level annotation cost is almost equivalent to video-level annotations (saving 6 times the cost compared to frame-level annotations), and at the same time provides more abundant guiding information, which has received extensive attention in the academic and industrial fields.
[0003] Existing PTAL algorithms usually follow the multi-instance learning (MIL) framework, generating segment-level class activation scores (CAS) through classification supervision, and then thresholding to select action proposals. CAS represents the probability of an action occurring. It can be seen that the quality of CAS determines the upper limit of the model performance. However, due to the sparsity of point labels, it is impossible to guide the complete learning of action boundaries. Coupled with the differences in classification and localization optimization objectives, the model tends to only focus on highly discriminative action regions, resulting in incomplete action proposals. Although some studies have tried to alleviate label sparsity by mining pseudo-action points or background points, these points are usually discontinuous and difficult to effectively capture the global information and temporal structure of action instances, as shown in Figure 1 (a) of, especially more difficult when dealing with complex and dense action instances. In addition, existing studies usually use methods such as self-attention mechanisms and transformers to model the temporal relationships between segments to capture more complete action information. However, while these methods focus on the relationships between segments, they do not specifically address the adverse effects of fuzzy segments on the overall features. Fuzzy segments may introduce interfering information during the feature propagation process, thereby weakening the expression of discriminative features.
[0004] Combined with the above limitations, starting from the inherent reliability of point labels, the present invention introduces a hierarchical learning strategy from feature discriminative learning to boundary sensitivity enhancement, forming a collaborative optimization that squeezes the boundary from both positive and negative directions. The first level focuses on enhancing the discriminability of positive segments. By suppressing fuzzy segments and strengthening the transmission of discriminative features, high-quality boundary pseudo-labels are generated to provide reliable guidance for subsequent optimization. The second level, with the constraint of pseudo-labels, introduces a boundary-sensitive temporal modeling mechanism to refine the expression of the action boundary region in the reverse direction, ultimately achieving precise positioning of the action boundary and improvement of integrity, as shown in Figure 1 the (b) of Summary of the Invention
[0005] The present invention proposes a collaborative algorithm for action discriminability and boundary sensitivity (SADBS), adopting a collaborative optimization strategy that squeezes the boundary from both positive and negative directions. First, a discriminative-guided propagation mechanism is introduced to explicitly model the semantic information of discriminative segments and fuzzy segments. This method promotes the knowledge propagation of discriminative information to the fuzzy region and inhibits the noise diffusion in the fuzzy region. At the same time, combined with the center contrast loss and feature consistency loss constraints, the discriminability of segment features is enhanced to generate high-quality class activation scores (CAS). Further, to overcome the limitations of CAS in boundary modeling, the present invention designs a boundary-sensitive temporal modeling mechanism to enhance the sensitivity of the boundary and the robustness of proposal expression. By orderly modeling action features and strengthening their structural expression, and at the same time combining disordered processing to capture potential associations across time periods. On this basis, the feature discriminative learning of segments is realized to the enhancement of boundary sensitivity of proposals, thereby improving action localization.
[0006] To achieve the above object, the present invention adopts the following technical solution as a point-supervised temporal action localization method based on the collaboration of action discriminability and boundary sensitivity. The implementation steps of this method are as follows:
[0007] Step 1: Construct a feature extraction module. First, divide the input unclipped video into non-overlapping segments, and apply the two-stream network I3D to extract the scene feature RGB and motion feature Flow, and then stack them along the temporal dimension to construct a video feature representation;
[0008] Step 2: Segment discriminative learning. According to the input video features, first divide the video segments into discriminative segments and fuzzy segments, and model the relationship between the two through a specific connection strategy, so that the fuzzy segments can receive supplementary information from the discriminative segments, but cannot spread their own information back to the discriminative segments. At the same time, combine the feature consistency loss to constrain the coherence of feature expression, and strengthen the difference between different categories through the center contrast loss to further improve the discriminability of the overall features.
[0009] Step 3: Proposal Integrity Refinement. Discriminative learning effectively alleviates the impact of ambiguous information on the discriminability of overall features by enhancing the feature representation ability in ambiguous regions. However, it is difficult to comprehensively capture the complete time range of an action solely relying on discriminative learning, especially in accurately locating the boundary regions. To further strengthen the boundary features and optimize the action integrity modeling, the present invention first applies differential convolution to the temporal domain and designs a boundary-sensitive temporal modeling mechanism. This mechanism effectively combines differential composite convolution and temporal perturbation strategy. Differential composite convolution not only continues its advantage in capturing local gradient information but also captures the evolution of action boundaries by adapting to the temporal dynamic features of the video, significantly enhancing the model's sensitivity to the start and end positions of actions. The temporal perturbation strategy, on the other hand, breaks the inherent temporal constraints in features and encourages the model to capture potential correlations across time periods. This complementary processing method improves the temporal integrity of action proposals and the robustness of feature representation.
[0010] Steps 2 and 3 implement a hierarchical learning strategy from segment discriminative learning to proposal integrity refinement, improving the accuracy of action boundary localization and the integrity of proposal expression under point supervision.
[0011] Step 4: Present the joint optimization and model localization inference process. Brief Description of the Drawings
[0012] Figure 1 It is a comparison diagram between the existing method and the method proposed by the present invention.
[0013] Figure 2 It is the overall framework diagram of the network of the present invention.
[0014] Figure 3 Schematic diagram of the boundary-sensitive temporal modeling mechanism.
[0015] Figure 4 It is the result diagram of qualitative visualization analysis. Detailed Description of the Specific Embodiment
[0016] The following details the specific implementation of the present invention with reference to the accompanying drawings.
[0017] The technical solution adopted by the present invention is a point - supervised temporal action localization method based on the collaboration of action discriminability and boundary sensitivity. By introducing a hierarchical learning strategy from segment discriminative learning to proposal integrity refinement, action localization is improved. The system modules for implementing this method include a feature extraction module, a segment discriminative learning module SDL, and a proposal integrity refinement module PIR. The feature extraction module is used for feature extraction and serves as the input data for subsequent processes. SDL first groups video segments and then constructs a graph through a discriminative - guided propagation mechanism for knowledge propagation, realizing the effective transfer of discriminative information to fuzzy regions while suppressing the diffusion of fuzzy information. Combining the feature consistency loss and the center contrast loss constraints further optimizes the feature distribution and generates high - quality pseudo - labels. PIR uses the pseudo - labels as supervision and introduces a boundary - sensitive temporal modeling mechanism to focus on the boundary optimization and integrity modeling of proposals. By complementing the structural reinforcement of the ordered processing of action features with the flexibility of the disordered processing between segments, the robustness of proposal expression is enhanced. Finally, the present invention presents the joint loss optimization and action inference process.
[0018] The overall framework diagram of the technical solution of the present invention is as Figure 2 shown. Further, the feature extraction module is trained on a dataset containing daily activities and sports events. Given an untrimmed video V with T segments, a single timestamp is annotated for each action instance, obtaining a point set where N is the total number of instances in the video, t n , y n represent the annotated timestamp and action category respectively. y n ∈{1,2,...,C}, C indicates the sum of all action classes, and the present invention also adds the (C + 1) - th class to represent the background class.
[0019] Following the conventional PTAL algorithm for processing, first, the pre - trained I3D is used to extract scene features X R ∈R T×D and motion features X F ∈R T×D from RGB and Flow respectively, and then the two are concatenated along the channel dimension to form the video feature representation X∈R T×2D , where D is the feature dimension, and these are used as the input data for subsequent processes.
[0020] The implementation process of the segment discriminative learning module (SDL) is as follows:
[0021] 1) Graph construction
[0022] The present invention constructs a graph structure G=(V, E, A) for video segments, where V represents the node set of the entire video segment, E represents the edge set between nodes, and A is its adjacency matrix. Based on the difference in attention values, the node set V will be divided into a discriminative action segment set V a , a discriminative background segment set V b , and a fuzzy segment set V m . Since E and A are in one-to-one correspondence, they are abbreviated as G=(V, A) in the following text.
[0023] First, the features X R and X F are respectively input into the attention module to calculate the segment-level attention weights, which includes a convolutional layer and RELU, and M R and M F are obtained respectively.
[0024] Then, by jointly considering M R and M F , and taking the action weight of the segment as the judgment basis, V is divided into V a , V b , and V m . Specifically, set thresholds μ a and μ b . For each time point t, t∈{1, 2,.., T}, if and , it is regarded as an action segment. If and , it is regarded as a background segment. Otherwise, it is used as a fuzzy segment. It is expressed as follows:
[0025]
[0026] On this basis, three subgraphs are defined, which are respectively used for information interaction and propagation between different segments. The undirected action graph is composed of V a and its symmetric adjacency matrix A a , denoted as G a =(V a , A a ). The edge weight of the adjacency matrix A a (v i , v j ) is defined by the cosine similarity between features. It should be noted that in order to ensure the consistency of video time series, in the process of graph modeling of the present invention, the size of the calculated matrix always matches the number of video segments.
[0027]
[0028] Similarly, the undirected background graph is G b =(V b , Ab ), where A b Adopts the same calculation method as A a .
[0029] Different from this, the directed fuzzy graph G m Is composed of V and its asymmetric adjacency matrix A m , and the direction of the edge always goes from V a And V b Points to V m . This structure is used to transfer discriminant information to the fuzzy segment to prevent the reverse propagation of fuzzy information to the discriminant node, thereby enhancing the feature discriminability of the fuzzy segment, that is, G m =(V, A m ). For G m , its A m (v i , v j ) is defined as follows:
[0030]
[0031] After completing the subgraph partitioning and adjacency matrix construction, knowledge learning based on subgraph propagation will be carried out, aiming to utilize the characteristics of different subgraphs to achieve the effective propagation of discriminant information and enhance the representation ability of the overall features.
[0032] 2) Knowledge learning based on subgraph propagation
[0033] For G a And G b , information interaction is carried out by using the internal structure and feature correlation of the discriminant segment. Specifically, the present invention directly performs graph averaging and graph convolution modeling, and residual connection to obtain enhanced features. Taking the action graph as an example, the calculation is as follows:
[0034]
[0035] Among them, Is the output result of the L-layer GCN.
[0036] For the fuzzy graph, since the fuzzy information itself lacks a clear structure, GCN modeling is not applied alone, and directly use And To aggregate, the graph averaging calculation is the same as above.
[0037]
[0038] By using the discriminant feature as a guide, the fuzzy region can be better integrated into the learning process without amplifying the uncertainty. After the segments are modeled, they are arranged in chronological order, and the original features are connected residually as the subsequent input, denoted as X e .
[0039] Enhanced feature X e is then input into the segment-level classifier to obtain the class activation sequence (CAS) S ∈ R T×(C+1) , and the convolutional layer calculates the action score K ∈ R T , and the two are element-wise multiplied to obtain the final score P ∈ R T×(C+1) , and finally the video-level prediction P vid ∈ R C+1 .
[0040] Pseudo-label mining. To alleviate the sparsity of labels, based on the feature that each action instance contains a point annotation and adjacent point annotations are in different action instances, according to the point annotation and the action score, by setting a threshold λ act , the segments with scores greater than the threshold near the point annotation are marked as pseudo-action segments, which share the same action category as the point annotation. Conversely, the segments located in the middle of two annotation points and having a score lower than the threshold λ bkg are marked as pseudo-backgrounds, and Λ + ={t i;c |i ∈ [1, N act , c ∈ [1, C]} and Λ - ={t j;C+1 |j ∈ [1, N bkg} are obtained respectively. The present invention uses these pseudo-segments as additional supervision to improve the model's ability to separate actions and backgrounds.
[0041]
[0042] where N act and N bkg respectively represent the total number of pseudo-action and pseudo-background segments, λ1 is a parameter, FF(·) represents a focusing function, and the cross-entropy classification loss L vid at the video level.
[0043] Feature consistency loss. To prevent the over-smoothing of GCN modeling and ensure the difference of highly discriminative features. The present invention adds constraints between the graph convolutional features and the average features of the action graph and the background graph respectively to encourage the separability between features.
[0044]
[0045] Center contrast loss. The center contrast loss calculates the similarity between the segment features in the pseudo-label set (Λ + and Λ - ) and their corresponding class centers, improves the aggregation of intra-class features, and increases the difference between classes. Taking the action center contrast loss as an example, the calculation is as follows:
[0046]
[0047] Among them, cos(·) represents the cosine distance between two vectors, and γ i is the binary label of the pseudo-action point, which is used to indicate whether it belongs to the action class c.
[0048] Proposal generation. The present invention first screens the action classes to be located on P vid , and then adopts a multi-threshold strategy to select class-specific activation score segments to generate candidate proposals for each predicted class, denoted as Props = {(s n , e n , c n , p n )}. The reliability of each action proposal is measured by the OIC (outer-inner-contrast) score, denoted as p n . A lower reliability score indicates an incomplete or over-complete prediction.
[0049] In addition, the candidate proposals are divided into reliable proposals (RP) according to the confidence score, that is, for each point annotation in the video, the proposal contains the point annotation and has the highest confidence score; and positive proposals (PP), that is, all the remaining candidate proposals. To achieve positive and negative sample balance, the present invention groups the segments with action scores lower than a predefined threshold to generate negative proposals (NP). At the same time, to improve the proposal recall rate, the present invention expands the boundaries of the positive proposals, defines all the segments detected by the proposals as the central region (R c ), and expands the proposal boundaries according to a preset ratio (ε = 0.25) to form the start region (R s ) and the end region (R e ). The reliable proposal RP with the highest confidence will be selected as the pseudo-label to further optimize the boundary and provide guidance for the subsequent refinement of the positive proposal integrity.
[0050] Furthermore, the implementation process of the proposal integrity refinement module (PIR) is as follows:
[0051] Since fragment-level learning lacks clear boundary information, it is often difficult to effectively guide the integrity clues of actions, thus affecting the accuracy of action localization. To solve this deficiency, the present invention introduces a boundary-sensitive temporal modeling mechanism, combines differential composite convolution (DCC) and temporal perturbation operation (TDO) for dual processing to strengthen boundary information extraction, and improves the accuracy of action localization, as shown in Figure 3 part (a) of.
[0052] The differential composite convolution (DCC) parallelly deploys 4 convolutional layers (three differential convolutions and one ordinary convolution). The differential convolutions include the central difference convolution (CDC), the angular difference convolution (ADC), and the horizontal difference convolution (HDC), which calculate the gradient information for selected feature pairs from different directions to capture the changes of features at different angles. The ordinary convolution is used to obtain intensity-level information, while the differential convolutions focus on enhancing the gradient-level information. To avoid introducing additional computational amount and computational cost, the present invention uses a reparameterization technique, as shown in part (b) of Figure 3 , adding multiple convolutional kernels of DCC at corresponding positions, so that the gradient prior is effectively encoded into the ordinary convolution to obtain an equivalent kernel that can produce the same output. Through this design, DCC can effectively encode fine details, especially boundaries, and improve the modeling ability of complex features. The calculation is as follows (the bias is omitted for simplicity):
[0053]
[0054] where DEConv(·) represents the differential composite convolution operation proposed by the present invention, and K i=1:4 represent VC, CDC, HDC, and ADC respectively. * represents the convolution operation, and K all represents the equivalent kernel.
[0055] X e obtains X after passing through Relu, residual connection, and convolutional layers o . Then, they are placed in sequence to calculate the attention weights in the temporal dimension and the feature dimension. The temporal attention W T ∈R T×1 is used to capture temporal changes, especially boundary sequence information. The feature attention W F ∈R 1×D allocates weights in the feature dimension to highlight the responses of important features. The calculation is as follows:
[0056]
[0057] where C k×k represents the convolutional layer with a k×k convolutional kernel, and [·,·] represents the concatenation operation. and represent the global average pooling and global maximum pooling features respectively. Then, W T and W F are fused through the broadcast rule to obtain the rough attention W cos .
[0058] Then, the input features are introduced as guidance to refine the weights, and the original temporal order is broken through the temporal shuffle operation, thereby introducing certain perturbations for modeling temporal features. This helps the model capture the potential connections between different segments more flexibly and improve the robustness of action expression, as shown in Figure 3 part (c).
[0059] F = sigmoid(CS([X o , W cos )) * X o + X o (11)
[0060] where CS(·) represents the temporal shuffle operation, which is implemented by group convolution.
[0061] According to the candidate proposal regions R s , R c and R e , the corresponding region features are mapped to F, and then and are concatenated along the temporal dimension and fed into the integrity score head (ISH) to predict the integrity score P comp .
[0062]
[0063] where respectively represent the maximum pooling features in the time dimension. This proposal-level integrity score process is supervised by the matching RP:
[0064]
[0065] where N p , N n are the total numbers of positive and negative proposals respectively, and H IoU represents the IoU between the proposal and the RP.
[0066] To obtain more accurate proposal boundaries, the present invention feeds into the boundary regression head (BRH) to predict the offsets of the start time and the end time, and then obtain the refined boundaries.
[0067]
[0068] where ω p = I e - I s . This boundary regression process is constrained by reliable proposals and calculated as follows:
[0069]
[0070] Among them, R comp represents the IoU between the optimized proposal and the RP.
[0071] Overall optimization function. During the training phase, the model is guided by multiple loss functions to ensure comprehensive optimization. These loss functions include the video classification loss L vid , the point-level loss L point , the feature consistency loss L fc , the center contrast loss L center , the proposal classification loss L cls and the regression loss L reg . This overall loss function effectively balances the classification accuracy, boundary accuracy, and feature quality. The calculation is as follows:
[0072] L total = λ2L vid + L point + L fc + L center + L cls + L reg (16)
[0073] Among them, λ2 is a learnable parameter.
[0074] Model inference. During the inference phase, first, the initial action proposals are generated in the proposal generation stage, and the proposal feature representation is refined through a boundary-sensitive temporal modeling mechanism. Subsequently, they are input into the BRH and ISH to generate more refined proposals {(s' n , e' n , c n , p n + p comp )}. Finally, the high-overlap proposals are removed using NMS to obtain the final localization results.
[0075] Experimental part
[0076] Experimental datasets: The method of the present invention is evaluated using the publicly available and highly challenging THUMOS14 and ActivityNet1.3 datasets. THUMOS14 covers 413 untrimmed videos, involving 20 action classes. The model is trained on 200 validation videos and evaluated on 213 test videos. ActivityNet1.3 contains 19,994 untrimmed videos from 200 action classes, which are divided into a training set, a validation set, and a test set in a ratio of 2:1:1. The present invention is trained on the training set and tested on the validation set.
[0077] Evaluation Metrics: Following the standard evaluation protocol, the mean Average Precision (mAP) at different Intersection over Union (IoU) thresholds is used as the metric to evaluate the performance of point-supervised temporal action localization, denoted as mAP@IoU. Specifically, in the ablation experiments, the detection evaluation metric is initialized as follows: on the THUMOS14 dataset, IoU is set to [0.1:0.1:0.7], and on the ActivityNet1.3 dataset, IoU is set to [0.5:0.25:0.95].
[0078] Experimental Settings: In the present invention, the unclipped video is divided into 16-frame segments, and the I3D network pre-trained on the Kinetics dataset is used to extract RGB and Flow features, where the Flow maps are generated using the TV-L1 algorithm. For fairness, no fine-tuning operations are introduced to the I3D network in the present invention.
[0079] The model training adopts a two-stage approach. In the first stage on THOUMS14, it is set to 2000 epochs, while on ActivityNet1.3 it is 5000 epochs, and in the second stage, it is 30 epochs for both. The training batch sizes are set to 16 and 512 respectively. The Adam optimizer is used for model optimization, with the learning rate set to 0.0001 on THOUMS14 and 0.0005 on ActivityNet1.3. The feature dimension D is set to 1024, μ a and μ b is set to 0.7, the parameters λ1 and λ2 are 0.5, λ act and λ bkg are set to 0.1 and 0.95 respectively. Finally, to remove highly overlapping action proposals, the IoU of non-maximum suppression (NMS) is set to 0.7. The entire model is implemented using PyTorch 1.7.
[0080] The present invention evaluates the SADBS method on the THUMOS14 and ActivityNet1.3 datasets and demonstrates advanced performance compared with the state-of-the-art localization methods. Secondly, the present invention conducts ablation studies on different components, loss functions, etc., and the specific data analysis is as follows.
[0081] Table 1: Contribution of the proposed algorithm components to the model in ablation on the THUMOS14 dataset. + indicates the gain of each setting compared to the baseline.
[0082]
[0083]
[0084] Effectiveness of Ablating Each Module. The present invention conducts a series of ablation studies on the proposed SADBS (as shown in Tables 1 and 2), and deeply analyzes the contributions of different components relative to the baseline setting. In the experiment, "BTMM" represents the boundary-sensitive timing modeling mechanism, "BIH" indicates the boundary regression head and integrity classification head, and "DKPM" indicates the discriminative guidance propagation mechanism, with the one without the above components as the Baseline.
[0085] As can be seen from Table 1, compared with the baseline, the SADBS proposed by the present invention achieves the best performance in various metrics. This indicates that each component is necessary for achieving the best performance, and each component is effective. Specifically, on the basis of the baseline, the present invention introduces BTMM and BIH. Although the model can capture the boundary information of action instances more meticulously and improve the accuracy of action expression, due to the low quality of the proposed pseudo-labels, the performance gain is limited, with an average gain of 1.2%. On this basis, the DKPM module is further introduced, and the graph knowledge propagation mechanism significantly enhances the feature discriminability, thus generating high-quality pseudo-labels. The combination of these three modules enables the model performance to reach the maximum gain, with an average increase of 6.2%, ensuring the integrity and accuracy of action localization.
[0086] Table 2: Ablation of the contributions of the proposed loss function, differential composite convolution (DCC), and temporal perturbation operation (TDO) components to the model on THUMOS14.
[0087]
[0088]
[0089] Effectiveness of Ablating DCC and TDO. The last 3 rows of Table 2 are the ablation of the ordered processing (differential composite convolution) and unordered processing (temporal shuffle operation) in the temporal harmonization mechanism. The experiment shows that the combination of the ordered and unordered processing methods can take into account both the boundary precision and the flexibility of temporal information. The differential composite convolution strengthens the action boundary recognition ability through ordered processing, while the temporal shuffle operation increases the model's expression ability in complex time series through unordered processing. The combination of the two makes the model perform more comprehensively in capturing action details and localization accuracy, improving the robustness of the proposal.
[0090] Ablation of L fc and L center Effectiveness of the Loss. Table 2 deeply explores the influence of the feature consistency loss L fc and the center contrast loss L center on the model performance. Specifically, L fc effectively avoids the over-smoothing of discriminative features during the knowledge propagation process, significantly improving the stability and reliability of the features. And Lcenter By constraining the consistency of the distribution of fragments with the same attributes in the feature space, the inter-class gap is further enlarged, thereby enhancing the discriminative ability of the model and enabling the overall performance to reach 57.1%. The synergistic effect of these two loss functions enables the model to maintain higher discriminability during the sub-graph knowledge propagation process, generate high-quality pseudo-labels, and ultimately achieve a significant improvement in action localization performance, with an overall gain of 5.9%.
[0091] Ablation of the effectiveness of differential convolution. The present invention explores the impact of parallel convolution design on the model in rows 5 and 6 of Table 2, where VC represents ordinary convolution. It can be seen from the experimental results that when two identical ordinary convolutions (VC) are deployed in parallel, the average precision of the model is 60.0% @AVG. Although there is a certain improvement compared to the Baseline (54.4%), due to the extraction of redundant features between convolutional layers, the training process is relatively difficult and the performance improvement is limited. However, with the design of detail-enhanced convolution (DCC), by parallelly deploying ordinary convolution and three differential convolutions, it can more effectively capture the gradient information and boundary details in the features, significantly enhancing the expression ability and boundary modeling ability of the model, and finally achieving an average precision of 60.5%. Qualitative visual analysis. To more intuitively verify the effectiveness of the method proposed in the present invention, Figure 4 The local regions of the "hockey penalty shot" and "javelin throw" actions and their CAS results are visually displayed. Compared with the baseline method, the CAS distribution generated by the SADBS model proposed in the present invention is more abundant, and it can more accurately locate the action boundaries in the video. Specifically, for the "hockey penalty shot" action, SADBS introduces a boundary-sensitive temporal modeling mechanism to help the model better capture action details and boundary information, making up for the problem of ignoring low-discriminative features caused by sparse supervision and improving the integrity of action localization. In the "javelin throw" action, SADBS constructs graph knowledge propagation to enhance the distinctiveness between different features, generates high-quality CAS, and thus improves the accuracy of action localization.
Claims
1. A point-supervised temporal action localization method based on the collaboration of action discriminability and boundary sensitivity, characterized in that The implementation steps of this method are as follows: Step 1: Construct a feature extraction module; divide the input unclipped video into non-overlapping segments, apply the two-stream network I3D to extract the scene feature RGB and the motion feature Flow, and then stack them along the temporal dimension to construct the video feature representation; Step 2: Segment discriminative learning; according to the input video features, divide the video segments into discriminative segments and ambiguous segments, and model the relationship between the two through a specific connection strategy, so that the ambiguous segments can receive complementary information from the discriminative segments without propagating their own information back to the discriminative segments; Combine the feature consistency loss to constrain the coherence of the feature expression, and strengthen the difference between different categories through the center contrast loss to improve the discriminability of the overall features; Step 3: Proposal integrity refinement; discriminative learning alleviates the impact of ambiguous information on the discriminability of the overall features by enhancing the feature expression ability of the ambiguous regions; To strengthen the boundary features and optimize the action integrity modeling, the differential convolution is first applied to the temporal domain, and a boundary-sensitive temporal modeling mechanism is designed; this temporal modeling mechanism combines the differential composite convolution and the temporal perturbation strategy to capture the evolution of the action boundary by adapting to the temporal dynamic features of the video; Step 4: Give the joint optimization and model localization inference process.
2. The point-supervised temporal action localization method based on the collaboration of action discriminability and boundary sensitivity according to claim 1, characterized in that The system modules for implementing this method include a feature extraction module, a segment discriminative learning module SDL, and a proposal integrity refinement module PIR; the feature extraction module is used for feature extraction and serves as the input data for subsequent links; SDL first groups the video segments, and then constructs a graph through the discriminative guidance propagation mechanism for knowledge propagation, realizes the effective transmission of discriminative information to the ambiguous regions, and at the same time suppresses the diffusion of ambiguous information; Combine the feature consistency loss and the center contrast loss constraints to further optimize the feature distribution and generate high-quality pseudo-labels; PIR uses the pseudo-labels as supervision and introduces a boundary-sensitive temporal modeling mechanism to focus on the boundary optimization and integrity modeling of the proposals; by complementing the structural reinforcement of the ordered processing of action features and the flexibility of the disordered processing between segments, the robustness of the proposal expression is improved; finally, joint loss optimization and action inference are performed.
3. A point-supervised temporal action localization method based on the collaboration of action discriminability and boundary sensitivity according to claim 2, characterized in that The feature extraction module is trained on a dataset containing daily activities and sports events; given an unclipped video V with T segments, it annotates a single timestamp for each action instance to obtain a point set where N is the total number of instances in the video, t n , y n represent the annotated timestamp and action class respectively; y n ∈ {1, 2,..., C}, where C indicates the sum of all action classes, and the (C + 1)-th class is added to represent the background class; For the PTAL algorithm processing, first, the pre-trained I3D is used to extract the scene feature \(X^{(rgb)}\in\mathbb{R}^{D^{(rgb)}}\) from RGB and the motion feature \(X^{(flow)}\in\mathbb{R}^{D^{(flow)}}\) from Flow respectively. Then, the two are concatenated along the channel dimension to form the video feature representation \(X\in\mathbb{R}^{D}\), where \(D\) is the feature dimension and serves as the input data for subsequent steps. R ∈R T×D and the motion feature \(X^{(flow)}\) F ∈R T×D , and then the two are concatenated along the channel dimension to form the video feature representation \(X\in\mathbb{R}^{D}\) T×2D , where \(D\) is the feature dimension and serves as the input data for subsequent steps.
4. A point-supervised temporal action localization method based on the collaboration of action discriminability and boundary sensitivity according to claim 2, characterized in that The implementation process of the Segment Discriminative Learning module SDL is as follows: 1) Graph construction; construct a graph structure G=(V, E, A) for a video segment, where V represents the node set of the entire video segment, E represents the edge set between nodes, and A is its adjacency matrix; based on the difference in attention values, the node set V will be divided into a discriminative action segment set V a , a discriminative background segment set V b and a fuzzy segment set V m ; E and A correspond one by one, briefly denoted as G=(V, A); First, input feature X R and X F into the attention module respectively to calculate the segment-level attention weights, which includes a convolutional layer and RELU, and obtain M R and M F ; Then, by jointly considering M R and M F , based on the action weights of the segments, divide V into V a , V b and V m ; specifically, set thresholds μ a and μ b . For each time point t, t ∈ {1, 2,.., T}, if and then it is regarded as an action segment. If and are regarded as background segments. Otherwise, it is regarded as a fuzzy segment; it is expressed as follows: Define three subgraphs for information interaction and dissemination between different segments respectively; the undirected action graph consists of V a and its symmetric adjacency matrix A a to form, denoted as G a =(V a , A a ); the edge weight of the adjacency matrix A a (v i , v j ) is defined by the cosine similarity between features; to ensure the consistency of video timings, during the graph modeling process, the size of the calculated matrix always matches the number of video segments; The undirected background graph is G b =(V b , A b ), where A b adopts the same calculation method as A a ; Directed fuzzy graph G m is composed of V and its asymmetric adjacency matrix A m , and the direction of the edges always goes from V a to V b ; this structure is used to transfer discriminative information to the fuzzy segments to prevent the reverse propagation of fuzzy information to the discriminative nodes, thereby enhancing the feature discriminability of the fuzzy segments, that is, G m m = (V, A m ); for G m , its A m (v i , v j ) edge weights are defined as follows: After completing the subgraph division and adjacency matrix construction, perform knowledge learning based on subgraph propagation, utilize the characteristics of different subgraphs to realize the effective propagation of discriminative information, and enhance the representation ability of the overall features; 2) Knowledge learning based on subgraph propagation; For G a and G b , information interaction is carried out by using the internal structure and feature correlation of discriminative segments; graph averaging and graph convolution modeling are performed, and residual connection is used to obtain enhanced features; The action graph is calculated as follows: Among them, is the output result of the L-layer GCN; For the fuzzy graph, directly use and to de-aggregate. The graph average calculation is the same as above; By using discriminative features as a guide; the segments are modeled and arranged chronologically, and the original features are residually connected as subsequent inputs, denoted as X e ; Enhanced Feature X e is then input into the segment-level classifier to obtain the class activation sequence S ∈ R T×(C+1) , and the convolutional layer calculates the action score K ∈ R of the segment T . The two are element-wise multiplied to obtain the final score P ∈ R T×(C+1) . Finally, the video-level prediction P vid ∈ R C+1 is obtained through the top-k strategy 5. A point-supervised temporal action localization method based on the collaboration of action discriminability and boundary sensitivity according to claim 2, characterized in that The process of pseudo-label mining is as follows: Based on the fact that each action instance contains a point annotation and adjacent point annotations are in different action instances, according to the point annotation and action score, by setting a threshold λ act , the segments with scores greater than the threshold near the point annotation are marked as pseudo-action segments, which share the same action category as the point annotation; conversely, the segments located in the middle of two annotation points and with a score lower than the threshold λ bkg are marked as pseudo-backgrounds, and Λ + = {t i;c |i ∈ [1, N act , c ∈ [1, C]} and Λ - = {t j;C+1 |j ∈ [1, N bkg} are obtained respectively; using these pseudo-segments as additional supervision to improve the model's ability to separate actions and backgrounds; Among them, N act and N bkg respectively represent the total number of pseudo-action and pseudo-background segments, λ1 is a parameter, FF(·) represents the focusing function, and the video-level cross-entropy classification loss L vid .
6. A point-supervised temporal action localization method based on the collaboration of action discriminability and boundary sensitivity according to claim 2, characterized in that In the feature consistency loss, add constraints between the graph convolution features and the average features of the action graph and the background graph respectively to encourage the separability between features; 7. A point-supervised temporal action localization method based on the collaboration of action discriminability and boundary sensitivity according to claim 2, characterized in that The center contrast loss improves the aggregation of intra-class features and increases the difference between classes by calculating the similarity between the segment features within the pseudo-label set and their corresponding class centers; the action center contrast loss is calculated as follows: where cos(·) represents the cosine distance between two vectors, and γ i is the binary label of the pseudo-action point, which is used to indicate whether it belongs to the action class c.
8. The method for localizing point-supervised temporal actions based on the coordination of action discriminability and boundary sensitivity according to claim 2, characterized in that: Proposal generation; First, filter the action categories to be located on P vid , and then adopt a multi-threshold strategy to select class-specific activation score segments to generate candidate proposals for each predicted class, denoted as Props = {(s n , e n , c n , p n )}; The reliability of each action proposal is measured by the OIC score, denoted as p n ; According to the confidence score, divide the candidate proposals into reliable proposals RP, that is, for each point annotation in the video, this reliable proposal RP contains this point annotation and has the highest confidence score; And actively propose PP, that is, all remaining candidate proposals; to achieve positive and negative sample balance, group the segments with action scores lower than the predefined threshold to generate negative proposals NP; to improve the proposal recall rate, expand the boundaries of the positive proposals, and define all segments detected by the proposals as the central region R c , and expand the proposal boundary according to the preset ratio ε = 0.25 to form the start region R s and the end region R e .
9. A point-supervised temporal action localization method based on the collaboration of action discriminability and boundary sensitivity according to claim 2, characterized in that The implementation process of the said proposal integrity refinement module PIR is as follows: Introduce a boundary-sensitive temporal modeling mechanism, and combine the differential composite convolution (DCC) and the temporal disturbance operation (TDO) for dual processing to strengthen the extraction of boundary information and improve the accuracy of action localization; The differential composite convolution (DCC) deploys 4 convolutional layers in parallel. The 4 convolutional layers include three differential convolutions and one ordinary convolution. The differential convolutions include the central difference convolution (CDC), the angular difference convolution (ADC), and the horizontal difference convolution (HDC), which calculate the gradient information for the selected feature pairs from different directions to capture the changes of features at different angles. The ordinary convolution is used to obtain intensity-level information, while the differential convolutions focus on enhancing gradient-level information. The reparameterization technique is used to add multiple convolutional kernels of DCC at corresponding positions, and the calculation is as follows: Among them, DEConv(·) represents the differential composite convolution operation, and K i=1:4 respectively represent VC, CDC, HDC, ADC; * represents the convolution operation, and K all represents the equivalent kernel; X e After passing through Relu, residual connection, and convolutional layers, X is obtained o ; They are placed in sequence to calculate the attention weights in the temporal dimension and feature dimension. Temporal attention W T ∈R T×1 is used to capture temporal changes, especially boundary sequence information; Feature attention W F ∈R 1 ×D Performs weight assignment in the feature dimension to highlight the responses of important features; The calculation is as follows: Among them, C k×k represents a convolutional layer with a k×k convolutional kernel, and [·,·] represents a concatenation operation; and represent global average pooling and global max pooling features respectively; then W T and W F are fused through the broadcast rule to obtain the rough attention W cos ; Then, introduce the input features as guidance to refine the weights, and break the sequentiality of the original time series through the temporal shuffle operation to introduce a certain disturbance for the temporal feature modeling; F = sigmoid(CS([X o , W cos )) * X o + X o (11) Among them, CS(·) represents the temporal shuffle operation, which is implemented by group convolution; According to the candidate proposed region R s , R c and R e are mapped to F to obtain the corresponding region features, and then F Rs , and are concatenated in the time series dimension and fed into the integrity score head ISH to predict the integrity score P comp ; Among them, respectively represent the maximum pooling features in the time dimension; the proposed level integrity scoring process is supervised by the matching RP: Among them, N p , N n are the total numbers of positive proposals and negative proposals respectively, and H IoU represents the IoU between the proposal and the RP; To obtain a more accurate proposal boundary, it is input into the boundary regression head BRH to predict the offsets of the start time and the end time, and then the refined boundary is obtained; where ω p = I e - I s ; the boundary regression process is constrained by a reliable proposal and calculated as follows: where R comp represents the IoU between the optimized proposal and the RP; During the training phase, the model is guided by multiple loss functions, including video classification loss L vid , point-level loss L point , feature consistency loss L fc , center contrast loss L center , proposal classification loss L cls and regression loss L reg ; The overall loss function effectively balances classification accuracy, boundary accuracy, and feature quality; The calculation is as follows: L total = λ2L vid + L point + L fc + L center + L cls + L reg (16) Among them, λ2 is a learnable parameter.
10. A point-supervised temporal action localization method based on the collaboration of action discriminability and boundary sensitivity according to claim 2, characterized in that Model inference; in the inference stage, first, the proposal generation stage generates initial action proposals, and refines the proposal feature representation through a boundary-sensitive temporal modeling mechanism. Subsequently, the refined proposals are input into BRH and ISH to generate more refined proposals {(s' n ,e' n ,c n ,p n +p comp )); finally, NMS is used to remove highly overlapping proposals to obtain the final localization results.