Point supervision time sequence action positioning method and system based on dual-stage neural network

Through the two-stage neural network method, candidate proposals are generated through frame-level prototype learning, and then positioning boundaries are refined through instance-level boundaries learning, which solves the shortcomings of the existing model in distinguishing backgrounds and actions and extracting global features, and achieves higher positioning accuracy.

CN119992429AInactive Publication Date: 2025-05-13HANGZHOU DIANZI UNIVERSTIY INFORMATION ENG SCHOOL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510471124.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing point-supervised timing action positioning model has shortcomings in distinguishing backgrounds and actions and extracting global action features, making it difficult to achieve accurate positioning.

Method used

Using a method based on a two-stage neural network, candidate proposals are first generated through frame-level prototype learning, and then position boundaries are refined through instance-level boundary learning, and time series features are extracted using the bidirectional Mamba module, and boundary correction is performed in combination with candidate proposals.

Benefits of technology

Effectively distinguish the background and actions in the video, extract the global action characteristics of the video, and improve the accuracy of timing action positioning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992429A_ABST
    Figure CN119992429A_ABST
Patent Text Reader

Abstract

The invention discloses a point supervision time sequence action positioning method and system based on a dual-stage neural network. The method comprises the following steps: firstly, extracting video features of each action video through an I3D video feature extraction network for a time sequence action positioning data set of point supervision labeling; carrying out first-stage frame-level prototype learning on the candidate proposal generation module, and carrying out second-stage instance-level boundary learning on the boundary positioning module; and finally, aiming at a target action video, extracting video features of the target action video through an I3D video feature extraction network, inputting the video features into a learned candidate proposal generation module, generating all candidate proposals, inputting the generated candidate proposals into a learned boundary positioning module, calculating proposal scores of all the obtained corrected proposals, and executing a soft-NMS algorithm to obtain a final proposal. The method can effectively distinguish the background and the action in the video, and extracts the global action characteristics of the video at the same time. Human body time sequence action positioning is achieved, and the positioning accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision and deep learning, and specifically relates to a point-supervised temporal action positioning method and system based on a two-stage neural network. Background Art

[0002] The purpose of Temporal Action Localization (TAL) is to locate the start and end time of an action in an uncut video and identify the category of the action at the same time. Temporal action localization is a downstream task of the video understanding task, which is one of the important research directions of deep learning. The temporal action localization task has high research value in actual production and life, such as intelligent monitoring, danger warning and other human-computer interaction fields.

[0003] The earliest TAL task was based on fully supervised learning. However, the fully supervised method requires manual annotation of the start and end time of the action, which results in a large amount of annotation work for the fully supervised method. In order to solve the problem of heavy workload, researchers used weak supervision methods in the TAL task. Weak supervision methods that are trained only with video-level labels can solve this problem. In recent years, point supervision methods have been further proposed to balance model performance and annotation workload. Point supervision methods in the video field add annotation of one frame of action on the basis of the original weak supervision method.

[0004] However, existing point-supervised temporal action localization (PSTAL) models face the following two problems: (1) How to effectively distinguish between background and action. For TAL models, the ability to distinguish between background and action is the key to accurately localizing action. Due to insufficient annotations of action instances, weakly supervised and point-supervised methods are naturally weaker than fully supervised methods in distinguishing background. (2) How to extract global action features. Most existing TAL models use video features extracted by I3D networks, while convolutional models can only capture local features and it is difficult to establish effective connections between global videos or video frames that are far apart. Summary of the invention

[0005] The purpose of the present invention is to address the above problems and to propose a point-supervised sequential action localization method based on a two-stage neural network.

[0006] To achieve the above object, the technical solution adopted by the present invention is:

[0007] In a first aspect, the present invention provides a point-supervised temporal action localization method based on a two-stage neural network, which comprises:

[0008] S1. For the point-supervised labeled temporal action localization dataset, the video features of each action video are extracted through the I3D video feature extraction network;

[0009] S2. Perform the first stage of frame-level prototype learning on the candidate proposal generation module; the candidate proposal generation module includes a prototype learning module and a three-branch attention module, wherein the prototype learning module stores the class prototype of each action category, and the video features of each action video are fused with the class prototype through the attention mechanism, and the fused features are input into the three-branch attention module and the classifier to generate an attention score and a class activation sequence respectively, and finally the class activation sequence is weightedly fused according to the attention score of each branch to generate a proposal activation sequence, and the candidate proposal and background proposal are obtained through threshold extraction;

[0010] S3. For the temporal action positioning data set, the candidate proposal generation module obtained by the first stage training is used to generate candidate proposals and background proposals corresponding to each action video, and the candidate proposals are divided into prototype proposals and ordinary proposals. The prototype proposals are used as pseudo labels to perform the second stage of instance-level boundary learning on the boundary positioning module; the boundary positioning module includes a bidirectional Mamba module and a boundary positioning module; wherein the bidirectional Mamba module is used to extract time series features from the video features, the boundary positioning module first extends the time boundary of each candidate proposal, and then extracts the features of the extended part from the time series features and maps them into boundary correction amounts through convolution, and re-corrects the time boundary of the candidate proposal to obtain a corrected proposal;

[0011] S4. For the target action video, the video features are extracted through the I3D video feature extraction network and then input into the learned candidate proposal generation module. After all candidate proposals are generated, they are input into the learned boundary positioning module. The proposal scores of all the corrected proposals are calculated and the soft-NMS algorithm is executed to obtain the final proposal.

[0012] As a preferred embodiment of the first aspect above, a prototype memory is set in the prototype learning module to store a class prototype for each action category, and the class prototype is initialized by the action frame annotated by point supervision before the first stage training; the method for fusing in the prototype learning module to obtain the fusion feature is: projecting the video features of the action video as a query through the first linear layer, splicing the video features of the action video with the class prototype in the prototype memory and projecting them as keys and values ​​through the second linear layer and the third linear layer respectively, and then generating an attention score for the query, the key and the value through the attention mechanism, and after the attention score and the video features of the action video are residually connected, they are continuously input into a feedforward neural network (FFL) using residual connection to obtain the fusion feature.

[0013] As a preferred embodiment of the above-mentioned first aspect, the three-branch attention module includes an instance attention branch, a context attention branch and a background attention branch, each branch takes the fused feature as input, passes through the convolution layer and then passes through the Softmax layer to calculate the attention score; the instance attention score, context attention score and background attention score output by the three branches are multiplied by the class activation sequence respectively to obtain the instance proposal activation sequence, the context proposal activation sequence and the background proposal activation sequence; all candidate proposals for each action category are extracted from the instance proposal activation sequence through a preset multiple first thresholds, and all background proposals are extracted from the background proposal activation sequence through a preset multiple second thresholds.

[0014] As a preferred embodiment of the first aspect, the loss function used for the first stage frame-level prototype learning of the candidate proposal generation module is a weighted sum of the fragment perception loss, the prototype optimization loss and the three-branch classification loss;

[0015] The segment-aware loss is the sum of the focus loss of the instance proposal activation sequence and the focus loss of the background proposal activation sequence;

[0016] The prototype optimization loss is the sum of the prototype contrast loss between action categories and the prototype contrast loss between action and background;

[0017] The three-branch classification loss is the sum of the cross entropy losses between the proposal activation sequences corresponding to the three attention branches and the true labeled activation sequences.

[0018] As a preferred embodiment of the above-mentioned first aspect, in the boundary positioning module, the original time boundary of each candidate proposal is forward extended and backward extended, and then the sequence features of the forward extended period are extracted from the time series features as the start features, the sequence features within the original time boundary are extracted as the proposal features, and the sequence features of the backward extended period are extracted as the end features. The start features and the end features are respectively mapped into the start time correction amount and the end time correction amount through different convolutional layers, and the expanded time boundary of the candidate proposal is re-corrected to obtain the corrected proposal.

[0019] As a preferred embodiment of the first aspect, a loss function used for performing the second-stage instance-level boundary learning on the boundary positioning module is a weighted sum of a contrast score loss and a regression loss;

[0020] The comparison score loss is the average SmoothL1 loss between the integrity scores and the intersection scores of all common proposals and background proposals; the integrity score of each proposal is obtained by maximum pooling the start feature, proposal feature and end feature of the proposal in the time dimension, and then concatenating the difference between the pooled proposal feature and the pooled start feature, the pooled proposal feature, and the difference between the pooled proposal feature and the pooled end feature to form a concatenated feature. The concatenated feature is input into the score head built based on the fully connected layer to calculate the integrity score; the intersection score of each proposal is the intersection ratio between the extended time period of the proposal and the extended time period of the prototype proposal;

[0021] The regression loss is the average of the SmoothL1 losses between the intersection scores of all normal proposals and the value 1.

[0022] As a preferred embodiment of the first aspect, the method for calculating the proposal score for each revised proposal obtained is: calculating the OIC score and the completeness score of the revised proposal, and adding the two together as the final proposal score.

[0023] In a second aspect, the present invention provides a point-supervised temporal action positioning system based on a two-stage neural network, comprising:

[0024] A target video input module is used for the user to specify the target action video that needs to be time-series action positioned;

[0025] A temporal action localization module is used to perform temporal action localization on the target action video according to the point-supervised temporal action localization method based on a two-stage neural network as described in any one of the schemes of the first aspect above, and to locate the action segment in the target action video according to the time boundary in the final proposal.

[0026] In a third aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the point-supervised temporal action localization method based on a two-stage neural network as described in any one of the schemes of the first aspect above is implemented.

[0027] In a fourth aspect, the present invention provides a computer electronic device comprising a memory and a processor;

[0028] The memory is used to store computer programs;

[0029] The processor is used to implement the point-supervised temporal action positioning method based on a two-stage neural network as described in any one of the solutions of the first aspect above when executing the computer program.

[0030] Compared with the prior art, the present invention has the following beneficial effects:

[0031] The present invention proposes a point-supervised temporal action localization method based on a two-stage neural network of Mamba, which adopts a two-stage learning strategy to optimize the point-supervised temporal action localization effect. The first stage of the present invention is a frame-level prototype learning stage, in which a class prototype is constructed for each class, and the learning of frame-level information is realized by maintaining the class prototype. And the three-branch attention module enhances the ability of the model to distinguish background and action, and provides an effective guarantee for the construction of the class prototype. The frame-level prototype learning stage generates candidate proposals and provides them to the instance-level boundary learning stage. The second stage of the present invention is an instance-level boundary learning stage, in which bidirectional Mamba is used to re-extract features, further extract the temporal information in the video features, and refine the boundaries of positioning by combining the candidate proposals and the bidirectional Mamba features to achieve the best positioning effect. In addition, the present invention also sets loss functions for the two stages respectively, and proves the effectiveness of the loss function used in the present invention through experiments. The method of the present invention can effectively distinguish the background and action in the video, and extract the action features of the video global at the same time. In order to realize the temporal action localization of the human body and improve the positioning accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 It is a schematic diagram of the steps of the point-supervised temporal action localization method based on a two-stage neural network;

[0033] Figure 2 A schematic diagram of a two-stage training framework of the method of the present invention;

[0034] Figure 3 This is a structural diagram of the first stage of the frame-level prototype learning stage of the present invention;

[0035] Figure 4 It is a structural diagram of the prototype learning module of the present invention;

[0036] Figure 5 It is a structural diagram of the instance-level boundary learning stage of the second stage of the present invention;

[0037] Figure 6 It is a schematic diagram of the structure of the bidirectional Mamba module of the present invention;

[0038] Figure 7 It is a schematic diagram of the module composition of the point-supervised sequential action localization system based on a two-stage neural network;

[0039] Figure 8 It is a schematic diagram of the structure of computer electronic equipment. DETAILED DESCRIPTION

[0040] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0041] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application.

[0042] In the description of the present invention, it should be understood that the terms "first" and "second" are only used for the purpose of distinguishing descriptions, and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features.

[0043] like Figure 1 As shown, in a preferred embodiment of the present invention, a point-supervised sequential action positioning method based on a two-stage neural network is provided, which includes steps S1 to S4. The specific implementation of each step is described in detail below.

[0044] S1. For the point-supervised labeled temporal action localization dataset, the video features of each action video are extracted through the I3D video feature extraction network.

[0045] In an embodiment of the present invention, the temporal action localization dataset may be any dataset that can be used for a temporal action localization task, such as the Thumos14 dataset.

[0046] In addition, it should be noted that the above-mentioned I3D video feature extraction network is a deep learning model for video action recognition. Its main purpose is to capture the spatiotemporal features in the video, including spatial features (objects and scenes in the image) and temporal features (changes of objects and scenes over time). The core idea of ​​I3D is to expand the traditional two-dimensional convolutional neural network (2DCNN) architecture to three-dimensional video data, so that the convolution kernel has a receptive field not only in the image height and width, but also in the temporal dimension. The I3D video feature extraction network belongs to the prior art and will not be described in detail. The I3D video feature extraction network needs to be trained in advance so that it can accurately extract the original features of the video. In addition, before using the I3D video feature extraction network to extract features from the action video, the action video needs to be segmented in advance by sampling to form a series of segments. The I3D video feature extraction network needs to encode features for each segment to finally form a complete video feature. When segmenting the action video, the video segment sampling rate can be set to sample once every 16 frames, and the dimension of each segment is 1024.

[0047] Based on the above-mentioned temporal action localization dataset, the present invention needs to perform two-stage training on the model framework of temporal action localization, wherein the first stage is the frame-level prototype learning stage, and the second stage is the instance-level boundary learning stage. The overall training framework is as follows: Figure 2 As shown in the figure, the specific training process of the two stages is introduced through steps S2 and S3 respectively.

[0048] S2. Perform the first stage of frame-level prototype learning on the candidate proposal generation module.

[0049] like Figure 3 As shown, the candidate proposal generation module includes a prototype learning module and a three-branch attention module, wherein the prototype learning module stores the class prototype of each action category, and the video feature X of each action video is related to the class prototype (i.e., ) are fused through the attention mechanism, and the fused features are input into the three-branch attention module and the classifier (implemented by sampling fully connected layer) to generate attention scores and class activation sequences respectively. Finally, the class activation sequences are weightedly fused according to the attention scores of each branch to generate proposal activation sequences, and candidate proposals and background proposals are obtained through threshold extraction.

[0050] In the embodiment of the present invention, the structure inside the prototype learning module is as follows: Figure 4 As shown in Figure 2. In the prototype learning module, a prototype memory is set to store a class prototype for each action category, and the class prototype is initialized by the action frame annotated by point supervision before the first stage of training. For each action category c, the class prototype stored in the prototype memory is recorded as Since the present invention adopts point supervision training, it is necessary to use the point annotations in the point supervision data to initialize the class prototype before the training starts. Specifically, for each class c, all point annotations of the corresponding categories in the action video of the training data are obtained, and then they are normalized. Each category uses the normalized point annotations as the class prototype.

[0051] In addition, see Figure 4 As shown, the method for obtaining the fusion feature by fusion in the prototype learning module is: the video feature X of the action video is passed through the first linear layer Projected as a query, the video feature X of the action video is compared with the class prototype in the prototype memory (i.e. ) are concatenated and then pass through the second linear layer and the third linear layer The query, key and value are projected into a key and a value, and then the query, key and value are used to generate an attention score A through an attention mechanism. After the attention score A is residually connected with the video feature X of the action video, it is input into a feedforward neural network (FFL) using a residual connection to obtain the desired output fusion feature. .

[0052] The fusion feature On the one hand, it needs to be input into the three-branch attention module, and each attention branch outputs the attention score. On the other hand, it needs to be input into the classifier built based on the fully connected layer to obtain the class activation sequence CAS (Class Activation Sequence). CAS is a sequence used to represent the activation degree of the action category of the video clip in weakly supervised temporal action localization. It records the probability score (i.e. confidence) sequence for each clip in the video (a clip consisting of a series of consecutive frames) belonging to a certain action category.

[0053] In the embodiment of the present invention, see Figure 3 As shown, the above three-branch attention module includes instance attention branch, context attention branch and background attention branch, each of which is based on the fusion feature As input, it passes through the convolution layer and then the Softmax layer to calculate the attention score. The attention scores output by the instance attention branch, context attention branch, and background attention branch are recorded as , , . Instance attention scores output by the three branches , contextual attention score , background attention score Multiply them with the aforementioned class activation sequence CAS to get the instance proposal activation sequence , context proposal activation sequence , background proposal activation sequence . Activation sequence from instance proposal In the above method, all candidate proposals for each action category are extracted by using multiple preset first thresholds, and the sequence of candidate proposals activated from the background All background proposals are extracted by using multiple preset second thresholds.

[0054] It should be noted that the proposals in the present invention are candidate action segments generated in the temporal action localization task, indicating the time interval in the video that may contain the action. Each proposal consists of a start time, an end time, and may also carry its confidence score.

[0055] It should be noted that the above activation sequence from instance proposal Extract all candidate proposals for each action category by combining multiple preset first thresholds And click the annotation tag to achieve it. Because the instance proposal activation sequence The probability score of each segment in the video belonging to an action segment is recorded in , the probability score can exceed this first threshold The time interval corresponding to the segment containing the point annotation label is taken as the candidate proposal of the action instance. The candidate proposals extracted are summarized as the final candidate proposal set under this action category and input into the next stage for frame-level prototype learning. The multiple first thresholds in the present invention can be optimized according to the actual effect.

[0056] Similarly, the above activation sequence from background proposal To extract background proposals without motion, it is necessary to combine multiple preset second thresholds to achieve this. Because the background proposal activation sequence The probability score of each clip in the video belonging to the background is recorded in , the probability score can exceed this second threshold The time interval corresponding to the segment with the second threshold is taken as the candidate proposal. The candidate proposals extracted are summarized as the final background proposal set and input into the next stage for frame-level prototype learning. The multiple second thresholds in the present invention can be optimized according to the actual effect.

[0057] In an embodiment of the present invention, the loss function used for the first stage of frame-level prototype learning of the candidate proposal generation module is the weighted sum of the segment perception loss, the prototype optimization loss and the three-branch classification loss. The segment perception loss is the sum of the focus loss of the instance proposal activation sequence and the focus loss of the background proposal activation sequence; the prototype optimization loss is the sum of the prototype-like contrast loss between action categories and the prototype-like contrast loss between action and background; the three-branch classification loss is the sum of the cross entropy loss between the proposal activation sequence corresponding to each of the three attention branches and the real labeled activation sequence.

[0058] In an embodiment of the present invention, the calculation process of the total loss function used in the frame-level prototype learning in the first stage is expressed by the following formula:

[0059] 1) Fragment-aware loss. This paper uses point annotations to optimize the model’s ability to recognize action fragments and activate the sequence of instance proposals. The surrounding segments of the point annotations above a given threshold are marked as pseudo action instance segments. ,all Pseudo-action example snippet Positive samples for training . The background proposal activation sequence of the video The fragments above a given threshold in the image are marked as pseudo background instance fragments. ,all Pseudo background instance fragment Negative samples for training Two types of pseudo instances are used as positive and negative samples to supervise model training. The calculation method of fragment-aware loss is as follows:

[0060]

[0061] in, and is the number of pseudo-action instances and pseudo-background instances, is the focus loss function.

[0062] 2) Prototype Optimization Loss. In order to optimize the prototype clip, the present invention sets a contrast loss to separate the background clip from the action clip, and performs contrast loss on the separated different action clips. The calculation formula of the prototype optimization loss is as follows:

[0063]

[0064] in, Is the action category Pseudo action example snippet The features of (obtained by the I3D video feature extraction network by encoding the features of the pseudo-action instance fragment), Action Category The class prototype, Represents the total number of action categories. Function express The formula is: is the temperature hyperparameter. Represents the features of the pseudo background instance fragment (obtained by feature encoding the pseudo background instance fragment by the I3D video feature extraction network).

[0065] 3) Three-branch classification loss. The present invention uses three attention branches to improve the model to distinguish between background instances and action instances. For each attention branch, we need to compare them with the corresponding real video action probability distribution. The cross-comparison loss is performed between the three attention branches. , , The calculation formula is as follows:

[0066]

[0067]

[0068]

[0069] Where: , , They represent the activation sequences of the true annotations corresponding to the instance attention branch, the context attention branch, and the background attention branch, respectively. The value of the corresponding action class in the activation sequence is defined as 1, and the background class is defined as 0. The value corresponding to the action class in the activation sequence is defined as 1, and the background class is defined as 1. The value of the corresponding action class in the activation sequence is defined as 0, and the background class is defined as 1. , , They represent the proposal activation sequences corresponding to the instance attention branch, the context attention branch, and the background attention branch, respectively. When represents the background class, When represents C action classes.

[0070] Add the losses of the three attention branches together to get the final three-branch classification loss:

[0071]

[0072] Finally, the three loss functions are combined to form the total loss function of the frame-level prototype learning stage:

[0073]

[0074] in, , and is a pre-set weight hyperparameter.

[0075] S3. For the temporal action localization dataset, the candidate proposal generation module obtained by the first stage training is used to generate candidate proposals and background proposals corresponding to each action video, and the candidate proposals are divided into prototype proposals and ordinary proposals. The prototype proposals are used as pseudo-labels to perform the second stage of instance-level boundary learning on the boundary localization module.

[0076] like Figure 5 As shown, the boundary positioning module includes a bidirectional Mamba module and a boundary positioning module. The bidirectional Mamba module is used to extract time series features from the video feature X, and the boundary positioning module first expands the time boundary of each candidate proposal, and then extracts the features of the expanded part from the time series features and maps them into a boundary correction amount through convolution, and re-corrects the time boundary of the candidate proposal to obtain a corrected proposal.

[0077] It should be noted that the bidirectional Mamba module is built based on the bidirectional Mamba network, which is an existing technology. It is a deep learning architecture based on the bidirectional state-space model (SSM) that aims to capture long-distance dependencies in sequence data (such as images, videos or text) through information flow in both forward and reverse directions.

[0078] Although the bidirectional Mamba module itself belongs to the prior art, in the present invention, in order to facilitate understanding, it is further combined with Figure 6 , showing the data processing process inside the bidirectional Mamba module. Its input is the video feature X of each action video extracted by the I3D video feature extraction network. The input video feature X is first separated from the forward and reverse features by different linear layers Linear and convolutional layers Conv1D. Each video frame is processed step by step to extract local and global feature information, and form forward and reverse feature sequences. Then they are input into the SSM with shared parameters for bidirectional scanning, and the states in the video sequence are iteratively updated and modeled according to the state update equation and observation equation. The bias formed by the linear layer and the activation layer (using the SiLu activation function) is used to weaken the specific state and connect the features. Finally, the sinusoidal gating mechanism is used to adjust the output feature dimension to be consistent with the input feature dimension, and the time series features extracted from the video feature X are output.

[0079] It should also be noted that during the second stage of training, the candidate proposals obtained in the first stage need to be graded to form prototype proposals and ordinary proposals, and at the same time, combined with background proposals to form three levels of proposals to overall optimize the boundary positioning module.

[0080] The prototype proposals need to be selected from the candidate proposals as pseudo labels for the temporal action localization task. The prototype proposals can be selected based on the OIC scores of all candidate proposals. For each proposal, the specific calculation method of its OIC (Outer-Inner-Contrastive) score belongs to the existing technology. For each candidate proposal, the method for calculating the OIC score is: the original time region [t 1 ,t 2 ] to expand the time boundary, and the expanded time area is [t 1 -γ(t 2 -t 1 ), t 2 +γ(t 2 -t 1 )], and then calculate the original time zone of the candidate proposal in the instance proposal activation sequence The average activation value A on 0 , and the extended time region (the part without the original time region) in the instance proposal activation sequence The average activation value A on 1 , and then the difference A between the two 1 -A 0 As the OIC score. In an embodiment of the present invention, the candidate proposal with the highest OIC score and containing the point annotation label in the candidate proposal set obtained in the first stage can be defined as the prototype proposal, and the remaining candidate proposals are all defined as ordinary proposals. The prototype proposal is used as a pseudo-label, and the ordinary proposal is used as the main data for model training in the second stage. Through the learning in the second stage, the ordinary proposal is made to approach the characteristics of the prototype proposal. The role of the background proposal is to serve as a negative sample for comparative learning with the ordinary proposal.

[0081] In an embodiment of the present invention, the specific process of performing boundary correction on candidate proposals in the above-mentioned boundary positioning module is as follows: the original time boundary of each candidate proposal is forward extended and backward extended, and then the sequence features of the forward extension period are extracted from the time series features as the start feature, the sequence features within the original time boundary are extracted as the proposal feature, and the sequence features of the backward extension period are extracted as the end feature. The start feature and the end feature are respectively mapped into the start time correction amount and the end time correction amount through different convolutional layers, and the time boundary of the candidate proposal after expansion is re-corrected to obtain the corrected proposal.

[0082] It should be noted that each candidate proposal corresponds to the start time t of a video clip. 1 and end time t 2 Therefore, when expanding the original time boundary of each candidate proposal, it is necessary to expand it in both directions along the timeline, that is, the start time boundary needs to be expanded forward and the end time boundary needs to be expanded backward. When expanding the boundary, an expansion ratio γ can be preset. Based on the set ratio γ, the original boundary of the candidate proposal [t 1 ,t 2 ], and obtain the starting time boundary t1-γ(t 2 -t 1 ) and the extended end time boundary t 2 +γ(t 2 -t 1 ). Then, since the time series features extracted by the bidirectional Mamba module contain features corresponding to different moments, the start feature (i.e. [t 1 -γ(t 2 -t 1 ),t 1 ] time period), proposal features (i.e. [t 1 ,t 2 ] period) and the end feature (i.e. [t 2 , t 2 +γ(t 2 -t 1 )] The features corresponding to the time period), the start feature, proposal feature, and end feature are pooled separately in the time dimension to obtain the start feature after pooling , Proposal features after pooling And the end feature after pooling ,Will After the convolution operation (Conv1d), the start time correction and the end time correction are obtained, and then the original proposal boundary [t 1 ,t 2 ] to make amendments and obtain amendment proposals.

[0083] Continue to see Figure 5 As shown, in order to perform the second stage instance-level boundary learning for the boundary positioning module, the present invention introduces an additional score head to calculate the loss. Specifically, in an embodiment of the present invention, the loss function used for the second stage instance-level boundary learning for the boundary positioning module is the contrast score loss and regression loss The weighted sum of , where:

[0084] The above contrast score loss The completeness score of all common proposals and background proposals Intersection score The average SmoothL1 loss between . The completeness score of each proposal After the start feature, proposal feature and end feature of the proposal are pooled in the time dimension, the difference between the pooled proposal feature and the start feature after pooling, the pooled proposal feature, and the difference between the pooled proposal feature and the end feature after pooling are concatenated to form a concatenated feature. The concatenated feature is input into the score head built based on the fully connected layer to calculate the completeness score; the intersection score of each proposal It is the union-intersection ratio between the extended time period of the proposal and the extended time period of the prototype proposal.

[0085] The above regression loss It is the average SmoothL1 loss between the intersection score and the value 1 of all normal proposals.

[0086] In an embodiment of the present invention, the calculation process of the total loss function in the second stage can be expressed as follows through the formula:

[0087] First, the integrity score of each proposal is calculated, and then the average is taken to get the integrity loss. The calculation method is:

[0088]

[0089] The fractional head Use fully connected layers.

[0090] Then, the intersection over union (IoU) between the prototype proposal and each common proposal and the background proposal is calculated, which is recorded as the intersection over union score. . Then the calculation formula for the contrast score loss is as follows:

[0091]

[0092] Where N is the total number of normal proposals and background proposals. is a loss function, and the loss function used in the present invention is SmoothL1.

[0093] In order to refine the action positioning boundary in the present invention, a boundary correction module is used to predict the boundary offset to correct the final positioning boundary. The offset output by the boundary correction module is defined as the start time correction amount and end time correction , from which the IoU between the refined positioning result of each common proposal and the prototype proposal can be calculated, denoted as . Then the regression loss is calculated as follows:

[0094]

[0095] In the formula is the number of common proposals. is a loss function, and the loss function used in the present invention is SmoothL1.

[0096] Finally, by setting the weight hyperparameters Combination and The two loss functions can be used to obtain the total loss of instance-level learning in the second stage. The formula is as follows:

[0097]

[0098] The above weight hyperparameters It can be optimized according to actual situation.

[0099] S4. For the target action video, the video features are extracted through the I3D video feature extraction network and then input into the learned candidate proposal generation module. After all candidate proposals are generated, they are input into the learned boundary positioning module. The proposal scores of all the corrected proposals are calculated and the soft-NMS algorithm is executed to obtain the final proposal.

[0100] In an embodiment of the present invention, the method for calculating the proposal score for each revised proposal is as follows: calculating the OIC score of the revised proposal and the integrity score (From the aforementioned fractional head Calculate), and add the two together as the final proposal score. Based on the proposal score and the corrected proposal generated by the boundary correction module, soft-NMS can be used to deduplicate all candidate proposals in each action category to obtain the final proposal, which is the action segment positioning result.

[0101] It should be noted that soft-NMS is an improved NMS (Non-Maximum Suppression) algorithm, which is mainly used to solve the problem of detection frame overlap in target detection. The basic idea of ​​NMS is to select the detection frame with the highest score and suppress other detection frames with high overlap. In the present invention, the score corresponds to the comprehensive proposal score, and the overlap between detection frames is the overlap of the time region of the proposal.

[0102] It should be noted that the point-supervised temporal action positioning method based on a two-stage neural network described in S1 to S4 above can actually be implemented in the form of a computer program or a functional module.

[0103] Therefore, based on the same inventive concept, in another embodiment of the present invention, a point-supervised sequential action positioning system based on a two-stage neural network is provided. Figure 7 As shown, the system includes:

[0104] A target video input module is used for the user to specify the target action video that needs to be time-series action positioned;

[0105] The temporal action localization module is used to perform temporal action localization on the target action video according to the point-supervised temporal action localization method based on a two-stage neural network described in S1 to S4 in the above embodiments, and locate the action segment in the target action video according to the time boundary in the final proposal.

[0106] Similarly, based on the same inventive concept, another embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the point-supervised temporal action positioning method based on a two-stage neural network described in S1 to S4 in the above embodiments is implemented.

[0107] Similarly, based on the same inventive concept, another embodiment of the present invention provides a computer electronic device, such as Figure 8 As shown, it includes a memory and a processor;

[0108] The memory is used to store computer programs;

[0109] The processor is used to implement the point-supervised temporal action positioning method based on a two-stage neural network described in S1 to S4 in the above embodiments when executing the computer program.

[0110] The present invention will further demonstrate the detailed implementation process and technical effects of the point-supervised temporal action localization method based on a two-stage neural network shown in the above steps S1 to S4 on a specific data set through a specific embodiment, so as to facilitate understanding of the essence of the present invention.

[0111] Example

[0112] In this embodiment, the specific method steps are the same as the point-supervised temporal action positioning method based on the two-stage neural network shown in the above-mentioned steps S1 to S4, and will not be repeated here. The specific data set, some specific parameter settings and implementation results of this embodiment are mainly demonstrated.

[0113] The temporal action localization dataset used in this embodiment is the public Thumos14 dataset.

[0114] The training parameters in this embodiment are as follows: the number of sampled segments is set to 320. The training cycle is 10000 and the batch size is 10. The Adam optimizer is used and the weight decay is set to 1e-3 and the learning rate is set to 1e-4. Regarding the threshold of the model, in the first stage, the loss function , and They are set to 0.5, 1, and 1 respectively. For a series of first thresholds for generating candidate proposals , which ranges from 0 to 0.25 with an interval of 0.05, so the first threshold used to extract candidate proposals There are 6 in total; similar settings are used to generate the second threshold for background proposals , which ranges from 0 to 0.1 with an interval of 0.01, so the second threshold for extracting background proposals There are 11 in total. In the second stage, the temperature coefficient Set to 0.1, Set to 1.

[0115] Table 1 lists the IoU parameters of all compared methods at different thresholds. The IoU range is 0.1 to 0.7 with an interval of 0.1, and the average values ​​of IoU in the three ranges of 0.1 to 0.5, 0.3 to 0.7, and 0.1 to 0.7 are given. Although MDSM is slightly inferior to existing methods at IoU=0.1 and IoU=0.2, it outperforms all existing methods at the three average values, among which the (0.3:0.7) average value has the most significant improvement over the state-of-the-art methods at high thresholds.

[0116] Table 2 shows the effect of MDSM on the ActivityNet 1.3 dataset and the comparison with the most advanced methods. On this dataset, the present invention only compares it with the method using the point supervision method. The localization effect of the model on this dataset is generally measured on ActivityNet1.3, using three IoU thresholds of 0.5, 0.75 and 0.95 and their average. In all four indicators, MDSM outperforms the current methods.

[0117] Table 1 Comparison of experimental results on the Thumos14 dataset

[0118]

[0119] Table 2 Comparison of experimental results on ActivityNet 1.3 dataset

[0120]

[0121] In addition, the present invention conducts ablation experiments on the two main modules of the frame-level prototype learning stage. The baseline used for comparison in Table 3 is that the I3D feature skips the prototype learning module and directly inputs the classifier to output CAS and the attention score at the same time. The attention CAS weight is created by fusing CAS and attention score, and then the attention CAS weight is used to generate candidate proposals.

[0122] Table 3 Ablation experiments of the main modules in the frame-level prototype learning stage

[0123]

[0124] It is obvious that using these two modules together can achieve the best results; on the contrary, if these two modules are not used, the model will perform poorly. It can also be found that the prototype-aware module can improve the performance of the model, because the prototype-aware model stores a class prototype for each class, thereby improving the frame-level perception ability. At the same time, the three-branch attention module can assist the prototype-aware module and improve the model's ability to distinguish between background and action.

[0125] In the candidate solutions of the frame-level prototype learning stage, the present invention uses multiple first thresholds to generate candidate solutions. In addition, the present invention also uses different thresholds to generate candidate solutions to obtain the best recognition results. Table 4 lists the experimental results. By testing the thresholds of the candidate solutions, the present invention finally selects the maximum threshold of 0.25 and the spacing of 0.05 to achieve the best model effect.

[0126] Table 4 Experiments on the maximum threshold for generating proposals

[0127]

[0128] Multiple loss functions are set in the frame-level prototype learning module to optimize the model parameters. In order to show the effectiveness of each loss function, an ablation test is performed and the results are shown in Table 5 (the symbol “ " indicates that the corresponding loss item exists, and the symbol "-" indicates that the corresponding loss item is deleted).

[0129] Table 5 Ablation experiment of loss function

[0130]

[0131] The experiment of this embodiment thus proves the effectiveness of the loss function used in the present invention. Among them, the prototype optimization loss has the greatest improvement on the model, while in the instance learning stage, the regression learning loss is more important than the comparison score loss. and They can complement each other and make the model have stronger recognition ability.

[0132] The above-described embodiments are only some preferred implementations of the present invention, but are not intended to limit the present invention. A person skilled in the relevant technical field may make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, any technical solution obtained by equivalent replacement or equivalent transformation falls within the protection scope of the present invention.

Claims

1. A point-supervised temporal action localization method based on a two-stage neural network, characterized in that: include: S1, for the point-supervised labeled temporal action localization dataset, the video features of each action video are extracted through the I3D video feature extraction network; S2. Perform the first stage of frame-level prototype learning on the candidate proposal generation module; the candidate proposal generation module includes a prototype learning module and a three-branch attention module, wherein the prototype learning module stores the class prototype of each action category, and the video features of each action video are fused with the class prototype through the attention mechanism, and the fused features are input into the three-branch attention module and the classifier to generate an attention score and a class activation sequence respectively, and finally the class activation sequence is weightedly fused according to the attention score of each branch to generate a proposal activation sequence, and the candidate proposal and background proposal are obtained through threshold extraction; S3. For the temporal action positioning data set, the candidate proposal generation module obtained by the first stage training is used to generate candidate proposals and background proposals corresponding to each action video, and the candidate proposals are divided into prototype proposals and ordinary proposals. The prototype proposals are used as pseudo labels to perform the second stage of instance-level boundary learning on the boundary positioning module; the boundary positioning module includes a bidirectional Mamba module and a boundary positioning module; wherein the bidirectional Mamba module is used to extract time series features from the video features, the boundary positioning module first extends the time boundary of each candidate proposal, and then extracts the features of the extended part from the time series features and maps them into boundary correction amounts through convolution, and re-corrects the time boundary of the candidate proposal to obtain a corrected proposal; S4. For the target action video, the video features are extracted through the I3D video feature extraction network and then input into the learned candidate proposal generation module. After all candidate proposals are generated, they are input into the learned boundary positioning module. The proposal scores of all the corrected proposals are calculated and the soft-NMS algorithm is executed to obtain the final proposal.

2. The point-supervised temporal action positioning method based on a two-stage neural network as claimed in claim 1 is characterized in that: The prototype learning module is provided with a prototype memory to store a class prototype for each action category, and the class prototype is initialized by the action frame annotated by point supervision before the first stage training; The method for fusing and obtaining the fused features in the prototype learning module is as follows: projecting the video features of the action video as a query through the first linear layer, splicing the video features of the action video with the class prototype in the prototype memory, and projecting them as keys and values ​​through the second linear layer and the third linear layer respectively, and then generating an attention score through the attention mechanism for the query, the key and the value, performing a residual connection on the attention score and the video features of the action video, and then continuing to input them into a feedforward neural network using a residual connection to obtain the fused features.

3. The point-supervised temporal action positioning method based on a two-stage neural network as claimed in claim 1, characterized in that: The three-branch attention module includes an instance attention branch, a context attention branch, and a background attention branch. Each branch takes the fused feature as input, passes through a convolutional layer, and then passes through a Softmax layer to calculate an attention score. The instance attention score, context attention score, and background attention score output by the three branches are multiplied by the class activation sequence to obtain the instance proposal activation sequence, context proposal activation sequence, and background proposal activation sequence; All candidate proposals for each action category are extracted from the instance proposal activation sequence by using multiple preset first thresholds, and all background proposals are extracted from the background proposal activation sequence by using multiple preset second thresholds.

4. The method for positioning point-supervised sequential actions based on a two-stage neural network as claimed in claim 3, characterized in that: The loss function used for the first stage of frame-level prototype learning of the candidate proposal generation module is the weighted sum of the fragment perception loss, prototype optimization loss and three-branch classification loss; The segment-aware loss is the sum of the focus loss of the instance proposal activation sequence and the focus loss of the background proposal activation sequence; The prototype optimization loss is the sum of the prototype contrast loss between action categories and the prototype contrast loss between action and background; The three-branch classification loss is the sum of the cross entropy losses between the proposal activation sequences corresponding to the three attention branches and the true labeled activation sequences.

5. The method for positioning point-supervised sequential actions based on a two-stage neural network as claimed in claim 1, characterized in that: In the boundary positioning module, the original time boundary of each candidate proposal is extended forward and backward, and then the sequence features of the forward extension period are extracted from the time series features as the start feature, the sequence features within the original time boundary are extracted as the proposal feature, and the sequence features of the backward extension period are extracted as the end feature. The start feature and the end feature are respectively mapped into the start time correction amount and the end time correction amount through different convolutional layers, and the expanded time boundary of the candidate proposal is re-corrected to obtain the corrected proposal.

6. The method for positioning point-supervised sequential actions based on a two-stage neural network as claimed in claim 1, characterized in that: The loss function used for the second stage instance-level boundary learning of the boundary localization module is the weighted sum of the contrast score loss and the regression loss; The comparison score loss is the average SmoothL1 loss between the integrity scores and the intersection scores of all common proposals and background proposals; the integrity score of each proposal is obtained by maximum pooling the start feature, proposal feature and end feature of the proposal in the time dimension, and then concatenating the difference between the pooled proposal feature and the pooled start feature, the pooled proposal feature, and the difference between the pooled proposal feature and the pooled end feature to form a concatenated feature. The concatenated feature is input into the score head built based on the fully connected layer to calculate the integrity score; the intersection score of each proposal is the intersection ratio between the extended time period of the proposal and the extended time period of the prototype proposal; The regression loss is the average of the SmoothL1 losses between the intersection scores of all normal proposals and the value 1.

7. The method for positioning point-supervised sequential actions based on a two-stage neural network as claimed in claim 6, characterized in that: The method for calculating the proposal score for each revised proposal obtained is: calculating the OIC score and the completeness score of the revised proposal, and adding the two together as the final proposal score.

8. A point-supervised sequential action positioning system based on a two-stage neural network, characterized in that: include: A target video input module is used for the user to specify the target action video that needs to be time-series action positioned; A temporal action localization module is used to perform temporal action localization on the target action video according to the point-supervised temporal action localization method based on a two-stage neural network as described in any one of claims 1 to 7, and locate the action segment in the target action video according to the time boundary in the final proposal.

9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by the processor, the point-supervised temporal action positioning method based on a two-stage neural network as described in any one of claims 1 to 7 is implemented.

10. A computer electronic device, characterized in that: including memory and processor; The memory is used to store computer programs; The processor is used to implement the point-supervised temporal action positioning method based on a two-stage neural network as described in any one of claims 1 to 7 when executing the computer program.

Citation Information

Patent Citations

  • Universal category multi-target tracking method fusing prototype learning mechanism

    CN118351138A

  • Continuous video human body behavior positioning method based on context information

    CN119559696A