Weakly Supervised Action Segmentation Method Based on Timestamp
By introducing a weak supervision method based on timestamps in action detection, generating pseudo-labels and optimizing action boundaries, the problem of action detection in the prior art relying on fully annotated videos is solved, and efficient action segmentation performance and the effect of reducing training costs is achieved.
Patent Information
- Application Number
- CN202310806223.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-03
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2043-07-03
AI Technical Summary
Most existing action detection methods rely on fully annotated videos, resulting in high training costs, time-consuming and risk of labeling errors.
A weakly supervised action segmentation method based on timestamps is proposed, and all video frames are trained by generating pseudo-labels, frame set changes are optimized using action timing relationships to estimate action boundaries, and boundary optimization losses are proposed based on energy functions and confidence constraints.
The action segmentation performance comparable to the fully supervised method is achieved, the complete information of the action is effectively learned, and the training cost and labeling error risk is reduced.
Smart Images

Figure CN116863374B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of pattern recognition and computer vision, and particularly relates to a weakly supervised action segmentation method based on timestamps. Background Art
[0002] Action detection has been a hot research topic in recent years. In the fields of driverless, security monitoring, transportation, human-computer interaction systems, etc., the application of action segmentation is becoming more and more extensive. Most of the existing state-of-the-art action detection methods still rely on fully annotated videos, that is, it is necessary to annotate the labels frame by frame for the video. However, this level of fully supervised training is very expensive and time-consuming, and the action boundaries are usually very similar, and there are also differences in different label annotations, which increases the risk of annotation errors caused by subjective factors of the annotators. Summary of the Invention
[0003] In order to solve the problems existing in the prior art, the present invention applies the timestamp supervision method to the action segmentation task, further analyzes and optimizes the weakly supervised action segmentation method, and thus proposes a weakly supervised action method based on timestamps. All video frames are trained by generating pseudo-labels to obtain complete action information. For timestamp training, the action boundary is estimated by optimizing the frame set change based on the action temporal relationship, and the timestamp labels are assigned to the corresponding frames to generate pseudo-labels to complete the model training. For the boundary prediction problem in timestamp supervision, a boundary optimization loss is proposed based on the constraint of the energy function on the video frame confidence to ensure that the complete information of the action can be learned during the training process. The model trained with timestamp annotations achieves performance comparable to that of the fully supervised method. It can effectively segment video actions.
[0004] The technical solution specifically adopted by the present invention to solve its technical problems is:
[0005] A weakly supervised action segmentation method based on timestamps, all video frames are trained by generating pseudo-labels to obtain complete action information: for timestamp training, the action boundary is estimated by optimizing the frame set change based on the action temporal relationship, and the timestamp labels are assigned to the corresponding frames to generate pseudo-labels to complete the model training; for the boundary prediction problem in timestamp supervision, a boundary optimization loss is adopted based on the constraint of the energy function on the video frame confidence to ensure that the complete information of the action can be learned during the training process.
[0006] Further, it includes the following steps:
[0007] Step S1: Extract video features from the input video through the I3D network;
[0008] Step S2: Input the timestamp labels and the extracted video features into the Asformer action segmentation model to extract sequence features;
[0009] Step S3: Calculate the energy function estimation boundary based on the sequence features and timestamp labels to generate pseudo-labels for auxiliary training;
[0010] Step S4: Calculate the boundary loss based on the pseudo-labels to evaluate and optimize the video segmentation effect;
[0011] Step S5: Generate the video action segmentation sequence.
[0012] Furthermore, in Step S1, the I3D network is used to extract features from the video data, obtaining a feature sequence X ∈ R of size N×d N×d , where N is the video length and d is the feature dimension, which is used in the training process.
[0013] Furthermore, in Step S2:
[0014] The timestamp label means that for a given training video X, which contains N frames and K action segments, where N < T, there is randomly only one timestamp label in each action segment, that is, only K frames in the entire video segment are labeled. where frame s i belongs to the i-th action segment.
[0015] Furthermore, Step S3 specifically includes the following steps:
[0016] Step S31: Calculate the energy function to detect the action change between two timestamps s i and s i+1 , that is, find the time that minimizes the following timestamp energy function:
[0017]
[0018]
[0019]
[0020] where d(·,·) is the Euclidean distance, h s is the output of the action segmentation model at time s, c i is the average value of the outputs between the first timestamp s i and the estimate , c i+1 is the average value of the outputs between the estimate and the second timestamp s i+1 ; that is, find the time that divides the frames between the two timestamps into two segments such that the distance between the frames and their corresponding segment centers is minimized;
[0021] Step S32: By calculating the forward estimate and backward estimation The average value of The specific formula is as follows:
[0022]
[0023]
[0024]
[0025] Among them, the forward estimate For use and i+1 The frames between Backward Estimation To use i+1 and The frames between
[0026] Step S33: until The estimated value and the next timestamp s i The frames between are assigned action labels
[0027] Furthermore, in step S4, the boundary loss function includes two parts. The first part encourages the model to make higher probability predictions for regions with lower confidence so that these regions are surrounded by regions with higher confidence. The second part suppresses abnormal frames with high confidence but far from the timestamp. The specific formula is as follows:
[0028]
[0029]
[0030] in, The video action of the sth frame is The probability of s'=2(s N -s 1 ) is the number of frames considered by the loss function.
[0031] Furthermore, in step S5, the classification result of each frame is output through the fully connected layer and the video action segmentation result is generated.
[0032] Compared with the prior art, the present invention and its preferred embodiment have the following beneficial effects:
[0033] 1. The present invention proposes a timestamp supervision training method. According to the change of actions, the video frames before the action boundary can be marked as the actions of the previous timestamp, and the frames after the action boundary can be marked as the actions of the next timestamp. Therefore, based on the action temporal relationship, the action boundary is estimated by optimizing the change of the frame set, and the timestamp labels are assigned to the corresponding frames to generate pseudo-labels to complete the model training.
[0034] 2. Based on the constraint of the energy function on the video frame confidence, the present invention proposes a boundary optimization loss to ensure that the complete information of the action can be learned during the training process. The model learns the complete video sequence of the action during the training process, rather than just the video frames near the timestamp. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] The present invention will be further described in detail below in conjunction with the drawings and specific embodiments:
[0036] Figure 1 It is a schematic diagram of the working process and principle of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0037] To make the features and advantages of this patent more obvious and understandable, specific embodiments are given below for detailed description as follows:
[0038] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs.
[0039] As Figure 1 shown, this embodiment provides a weakly supervised action method based on timestamps, which specifically includes the following steps:
[0040] Step S1: Extract video features from the input video through the I3D network;
[0041] Step S2: Input the timestamp labels and the extracted video features into the Asformer action segmentation model to extract sequence features;
[0042] Step S3: Calculate the energy function according to the sequence features and timestamp labels to estimate the boundary and generate pseudo-labels for auxiliary training;
[0043] Step S4: Calculate the boundary loss according to the pseudo-labels to evaluate and optimize the video segmentation effect;
[0044] Step S5: Generate a video action segmentation sequence.
[0045] In this embodiment, in step S1:
[0046] Extract features from video data using the I3D network to obtain a feature sequence X ∈ R of size N×d N×d , where N is the video length and d is the feature dimension, which is used in the training process;
[0047] Among them, the I3D network adopted is a mainstream neural network model for video sequence feature modeling in action recognition, and it has excellent performance in video sequence feature extraction.
[0048] In this embodiment, step S2 inputs the timestamp label and the extracted video features into the Asformer action segmentation model to extract sequence features;
[0049] Among them, the timestamp label means that given a training video X, which contains N frames and K action segments, where N < T, and there is only one timestamp label randomly in each action segment, that is, only K frames in the entire video segment are labeled. Among them, frame s i belongs to the i-th action segment.
[0050] The adopted Asformer action segmentation model applies Transformer to action segmentation. This method combines the sparse attention mechanism and one-dimensional convolution to model the long-term temporal dependence relationship, showing excellent segmentation effect.
[0051] In this embodiment, step S3 specifically includes the following steps:
[0052] Step S31: Calculate the energy function to detect the action change between two timestamps s i and s i+1 , that is, find the time that minimizes the following timestamp energy function:
[0053]
[0054]
[0055]
[0056] Among them, d(·,·) is the Euclidean distance, h s is the output of the action segmentation model at time s, c i is the average value of the output between the first timestamp s i and the estimate , and c i+1 is the average value of the output between the estimate and the second timestamp s i+1 . That is, find the time that divides the frames between two timestamps into two segments.Minimize the distance between the frame and its corresponding segment center.
[0057] Step S32: Obtain a more accurate estimation result by calculating the average value of the forward estimation and the backward estimation The specific formula is as follows: Specifically, the formula is as follows:
[0058]
[0059]
[0060]
[0061] Among them, the forward estimation is to estimate using the frames between i+1 and s The backward estimation is to estimate i+1 using the frames between s and
[0062] Step S33: Until the frames between the estimation value and the next timestamp s will be assigned action labels i In this embodiment, step S4 calculates the boundary loss function based on the confidence of each frame.
[0063] Among them, the boundary loss function includes two parts. The first part encourages the model to make higher probability predictions for regions with lower confidence, so that these regions are surrounded by regions with high confidence. The second part suppresses those abnormal frames with very high confidence but far from the timestamp, because these frames are not supported by regions with high confidence. The specific formula is as follows:
[0064] Among them,
[0065]
[0066]
[0067] Among them, refers to the probability that the action of the s-th frame of the video is S' = 2(s N - s 1 ) is the number of frames considered by the loss function.
[0068] In this embodiment, step S5 outputs the classification result of each frame through a fully connected layer and generates the video action segmentation result.
[0069] In this embodiment, a weakly supervised action segmentation task model based on timestamps is designed, and all video frames are trained by generating pseudo-labels to obtain complete action information. For timestamp training, based on the action temporal relationship, the action boundaries are estimated by optimizing the change of the frame set, and the labels of the timestamps are assigned to the corresponding frames to generate pseudo-labels to complete the model training. Aiming at the boundary prediction problem in timestamp supervision, a boundary optimization loss is proposed based on the constraint of the confidence of video frames by the energy function to ensure that the complete information of the action can be learned during the training process.
[0070] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0071] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0072] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means realizes the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0073] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0074] This patent is not limited to the above-mentioned best implementation mode. Anyone inspired by this patent can derive various other forms of timestamp-based weakly supervised action segmentation methods. All equivalent changes and modifications made according to the scope of the patent application of this invention shall fall within the scope covered by this patent.
Claims
1. A weakly-supervised action segmentation method based on timestamps, characterized in that: It includes the following steps; Step S1: Extract video features from the input video through the I3D network; Step S2: Input the timestamp labels and the extracted video features into the Asformer action segmentation model to extract sequence features; Step S3: Calculate the energy function according to the sequence features and timestamp labels to estimate the boundaries and generate pseudo-labels for auxiliary training; Step S4: Calculate the boundary loss according to the pseudo-labels to evaluate and optimize the video segmentation effect; Step S5: Generate a video action segmentation sequence; Step S3 specifically includes the following steps: Step S31: Calculate the energy function to detect the action change between two timestamps s i and s i+1 That is, find the time that minimizes the following timestamp energy function: where d(·,·) is the Euclidean distance, h s is the output of the action segmentation model at time s, c i is the first timestamp s i and the average value of the outputs between and the estimate i+1 is the average value of the outputs between the estimate and the second timestamp s i+1 ; that is, find the time that divides the frames between the two timestamps into two segments such that the distance between the frame and the center of its corresponding segment is minimized; Step S32: Obtain a more accurate estimation result by calculating the average of the forward estimation and the backward estimation Specific formula is as follows: Specific formula is as follows: Among them, forward estimation is to use and the frames between s i+1 to estimate Backward estimation is to use s i+1 and the frames between them to estimate Step S33: Until the frames between the estimated value and the next timestamp s i are assigned action labels In step S4, the boundary loss function includes two parts. The first part encourages the model to make higher probability predictions for regions with lower confidence, so that these regions are surrounded by regions with high confidence. The second part suppresses the abnormal frames with high confidence but far from the timestamps. The specific formula is as follows: Among them, refers to the probability that the video action of the s-th frame is , and S' = 2(s N - s 1 ) is the number of frames considered by the loss function.
2. The weakly-supervised action segmentation method based on timestamps according to claim 1, characterized in that: In step S1, the I3D network is used to extract features from the video data, obtaining a feature sequence X ∈ R of size N×d N×d , where N is the video length and d is the feature dimension.
3. The weakly-supervised action segmentation method based on timestamps according to claim 2, characterized in that: In step S2: The timestamp label means that for a given training video X, which contains N frames and K action segments, where N < T, there is randomly only one timestamp label in each action segment, that is, only K frames in the entire video segment are labeled. where represents the timestamp s i belongs to the i-th action segment.
4. The weakly-supervised action segmentation method based on timestamps according to claim 1, characterized in that: In step S5, the classification result of each frame is output through a fully-connected layer and a video action segmentation result is generated.
Citation Information
Patent Citations
Time sequence action detection method based on high-precision boundary prediction and computer equipment
CN115588230A
KR20220040063A