Error detection method and system in programmed task video
By introducing Gaussian hybrid model and causal expansion convolution module in first-person procedural task videos, the problems of large intra-class variance and lack of causality in timing are solved, more accurate error detection is achieved, and the effects of intelligent monitoring and task automation are improved.
Patent Information
- Application Number
- CN202510304516.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-08-08
AI Technical Summary
The prior art is difficult to accurately identify operational errors in first-person procedural task videos, especially due to insufficient accuracy in error detection due to large intra-class variance, small inter-class distinction and lack of time-sequence causal relationships.
Gaussian mixed model (GMM) is used to learn the probability distribution for each action category, and a causal expansion convolution module (CDC) is introduced to enhance timing consistency, and error detection is performed in combination with the timing action segmentation model.
Improves the accuracy and robustness of error detection, can better identify error actions in video, reduce noise interference, and provide more reliable intelligent monitoring and task automation support.
Smart Images

Figure CN120451851A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of error detection in first-person perspective videos. Specifically, it relates to a method and system for error detection in first-person perspective procedural task videos based on probabilistic embedding and causal constraints, especially a technology for identifying operational deviations in application scenarios such as industrial process automation and intelligent monitoring. Background Art
[0002] With the rapid development of computer vision technology in recent years, the analysis of first-person perspective video has become a significant research area, with applications spanning diverse fields such as autonomous driving, intelligent security, and industrial process automation. Accurately identifying operational errors during procedural tasks is crucial in these applications, providing critical support for intelligent monitoring and task automation. For example, in industrial environments, the ability to detect operator errors in real time can effectively reduce production accidents and improve efficiency.
[0003] Traditional methods, such as those based on anomaly detection, primarily learn normal behavior patterns and identify deviations from these patterns as anomalies. These methods have achieved good results in some controlled environments. However, when processing complex task videos from a first-person perspective, these methods often struggle to capture the complex temporal dynamics and diverse action patterns. For example, in procedural tasks, actions often have specific sequences and causal relationships, and traditional anomaly detection methods struggle to model these complex temporal dependencies, resulting in limited error detection performance.
[0004] To address these challenges, existing research attempts to use a temporal action segmentation model as the backbone network for action segmentation, followed by a K-means clustering algorithm to learn prototypes for each action category. This approach allows the model to learn multiple prototypes for each action category, which serve as references for detecting deviations from normal task execution. While this approach lays the foundation for error detection in first-person perspective procedural task videos, it still has some significant limitations: (1) Limitations of prototype representation: Existing methods often ignore the inherent characteristics of the data. For example, the feature differences between frames within the same action category may be large (i.e., large intra-class variance), while the feature differences between different action categories may be small (i.e., small inter-class discrimination). This causes the learned prototypes to be close to each other in the feature space, making it difficult to distinguish different action categories or detect erroneous behaviors. Due to the large intra-class variance and small inter-class discrimination, prototypes of different actions are easily confused, which affects the accuracy of error detection.
[0005] (2) Lack of causal consistency in temporal modeling: Among existing methods, temporal action segmentation (TAS) models generally lack modeling of temporal causal consistency. In procedural task videos, the occurrence of actions often has a specific causal relationship. For example, the correct execution of a certain step is a prerequisite for the execution of the next step. The TAS model fails to effectively capture this causal relationship, which will lead to incorrect segmentation and affect subsequent error detection. If the TAS model fails to correctly capture the temporal causal relationship of the action, then the incorrect segmentation will lead to incorrect error detection results.
[0006] In summary, to overcome the shortcomings of existing methods for error detection in first-person perspective procedural task videos, this paper proposes a novel error detection method based on probabilistic embedding and causal constraints. This method effectively captures the internal variations and distributional characteristics of actions by learning a Gaussian mixture model (GMM) for each action category to model its probability distribution. Furthermore, to better model the temporal causal relationships of actions in procedural tasks, this paper introduces a causal dilated convolution module into the temporal action segmentation model to enhance the model's perception of temporal consistency. This approach enables more accurate identification of erroneous actions in videos and improves the robustness of error detection, providing more reliable support for applications such as intelligent monitoring and task automation. Summary of the Invention
[0007] The purpose of this invention is to address the shortcomings of existing error detection methods, build a unified model framework, model the underlying statistical characteristics of normal actions through Gaussian mixture models, and achieve accurate and robust error detection in egocentric program tasks.
[0008] The specific implementation contents of the present invention are as follows: The present invention proposes a method for error detection in a programmed task video recorded from a first-person perspective, which comprises the following steps: Step 1: Select a training dataset; Step 2: Use a pre-trained visual feature extraction model, such as a 3D deep convolutional neural network, to extract frame-level features of the input video data, providing high-quality visual input features for the subsequent temporal action segmentation model; Step 3: Using the frame-level visual features extracted in step 2 and the true labels of temporal actions, a temporal action segmentation model with a causal dilated convolutional module is trained to accurately predict the action category of each frame. Step 4: Use the AdamW optimizer with weight decay to optimize the model parameters. Adjust the corresponding hyperparameters according to the different training data sets to improve the performance and robustness of the model. Step 5: For each action category in the training set, use the trained temporal action segmentation model to extract the intermediate-level features of all frames in the corresponding category. Using these features as input, the Expectation Maximization algorithm is used to train a Gaussian mixture model for the corresponding action category to learn the distribution pattern of features in each category. Step 6: In the inference phase, we first use the trained temporal action segmentation model to perform temporal action segmentation on the training set videos to obtain the action category for each frame. Then, we input the intermediate layer features of each frame into the trained Gaussian mixture model corresponding to the action category and calculate its log-likelihood value; Step 7: Based on all the log-likelihood values of each category obtained in step 6, take the minimum and median of the log-likelihood values of all frames, and use them as boundaries to generate multiple threshold ranges. These thresholds are then used to determine whether each frame is abnormal; Step 8: For each test video to be inferred, we first use the temporal action segmentation model to segment each action segment. We then input the intermediate layer features of each frame into the trained Gaussian mixture model (GMM) corresponding to the action category and calculate its log-likelihood value.
[0009] Step 9: We aggregate the frame-level log-likelihood values obtained in Step 8 into segments. To reduce the impact of outliers at individual moments on the judgment results, we smooth the Gaussian mixture model predictions within each action segment using a one-dimensional Gaussian filter, thereby generating a more stable segment-level log-likelihood. Step 10: We compare the log-likelihood obtained in step 9 with the threshold after smoothing. Frames with a value greater than the threshold are considered normal, and frames with a value less than the threshold are considered error frames. Finally, we determine whether the entire segment is an error based on the accuracy of the majority of frames within each segment.
[0010] The present invention also relates to an error detection system in a programmed task video, wherein the programmed task video is recorded from a first-person perspective, and the error detection system specifically comprises: An acquisition module acquires programmed task videos recorded from a first-person perspective; The temporal action segmentation model is used to obtain the intermediate layer features of each frame in the test video and obtain the predicted action category of each frame; A log-likelihood value calculation unit, configured to input the intermediate layer features of each frame into a Gaussian mixture model corresponding to the action category, and calculate a frame-level log-likelihood value of the action category; An error detection unit is used to determine whether the programmed task video is abnormal based on the log-likelihood value.
[0011] Compared with the prior art, the present invention has the following advantages and beneficial effects: By introducing a causal dilated convolutional module into the temporal action segmentation model, our method can more effectively capture the temporal dependencies and causal relationships of actions in videos. This enables the model to better understand the order in which actions occur, thereby improving the recognition accuracy of action clips and reducing misjudgments caused by temporal confusion.
[0012] 2. Traditional prototype learning methods rely on fixed cluster centers and are difficult to adapt to situations with large intra-class variations. Our method uses a Gaussian mixture model (GMM) to learn a probability distribution for each action category, enabling a more detailed characterization of normal action behavior patterns. By calculating the log-likelihood of each frame, we can more effectively identify frames that deviate from normal patterns, thereby improving the accuracy and robustness of anomaly detection. In addition, we use the minimum and median log-likelihood values on the training set as bounds, which better fits the characteristics of the dataset and further improves detection performance.
[0013] 3. By applying Gaussian smoothing to the GMM log-likelihood, our method effectively reduces noise interference and produces more stable and consistent segment-level anomaly assessment results. This smoothing operation makes the judgment of the entire action segment more reliable and avoids misjudgments caused by abnormal fluctuations in individual frames. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 Flowchart for the implementation of the present invention; Figure 2 A schematic diagram of a framework for error detection based on probability embedding and causal constraints proposed by the present invention; Figure 3 The figure shows the effect of video error detection performed by the present invention on the three subtasks (Tea, Coffee, Pinwheels) of the EgoPER dataset. DETAILED DESCRIPTION
[0015] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. That is, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments.
[0016] Temporal Action Segmentation Model Temporal action segmentation models are deep learning models based on time series. They aim to segment long videos into action steps and predict the corresponding action category for each frame. These models simultaneously model temporal dependencies and local action features and are often used for the automatic analysis of structured processes, such as cooking steps and surgical procedures.
[0017] Common models that can be used in the present invention include: MS-TCN (Multi-Stage Temporal Convolutional Network) AF (ActionFormer).
[0018] The specific temporal action segmentation model provided by the present invention may also be used.
[0019] Intermediate layer features The intermediate layer features are high-dimensional vector representations of the output of a hidden layer during the forward propagation of a deep learning model (such as the activation values of a fully connected layer or a convolutional layer), which contain the nonlinear and high-level semantic information of the input data.
[0020] Gaussian mixture model A Gaussian mixture model is a probabilistic generative model that fits complex data distributions through a linear combination of multiple Gaussian distributions. In anomaly detection, the features of each action category are modeled as a separate GMM. The features of normal samples follow this distribution, while the features of abnormal samples deviate from this distribution.
[0021] 3D Deep Convolutional Neural Network Three-dimensional deep convolutional neural networks (3D-CNNs), specifically the I3D (Inflated 3D ConvNet) model, are deep learning models specifically designed for processing spatiotemporal data, such as videos. I3D's core principle is that it leverages established 2D-CNN architectures (such as Inception-V1), pre-trained on large-scale datasets like ImageNet and Kinetics, and extends them to three dimensions through an "inflation" operation.
[0022] AdamW optimizer with weight decay AdamW is an improved version of the Adam optimizer. By explicitly decoupling the weight decay term from the gradient update in the loss function, it solves the incompatibility problem between weight decay (L2 regularization) and the adaptive learning rate mechanism in the traditional Adam optimizer, thereby improving the stability and generalization ability of model training.
[0023] like Figure 1 、 2As shown in FIG3 , the implementation of the present invention can be divided into two stages, namely, model training and positioning using the model.
[0024] Step 1: Select the training data set; Step 2: Use a pre-trained visual feature extraction model, such as a 3D deep convolutional neural network, to extract frame-level features of the input video data, providing high-quality visual input features for the subsequent temporal action segmentation model; Step 3: Using the frame-level visual features extracted in step 2 and the true labels of temporal actions, a temporal action segmentation model with a causal dilated convolutional module is trained to accurately predict the action category of each frame. Step 4: Use the AdamW optimizer with weight decay to optimize the model parameters. Adjust the corresponding hyperparameters according to the different training data sets to improve the performance and robustness of the model. Step 5: For each action category in the training set, use the trained temporal action segmentation model to extract intermediate-level features from all frames of the corresponding category. Using these features as input, the Expectation Maximization algorithm is used to train a Gaussian mixture model for the corresponding action category to learn the distribution patterns of features for each category. Step 6: In the inference phase, we first use the trained temporal action segmentation model to perform temporal action segmentation on the training set videos to obtain the action category for each frame. Then, we input the intermediate layer features of each frame into the trained Gaussian mixture model corresponding to the action category and calculate its log-likelihood value; Step 7: Based on all the log-likelihood values of each category obtained in step 6, take the minimum and median of the log-likelihood values of all frames, and use them as boundaries to generate multiple threshold ranges. These thresholds are then used to determine whether each frame is abnormal; Step 8: For each test video to be inferred, we first use the temporal action segmentation model to segment each action segment. We then input the intermediate layer features of each frame into the trained Gaussian mixture model for the corresponding action category and calculate its log-likelihood.
[0025] Step 9: We aggregate the frame-level log-likelihood values obtained in Step 8 into segments. To reduce the impact of outliers at individual moments on the judgment results, we smooth the Gaussian mixture model predictions within each action segment using a one-dimensional Gaussian filter, thereby generating a more stable segment-level log-likelihood. Step 10: We compare the log-likelihood obtained in step 9 with the threshold after smoothing. Frames larger than the threshold are considered normal, and frames otherwise are considered erroneous. Finally, we determine whether the entire segment is an erroneous action based on the correctness of the majority of frames within each segment. As a preferred technical solution, step 2 includes in more detail: when extracting visual features, using the deep convolutional neural network model I3D to obtain the features of each frame: , represents the d-dimensional feature vector of the t-th frame in the n-th video.
[0026] As a preferred technical solution, step 3 includes in more detail: Step 3.1: Build a backbone network to input the frame-level feature sequence of each video ( ) is mapped to a refined feature sequence ( ). It can be expressed as: in Represents the backbone feature extractor based on deep neural network.
[0027] Step 3.2: Use the causal dilated convolution (CDC) module to convolution the output of the backbone network ( ) to perform time series modeling and maintain causal consistency. The core operation formula of the CDC module is: Where k is the convolution kernel size, d is the expansion coefficient, The weight parameters can be learned and temporal causality is guaranteed (the output at the current moment only depends on the input at the historical moment).
[0028] Then the output of the CDC module is connected to the output of the backbone network through residual connection: Step 3.3: The features processed by CDC are passed through the classification head network ( ) is mapped to action category prediction: in is the predicted action category.
[0029] As a preferred technical solution, step 4 includes more details: According to the learning goal, we use cross entropy loss as the task loss, which can be expressed by the following formula: in is the cross entropy loss; is the total number of frames; is the label (0 or 1) indicating whether the i-th frame belongs to the j-th category; Predict the probability that the i-th frame belongs to the j-th class for the model.
[0030] As a preferred technical solution, step 5 includes in more detail: Step 5.1: Randomly initialize the parameters of GMM, including the mixing coefficient ( ), mean vector ( ), and the covariance matrix ( ) (where k is the index of the Gaussian component, and there are K Gaussian components in total).
[0031] Step 5.2: Calculate the posterior probability that each data point belongs to each Gaussian component ( ), the formula is as follows: in,( ) is the intermediate layer feature of the t-th frame.
[0032] Step 5.3: Update the parameters of GMM according to the posterior probability calculated in step E. The formula is as follows: Step 5.4: Repeat steps 5.2 and 5.3 until the parameters of the Gaussian mixture model converge or the preset maximum number of iterations is reached.
[0033] As a preferred technical solution, step 7 includes in more detail: first, using the following formula for all log-likelihoods of action category j in the training set: in is the intermediate layer feature of the t-th frame of the video, are the parameters of the Gaussian mixture model of the corresponding category, express The probability density under the k-th Gaussian component.
[0034] For each action category, we calculate the minimum log-likelihood value of all frames of that category in the training set and median Based on the minimum and median , generating multiple threshold ranges For example, 41 equally spaced thresholds can be generated: .
[0035] As a preferred technical solution, step 8 includes in more detail: Step 8.1: Use the trained temporal action segmentation model to perform action segmentation on the test video, obtain the intermediate layer features of each frame in the test video, and obtain the predicted action category for each frame.
[0036] Step 8.2: The intermediate layer features of each frame Input it into the Gaussian mixture model of the corresponding predicted action category and calculate its log likelihood value. The formula is as follows: As a preferred technical solution, step 9 includes in more detail: Step 9.1: Based on the results of temporal action segmentation, segment the test video into action segments.
[0037] Step 9.2: For each action segment, the log-likelihood value sequence , use a one-dimensional Gaussian filter for smoothing. The kernel function of the Gaussian filter is: in is the standard deviation of the Gaussian filter.
[0038] Given a one-dimensional Gaussian filter acting on a time series with a sliding window width of an odd number W, the smoothed log-likelihood value sequence is ,in: As a preferred technical solution, step 10 includes in more detail: Step 10.1: Smooth the log-likelihood of each frame The threshold range generated in step 7 If there is a threshold such that If the value is lower than this threshold, the frame is considered an error frame.
[0039] Step 10.2: For each action segment, count the number of error frames and calculate the error frame ratio. If the error frame ratio exceeds a predefined threshold (50%), all frames in the segment are considered error-free; otherwise, all frames in the segment are considered normal.
[0040] The present invention also relates to an error detection system in a programmed task video, wherein the programmed task video is recorded from a first-person perspective, and the error detection system specifically comprises: An acquisition module acquires programmed task videos recorded from a first-person perspective; The temporal action segmentation model is used to obtain the intermediate layer features of each frame in the test video and obtain the predicted action category of each frame; A log-likelihood value calculation unit, configured to input the intermediate layer features of each frame into a Gaussian mixture model corresponding to the action category, and calculate a frame-level log-likelihood value of the action category; An error detection unit is used to determine whether the programmed task video is abnormal based on the log-likelihood value.
[0041] To test the model's error detection performance, we use the error detection accuracy (EDA), defined as De / GTe, where De and GTe represent the number of correctly predicted segments and the total number of segments in all test videos, respectively. Furthermore, we use the Area Under the Curve (AUC) metric based on frame-level error predictions—specifically, the micro-averaged AUC—to evaluate frame-level error detection capabilities.
[0042] Below, the effects and advantages of the present invention are more intuitively demonstrated: The first scenario example: Model training and testing were performed on the EgoPER dataset, and the results were compared with those of previous methods. Using the EDA and AUC metrics, the comparative results for five tasks on the EgoPER dataset are shown in Table 1: Table 1. Average false detection results of different methods for each task and all tasks on the EGOPER dataset.
[0043] The second scenario example: Model training and testing were performed on the HoloAssist dataset, and the results were compared with previous methods. The EDA and AUC metrics were used, with verbs and nouns as action categories, respectively. The comparative results for the five tasks on the HoloAssist dataset are shown in Table 2. Table 2. Error detection results on the HoloAssist dataset The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for detecting errors in a programmed task video, wherein the programmed task video is recorded from a first-person perspective, characterized in that: The error detection method specifically comprises the following steps: Obtain a video of a programmed task recorded from a first-person perspective; Based on the temporal action segmentation model, the intermediate layer features of each frame in the test video are obtained, and the predicted action category of each frame is obtained; Inputting the intermediate layer features of each frame into the Gaussian mixture model of the corresponding action category to calculate the frame-level log-likelihood value of the action category; Whether the programmed task video is abnormal is determined based on the log-likelihood value.
2. The error detection method in a programmed task video according to claim 1, characterized in that: The method of extracting the programmed task video and segmenting it into one or more action category videos by using a temporal action segmentation model comprises the following steps: Select a training dataset; Use a pre-trained 3D deep convolutional neural network (DCNN) to extract frame-level features from the input video data, providing high-quality visual input features for the subsequent temporal action segmentation model. Using the extracted frame-level visual features and the true labels of temporal actions, a temporal action segmentation model with a causal dilated convolution module is trained to predict the action category of each frame. The AdamW optimizer with weight decay is used to optimize the model parameters. The corresponding hyperparameters are adjusted according to the different training data sets to obtain a trained temporal action segmentation model. For each action category in the training set, use the trained temporal action segmentation model to extract the intermediate layer features of all frames of the corresponding category; In the inference stage, the trained temporal action segmentation model is used to perform temporal action segmentation on the videos of the training set to obtain the action category of each frame and the intermediate layer features of each frame.
3. The error detection method in a programmed task video according to claim 2, characterized in that: Input the intermediate layer features of each frame into the Gaussian mixture model of the corresponding action category to calculate the frame-level log-likelihood value of the action category, including: Using the intermediate layer features of all frames as input, the Expectation Maximization algorithm is used to train the Gaussian mixture model corresponding to the action category. This is used to learn the distribution law of the features of each category and obtain the trained Gaussian mixture model corresponding to the action category. The intermediate layer features of each frame are input into the trained Gaussian mixture model of the corresponding action category to calculate the frame-level log-likelihood value.
4. The error detection method in a programmed task video according to claim 1, characterized in that: Judging whether the programmed task video is abnormal based on the log-likelihood value includes: Input the intermediate layer features of each frame in the training set into the trained Gaussian mixture model of the corresponding action category and calculate its log-likelihood value; according to all the log-likelihood values of each action category, take the minimum and median of the log-likelihood values of all frames respectively, and use this as the boundary to generate the threshold range ; The threshold range Used to determine whether each frame is abnormal; The frame-level log-likelihood values are integrated into segments; the Gaussian mixture model prediction values in each action segment are smoothed using a one-dimensional Gaussian filter to obtain a more stable segment-level log-likelihood; Combine the obtained log-likelihood after smoothing with the threshold range If the frame value is greater than the threshold, it is considered a normal frame, otherwise it is considered an error frame. Finally, based on the correctness of the majority of frames in each segment, it is determined whether the segment as a whole is an error action.
5. The error detection method in a programmed task video according to claim 4, characterized in that: The formula for calculating the log-likelihood value is: ,in is the intermediate layer feature of the t-th frame of the video, are the parameters of the Gaussian mixture model of the corresponding category, express The probability density under the k-th Gaussian component.
6. The error detection method in a programmed task video according to claim 5, characterized in that: The frame-level log-likelihood values are integrated into segments; the Gaussian mixture model prediction values within each action segment are smoothed using a one-dimensional Gaussian filter to obtain a more stable segment-level log-likelihood, including: According to the results of temporal action segmentation, the test video is divided into action segments; For each action segment, the log-likelihood value sequence , use a one-dimensional Gaussian filter for smoothing, the kernel function of the Gaussian filter is: ,in is the standard deviation of the Gaussian filter; Given a one-dimensional Gaussian filter acting on a time series with a sliding window width of an odd number W, the smoothed log-likelihood value sequence is ,in: 。 7. The error detection method in a programmed task video according to claim 1, characterized in that: Combine the obtained log-likelihood after smoothing with the threshold range If the frame size is greater than the threshold, it is considered a normal frame, otherwise it is considered an error frame. Finally, based on the correctness of the majority of frames in each segment, it is determined whether the segment as a whole is an error action, including: The log-likelihood value of each frame after smoothing With threshold range Compare; if there is a threshold such that Below threshold range , then the frame is considered an error frame; For each action segment, the number of error frames is counted and the ratio of error frames is calculated. If the ratio of error frames exceeds a predefined threshold of 50%, all frames in the segment are judged to be error; otherwise, all frames in the segment are judged to be normal.
8. A system for detecting errors in a programmed task video, wherein the programmed task video is recorded from a first-person perspective, characterized in that: The error detection system specifically includes: An acquisition module acquires programmed task videos recorded from a first-person perspective; The temporal action segmentation model is used to obtain the intermediate layer features of each frame in the test video and obtain the predicted action category of each frame; A log-likelihood value calculation unit, configured to input the intermediate layer features of each frame into a Gaussian mixture model corresponding to the action category, and calculate a frame-level log-likelihood value of the action category; An error detection unit is used to determine whether the programmed task video is abnormal based on the log-likelihood value.
Citation Information
Cited By
Lightweight few-sample man-machine interaction action recognition method, system and equipment
CN121096028A