A self-supervised action procedure anomaly detection method and system based on contrast learning
Through a self-supervised action process anomaly detection method based on contrastive learning, video representation algorithms are used to automatically generate pseudo labels and train action detection models. This solves the problems of existing technologies such as dependence on large-scale labeled data and insufficient model generalization ability. It achieves rapid identification of anomalies such as missing actions, incorrect sequences, or redundant actions, and builds a robust action process anomaly detection system.
Patent Information
- Application Number
- CN202510207615.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-02-25
AI Technical Summary
Existing methods for detecting anomalies in action processes rely on large-scale labeled data, and the model has insufficient generalization capabilities, making it difficult to quickly identify anomalies such as missing actions, incorrect sequences, or redundant actions.
A self-supervised action process anomaly detection method based on contrastive learning is adopted. Pseudo labels are automatically generated through video representation algorithm. The pseudo labels are used to train the action detection model, including video feature extraction, dynamic time warping and frame-segment contrastive learning, to build a robust action process anomaly detection system.
It reduces the dependence on large-scale labeled data, and realizes effective technical means by automatically generating technical means from unlabeled videos. It can quickly identify abnormal situations such as missing actions, wrong sequences or redundant actions, and build a robust action process anomaly detection system, which solves the technical problems existing in the existing technology, realizes the dependence on labeled data, and improves the generalization ability and recognition speed of the model.
Smart Images

Figure CN120148108B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of operation process action anomaly detection in industrial control, and in particular to a self-supervised operation process anomaly detection method based on contrastive learning. Background Art
[0002] With the rapid development of intelligent manufacturing, automated production, and intelligent monitoring technologies, process monitoring and anomaly detection have become increasingly important in a variety of fields, including industrial production, video surveillance, and medical diagnosis. In these fields, timely and effective detection of process anomalies not only improves production efficiency but also reduces potential safety hazards and resource waste. Traditional anomaly detection methods typically rely on supervised learning, requiring large amounts of labeled anomaly data for training. However, in real-world scenarios, anomaly data is often difficult to obtain or extremely expensive to label. This is especially true for complex process flows, where manual labeling is both expensive and time-consuming. Consequently, the application of previous learning methods in anomaly detection has been limited. Existing methods for identifying anomaly in industrial control workflows often require large amounts of manually labeled data for model training. In real-world scenarios, these methods are unable to quickly identify anomalies such as missing actions, incorrect sequences, or redundant actions, and thus fail to ensure process efficiency and safety.
[0003] Patent document CN 118470014 A (authorized invention patent) discloses an industrial anomaly detection method and system, which includes: acquiring an image to be detected for anomalies; detecting and locating anomalies based on the acquired image and a pretrained anomaly detection model; the anomaly detection model includes a backbone network, a pooling layer, a cascaded stream, a dual substream, and a constant stream; the cascaded stream contains several sequentially arranged stream blocks. This invention makes the features of abnormal data more distinct from those of normal data, thereby improving the effectiveness of locating anomalies. The comparison object is a single image, and this prior art does not mention anomaly detection in operational or work processes within industrial control.
[0004] The existing patent document with the document number CN118115935A discloses a method and device for identifying abnormal behavior, which includes: obtaining M target video clips with detection areas from N video clips containing the complete work process of the operator, where N and M are both positive integers greater than 1, and M is less than or equal to N; based on the M target video clips, obtaining a list of action behaviors corresponding to the complete work process of the operator; based on the action behavior list and the standard work process, identifying the abnormal action behaviors of the operator. The device executes the method. This existing technology identifies abnormal behavior in video clips containing the complete work process of the operator, standardizes the action behaviors of the operator, ensures the personal safety of the operator, avoids irreparable losses to the factory, and improves the work efficiency of the operator. In this existing technology, the complete video needs to be collected, decomposed, and then trained separately. The model for abnormal judgment of video clips requires a large amount of video annotation during training.
[0005] Therefore, it is imperative to provide a method that can reduce the dependence on large-scale manually labeled data, save data labeling costs, and quickly identify industrial control process anomalies. Summary of the Invention
[0006] The technical problems to be solved by the present invention are:
[0007] The present invention discloses a self-supervised action process anomaly detection method and system based on contrastive learning, aiming to solve the problems that existing action process anomaly detection methods rely on large-scale labeled data and have insufficient model generalization ability.
[0008] The technical solution adopted by the present invention to solve the above technical problems is:
[0009] First, a self-supervised action flow anomaly detection method based on contrastive learning is provided. The method uses an innovative video representation algorithm to automatically generate pseudo labels from a large number of unlabeled videos. This data is then used to train an efficient action detection framework to build a robust action flow anomaly detection system. Specifically, the method generates pseudo labels by training a video representation learning model; the pseudo labels are then used to train an action detection model to implement action flow anomaly detection. The method mainly includes the following steps:
[0010] S1: Collect video sequences of normal action processes;
[0011] S2: Use the video representation learning model to extract features and perform fine-grained learning on the acquired video to identify the characteristics of each process action;
[0012] S3: Manually mark the sub-action boundaries of a standard action process video and determine the time interval of each sub-action stage;
[0013] S4: Apply dynamic time warping algorithm to find the optimal time alignment path P between the unlabeled video and the labeled standard procedure video;
[0014] S5: According to the time alignment path, map the labels of the standard procedure video to the unlabeled video to generate pseudo labels, and then obtain a series of labeled videos;
[0015] S6: Train the action detection model using the labeled videos as a dataset, so that the action detection model can predict the category of each frame of action;
[0016] S7: Real-time capture camera stream data, use the trained action detection model for detection, compare the detected action sequence with the standard action procedure template, if the sequence completely matches, it is determined as normal procedure; if the action is missing, the order is wrong or there is redundant action, it is determined as abnormal procedure.
[0017] In combination with the first aspect, further, the function implementation process of the video representation learning model of step 2 is:
[0018] View generation and enhancement: For the same video, a certain number of frames are randomly cropped to generate two videos. Data augmentation is performed on the two obtained videos, including standardization and size adjustment, brightness, contrast adjustment, color jittering, and horizontal flipping, etc. to enhance the robustness of the model to different scenes, generating two different views V1 and V2 of the same video;
[0019] Video encoder: encode the preprocessed video data, convert the input multiple time frames into embedding space representation through the video encoder to effectively capture the temporal and spatial information in the video;
[0020] Frame embedding generation: each frame in the two views is input into the encoder for feature extraction to generate 128-dimensional frame-level feature representation;
[0021] Frame-frame contrast learning: construct positive and negative sample pairs through the time distance of frames, frames with close time distance as positive samples, and frames with far time distance as negative samples; maximize the similarity of positive sample frame embeddings of different views, and minimize the similarity of negative sample frame embeddings, construct the loss function as follows:
[0022]
[0023] where w(i,j) represents the allocation weight of the positive sample pair (i,j), exp represents the exponential function, t i and t j respectively represent the time step of frame i and j, L framerepresents the frame contrast loss, K represents the number of frames contained in a single view, log represents the logarithmic function with the natural base e as the base, Σ represents the summation function, corr(i,j) represents the normalized similarity of the positive sample pair (i,j), τ is the temperature parameter, and ε is a very small number to prevent division by zero. It represents the cosine similarity between the feature vector of the i-th frame of view V1 and the feature vector of the j-th frame of view V2.
[0024] Clip-to-clip contrastive learning: The video sequence is divided into non-overlapping temporal segments, each of which is obtained by averaging the embeddings of a fixed number of frames. The generated clip features are normalized, similar to frame-to-frame contrast, with temporally closer segments used as positive samples and more distant segments as negative samples. The similarity of the frame embeddings of the positive sample segments is maximized, while the similarity of the embeddings of the negative sample segments is minimized.
[0025] Reverse Optimization Encoder: By calculating the loss function of frame-frame and clip-clip comparative learning, backward gradient propagation is performed to update the encoder weight values to better focus on the fine-grained features of the action in the video.
[0026] In combination with the first aspect, further, the encoder of the video representation learning model described in step 2 includes:
[0027] Spatial Feature Extraction Module: Based on the ResNet50 backbone network, it is used to extract spatial features of each frame. Adaptive max pooling is used to reduce the feature dimension and represent each frame as a fixed-size feature vector.
[0028] Temporal dynamics capture module: uses a one-dimensional temporal convolution layer to process the frame-level feature vector sequence; captures the short-term temporal dependencies between frames and outputs a feature sequence containing temporal dynamic information.
[0029] Long-term dependency modeling module: This module includes two Transformer encoder layers, each of which contains a multi-head self-attention mechanism and a feedforward network. This module is used to capture long-term temporal dependencies in videos and globally model temporal information.
[0030] Feature Mapping and Normalization Module: Uses a multi-layer perceptron (MLP) to map the features of long videos to a fixed 128-dimensional vector space; L2 normalization is used to ensure feature stability and applicability for contrastive learning.
[0031] In combination with the first aspect, further, the ResNet50 backbone network includes:
[0032] Multiple convolutional layers to extract spatial features from input video frames;
[0033] A specific Conv3c layer for extracting high-level features;
[0034] Adaptive max pooling layer to reduce the dimension of feature maps;
[0035] Flattening operation converts the pooled features into a feature vector for each frame.
[0036] In combination with the first aspect, further, the one-dimensional temporal convolution layer includes:
[0037] Multiple one-dimensional convolution kernels are used to perform convolution operations on the inter-frame feature vector sequence;
[0038] Activation function, used to introduce nonlinearity;
[0039] Normalization layer, used to stabilize the training process.
[0040] In combination with the first aspect, further, the Transformer encoder includes:
[0041] A multi-head self-attention layer, used to calculate the dependencies between different positions in the input sequence;
[0042] A feedforward network to further process the output of the self-attention layer;
[0043] Residual connections and layer normalization are used to enhance gradient flow and stabilize training.
[0044] In combination with the first aspect, further, the multilayer perceptron (MLP) includes:
[0045] Multiple fully connected layers are used to map the features output by the Transformer encoder into a 128-dimensional space;
[0046] Activation function, used to introduce nonlinearity;
[0047] The L2 normalization layer is used to normalize the output features for contrastive learning.
[0048] Combined with the first aspect, further, the process of labeling a standard video in step 3 is as follows: using LabelVideo Tools software, label the time interval of each action stage and generate a label sequence of the standard process video Where, Indicates that the i-th frame belongs to the j-th category.
[0049] In combination with the first aspect, further, the process of finding the path P in step 4 is as follows:
[0050] Use the encoder to extract the features of each frame of the labeled video and the unlabeled video to obtain the sequence and
[0051] Calculate the frame feature similarity matrix between the target video and the standard process video:
[0052] Here, sim(·) represents the similarity of the frame embedding vectors.
[0053] Apply the dynamic time warping algorithm to find the optimal time alignment path P = {(i1, j1), (i2, j1), ..., (i k ,j k )}.
[0054] Among them, (i m ,j n ) indicates that the mth frame of the unlabeled video should be aligned with the nth frame of the labeled video.
[0055] Combined with the first aspect, further, the process of generating pseudo labels by label mapping in step 5 is as follows: according to the path P, the label of the marked video is mapped to the target video
[0056] Where, L target [i] represents the label of the i-th frame of the target video, L standard [j] represents the label of the jth frame of the annotated video.
[0057] In combination with the first aspect, further, the action detection model training implementation process in step 6 is:
[0058] Use the 2D spatial backbone network ResNet50 to extract feature vectors from multiple frames of the video;
[0059] Encode the feature vector using a temporal attention-based encoder Longformer network, where the temporal attention-based encoder processes the feature vector sequentially to capture the temporal dependencies between frames;
[0060] Combine the encoded feature vector with the position encoding information;
[0061] A multi-layer perceptron is used to generate action classification predictions based on the combined encoded feature vector and positional encoding information.
[0062] The classification loss is calculated as follows:
[0063] Where C is the number of action flow categories, is the one-hot encoding of the true label, and is the probability predicted by the model.
[0064] The classification loss is calculated and the back gradient propagation is used to optimize the encoder weights.
[0065] In a second aspect, a self-supervised action procedure anomaly detection system based on contrast learning is provided, which has program modules corresponding to the steps, and when running, the steps of the self-supervised action procedure anomaly detection method based on contrast learning are executed, including:
[0066] a data input module configured to receive normal video sequences and real-time camera stream data;
[0067] a video feature learning module configured to extract spatio-temporal features of the video;
[0068] a dynamic alignment and pseudo-labeling module configured to generate pseudo-labeled data;
[0069] an action detection model configured to predict action categories of the video sequences;
[0070] an anomaly detection module configured to compare the action sequences with standard action templates and mark anomalies.
[0071] A computer-readable storage medium storing a computer program, the computer program being configured to implement the steps of the self-supervised action procedure anomaly detection method based on contrast learning when called by a processor.
[0072] A computer device including a memory and a processor, the memory storing a computer program, and when the processor runs the computer program stored in the memory, the processor executes the steps of the self-supervised action procedure anomaly detection method based on contrast learning.
[0073] The present application has the following advantages:
[0074] The self-supervised action procedure anomaly detection method based on contrast learning can automatically generate pseudo-labels from a large number of unlabeled videos through an innovative video feature algorithm, and then train an efficient action detection framework using these data to build a robust action procedure anomaly detection system, thereby solving the problems of dependence on a large amount of labeled data and insufficient model generalization ability in existing methods, solving the problem that the application of the past learning method in anomaly detection is limited, and solving the problem that the past action recognition model for anomaly detection relies on a large amount of manually labeled video data and is difficult to promote and use. The present application only needs to label the sub-action boundaries of an action procedure and a certain amount of unlabeled normal action procedure videos, and can train an efficient and robust action procedure anomaly detection model to quickly identify action missing, sequence error or redundant action anomalies in actual scenarios.
[0075] This invention utilizes an innovative self-supervised learning method to generate pseudo-annotations from large amounts of unlabeled video, significantly reducing reliance on large-scale manually labeled data and thus significantly saving data annotation costs. The constructed real-time action flow anomaly detection system can rapidly capture and analyze camera stream data, quickly identifying anomalies such as missing actions, incorrect sequences, or redundant actions in real-world scenarios, ensuring the efficiency and safety of industrial control processes.
[0076] This paper proposes a self-supervised action process anomaly detection system based on contrastive learning. The system can generate pseudo-annotated data from unlabeled action process videos, thereby significantly reducing the labeling cost. The core technologies include a video representation model based on an encoder composed of ResNet50, a temporal convolutional layer, and a Transformer, as well as action annotation generation and mapping technology combined with a dynamic time warping algorithm. The main functions of the system are: real-time capture of camera stream data, identification and comparison of action processes through a trained action detection model, and marking as anomalies if missing, incorrectly sequenced, or redundant actions are detected. This method is suitable for scenarios such as industrial production, video surveillance, and medical diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] Figure 1 It is the overall process framework diagram (flow chart of the self-supervised action process anomaly detection method based on contrastive learning according to the present invention);
[0078] Figure 2 Schematic diagram of the video representation learning process (functional composition and flow diagram of the video representation learning model);
[0079] Figure 3 This is the architecture diagram of the video representation learning model (schematic diagram of the composition and process of the video encoder in the video representation learning model);
[0080] Figure 4 Schematic diagram for action annotation generation (diagram of the process of mapping the labels of standard process videos to unlabeled videos by time alignment paths); Figures (a) and (b) show an example of the alignment path, and Figure (c) shows an example of label mapping, with different colors representing different action labels;
[0081] Figure 5 This is the architecture diagram of the motion anomaly detection model (composition and flow diagram of the motion video encoder);
[0082] The English words in the drawings have the commonly known meanings in the art. DETAILED DESCRIPTION
[0083] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the following will be combined with the appended drawings of the embodiments of the present invention. Figure 1-5, the technical solutions in the embodiments of the present invention are described clearly and completely. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments.
[0084] Implementation method 1, see Figure 1 This paper describes the implementation method. The present invention proposes a self-supervised action process anomaly detection method based on contrastive learning. The core of this method is to train a video representation learning model to generate pseudo labels, thereby meeting the training needs of the action detection model. By combining these two models, the system can perform real-time analysis of the monitoring video stream in the scene, thereby effectively determining whether an abnormality occurs in the action process. Specifically, Figure 1 As shown, the following steps are included:
[0085] S1: Collecting video sequences of normal action processes: The system first uses high-definition cameras installed at key locations in the scene to collect videos of normal action processes, providing raw data for subsequent model training;
[0086] S2: Use the video representation learning model to extract features and perform fine-grained learning on the acquired video to identify the characteristics of each process action, laying the foundation for generating pseudo labels;
[0087] S3: Manually mark the sub-action boundaries of a standard action process video and determine the time interval of each sub-action stage;
[0088] S4: Apply the dynamic time warping algorithm to find the optimal time alignment path P between the unlabeled video and the labeled standard process video;
[0089] S5: Map the labels of the standard process video to the unlabeled video according to the time alignment path, generate pseudo labels, and then obtain a series of labeled videos;
[0090] S6: Using the labeled video as a data set to train an action detection model, so that the action detection model can predict the category of each frame of action;
[0091] S7: Capture camera stream data in real time, use the trained motion detection model to detect it, and compare the detected motion sequence with the standard motion process template. If the sequence matches completely, it is judged as a normal process; if the action is missing, the order is wrong, or there are redundant actions, it is judged as an abnormal process.
[0092] Implementation method 2, see Figure 2 In the self-supervised action process anomaly detection method based on contrastive learning described in this embodiment, the specific implementation process of the video representation learning model based on contrastive learning is as follows:
[0093] View generation and enhancement: Randomly crop a certain number of frames from the same video to generate two videos. Data enhancement is performed on the two videos, including normalization and resizing, brightness and contrast adjustment, color dithering, and horizontal flipping, to enhance the model's robustness to different scenarios. Two different views, V1 and V2, of the same video are generated.
[0094] Video encoder: Encodes the preprocessed video data and converts the input multiple time frames into embedded spatial representations through the video encoder to effectively capture the temporal and spatial information in the video;
[0095] Frame embedding generation: Each frame in the two views is passed to the encoder for feature extraction to generate a 128-dimensional frame-level feature representation;
[0096] Frame-to-frame comparative learning: construct positive and negative sample pairs based on the temporal distance of frames, with frames with closer temporal distance as positive samples and frames with farther temporal distance as negative samples. Maximize the similarity of positive sample frame embeddings across different views while minimizing the similarity of negative sample frame embeddings. The loss function is constructed as follows:
[0097]
[0098] Among them, w(i,j) represents the weight assigned to the positive sample pair (i,j), exp represents the exponential function, and t i and t j denote the time steps of frames i and j respectively, L frame represents the frame contrast loss, K represents the number of frames contained in a single view, log represents the logarithmic function with the natural base e as the base, Σ represents the summation function, corr(i,j) represents the normalized similarity of the positive sample pair (i,j), τ is the temperature parameter, and ε is a very small number to prevent division by zero. It represents the cosine similarity between the feature vector of the i-th frame of view V1 and the feature vector of the j-th frame of view V2.
[0099] Clip-to-clip contrastive learning: The video sequence is divided into non-overlapping temporal segments, each of which is obtained by averaging the embeddings of a fixed number of frames. The generated clip features are normalized, similar to frame-to-frame contrast, with temporally closer segments used as positive samples and more distant segments as negative samples. The similarity of the frame embeddings of the positive sample segments is maximized, while the similarity of the embeddings of the negative sample segments is minimized.
[0100] Reverse Optimization Encoder: By calculating the loss function of frame-frame and clip-clip comparative learning, backward gradient propagation is performed to update the encoder weight values to better focus on the fine-grained features of the action in the video.
[0101] Implementation 3: This implementation further limits the encoder described in Implementation 2. The encoder is composed of the following parts:
[0102] Spatial feature extraction module:
[0103] The backbone network based on ResNet50 is used to extract spatial features of each frame. It includes:
[0104] Multiple convolutional layers to extract spatial features from input video frames;
[0105] A specific Conv3c layer for extracting high-level features;
[0106] Adaptive max pooling layer to reduce the dimension of feature maps;
[0107] Flattening operation converts the pooled features into a feature vector for each frame.
[0108] Time dynamic capture module:
[0109] A one-dimensional temporal convolutional layer is used to process the frame-level feature vector sequence. It captures the short-term temporal dependencies between frames and outputs a feature sequence containing temporal dynamic information. This includes:
[0110] Multiple one-dimensional convolution kernels are used to perform convolution operations on the inter-frame feature vector sequence;
[0111] Activation function, used to introduce nonlinearity;
[0112] Normalization layer, used to stabilize the training process.
[0113] Long-term dependency modeling module:
[0114] This module is used to capture long-term temporal dependencies in videos and perform global modeling of temporal information. It includes:
[0115] A multi-head self-attention layer, used to calculate the dependencies between different positions in the input sequence;
[0116] A feedforward network to further process the output of the self-attention layer;
[0117] Residual connections and layer normalization are used to enhance gradient flow and stabilize training.
[0118] Feature mapping and normalization module:
[0119] Map the features of long videos to a fixed 128-dimensional vector space and ensure the stability of the features and the applicability of contrastive learning. This includes:
[0120] Multiple fully connected layers are used to map the features output by the Transformer encoder into a 128-dimensional space;
[0121] Activation function, used to introduce nonlinearity;
[0122] The L2 normalization layer is used to normalize the output features for contrastive learning.
[0123] Implementation method 4, see Figure 4 To illustrate this embodiment, the specific process of generating pseudo labels in this embodiment is as follows:
[0124] Use Label Video Tools software to manually divide a standard action process video into action segments and mark the time interval of each action stage. Generate a label sequence for the standard process video
[0125] Where, l i j Indicates that the i-th frame belongs to the j-th category.
[0126] Contains labels for a standard flow video from the first to the last frame.
[0127] Feature extraction of standard videos and unlabeled videos:
[0128] The encoder is used to extract the features of each frame of the labeled video and the unlabeled video to obtain the sequence
[0129]
[0130] Calculate the frame feature similarity matrix between the target video and the standard process video:
[0131] Here, sim(·) represents the similarity of the frame embedding vectors.
[0132] Apply the dynamic time warping algorithm to find the optimal time alignment path P = {(i1, j1), (i2, j1), ..., (i k ,j k )}.
[0133] Among them, (i m ,j n ) indicates that the mth frame of the unlabeled video should be aligned with the nth frame of the labeled video.
[0134] According to the path P, the labels of the labeled video are mapped to the target video
[0135] Where, L target [i] represents the label of the i-th frame of the target video, L standard [j] represents the label of the jth frame of the annotated video.
[0136] Embodiment five, see Figure 5 To illustrate the embodiment, the action detection model training implementation process described in the embodiment is as follows:
[0137] A two-dimensional space backbone network, ResNet50 convolutional neural network, is used to extract feature vectors from multiple frames of the video;
[0138] A time attention-based encoder, Longformer network, is used to encode the feature vectors, wherein the time attention-based encoder sequentially processes the feature vectors to capture the temporal dependencies between the frames;
[0139] The encoded feature vectors are combined with position encoding information;
[0140] A multi-layer perceptron is used to generate action classification predictions based on the combined encoded feature vectors and position encoding information.
[0141] The classification loss is calculated as follows:
[0142] Where C is the number of action procedure categories, is the one-hot encoding of the true label, and is the probability predicted by the model.
[0143] The classification loss value is calculated, and the weights of the encoder are optimized by backpropagation of the gradient.
[0144] Embodiment six, this embodiment is a further limitation of the method for determining whether an action procedure is abnormal according to embodiment one.
[0145] The real-time action procedure anomaly detection process is as follows:
[0146] Real-time camera stream data is captured, input into the trained action detection model, and the detected action sequence is compared with the standard action procedure template. If the sequence is completely matched, it is determined to be a normal procedure. If there are missing actions, incorrect sequences, or extra actions, it is marked as abnormal and an alarm is recorded, so that when the operator has the above-mentioned abnormal behavior, the operator's action behavior can be standardized in a timely manner.
[0147] Embodiment seven, the computer device described in this embodiment comprises a memory and a processor, and the memory stores a computer program. When the processor runs the computer program stored in the memory, the processor executes the method described in any one of embodiments one to six.
[0148] Embodiment eight, the computer readable storage medium described in this embodiment is used to store a computer program, and the computer program executes the method described in any one of embodiments one to six.
[0149] Those skilled in the art will appreciate that embodiments of the present disclosure may be provided as methods, systems, or computer program products. Thus, the present disclosure may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0150] The present disclosure is described with reference to flowcharts and / or block diagrams of methods, devices (systems) and computer program products according to embodiments of the present disclosure. It should be understood that each process and / or block in the flowchart and / or block diagram and the combination of processes and / or blocks in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific way, so that the instructions stored in the computer-readable memory produce a product including the instruction device, which implements the function specified in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0151] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0152] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present disclosure and are not intended to limit its scope of protection. Although the present disclosure has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that after reading the present disclosure, those skilled in the art can still make various changes, modifications or equivalent substitutions to the specific implementation methods of the invention, but these changes, modifications or equivalent substitutions are all within the scope of protection of the disclosed claims.
Claims
1. A self-supervised action process anomaly detection method based on contrastive learning, characterized in that: The method generates pseudo labels by training a video representation learning model; and uses the pseudo labels to train an action detection model to detect action process anomalies, including the following steps: S1: Collect video sequences of normal action processes; S2: Use the video representation learning model to extract features and perform fine-grained learning on the acquired video to identify the characteristics of each process action; S3: Manually mark the sub-action boundaries of a standard action process video and determine the time interval of each sub-action stage; S4: Apply the dynamic time warping algorithm to find the optimal time alignment path P between the unlabeled video and the labeled standard process video; The process of finding the alignment path P is as follows: Use the encoder to extract the features of each frame of the labeled video and the unlabeled video to obtain the sequence and Calculate the frame feature similarity matrix between the target video and the standard process video: Where sim(·) represents the similarity of the frame embedding vectors; Apply the dynamic time warping algorithm to find the optimal time alignment path P = {(i1, j1), (i2, j1), ..., (i k ,j k )}; Among them, (i m ,j n ) indicates that the mth frame of the unlabeled video should be aligned with the nth frame of the labeled video; S5: Map the labels of the standard process video to the unlabeled video according to the time alignment path, generate pseudo labels, and then obtain a series of labeled videos; S6: Using the labeled video as a data set to train an action detection model, so that the action detection model can predict the category of each frame of action; S7: Capture camera stream data in real time, use the trained motion detection model to detect it, and compare the detected motion sequence with the standard motion process template. If the sequence matches completely, it is judged as a normal process; if the action is missing, the order is wrong, or there are redundant actions, it is judged as an abnormal process.
2. The self-supervised action process anomaly detection method based on contrastive learning according to claim 1 is characterized in that The functions of the video representation learning model described in step 2 include: View generation and enhancement: Randomly crop a certain number of frames from the same video to generate two videos. Data augmentation is performed on the two videos, including normalization and resizing, brightness and contrast adjustment, color jittering, and horizontal flipping, to enhance the model's robustness to different scenarios. The resulting two different views, V1 and V2, are generated from the same video. Video encoder: Encodes the preprocessed video data and converts the input multiple time frames into embedded spatial representations through the video encoder to effectively capture the temporal and spatial information in the video; Frame embedding generation: Each frame in the two views is passed to the video encoder for feature extraction to generate a 128-dimensional frame-level feature representation; Frame-to-frame comparative learning: construct positive and negative sample pairs based on the temporal distance of frames, with frames with closer temporal distance as positive samples and frames with farther temporal distance as negative samples. Maximize the similarity of positive sample frame embeddings across different views while minimizing the similarity of negative sample frame embeddings. The loss function is constructed as follows: Among them, w(i,j) represents the weight assigned to the positive sample pair (i,j), exp represents the exponential function, and t i and t j denote the time steps of frames i and j respectively, L frame represents the frame contrast loss, K represents the number of frames contained in a single view, log represents the logarithmic function with the natural base e as the base, Σ represents the summation function, corr(i,j) represents the normalized similarity of the positive sample pair (i,j), τ is the temperature parameter, and ε is a very small number to prevent division by zero. Denotes the cosine similarity between the feature vector of the i-th frame of view V1 and the feature vector of the j-th frame of view V2; Clip-to-clip contrastive learning: The video sequence is divided into non-overlapping temporal segments, each of which is obtained by averaging the embeddings of a fixed number of frames. The generated clip features are normalized, similar to frame-to-frame contrast, with temporally closer segments used as positive samples and more distant segments as negative samples. The similarity of the frame embeddings of the positive sample segments is maximized, while the similarity of the embeddings of the negative sample segments is minimized. Reverse Optimization Encoder: By calculating the loss function of frame-frame and clip-clip comparative learning, backward gradient propagation is performed to update the encoder weight values to better focus on the fine-grained features of the action in the video.
3. According to the self-supervised action flow anomaly detection method based on contrastive learning in claim 2, the video encoder in the video representation learning model includes: Spatial feature extraction module: Based on the ResNet50 backbone network, it is used to extract the spatial features of each frame; Use adaptive max pooling to reduce feature dimensions and represent each frame as a fixed-size feature vector; Temporal dynamics capture module: This module processes the frame-level feature vector sequence using a one-dimensional temporal convolution layer, captures the short-term temporal dependencies between frames, and outputs a feature sequence containing temporal dynamics information. Long-term dependency modeling module: This module includes two Transformer encoder layers, each of which contains a multi-head self-attention mechanism and a feedforward network. This module is used to capture long-term temporal dependencies in videos and perform global modeling of temporal information. Feature Mapping and Normalization Module: Uses a multi-layer perceptron (MLP) to map the features of long videos to a fixed 128-dimensional vector space; L2 normalization is used to ensure feature stability and applicability for contrastive learning.
4. The self-supervised action process anomaly detection method based on contrastive learning according to claim 3 is characterized in that: The ResNet50 backbone network in the spatial feature extraction module includes: Multiple convolutional layers to extract spatial features from input video frames; A specific Conv3c layer for extracting high-level features; Adaptive max pooling layer to reduce the dimension of feature maps; Flattening operation converts the pooled features into feature vectors for each frame; The one-dimensional temporal convolution layer in the temporal dynamic capture module includes: Multiple one-dimensional convolution kernels are used to perform convolution operations on the inter-frame feature vector sequence; Activation function, used to introduce nonlinearity; Normalization layer, used to stabilize the training process; The Transformer encoder described in the long-term dependency modeling module includes: A multi-head self-attention layer, used to calculate the dependencies between different positions in the input sequence; A feedforward network to further process the output of the self-attention layer; Residual connections and layer normalization to enhance gradient flow and stabilize training; The multi-layer perceptron MLP in the feature mapping and normalization module includes: Multiple fully connected layers are used to map the features output by the Transformer encoder into a 128-dimensional space; Activation function, used to introduce nonlinearity; The L2 normalization layer is used to normalize the output features for contrastive learning.
5. The self-supervised action process anomaly detection method based on contrastive learning according to claim 4 is characterized in that: The process of labeling a standard video in step 3 is as follows: Use the Label Video Tools software to label the time interval of each action stage and generate a label sequence for the standard process video. Where, Indicates that the i-th frame belongs to the j-th category.
6. The self-supervised action process anomaly detection method based on contrastive learning according to claim 5 is characterized in that: The process of generating pseudo labels by label mapping in step 5 is as follows: according to the path P, the labels of the labeled video are mapped to the target video. Where, L target [i] represents the label of the i-th frame of the target video, L standard [j] represents the label of the jth frame of the annotated video.
7. The self-supervised action process anomaly detection method based on contrastive learning according to claim 6, characterized in that: The action detection model training implementation process described in step 6 is as follows: Extract feature vectors from multiple frames of a video using a 2D spatial backbone network; encoding the feature vector using a temporal attention-based action video encoder, wherein the temporal attention-based action video encoder sequentially processes the feature vector to capture the temporal dependencies between frames; Combine the encoded feature vector with the position encoding information; Use a multi-layer perceptron to generate action classification predictions based on the combined encoded feature vector and positional encoding information; The classification loss is calculated as follows: Where C is the number of action flow categories, is the one-hot encoding of the true label, and is the probability predicted by the model; Calculate the classification loss value and use the back gradient propagation to optimize the encoder weight value; The two-dimensional spatial backbone network of the action video encoder is: ResNet50 convolutional neural network; The temporal network used by the temporal attention-based action video encoder is the Longformer network.
8. A self-supervised action process anomaly detection system based on contrastive learning, characterized in that: The system has a program module corresponding to the steps of any one of claims 1 to 7 above, and when running, executes the steps in the self-supervised action process anomaly detection method based on contrastive learning, which includes: A data input module is used to receive normal video sequences and real-time camera stream data; Video representation learning module, used to extract spatiotemporal features of videos; Dynamic alignment and pseudo-annotation module, used to generate pseudo-annotation data; Action detection model, used to predict action categories from video sequences; The anomaly detection module is used to compare action sequences with standard action templates and mark anomalies.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is configured to implement the steps of the self-supervised action process anomaly detection method based on contrastive learning according to any one of claims 1 to 7 when called by a processor.
10. A computer device, characterized in that: The method comprises a memory and a processor, wherein a computer program is stored in the memory. When the processor runs the computer program stored in the memory, the processor executes the steps of the self-supervised action process anomaly detection method based on contrastive learning as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Abnormal behavior recognition method and device
CN118115935A
Industrial anomaly detection method and system
CN118470014A
Video abnormal event detection method based on prompt learning and multi-scale time sequence fusion
CN118918506A
Weak supervision time sequence action positioning method and device based on semantic and significance knowledge cooperative propagation
CN119399825A