Electric power safety operation monitoring system
By performing cosine similarity screening and timing rearrangement of historical image frames, a high-quality training set is constructed, and the timing features are extracted using a three-dimensional convolutional model, the problem of insufficient real-time power operation video recognition in the existing technology is solved, and high-sensitivity power operation behavior recognition and safety monitoring are achieved.
Patent Information
- Application Number
- CN202510311454.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-29
AI Technical Summary
The existing power operation monitoring system cannot effectively identify the timing changes and behavioral continuity of operation behavior in real-time power operation videos, resulting in insufficient recognition sensitivity.
Using high-similarity screening and timing rearrangement mechanisms, a high-quality training set is constructed by calculating cosine similarity and timing rearrangement of historical image frames, and a three-dimensional convolutional model is used to extract the spatial and timing characteristics of image frames, and the model parameters are optimized in combination with multi-classified cross-entropy loss function.
It significantly improves the ability to identify power operation behavior, improves the model's sensitivity to timing changes and identification accuracy, and ensures the safety and efficiency of the operation site.
Smart Images

Figure CN120388314A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of power operation monitoring, and specifically to a power safety operation monitoring system. Background Art
[0002] With the complexity of power operation scenarios and the diversity of operation environments, the risks faced by power operation personnel during on-site operations are also increasing day by day. To ensure the safety of operation personnel, the existing power operation monitoring systems usually monitor operation behaviors based on real-time operation videos and identify power operation behaviors through intelligent technologies. The Chinese patent document with the patent publication number CN112989110A discloses a power operation intelligent video monitoring system, which realizes real-time warning and direct retrieval of violation behaviors by performing intelligent analysis and annotation on video stream data, thus greatly reducing the safety risks of power production. However, in the above-mentioned existing technologies, historical operation videos are not used to train the intelligent large model, and there is also a lack of a screening and optimization process for the training set during the training process. Therefore, it is impossible to achieve highly sensitive identification of operation behaviors in real-time power operation videos. Summary of the Invention
[0003] Aiming at the deficiencies of the existing technologies, the present invention provides a power safety operation monitoring system, which uses a behavior monitoring model to monitor real-time power operation videos, and the behavior monitoring model proposes a high similarity screening and time series rearrangement mechanism, solving the technical problem of insufficient sensitivity of the model to the time series changes and behavior continuity of real-time power operation videos in the existing technologies.
[0004] To achieve the above objectives, the present invention is realized through the following technical solutions:
[0005] A power safety operation monitoring system, comprising:
[0006] A real-time operation video acquisition module, configured to acquire real-time power operation videos and decompose them into first real-time image frames with N consecutive time series frame numbers;
[0007] An image frame preprocessing module, configured to preprocess the first real-time image frames to obtain second real-time image frames;
[0008] A label output module, configured to input the second real-time image frames into the trained operation monitoring model to obtain the predicted probability values of power operation behaviors in the second real-time image frames, and output the category corresponding to the maximum predicted probability value as the power operation behavior classification label;
[0009] An alarm strategy output module, configured to match the corresponding alarm strategy in a preset alarm strategy library according to the power operation behavior classification label and output it in real-time.
[0010] In some of these embodiments, the training process of the job monitoring model includes:
[0011] S1. Obtain an image frame training set; wherein, the training samples of the image frame training set are rearranged image frames, and the rearranged image frames represent historical image frames rearranged according to the sequential frame numbers.
[0012] S2. Input the rearranged image frames into a three-dimensional convolutional model for iterative training to obtain the job monitoring model; wherein, the job monitoring model can extract the spatial features and temporal features of the rearranged image frames.
[0013] In some of these embodiments, obtaining an image frame training set; wherein, the training samples of the image frame training set are rearranged image frames, and the rearranged image frames represent historical image frames rearranged according to the sequential frame numbers, includes:
[0014] S1-1. Obtain power operation videos within M historical time periods and decompose them into first historical image frames with N consecutive sequential frame numbers.
[0015] S1-2. Preprocess the first historical image frames to obtain second historical image frames.
[0016] S1-3. Calculate the cosine similarity of the second historical image frames to obtain the cosine similarity of the second historical image frames.
[0017] S1-4. Rearrange the second historical image frames with cosine similarity higher than the similarity threshold according to adjacent sequential frame numbers to obtain the rearranged image frames.
[0018] S1-5. Define the pixel values in the rearranged image frames as input features and the power operation behaviors as class labels to construct the training samples of the image frame training set.
[0019] S1-6. Aggregate the rearranged image frames to obtain the image frame training set.
[0020] In some of these embodiments, calculating the cosine similarity of the second historical image frames to obtain the cosine similarity of the second historical image frames includes:
[0021] S1-3-1. Extract the feature vectors of the second historical image frames.
[0022] The extraction of the feature vectors of the second historical image frames is:
[0023]
[0024] Wherein, represents the i-th second historical image frame, is the feature vector of the i-th second historical image frame, and FE represents feature extraction;
[0025] S1-3-2. According to the extracted feature vectors, calculate the cosine similarity between adjacent frame pairs; where the adjacent frame pairs represent adjacent second historical image frames;
[0026] The expression of the cosine similarity is:
[0027]
[0028] where, represents the cosine similarity between adjacent frame pairs, represents the inner product of the feature vectors of adjacent frame pairs, represents the product of the norm lengths of adjacent frame pairs.
[0029] S1-3-3. According to the cosine similarity between adjacent frame pairs, generate a set S of cosine similarities of adjacent frame pairs;
[0030] The expression of the set of cosine similarities is:
[0031] .
[0032] In some embodiments, rearranging the second historical image frames with cosine similarity higher than the similarity threshold according to adjacent time sequence frame numbers to obtain the rearranged image frames includes:
[0033] S1-4-1. Define a similarity threshold TS;
[0034] S1-4-2. Screen out adjacent frame pairs with cosine similarity higher than the similarity threshold from the set of cosine similarities to obtain screened adjacent image frames;
[0035] The expression for screening out adjacent frame pairs with cosine similarity higher than the similarity threshold from the set of cosine similarities is:
[0036]
[0037] where, represents the screened adjacent image frames, represents filtering out adjacent image frames smaller than the similarity threshold TS.
[0038] S1-4-3. Rearrange the screened adjacent image frames according to adjacent time sequence frame numbers to generate a set of the rearranged image frames;
[0039] The expression for rearranging the screened adjacent image frames according to adjacent time sequence frame numbers is:
[0040]
[0041] where, The screened adjacent image frames are denoted as, Reorder denotes reordering the screened adjacent image frames according to the adjacent time sequence frame numbers, and N denotes the number of image frames.
[0042] The expression for the set of the reordered image frames is:
[0043]
[0044] Where X represents the set of the reordered image frames, denotes the i-th reordered image frame after reordering according to the adjacent time sequence frame numbers;
[0045] In some embodiments, the reordered image frames are input into a three-dimensional convolutional model for iterative training to obtain the job monitoring model; wherein, the job monitoring model can extract the spatial features and temporal features of the reordered image frames, including:
[0046] S2-1, the three-dimensional convolutional model performs forward propagation on the image frame training set so that the output layer outputs the predicted probability value based on the pixel values in the reordered image frames;
[0047] The expression for the predicted probability value is:
[0048]
[0049] Where, denotes the predicted probability value, denotes the weight matrix of the fully connected layer of the model, denotes the flattened spatio-temporal feature vector, denotes the bias term, and Softmax represents converting the classification score vector of the fully connected layer into a normalized predicted probability value.
[0050] S2-2, converting the class label of the power operation behavior in the reordered image frames into a one-hot vector to obtain the true label value of the power operation behavior;
[0051] The expression for the true label value is:
[0052]
[0053] Where, denotes the true label value of the power operation behavior, denotes the class label of the power operation behavior, and OneHot represents converting the class label into a one-hot vector.
[0054] S2-3, calculating the multi-class cross-entropy loss of the predicted probability value and the true label value through the multi-class cross-entropy loss function;
[0055] The multi-class cross-entropy loss function is:
[0056]
[0057] Among them, represents the multi-class cross-entropy loss, represents the k-th true label value, represents the predicted probability value of the i-th rearranged image frame corresponding to the k-th class label, and K represents the total number of classes.
[0058] S2-4. According to the multi-class cross-entropy loss, perform the backpropagation algorithm to update the weight matrix and bias term of the fully connected layer layer by layer; among them, define the weight matrix and bias term of the fully connected layer as model parameters;
[0059] S2-5. When the total loss value of the multi-class cross-entropy is lower than the loss threshold, it is considered that the 3D convolutional model reaches the convergence condition;
[0060] The expression of the convergence condition is:
[0061]
[0062] Among them, represents the total loss value of the multi-class cross-entropy, characterizing the total loss of the image frame training set, represents the loss threshold.
[0063] S2-6. Define the model parameters at convergence as the optimal parameters, and use the 3D convolutional model with the optimal parameters as the trained job monitoring model;
[0064] The expression of the job monitoring model is:
[0065]
[0066] Among them, represents the job monitoring model, X represents the input rearranged image frame; h represents the spatio-temporal feature vector extracted by the 3D convolutional model, represents the optimal weight matrix of the fully connected layer, represents the optimal bias term of the fully connected layer.
[0067] In some of the embodiments, the 3D convolutional model performs forward propagation on the image frame training set so that the output layer outputs the predicted probability value based on the pixel values in the rearranged image frame, including:
[0068] S2-1-1. Perform a convolution operation on the rearranged image frame in the convolutional layer of the 3D convolutional model to obtain a first feature map with spatio-temporal features;
[0069] S2-1-2. After the convolution operation, perform a non-linear activation operation on the pixel values in the first feature map to obtain a second feature map;
[0070] S2-1-3. Perform a max pooling operation on the second feature map to obtain a third feature map;
[0071] S2-1-4. Perform convolution, non-linear activation, and max pooling on the third feature map n times to obtain a spatio-temporal feature map;
[0072] S2-1-5. Perform a flattening operation on the spatio-temporal feature map to obtain a flattened spatio-temporal feature vector;
[0073] S2-1-6. Input the flattened spatio-temporal feature vector into a fully connected layer, and the fully connected layer performs weighted processing on it to obtain a classification score vector based on the power operation behavior;
[0074] S2-1-7. Input the classification score vector based on the power operation behavior output by the fully connected layer into the Softmax layer of the classifier, and the Softmax layer performs normalization processing on it to obtain a predicted probability value based on the pixel values in the rearranged image frames.
[0075] In some of the embodiments, preprocessing the first real-time image frame to obtain a second real-time image frame includes:
[0076] Normalize the pixel values of the first real-time image frame and perform image scaling on the resolution of the first real-time image frame.
[0077] In some of the embodiments, the warning strategy includes:
[0078] Risk-free strategy: No warning information needs to be output, corresponding to the safe behavior label;
[0079] Low-risk strategy: Currently, it is a low-risk operation. The operator should stay alert and ensure that the protective equipment is complete, corresponding to the low-risk behavior label;
[0080] Medium-risk strategy: Currently, it is a medium-risk operation. The operator should carefully check the operation environment and tools to ensure strict compliance with the safety operation procedures;
[0081] High-risk strategy: Currently, it is a high-risk operation. Immediately start monitoring for high-risk operations. The operator must wear complete protective equipment and closely monitor the operation environment. If necessary, the operation can be suspended and additional safety measures can be taken.
[0082] The present invention provides a power safety operation monitoring system. By screening image frames through cosine similarity, adjacent frames with high similarity are selected. Then, through time series rearrangement, adjacent frame pairs are arranged in chronological order to ensure that the selected image frames are not only similar in spatial features but also coherent in time series features. Thus, a high-quality training set with guarantees in both time series and similarity is constructed, significantly improving the model's recognition ability of power operation behaviors. BRIEF DESCRIPTION OF THE DRAWINGS
[0083] Figure 1 It is a structural block diagram of a power safety operation monitoring system of the present invention;
[0084] Figure 2 It is a flowchart for obtaining an image frame training set in the present invention;
[0085] Figure 3 It is a flowchart for constructing rearranged image frames in the present invention;
[0086] Figure 4 It is a training flowchart of the operation monitoring model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0087] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0088] Embodiment 1: Please refer to Figure 1 , the present invention provides a power safety operation monitoring system, including:
[0089] A real-time operation video acquisition module, configured to acquire a real-time power operation video and decompose it into first real-time image frames with N consecutive time series frame numbers;
[0090] Specifically, the real-time operation video acquisition module in this embodiment can use an audio-video decoder to decode and decompose the power operation video through the audio-video decoder. Specifically, the audio-video decoder receives the power operation video as input, decodes the video, and obtains N consecutive image frames. At the same time, frame numbers are assigned to the decomposed image frames based on the continuous time series. For example, the first image frame can be marked as 0001, the second as 0002, and so on.
[0091] An image frame preprocessing module, configured to preprocess the first real-time image frames to obtain second real-time image frames;
[0092] A label output module, which is used to input the second real-time image frame into the trained job monitoring model, obtain the predicted probability value of the power operation behavior in the second real-time image frame, and output the category corresponding to the maximum predicted probability value as the power operation behavior classification label;
[0093] An alarm policy output module, which is used to match the corresponding alarm policy in the preset alarm policy library according to the power operation behavior classification label and output it in real time.
[0094] The power safety operation monitoring system of this embodiment obtains the real-time power operation video through the real-time operation video acquisition module; obtains the second real-time image frame through the image frame preprocessing module; obtains the power operation behavior classification label through the label output module, and obtains the corresponding alarm policy through the alarm policy output module; thus realizing the real-time monitoring and real-time alarm of the power operation site, ensuring the safety of operators and improving the operation efficiency.
[0095] Embodiment 2: Refer to Figures 2 to 4 , the technical solution of this Embodiment 2 different from that of Embodiment 1 is to disclose the training process of the job monitoring model in Embodiment 1, and this training process includes:
[0096] S1. Obtain an image frame training set; wherein, the training samples of the image frame training set are rearranged image frames, and the rearranged image frames represent historical image frames rearranged according to the time sequence frame numbers;
[0097] S2. Input the rearranged image frames into a three-dimensional convolutional model for iterative training to obtain a job monitoring model; wherein, the job monitoring model can extract the spatial features and temporal features of the rearranged image frames.
[0098] The job monitoring model of this embodiment is trained using a three-dimensional convolutional model, so it can capture the dimensional features of space and time in the video at the same time, making the model more focused on understanding continuous actions, and thus more accurately identifying different behaviors in power operations.
[0099] Further, step S1 specifically includes:
[0100] S1-1. Obtain the power operation videos within M historical time periods and decompose them into the first historical image frames with N consecutive time sequence frame numbers; wherein, the consecutive time sequence frame numbers mean that the power operation videos within the historical time periods are decomposed according to the time sequence, and each power operation video is decomposed into N first historical image frames, with a total of first historical image frames; wherein, each power operation video records the monitoring videos of several power operators performing power operations in different time periods.
[0101] S1-2. Preprocess the first historical image frame to obtain the second historical image frame. Specifically, the preprocessing operation in this embodiment includes normalizing the pixel values of the first historical image frame, rather than standardizing them. Because compared with standardization, normalization is more suitable for processing scene data with known and consistent pixel value ranges in this embodiment, while standardization is suitable for processing scene data with large differences in dimensions or data distributions closer to the normal distribution. The preprocessing also includes adjusting the resolution of the first historical image frame. For example, image scaling is performed on the image frame through an interpolation algorithm, that is, the resolution of the first historical image frame is adjusted to the required size using bilinear interpolation. In this embodiment, the specified resolution is 224*224.
[0102] S1-3. Calculate the cosine similarity of the second historical image frame to obtain the cosine similarity of the second historical image frame.
[0103] Further, step 1-3 also includes:
[0104] S1-3-1. Extract the feature vector of the second historical image frame. Among them, the feature vector can be the spatial feature extracted through a convolution operation or the histogram feature based on pixel values.
[0105] The expression for extracting the feature vector of the second historical image frame is:
[0106]
[0107] Among them, represents the i-th second historical image frame, is the feature vector of the i-th second historical image frame, and FE represents feature extraction.
[0108] S1-3-2. Calculate the cosine similarity between adjacent frame pairs according to the extracted feature vectors. The adjacent frame pairs represent adjacent second historical image frames.
[0109] The expression for cosine similarity is:
[0110]
[0111] represents the cosine similarity between adjacent frame pairs, represents the inner product of the feature vectors of adjacent frame pairs, represents the product of the norms of adjacent frame pairs;
[0112] S1-3-3. Generate a set S of cosine similarities of adjacent frame pairs according to the cosine similarities between adjacent frame pairs.
[0113] The expression for the set of cosine similarities is:
[0114] 。
[0115] In this embodiment, by performing feature extraction and cosine similarity calculation on historical image frames, adjacent frames with high similarity can be screened out, thereby maintaining the temporal sequence and similarity of historical image frames, and further ensuring that the model can better learn the variation law of power operation behaviors during learning.
[0116] S1-4. Rearrange the second historical image frames with cosine similarity higher than the similarity threshold according to adjacent time sequence frame numbers to obtain the rearranged image frames.
[0117] Further, step 1-4 further includes:
[0118] S1-4-1. Define the similarity threshold TS;
[0119] S1-4-2. Screen out adjacent frame pairs with cosine similarity higher than the similarity threshold from the cosine similarity set to obtain the screened adjacent image frames;
[0120] The expression for screening out adjacent frame pairs with cosine similarity higher than the similarity threshold from the cosine similarity set is:
[0121]
[0122] where represents the screened adjacent image frames, represents filtering out adjacent image frames with cosine similarity less than the similarity threshold TS.
[0123] S1-4-3. Rearrange the screened adjacent image frames according to adjacent time sequence frame numbers to generate a set of rearranged image frames.
[0124] The expression for rearranging the screened adjacent image frames according to adjacent time sequence frame numbers is:
[0125]
[0126] where represents the screened adjacent image frames, Reorder represents rearranging the screened adjacent image frames according to adjacent time sequence frame numbers, and N represents the number of image frames;
[0127] The expression for the set of rearranged image frames is:
[0128]
[0129] where X represents the set of rearranged image frames, represents the i-th rearranged image frame after rearrangement according to adjacent time sequence frame numbers;
[0130] Specifically, adjacent temporal frame numbers refer to: re - arranging the selected adjacent frames in chronological order to ensure that not only do the adjacent image frames after screening have high similarity, but also they are continuous in time sequence; the purpose of this process is to ensure that in the set of temporal frames, image frames with high similarity can appear in chronological order as much as possible, thereby improving the model's sensitivity to temporal changes.
[0131] The rearrangement of adjacent temporal frame numbers in this embodiment ensures that even the highly similar frames after screening retain the original time order, improving the model's sensitivity to action temporal changes and contributing to enhancing the recognition accuracy of power operation behaviors.
[0132] S1 - 5, Define the pixel values in the rearranged image frames as input features and the power operation behaviors as class labels, and construct training samples for the image frame training set;
[0133] In this embodiment, class labels can be defined for different power operation behaviors based on the power safety operation monitoring requirements. For example, the classification table of power operation behavior class labels is as follows:
[0134] Categories of electric power operation behaviors Risk levels Examples Safe behaviors 0 Operate with complete protective equipment Low-risk behaviors 1 Line inspection, equipment surface cleaning Medium-risk behaviors 2 Equipment maintenance, equipment grounding operation High-risk behaviors 3 Climbing heights, high-voltage equipment operation
[0135] S1 - 6, Aggregate the rearranged image frames to obtain the image frame training set.
[0136] The construction process of the image frame training set in this embodiment takes into account the temporal features and similarity of historical image frames. Through the cosine similarity calculation and rearrangement of historical image frames, it ensures that the training samples have good temporality and high similarity, which helps to improve the model's learning sensitivity to power operation behaviors.
[0137] Furthermore, step S2 specifically includes:
[0138] S2 - 1, The 3D convolutional model performs forward propagation on the image frame training set so that the output layer outputs the predicted probability values based on the pixel values in the rearranged image frames;
[0139] The expression of the predicted probability value is:
[0140]
[0141] Among them, represents the predicted probability value, represents the weight matrix of the model's fully - connected layer, represents the flattened spatio - temporal feature vector, represents the bias term, and Softmax represents converting the classification score vector of the fully - connected layer into a normalized predicted probability value.
[0142] S2-2. Convert the class label of the power operation behavior in the rearranged image frame into a one-hot vector to obtain the true label value of the power operation behavior;
[0143] The expression for the true label value is:
[0144]
[0145] where represents the true label value of the power operation behavior, represents the class label of the power operation behavior, and OneHot represents converting the class label into a one-hot vector;
[0146] S2-3. Calculate the multi-class cross-entropy loss between the predicted probability value and the true label value through the multi-class cross-entropy loss function;
[0147] The multi-class cross-entropy loss function is:
[0148]
[0149] where represents the multi-class cross-entropy loss, represents the k-th true label value, represents the predicted probability value of the i-th rearranged image frame corresponding to the k-th class label, and K represents the total number of classes.
[0150] S2-4. According to the multi-class cross-entropy loss, execute the backpropagation algorithm to update the weight matrix and bias term of the fully connected layer layer by layer; among them, define the weight matrix and bias term of the fully connected layer as model parameters.
[0151] Before starting backpropagation, to avoid gradient accumulation during iteration, then start backpropagating the error from the output layer.
[0152] The process of backpropagation is as follows:
[0153] In the fully connected layer, calculate the error gradient of the output layer with respect to the input of the fully connected layer according to the multi-class cross-entropy loss.
[0154] According to the error gradient and the current input data (flattened spatio-temporal feature vector), calculate the gradients of the weight matrix and bias term of the fully connected layer, and update these parameters according to the selected optimization algorithm (such as SGD or Adam) so that the model prediction can be closer to the true label during the next forward propagation.
[0155] S2-5. When the total loss value of the multi-class cross-entropy is lower than the loss threshold, it is considered that the 3D convolutional model reaches the convergence condition;
[0156] The expression for the convergence condition is:
[0157]
[0158] Among them, represents the total loss value of the multi-class cross-entropy, characterizing the total loss of the image frame training set. represents the loss threshold.
[0159] S2-6. Define the model parameters at convergence as the optimal parameters, and use the three-dimensional convolutional model with the optimal parameters as the trained job monitoring model.
[0160] The expression of the job monitoring model is:
[0161]
[0162] Among them, represents the job monitoring model, X represents the rearranged input image frame; h represents the spatio-temporal feature vector extracted by the three-dimensional convolutional model. represents the optimal weight matrix of the fully connected layer. represents the optimal bias term of the fully connected layer.
[0163] In this embodiment, the multi-class cross-entropy loss function is used to calculate the difference between the predicted probability value obtained by the model based on the rearranged image frame and the true label value converted into a one-hot vector; and the model parameters are optimized through the backpropagation algorithm to update the weights and bias terms, so that the difference in each round gradually decreases, and the model gradually converges during the training process.
[0164] Furthermore, step 2-1 further includes:
[0165] S2-1-1. Perform a convolution operation on the rearranged image frame in the convolutional layer of the three-dimensional convolutional model to obtain a first feature map with spatio-temporal features.
[0166] In this embodiment, the rearranged image frame passes through the convolutional layer of the three-dimensional convolutional model, and a convolution operation is performed on the pixel values of the rearranged image frame to extract its local spatial features and temporal features, obtaining the first layer of feature maps; the first layer of feature maps contains the local motion and spatial information of the images in each rearranged image frame, such as the action trajectories and operation areas in power operations.
[0167] S2-1-2. After the convolution operation, perform a non-linear activation operation on the pixel values in the first feature map to obtain a second feature map; among them, the non-linear activation specifically is to introduce the ReLU activation function, set the pixel values less than zero in the first feature map after convolution to 0, and keep the positive values unchanged, so as to obtain the second feature map.
[0168] S2-1-3. Perform max pooling operation on the second feature map to obtain a third feature map. Max pooling can reduce the size of the second feature map, extract key features in the rearranged image frame, and thus reduce the data dimension of the rearranged image frame.
[0169] S2-1-4. Perform convolution, non-linear activation, and max pooling on the third feature map for n times to obtain a spatio-temporal feature map.
[0170] S2-1-5. Perform a flattening operation on the spatio-temporal feature map to obtain a flattened spatio-temporal feature vector.
[0171] S2-1-6. Input the flattened spatio-temporal feature vector into a fully connected layer, and the fully connected layer performs weighted processing on it to obtain a classification score vector based on the power operation behavior.
[0172] S2-1-7. Input the classification score vector based on the power operation behavior output by the fully connected layer into the Softmax layer of the classifier. The Softmax layer performs normalization processing on it to obtain a predicted probability value based on the pixel values in the rearranged image frame.
[0173] During the forward propagation process, through a series of steps such as convolution operations, non-linear activation, and pooling operations, the model can gradually extract high-level spatio-temporal features (temporal features and spatial features) and finally obtain the predicted probability value.
[0174] In this embodiment, preprocessing the first real-time image frame to obtain a second real-time image frame includes:
[0175] Normalize the pixel values of the first real-time image frame and perform image scaling on the resolution of the first real-time image frame. Specifically, normalization and resolution adjustment are to unify the input data format for easy model processing, reduce noise interference, and improve the generalization ability and stability of the model.
[0176] In this embodiment, the warning strategy includes:
[0177] ① Risk-free strategy: No warning information needs to be output, corresponding to the safe behavior label.
[0178] ② Low-risk strategy: Currently, it is a low-risk operation. The operator should stay alert and ensure that the protective equipment is complete, corresponding to the low-risk behavior label.
[0179] ③ Medium-risk strategy: Currently, it is a medium-risk operation. The operator should carefully check the operation environment and tools to ensure strict compliance with safety operation procedures.
[0180] ④ High-risk strategy: Currently, it is a high-risk operation. Immediately initiate high-risk operation monitoring. The operators must wear complete protective equipment and closely monitor the operation environment. If necessary, the operation can be suspended and additional safety measures can be taken.
[0181] The alarm strategy library contains response strategies designed for different risk behaviors. These strategies can be in the form of voice announcements, earphone transmissions, etc., to timely remind the operators to pay attention to safety or take emergency measures when necessary, thus greatly reducing the possibility of accidents and ensuring the safety of the operation site.
[0182] In summary, a power safety operation monitoring system provided by the present invention can extract both temporal features and spatial features in real-time power operation videos by using a 3D convolutional model to train the rearranged image frames. Therefore, the system is more sensitive in processing continuous actions and improves the monitoring ability of power operation behaviors.
[0183] Specifically, the present invention calculates the cosine similarity of historical image frames and screens out adjacent frame pairs with similarity higher than the threshold, ensuring that the selected historical image frames are not only similar in spatial features but also have behavioral continuity, guaranteeing the high quality of training data and improving the generalization ability of the model; by temporally rearranging the selected adjacent frames, the present invention not only screens out similar image frames but also ensures that the features of the image frames have high similarity in the time series. Thus, the model can better learn the temporal variation law of operation behaviors and improve the recognition ability of behaviors with strong time dependence such as power operations.
[0184] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as infrared, wireless, microwave, etc.) means.
[0185] The computer-readable storage medium may be any available medium accessible by a computer or a data storage device such as a server or a data center that includes one or more collections of available media. The available media may be magnetic media (e.g., floppy disk, hard disk, magnetic tape), optical media (e.g., DVD), or semiconductor media. The semiconductor media may be a solid state drive.
[0186] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods may be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division of an underwater terrain change analysis system and method for a waterway. In actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other may be through some interfaces, and the indirect couplings or communication connections of the devices or units may be in electrical, mechanical, or other forms.
[0187] As described above, the above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
[0188] As described above, the above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application.
Claims
1. A power safety operation monitoring system, characterized in that, Including: A real-time operation video acquisition module, configured to acquire a real-time power operation video and decompose it into a first real-time image frame with N consecutive time sequence frame numbers; An image frame preprocessing module, configured to preprocess the first real-time image frame to obtain a second real-time image frame; A label output module, configured to input the second real-time image frame into a trained operation monitoring model to obtain a predicted probability value of the power operation behavior in the second real-time image frame, and output the category corresponding to the maximum predicted probability value as the power operation behavior classification label; An alarm policy output module, configured to match a corresponding alarm policy in a preset alarm policy library according to the power operation behavior classification label and output it in real time.
2. The power safety operation monitoring system according to claim 1, wherein The training process of the operation monitoring model includes: S1. Acquire an image frame training set; the training samples of the image frame training set are rearranged image frames, and the rearranged image frames represent historical image frames rearranged according to time sequence frame numbers; S2. Input the rearranged image frames into a three-dimensional convolutional model for iterative training to obtain the operation monitoring model; the operation monitoring model can extract the spatial features and time sequence features of the rearranged image frames.
3. The power safety operation monitoring system according to claim 2, characterized in that, The acquisition of the image frame training set includes: S1-1. Acquire power operation videos within M historical time periods and decompose them into first historical image frames with N consecutive time sequence frame numbers; S1-2. Preprocess the first historical image frames to obtain second historical image frames; S1-3. Calculate the cosine similarity of the second historical image frames to obtain the cosine similarity of the second historical image frames; S1-4. Rearrange the second historical image frames with cosine similarity higher than the similarity threshold according to adjacent time sequence frame numbers to obtain the rearranged image frames; S1-5. Define the pixel values in the rearranged image frames as input features and the power operation behavior as category labels to construct training samples of the image frame training set; S1-6. Aggregate the rearranged image frames to obtain the image frame training set.
4. An electric power safety operation monitoring system according to claim 3, characterized in that, The calculation of the cosine similarity of the second historical image frames includes: S1-3-1. Extract the feature vectors of the second historical image frames; The expression for extracting the feature vectors of the second historical image frames is: represents the i-th second historical image frame, is the feature vector of the i-th second historical image frame, where FE represents feature extraction; S1-3-2. Calculate the cosine similarity between adjacent frame pairs according to the extracted feature vectors; adjacent frame pairs represent adjacent second historical image frames; The expression for the cosine similarity is: Indicates the cosine similarity between adjacent frame pairs, Indicates the inner product of the feature vectors of adjacent frame pairs, Indicates the product of the magnitudes of adjacent frame pairs; S1-3-3. Generate a cosine similarity set S of adjacent frame pairs according to the cosine similarity between adjacent frame pairs; The expression for the cosine similarity set is: 。 5. A power safety operation monitoring system according to claim 4, characterized in that, The rearrangement of the second historical image frames with cosine similarity higher than the similarity threshold according to adjacent time sequence frame numbers includes: S1-4-1. Define a similarity threshold TS; S1-4-2. Screen out adjacent frame pairs higher than the similarity threshold from the cosine similarity set to obtain screened adjacent image frames; The expression for screening out adjacent frame pairs higher than the similarity threshold from the cosine similarity set is: Indicates adjacent image frames after screening, Indicates adjacent image frames with similarity less than the similarity threshold TS filtered out; S1-4-3. Rearrange the screened adjacent image frames according to adjacent time sequence frame numbers to generate a set of the rearranged image frames; The expression for rearranging the adjacent image frames after screening according to the adjacent time series frame numbers is as follows: Reorder represents rearranging the adjacent image frames after screening according to the adjacent time series frame numbers, and N represents the number of image frames; The expression for the set of rearranged image frames is as follows: X represents the set of rearranged image frames, denotes the i-th rearranged image frame after rearrangement by adjacent time sequence frame numbers.
6. An electric power safety operation monitoring system according to claim 5, characterized in that, Inputting the rearranged image frames into the 3D convolutional model for iterative training includes: S2-1, the 3D convolutional model performs forward propagation on the image frame training set so that the output layer outputs the predicted probability values based on the pixel values in the rearranged image frames; The expression for the predicted probability values is as follows: represents the predicted probability value, represents the weight matrix of the fully connected layer of the model, represents the flattened spatio-temporal feature vector, represents the bias term, and Softmax represents converting the classification score vector of the fully connected layer into a normalized predicted probability value; S2-2, converting the category labels of the power operation behaviors in the rearranged image frames into one-hot vectors to obtain the true label values of the power operation behaviors; The expression for the true label values is as follows: Represents the true label value of the power operation behavior, Represents the class label of the power operation behavior, and OneHot represents converting the class label into a one-hot vector; S2-3, calculating the multi-class cross-entropy loss between the predicted probability values and the true label values through the multi-class cross-entropy loss function; The multi-class cross-entropy loss function is as follows: Represents the multi-class cross-entropy loss, Represents the k-th true label value, Represents the predicted probability value of the k-th class label corresponding to the i-th rearranged image frame, where K represents the total number of classes; S2-4, according to the multi-class cross-entropy loss, execute the backpropagation algorithm to update the weight matrix and bias terms of the fully connected layer layer by layer; define the weight matrix and bias terms of the fully connected layer as model parameters; S2-5, when the total loss value of the multi-class cross-entropy is lower than the loss threshold, it is considered that the 3D convolutional model reaches the convergence condition; The expression for the convergence condition is as follows: Represents the total loss value of multi-class cross-entropy, characterizing the total loss of the image frame training set, Represents the loss threshold; S2-6, define the model parameters at convergence as the optimal parameters, and use the 3D convolutional model with the optimal parameters as the trained operation monitoring model; The expression for the operation monitoring model is as follows: represents the job monitoring model, X represents the rearranged input image frame; h represents the spatio-temporal feature vector extracted by the 3D convolutional model, represents the optimal weight matrix of the fully connected layer, represents the optimal bias term of the fully connected layer.
7. An electric power safety operation monitoring system according to claim 6, characterized in that, The 3D convolutional model performs forward propagation on the image frame training set, including: S2-1-1, perform a convolution operation on the rearranged image frames in the convolutional layer of the 3D convolutional model to obtain a first feature map with spatio-temporal features; the spatio-temporal features represent temporal features and spatial features; S2-1-2, after the convolution operation, perform a non-linear activation operation on the pixel values in the first feature map to obtain a second feature map; S2-1-3, perform a max pooling operation on the second feature map to obtain a third feature map; S2-1-4, perform convolution, non-linear activation, and max pooling n times on the third feature map to obtain a spatio-temporal feature map; S2-1-5, perform a flattening operation on the spatio-temporal feature map to obtain a flattened spatio-temporal feature vector; S2-1-6, input the flattened spatio-temporal feature vector into the fully connected layer, and the fully connected layer performs weighted processing on it to obtain a classification score vector based on the power operation behaviors; S2-1-7, input the classification score vector based on the power operation behaviors output by the fully connected layer into the Softmax layer of the classifier, and the Softmax layer performs normalization processing on it to obtain the predicted probability values based on the pixel values in the rearranged image frames.
8. An electric power safety operation monitoring system according to claim 7, characterized in that, Preprocessing the first real-time image frame includes: Normalizing the pixel values of the first real-time image frame and performing image scaling on the resolution of the first real-time image frame.
9. A power safety operation monitoring system according to claim 1, characterized in that, The alarm strategy includes: Risk-free strategy: No alarm information needs to be output, corresponding to the safe behavior label; Low-risk strategy: Currently, it is a low-risk operation, and the operator should stay alert and ensure that the protective equipment is complete, corresponding to the low-risk behavior label; Medium-risk strategy: Currently, it is a medium-risk operation. Operators should carefully check the operation environment and tools to ensure strict compliance with safety operation procedures; High-risk strategy: Currently, it is a high-risk operation. Immediately initiate high-risk operation monitoring. Operators must wear complete protective equipment and closely monitor the operation environment.
Citation Information
Patent Citations
Intelligent video monitoring system for power operation
CN112989110A
Human behavior identification method based on three-dimensional convolutional neural network and transfer learning model
CN107506740A
Method and device for obtaining candidate fragments from video and processing device
CN109977262A
Real-time intelligent video monitoring abnormal behavior analysis method based on slowfast double-frame rate
CN113743306A
Hand cleanliness detection method and system
CN116071687A