Power transmission line hot-line work behavior identification method and system based on deep learning

By building a live job behavior recognition system based on deep learning, the behavior of live jobs is identified using the space-time attention mechanism, which solves the problem of determining safe distances in live jobs and achieves higher work safety.

CN119992647APending Publication Date: 2025-05-13CHINA ELECTRIC POWER RESEARCH INSTITUTE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510024016.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

During live operation, how to effectively determine which safety distance constraint value is selected at different times to ensure the safety of the operators?

Method used

A method and system for identifying live-operated operations in transmission lines based on deep learning is constructed. By constructing a short video data set of typical behaviors of live-operated operations, a behavior recognition model based on the spatiotemporal attention mechanism is used to identify the behavior of live-operated workers, and the corresponding distance calculation and monitoring strategies are automatically triggered.

Benefits of technology

It realizes accurate identification of the behavior of live-operated workers, can automatically adjust the safety distance, improves the safety of the work, and solves the problem that the safety distance cannot be effectively determined in the prior art.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992647A_ABST
    Figure CN119992647A_ABST
Patent Text Reader

Abstract

The invention discloses a deep learning-based power transmission line hot-line work behavior identification method and system. The method comprises the steps of constructing a short video data set of hot-line work typical behaviors; preprocessing the short video to obtain a corresponding video frame sequence; dividing the video frame sequence into a training set and a test set, and marking a behavior category to which the video frame sequence in the training set belongs; constructing a live working behavior recognition model based on a space-time attention mechanism; training the model by using the video frame sequence in the training set; using the video frame sequence in the test set to test the model; when the recognition accuracy of the model on the behavior category to which each video frame sequence in the test set belongs reaches a preset threshold value, completing training and testing; and inputting the video frame sequence to be identified into the live-line work behavior identification model, and outputting a corresponding live-line work behavior category. The problem of effective identification of typical behaviors of live working personnel in equipotential entering and exiting and on-line working processes is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of live working safety protection, and in particular to a method and system for identifying live working behaviors on transmission lines based on deep learning. Background Art

[0002] Live working is a key technical means in power maintenance work, and it has always played an important role in ensuring the safety and stability of large power grids with ultra-high voltage and UHV as the backbone grid. In order to effectively reduce the risk of discharge, industry specifications have set limits on the safe distance between workers and infrastructure such as poles, towers, and conductors, as well as the combined gap constraints between workers and grounded bodies and live bodies during the process of entering and exiting the equipotential.

[0003] The constraint value of the safe clearance distance varies depending on the location of the worker. When monitoring the behavior of live workers, how to determine which safe distance constraint value to use at the current moment is the main problem. Ideally, behavior recognition is an effective way to identify the actual activities of live workers, and the subsequent distance calculation and monitoring strategies can be automatically triggered based on the recognition results.

[0004] At present, the safety control methods based on human motion recognition reported in the literature are mainly applied to on-site power outage maintenance operations and distribution network non-stop operations. Target detection models such as R-CNN and YOLO series have shown high accuracy and good generalization in scenarios such as helmet wearing recognition, short-sleeved shorts and illegal clothing recognition, and seat belt wearing recognition for operators. Deep learning algorithms such as NanoDet use the GhostBlock structure in the feature fusion network to complete up and down sampling, and achieve effective recognition of illegal behaviors such as not wearing insulating gloves, not wearing insulating clothes, not insulating shielding, and not wearing goggles. However, there is no research on the recognition of operator action behavior during live transmission line operations, especially during equipotential operations. Summary of the invention

[0005] In order to solve the problems existing in the prior art, the present invention provides a method for identifying live working behaviors of power transmission lines based on deep learning, comprising:

[0006] Construct a short video dataset of typical behaviors of live working; preprocess each short video in the short video dataset to obtain a corresponding video frame sequence; divide the video frame sequence into a training set and a test set, and annotate the behavior category to which each video frame sequence in the training set belongs;

[0007] Construct a live working behavior recognition model based on spatiotemporal attention mechanism;

[0008] The live working behavior recognition model is trained using the video frame sequence in the training set; the live working behavior recognition model is tested using the video frame sequence in the test set; when the recognition accuracy of the behavior category to which each video frame sequence in the test set belongs reaches a preset threshold through the live working behavior recognition model, the live working behavior recognition model completes training and testing;

[0009] The video frame sequence to be identified is input into the live working behavior identification model, and the corresponding live working behavior category is output.

[0010] Preferably, preprocessing each short video in the short video dataset to obtain a corresponding video frame sequence includes:

[0011] Video frames are extracted from the original short video of typical live working behaviors at a fixed sampling rate of 0.2s, and every 12 consecutive frames constitute a video frame sequence;

[0012] Use a sliding window with an interval of 1s to extract all video frame sequences;

[0013] Delete the video frame sequences that do not contain the operating personnel.

[0014] Preferably, the live working behavior recognition model based on the spatiotemporal attention mechanism includes:

[0015] Convolutional layer, max pooling layer, 4 consecutive and identical residual layers, global pooling layer and fully connected layer connected in sequence.

[0016] Preferably, the structure and specific steps of the residual layer are:

[0017] The feature vector of the input residual layer is input into the temporal attention module TAM, channel attention module CAM and spatial attention module SAM arranged in parallel respectively;

[0018] After the output combination of the three modules is added, it is outputted through the first convolutional layer, the second convolutional layer, and the third convolutional layer in sequence;

[0019] The feature vector of the input residual layer is downsampled, and the downsampled result is combined and added with the result of the output of the third convolutional layer to obtain the residual layer output.

[0020] Preferably, the time attention module TAM is used to identify important information in the time dimension during live working and assign corresponding weights to the important information. The specific steps are:

[0021] The input feature matrix Use global average pooling and global maximum pooling to perform combined pooling calculations to obtain the initial weight of each frame in the feature matrix; f, h, w, c represent the number of frames, height pixels, width pixels, and channels of the input feature, respectively;

[0022] Weighted calculation results of combined pooling By introducing trainable parameters And the nonlinear activation function ReLU, perform the expansion operation; the calculation process of the expansion operation is:

[0023] M te =σ(M tp ·w t1 )

[0024] Introducing trainable parameters And the nonlinear activation function Sigmoid, restore the learned frame weight features to a dimension equal to the number of frames, and complete the calculation results of the expansion operation The calculation process of the compression operation is:

[0025]

[0026] After the compression operation, the final output of TAM is the input feature X t Frame weight matrix in the time dimension

[0027] Preferably, the channel attention module CAM is used to improve the screening and recognition capability of the channel, and the specific steps are:

[0028] Introducing trainable parameters For the input feature X c Perform the first compression operation, retaining only the channel dimension of the feature matrix. The calculation process of the first compression operation is:

[0029] M cs1 =w c1 ·X c

[0030] The trainable parameters The result of the first compression operation Perform matrix multiplication, and then pass the nonlinear activation function ReLU to calculate the result of the first compression operation Perform an expansion operation. The calculation process of the expansion operation is:

[0031] M cs2 =σ(w c2 ·M cs1 )

[0032] The result of the calculation of the extended operation Perform a second compression operation to restore the learned frame weight features to the required dimensions and calculate the results of the expansion operation Perform the second compression operation. The calculation process of the second compression operation is:

[0033]

[0034] The final output of CAM after the second compression operation is the input feature X c The channel attention weight matrix

[0035] Preferably, the spatial attention module SAM is used to detect the local intensity and edge features of the image, and the specific steps are:

[0036] The softmax function is used to map the value of each pixel in the input feature map to the probability of the corresponding point, and the attention weight matrix of the image intensity is obtained;

[0037] Use the Sobel operator as the convolution kernel to perform two-dimensional convolution on the spatial dimension of the input feature map, normalize the two-dimensional convolution result to the maximum and minimum values, and obtain the spatial attention weight based on the image edge;

[0038] The attention weight based on image intensity and the attention weight based on image edge are averaged and fused, and the Sigmoid function is used for nonlinear activation to obtain the SAM weight.

[0039] The present invention also provides a transmission line live working behavior recognition system based on deep learning, comprising:

[0040] The training set and test set construction module is used to construct a short video data set of typical behaviors of live working; preprocess each short video in the short video data set to obtain a corresponding video frame sequence; divide the video frame sequence into a training set and a test set, and annotate the behavior category to which each video frame sequence in the training set belongs;

[0041] Recognition model building module, used to build a live working behavior recognition model based on spatiotemporal attention mechanism;

[0042] A training and testing module, used to train the live working behavior recognition model using the video frame sequence in the training set; and to test the live working behavior recognition model using the video frame sequence in the test set; when the recognition accuracy of the behavior category to which each video frame sequence in the test set belongs reaches a preset threshold through the live working behavior recognition model, the live working behavior recognition model completes training and testing;

[0043] The recognition module is used to input the video frame sequence to be recognized into the live working behavior recognition model and output the corresponding live working behavior category.

[0044] Preferably, the training set and test set construction modules include:

[0045] The video frame sequence construction submodule is used to extract video frames from the original short video of typical live working behaviors at a fixed sampling rate of 0.2s, and every 12 consecutive frames constitute a video frame sequence;

[0046] The video frame sequence extraction submodule is used to extract all video frame sequences using a sliding window with an interval of 1s;

[0047] The video frame sequence cleaning submodule is used to delete the video frame sequence whose video content does not contain the operator.

[0048] Preferably, the live working behavior recognition model based on the spatiotemporal attention mechanism includes:

[0049] Convolutional layer, max pooling layer, 4 consecutive and identical residual layers, global pooling layer and fully connected layer connected in sequence.

[0050] The present invention provides a method and system for identifying live working behaviors on power transmission lines based on deep learning. The constructed TAM can give greater weights to frames with larger amounts of information, thereby enhancing the accuracy of feature extraction. For the channel dimension of image features, the constructed CAM helps to improve the model's ability to screen and identify channels with larger amounts of information and more distinguishable features. For the length and width dimensions of image features, the constructed SAM can detect the local intensity and edge features of the image, and use the two as spatial dimensions to calculate weights. The above method and model can suppress information redundancy and sparse distribution of feature channels containing action information, thereby improving the accuracy and performance of the network as a whole. It solves the problem of effectively identifying typical behaviors of live working personnel when entering and exiting equipotential and working on the line. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 is an overall flow chart of the method of the embodiment of the present invention;

[0052] Figure 2 It is a screenshot of part of the data set used in the embodiment of the present invention;

[0053] Figure 3 It is a schematic diagram of the model structure of the method of the embodiment of the present invention;

[0054] Figure 4 is a diagram of ablation experiment results of the method according to an embodiment of the present invention;

[0055] Figure 5It is a schematic diagram of the structure of the system of the embodiment of the present invention. DETAILED DESCRIPTION

[0056] Many specific details are set forth in the following description to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of the present invention, so the present invention is not limited to the specific implementation disclosed below.

[0057] The present invention mainly aims at the recognition of human behavior in short videos of live working. In the embodiment, the method provided by the present invention is used to recognize 8 types of typical behavior short videos collected. Figure 2 The above 8 typical behaviors include entering equipotential, riding across the conductor, lifting, potential transfer, online operation, routing, crossing two and shorting three, and exiting equipotential.

[0058] Example 1

[0059] A method for identifying live working behaviors on power transmission lines based on deep learning is provided in a preferred embodiment of the present invention, and a schematic diagram of the model structure is shown in FIG. Figure 3 As shown in the figure, the stacked live working short video frames are used as the input network, and the output feature vector X is obtained through the convolution layer, and then the output feature vector X is obtained through the maximum pooling layer MaxPool. m .X m After passing through four residual structures (Res1 / 2 / 3 / 4+TAM+CAM+SAM), the final live working behavior recognition is completed through the global pooling layer AvgPool and the fully connected layer FC.

[0060] The specific steps of the present invention are as follows Figure 1 As shown, the following steps are included:

[0061] Step S101, constructing a short video dataset of typical behaviors of live working; preprocessing each short video in the short video dataset to obtain a corresponding video frame sequence; dividing the video frame sequence into a training set and a test set, and labeling the behavior category to which each video frame sequence in the training set belongs.

[0062] Step S102, constructing a live working behavior recognition model based on the spatiotemporal attention mechanism.

[0063] Step S103, use the video frame sequence in the training set to train the live working behavior recognition model; use the video frame sequence in the test set to test the live working behavior recognition model; when the recognition accuracy of the behavior category belonging to each video frame sequence in the test set reaches a preset threshold through the live working behavior recognition model, the live working behavior recognition model completes training and testing.

[0064] Step S104: input the video frame sequence to be identified into the live working behavior identification model, and output the corresponding live working behavior category.

[0065] Preferably, in step S101, preprocessing each short video in the short video dataset to obtain a corresponding video frame sequence includes:

[0066] 1) Extract video frames from the original short video of typical live working behaviors at a fixed sampling rate of 0.2s, and every 12 consecutive frames constitute a video frame sequence;

[0067] 2) Use a sliding window with an interval of 1s to extract all video frame sequences;

[0068] 3) Delete the video frame sequence whose video content does not contain the operator.

[0069] Preferably, in the step S102, the live working behavior recognition model (STAM-3DRN) based on the spatiotemporal attention mechanism is constructed, including: a convolutional layer, a maximum pooling layer, 4 consecutive and identical residual layers, a global pooling layer and a fully connected layer connected in sequence.

[0070] The structure and specific steps of the residual layer are as follows:

[0071] 1) Input the feature vector of the input residual layer into the temporal attention module TAM, channel attention module CAM and spatial attention module SAM arranged in parallel respectively;

[0072] 2) After adding the output combinations of the three modules, they are outputted through the first convolutional layer, the second convolutional layer, and the third convolutional layer in sequence;

[0073] 3) Perform a downsampling operation on the feature vector of the input residual layer, and add the downsampling result and the result of the output of the third convolutional layer together to obtain the residual layer output.

[0074] The time attention module TAM is used to identify important information in the time dimension during live working and assign corresponding weights to the important information. The specific steps are as follows:

[0075] The input feature matrix Use global average pooling and global maximum pooling to perform combined pooling calculations to obtain the initial weight of each frame in the feature matrix; f, h, w, c represent the number of frames, height pixels, width pixels, and channels of the input feature, respectively;

[0076] Weighted calculation results of combined pooling By introducing trainable parameters And the nonlinear activation function ReLU, perform the expansion operation; the calculation process of the expansion operation is:

[0077] M te =σ(M tp ·w t1 )

[0078] Introducing trainable parameters

[0079] And the nonlinear activation function Sigmoid, restore the learned frame weight features to a dimension equal to the number of frames, and complete the calculation results of the expansion operation The calculation process of the compression operation is:

[0080]

[0081] After the compression operation, the final output of TAM is the input feature X t Frame weight matrix in the time dimension

[0082] The channel attention module CAM is used to improve the screening and recognition capabilities of channels. The specific steps are as follows:

[0083] Introducing trainable parameters For the input feature X c Perform the first compression operation, retaining only the channel dimension of the feature matrix. The calculation process of the first compression operation is:

[0084] M cs1 =w c1 ·X c

[0085] The trainable parameters The result of the first compression operation Perform matrix multiplication, and then pass the nonlinear activation function ReLU to calculate the result of the first compression operation Perform an expansion operation. The calculation process of the expansion operation is:

[0086] M cs2 =σ(w c2 ·M cs1 )

[0087] The result of the calculation of the extended operation Perform a second compression operation to restore the learned frame weight features to the required dimensions and calculate the results of the expansion operation Perform the second compression operation. The calculation process of the second compression operation is:

[0088]

[0089] The final output of CAM after the second compression operation is the input feature X c The channel attention weight matrix

[0090] The spatial attention module SAM is used to detect the local intensity and edge features of the image. The specific steps are as follows:

[0091] The softmax function is used to map the value of each pixel in the input feature map to the probability of the corresponding point, and the attention weight matrix of the image intensity is obtained;

[0092] Use the Sobel operator as the convolution kernel to perform two-dimensional convolution on the spatial dimension of the input feature map, normalize the two-dimensional convolution result to the maximum and minimum values, and obtain the spatial attention weight based on the image edge;

[0093] The attention weight based on image intensity and the attention weight based on image edge are averaged and fused, and the Sigmoid function is used for nonlinear activation to obtain the SAM weight.

[0094] Example 2 The present invention mainly focuses on the recognition of human behavior in short videos of live working. In the example, the method provided by the present invention is used to recognize 8 types of typical behavior short videos collected. Figure 2 The above 8 typical behaviors include entering equipotential, riding across the conductor, lifting, potential transfer, online operation, routing, crossing two and shorting three, and exiting equipotential.

[0095] A method for identifying live working behaviors on power transmission lines based on deep learning is provided in a preferred embodiment of the present invention, and a schematic diagram of the model structure is shown in FIG. Figure 3 As shown in the figure, the stacked live working short video frames are used as the input network, and the output feature vector X is obtained through the convolution layer, and then the output feature vector X is obtained through the maximum pooling layer MaxPool. m .X m After passing through four residual structures (Res1 / 2 / 3 / 4+TAM+CAM+SAM), the final live working behavior recognition is completed through the global pooling layer AvgPool and the fully connected layer FC.

[0096] Preferably, the operation of preprocessing the short video of typical live working behaviors in step 1 is:

[0097] 1) Extract video frames from the original short video of typical live working behaviors at a fixed sampling rate of 0.2s, ensuring that every 12 consecutive frames constitute a video frame sequence;

[0098] 2) Use a sliding window with an interval of 1s to extract all video frame sequences;

[0099] 3) Delete the video frame sequence whose video content does not contain the operator.

[0100] A total of 100 short videos of typical operations of live transmission line work were selected, covering common entry and exit equipotential methods and online work methods. Through the above preprocessing method, a total of 100,000 training samples were obtained and divided into training set and validation set in a 2:1 ratio.

[0101] Preferably, the live working behavior recognition model based on spatiotemporal attention mechanism (STAM-3DRN) in step 2 includes a convolutional layer, a maximum pooling layer, 4 consecutive and identical residual layers, a global pooling layer and a fully connected layer connected in sequence.

[0102] The structure and specific steps of the residual layer are as follows:

[0103] 1) Input the feature vector of the input residual layer into the temporal attention module (TAM), channel attention module (CAM) and spatial attention module (SAM) arranged in parallel respectively;

[0104] 2) After adding the output combinations of the three modules, they are sequentially passed through the first convolutional layer, the second convolutional layer, and the third convolutional layer;

[0105] The temporal attention module (TAM) can identify important information in the time dimension during live working, and can give greater weight to frames with more temporal information, thereby enhancing the accuracy of feature extraction. TAM is divided into the following three steps:

[0106] 1) Input feature matrix Use global average pooling and global maximum pooling to perform combined pooling calculations to obtain the initial weights of each frame in the feature matrix. f, h, w, c represent the number of frames, height pixels, width pixels, and channels of the input feature, respectively;

[0107] 2) Weighted calculation results of combined pooling Perform extended operations. By introducing trainable parameters And the nonlinear activation function ReLU, enhance the information capture ability of the model. The calculation process of the expansion operation is:

[0108] M te =σ(M tp ·w t1 )

[0109] 3) Calculation results of step 2) Perform compression operations. By introducing trainable parameters

[0110] And the nonlinear activation function Sigmoid, restore the learned frame weight features to a dimension equal to the number of frames. The calculation process of the compression operation is:

[0111]

[0112] After the compression operation is completed, the final output of TAM is obtained, that is, the input feature X t Frame weight matrix in the time dimension

[0113] The channel attention module (CAM) can improve the model's ability to screen and identify channels with larger information and more distinguishable features. CAM is divided into the following three steps:

[0114] 1) Introducing trainable parameters For the input feature X c Perform the first compression operation to retain only the channel dimension of the feature matrix. The calculation process of the first compression operation is:

[0115] M cs1 =w c1 ·X c

[0116] 2) Calculation results of step 1) Perform an expansion operation. Perform matrix multiplication with the output of step 1) and then pass through the nonlinear activation function ReLU. The calculation process of the expansion operation is:

[0117] M cs2 =σ(w c2 ·M cs1 )

[0118] 3) Calculation results of step 2) Perform the second compression operation to restore the learned frame weight features to the required dimension. The calculation process of the second compression operation is:

[0119]

[0120] After the second compression operation, the final output of CAM is obtained, that is, the input feature X c The channel attention weight matrix

[0121] The spatial attention module (SAM) can detect the local intensity and edge features of the image. SAM is divided into the following three steps:

[0122] 1) Use the softmax function to map the value of each pixel in the input feature map to the probability of the corresponding point, and obtain the attention weight matrix of the image intensity;

[0123] 2) Use the Sobel operator as the convolution kernel to perform two-dimensional convolution on the spatial dimension of the input feature map, and then perform maximum and minimum value normalization to obtain the spatial attention weight based on the image edge;

[0124] 3) The attention weight based on image intensity and the attention weight based on image edge are averaged and fused, and the Sigmoid function is used for nonlinear activation to finally obtain the SAM weight.

[0125] After the pre-shot and selected short videos of live working are input into the constructed deep learning model, it passes through four consecutive residual structures (Res1 / 2 / 3 / 4+TAM+CAM+SAM) and then through the global pooling layer AvgPool and the fully connected layer FC, and finally the recognition results of the behavior actions of the live working personnel are obtained.

[0126] In a preferred embodiment, for a method for identifying live working behaviors on power transmission lines based on deep learning provided by the present invention, a 3D ResNet-50 network is used as the baseline model, and TAM, CAM, and SAM modules are added to perform ablation experiments. The experimental results are shown in Figure 2. Figure 4 shown.

[0127] Figure 4 The curves showing the changes in the accuracy of the four models as the number of elapsed times increases during the test process. The 3D ResNet network with all three types of attention modules added has the highest accuracy and the fastest convergence speed. The accuracy of the three networks, 3D ResNet+TAM, 3DResNet+CAM, and 3D ResNet+SAM, is relatively close. The difference in accuracy between the 3D ResNet+TAM / CAM / SAM network with a single module added and the 3D ResNet network with all three types of attention modules added is higher than the difference between them and the baseline network. It can be seen that when the three types of attention modules act separately, the degree of improvement in network performance is similar and not significant; when the three types of attention modules act together, the improvement in network performance is more significant. This shows that the three types of attention modules can jointly improve the behavior recognition effect of the model by parallel access, and there will be no conflict or interference between them.

[0128] In a preferred embodiment, a method for identifying live working behaviors on power transmission lines based on deep learning provided by the present invention was evaluated, and the recognition accuracy of the method for eight live working behavior categories was obtained, and it was compared with two popular deep learning models, 3D ResNet and 3D ResNeXt. The comparison results are shown in Table 1.

[0129] Table 1 Comparison of recognition accuracy of three models

[0130] category Number of samples 3DRN 3DRNX STAM-3DRN Entering the equipotential 11,869 0.803 0.815 0.898 Riding across the wire 5,112 0.758 0.799 0.865 promote 17,125 0.779 0.751 0.886 Potential transfer 6,089 0.831 0.825 0.912 Online homework 26,481 0.818 0.825 0.909 Routing 15,849 0.824 0.811 0.911 Cross-two-short-three method 6,980 0.834 0.827 0.925 Exit Equipotential 10,495 0.809 0.817 0.907 overall 100,000 0.808 0.807 0.903

[0131] As can be seen from the table above, among all action categories, STAM-3DRN has the highest accuracy in identifying the "cross two short three" method, which can reach 92.5%, an increase of about 9.1% compared to 3DRN and about 9.8% compared to 3DRNX. This is because the operator takes a leaning posture when performing the "cross two short three" operation on the insulator string, while the upper body of the operator is in an upright state when performing the other 7 types of actions. In addition, the background of the insulator string is significantly different from that of other split conductors or towers. The spatial attention module can adaptively increase the weight assignment; when the operator moves with the same hands and feet on the insulator string, the temporal attention module can effectively capture the instantaneous movement changes, thereby improving the recognition accuracy.

[0132] Based on the same inventive concept, the present invention also provides a transmission line live working behavior recognition system based on deep learning, such as Figure 5 As shown, including:

[0133] The training set and test set construction module 510 is used to construct a short video data set of typical behaviors of live working; pre-process each short video in the short video data set to obtain a corresponding video frame sequence; divide the video frame sequence into a training set and a test set, and annotate the behavior category to which each video frame sequence in the training set belongs;

[0134] A recognition model building module 520 is used to build a live working behavior recognition model based on a spatiotemporal attention mechanism;

[0135] The training and testing module 530 is used to train the live working behavior recognition model using the video frame sequence in the training set; and to test the live working behavior recognition model using the video frame sequence in the test set; when the recognition accuracy of the behavior category to which each video frame sequence in the test set belongs reaches a preset threshold through the live working behavior recognition model, the live working behavior recognition model completes training and testing;

[0136] The recognition module 540 is used to input the video frame sequence to be recognized into the live working behavior recognition model and output the corresponding live working behavior category.

[0137] Preferably, the training set and test set construction modules include:

[0138] The video frame sequence construction submodule is used to extract video frames from the original short video of typical live working behaviors at a fixed sampling rate of 0.2s, and every 12 consecutive frames constitute a video frame sequence;

[0139] The video frame sequence extraction submodule is used to extract all video frame sequences using a sliding window with an interval of 1s;

[0140] The video frame sequence cleaning submodule is used to delete the video frame sequence whose video content does not contain the operator.

[0141] Preferably, the live working behavior recognition model based on the spatiotemporal attention mechanism includes:

[0142] Convolutional layer, max pooling layer, 4 consecutive and identical residual layers, global pooling layer and fully connected layer connected in sequence.

[0143] The present invention proposes a method and system for identifying live working behaviors of power transmission lines based on deep learning. The constructed TAM can give greater weights to frames with larger amounts of information, thereby enhancing the accuracy of feature extraction. For the channel dimension of image features, the constructed CAM helps to improve the model's ability to screen and identify channels with larger amounts of information and more distinguishable features. For the length and width dimensions of image features, the constructed SAM can detect the local intensity and edge features of the image, and use the two as spatial dimensions to calculate weights. The above methods and models can suppress information redundancy and the sparse distribution of feature channels containing action information, thereby improving the accuracy and performance of the network as a whole.

[0144] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0145] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0146] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0147] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0148] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents, and any modifications or equivalents that do not depart from the spirit and scope of the present invention should be included in the scope of the claims of the present invention.

Claims

1. A method for identifying live working behaviors on power transmission lines based on deep learning, characterized in that: include: Construct a short video dataset of typical behaviors of live working; Preprocess each short video in the short video data set to obtain a corresponding video frame sequence; divide the video frame sequence into a training set and a test set, and annotate the behavior category to which each video frame sequence in the training set belongs; Construct a live working behavior recognition model based on spatiotemporal attention mechanism; The live working behavior recognition model is trained using the video frame sequence in the training set; the live working behavior recognition model is tested using the video frame sequence in the test set; when the recognition accuracy of the behavior category to which each video frame sequence in the test set belongs reaches a preset threshold through the live working behavior recognition model, the live working behavior recognition model completes training and testing; The video frame sequence to be identified is input into the live working behavior identification model, and the corresponding live working behavior category is output.

2. The method according to claim 1, characterized in that Preprocessing each short video in the short video dataset to obtain a corresponding video frame sequence includes: Video frames are extracted from the original short video of typical live working behaviors at a fixed sampling rate of 0.2s, and every 12 consecutive frames constitute a video frame sequence; Use a sliding window with an interval of 1s to extract all video frame sequences; Delete the video frame sequences that do not contain the operating personnel.

3. The method according to claim 1, characterized in that The live working behavior recognition model based on spatiotemporal attention mechanism includes: Convolutional layer, max pooling layer, 4 consecutive and identical residual layers, global pooling layer and fully connected layer connected in sequence.

4. The method according to claim 3, characterized in that The structure and specific steps of the residual layer are as follows: The feature vector of the input residual layer is input into the temporal attention module TAM, channel attention module CAM and spatial attention module SAM arranged in parallel respectively; After the output combination of the three modules is added, it is outputted through the first convolutional layer, the second convolutional layer, and the third convolutional layer in sequence; The feature vector of the input residual layer is downsampled, and the downsampled result is combined and added with the result of the output of the third convolutional layer to obtain the residual layer output.

5. The method according to claim 4, characterized in that The time attention module TAM is used to identify important information in the time dimension during live working and assign corresponding weights to the important information. The specific steps are as follows: The input feature matrix Use global average pooling and global maximum pooling to perform combined pooling calculations to obtain the initial weight of each frame in the feature matrix; f, h, w, c represent the number of frames, height pixels, width pixels, and channels of the input feature, respectively; Weighted calculation results of combined pooling By introducing trainable parameters And the nonlinear activation function ReLU, perform the expansion operation; the calculation process of the expansion operation is: M te =σ(M tp ·w t1 ) Introducing trainable parameters And the nonlinear activation function Sigmoid, restore the learned frame weight features to a dimension equal to the number of frames, and complete the calculation results of the expansion operation The calculation process of the compression operation is: After the compression operation, the final output of TAM is the input feature X t Frame weight matrix in the time dimension 6. The method according to claim 4, characterized in that The channel attention module CAM is used to improve the screening and recognition capabilities of channels. The specific steps are as follows: Introducing trainable parameters For the input feature X c Perform the first compression operation, retaining only the channel dimension of the feature matrix. The calculation process of the first compression operation is: M cs1 =w c1 ·X c The trainable parameters The result of the first compression operation Perform matrix multiplication, and then pass the nonlinear activation function ReLU to calculate the result of the first compression operation Perform an expansion operation. The calculation process of the expansion operation is: M cs2 =σ(w c2 ·M cs1 ) The result of the calculation of the expansion operation Perform a second compression operation to restore the learned frame weight features to the required dimensions and calculate the results of the expansion operation Perform the second compression operation. The calculation process of the second compression operation is: The final output of CAM after the second compression operation is the input feature X c The channel attention weight matrix 7. The method according to claim 4, characterized in that The spatial attention module SAM is used to detect the local intensity and edge features of the image. The specific steps are as follows: The softmax function is used to map the value of each pixel in the input feature map to the probability of the corresponding point, and the attention weight matrix of the image intensity is obtained; Use the Sobel operator as the convolution kernel to perform two-dimensional convolution on the spatial dimension of the input feature map, normalize the two-dimensional convolution result to the maximum and minimum values, and obtain the spatial attention weight based on the image edge; The attention weight based on image intensity and the attention weight based on image edge are averaged and fused, and the Sigmoid function is used for nonlinear activation to obtain the SAM weight.

8. A transmission line live working behavior recognition system based on deep learning, characterized in that: include: Training set and test set construction module, used to construct a short video dataset of typical behaviors of live working; Preprocess each short video in the short video data set to obtain a corresponding video frame sequence; divide the video frame sequence into a training set and a test set, and annotate the behavior category to which each video frame sequence in the training set belongs; Recognition model building module, used to build a live working behavior recognition model based on spatiotemporal attention mechanism; A training and testing module, used to train the live working behavior recognition model using the video frame sequence in the training set; and to test the live working behavior recognition model using the video frame sequence in the test set; when the recognition accuracy of the behavior category to which each video frame sequence in the test set belongs reaches a preset threshold through the live working behavior recognition model, the live working behavior recognition model completes training and testing; The recognition module is used to input the video frame sequence to be recognized into the live working behavior recognition model and output the corresponding live working behavior category.

9. The system according to claim 8, characterized in that The training set and test set building blocks include: The video frame sequence construction submodule is used to extract video frames from the original short video of typical live working behaviors at a fixed sampling rate of 0.2s, and every 12 consecutive frames constitute a video frame sequence; The video frame sequence extraction submodule is used to extract all video frame sequences using a sliding window with an interval of 1s; The video frame sequence cleaning submodule is used to delete the video frame sequence whose video content does not contain the operator.

10. The system according to claim 8, characterized in that The live working behavior recognition model based on spatiotemporal attention mechanism includes: Convolutional layer, max pooling layer, 4 consecutive and identical residual layers, global pooling layer and fully connected layer connected in sequence.