A power grid operator safety risk detection method and device
By fusing 2D and 3D posture features and utilizing a temporal decoupled adaptive graph convolutional network, the accuracy problem of temporal behavior recognition at power grid operation sites is solved, enabling refined and accurate detection of safety risks for power grid operators and reducing the probability of accidents.
Patent Information
- Application Number
- CN202511096246.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-08-06
AI Technical Summary
Existing technologies have difficulty accurately identifying temporal behaviors at power grid operation sites. A single video segment division method cannot cover the segmentation lengths of different behaviors, resulting in the inability to accurately identify certain behaviors with shorter durations or ignoring the temporal characteristics of the entire behavior.
By collecting videos of workers, extracting and fusing 2D and 3D posture features, and using temporal decoupled adaptive graph convolutional networks for behavior recognition, a 2D-3D spatial attention feature fusion network and cross-attention operation are designed, and temporal behavior recognition is performed by combining the changes in posture estimation results in multiple time periods.
It improves the accuracy of temporal behavior recognition, avoids posture estimation errors caused by occlusion and similar actions, ensures safety risk detection at power grid operation sites, reduces the probability of accidents, and protects the personal safety of operators.
Smart Images

Figure CN120599524B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of pose estimation and temporal behavior recognition, in particular to a power grid worker safety risk detection method and device. BACKGROUND
[0002] Power grid operation site accidents usually bring unavoidable personnel casualties and economic losses, and about 80% of power grid operation accidents are related to the illegal operation of workers. Therefore, on the one hand, the safety awareness and personal quality of the workers are improved, and on the other hand, the safety risk detection is strengthened to supervise the behavior, which can more effectively avoid the occurrence of power grid operation accidents and protect the safety of the workers.
[0003] Through video monitoring, the pose estimation and behavior recognition of the workers can be realized to detect the safety risk of the current operation task in real time, and when there is a safety risk such as improper operation and illegal behavior, a prompt or alarm is given to protect the personal safety of the power grid workers and reduce the probability of power plant operation site accidents. Human pose estimation based on video clips focuses on 2D and 3D directions: the 2D human pose estimation method has excellent performance and relatively simple model design, but is easily affected by the problems of occlusion and motion similarity, for example, the human abnormal behavior recognition method under the video monitoring of the substation based on pose estimation disclosed in Chinese patent publication No. CN116912930A is a 2D pose estimation method; the 3D human pose estimation method can provide more rich spatial information, but due to the limitation of input data, the model may be misjudged, resulting in incorrect pose estimation results. Using 2D pose as an intermediate representation, combined with the correlation of human key points between consecutive frames of video, the 3D pose can be inferred to achieve more accurate pose estimation results to improve the accuracy of temporal behavior recognition, for example, the video three-dimensional human pose estimation method and system based on multi-level supervised graph convolution disclosed in Chinese patent publication No. CN114694261A is to infer 3D pose through 2D pose.
[0004] However, the above scheme, when performing temporal behavior recognition, a single video clip division method is difficult to cover the division length of different behaviors, and too long video clips are easy to ignore some behaviors with short duration, while too short video clips make the model unable to capture the temporal features of the entire behavior, resulting in inaccurate recognition of such behaviors. Therefore, designing an accurate human pose estimation method and a method of temporal behavior recognition based on changes in multi-time period pose estimation results is of great significance for power grid worker safety risk detection. SUMMARY
[0005] The technical problem to be solved by the present application is how to provide a method of accurate temporal behavior recognition based on changes in multi-time period pose estimation results.
[0006] The present invention solves the above technical problems through the following technical means: a method for detecting safety risks of power grid operators, comprising:
[0007] S1. Video data collection of operators in different working scenes;
[0008] S2, extracting the 2D posture features and 3D posture features of each frame of the input video clip, and then fusing the 2D posture features and 3D posture features to obtain a fused posture feature;
[0009] S3, divide the obtained fusion posture features into layer, is an integer greater than or equal to 0. Each layer of the fused posture feature is used to perform temporal behavior recognition to obtain the corresponding temporal features for behavior recognition in the current video. Then, the two adjacent divided temporal features are fused to obtain a longer temporal feature for behavior recognition. This process is repeated. After that, the temporal features of the entire video are obtained, and behavior recognition is performed. The sequence composed of all behavior recognition results of the video is input into the risk assessment network to obtain the risk detection result. The risk assessment network is composed of three layers of fully connected networks in series.
[0010] Furthermore, S1 includes:
[0011] Collect videos of workers to build a dataset , divided into training set and test set , the videos in the dataset are divided into video clips and labeled with specific risk types. The training set is represented as ,in, Indicates that the training set contains Video clips, Represents the training set video clips, and the corresponding video clip types are marked as ,in , Indicates the number of new video segments each video is divided into, Represents the training set Type annotation of video clips, Indicates the The type label of the first new video segment of layer 0 divided by the video segments, Indicates the The type label of the first new video segment of the cth layer divided by the video segments, Indicates the The cth layer is divided into The type of new video clips is annotated; the test set is divided in the same way and is represented as , the type label of the first new video segment of the 0th layer divided by the th video segment of the test set, , denotes that the test set contains video segments, denotes the th video segment of the test set, denotes the type label of the th video segment of the test set, denotes the type label of the first new video segment of the 0th layer divided by the th video segment of the test set, denotes the type label of the first new video segment of the cth layer divided by the th video segment of the test set, denotes the type label of the th new video segment of the cth layer divided by the th video segment of the test set.
[0012] Further, the extraction manner of the 2D pose feature is:
[0013] Each frame of the input video is input into a 2D pose estimation network, and a 2D pose feature is output, and the working process of the 2D pose estimation network is to input each frame of the input video into a block embedding layer to obtain a block embedding result , then input into an encoder composed of multiple sequentially cascaded converters to obtain an encoding feature , the operation of each converter is represented as follows
[0014]
[0015] wherein, denotes the intermediate variable of the th converter, denotes the output of the th converter, denotes a layer normalization operation, denotes a multi-head self-attention operation, denotes a combined network comprising a first linear layer, a nonlinear activation function , Dropout and a second linear layer; then the encoding feature is input into a decoder to obtain a 2D pose feature , which is represented as follows:
[0016]
[0017] wherein, denotes a bilinear interpolation operation, a convolutional operation layer with a convolution kernel size of 3*3, linear transformation operation.
[0018] Further, the extraction manner of the 3D pose feature is:
[0019] The 2D pose feature is taken as an input of a 3D pose estimation network, and a 3D pose feature is output, the 3D pose estimation network being a sequentially connected feature extraction module, a space-time attention module and a channel attention module, the space-time attention module being composed of a plurality of time encoders and a plurality of space encoders connected in series.
[0020] Further, the processing process of the feature extraction module is:
[0021] Another block embedding layer is adopted to obtain a block embedding result , and then a convolution, a nonlinear activation function and a Dropout processing are performed to obtain a feature for 3D pose estimation in combination with position encoding information , and the calculation process is as follows:
[0022]
[0023] wherein, represents a batch normalization operation, represents a Dropout operation, represents a position embedding operation.
[0024] Further, the working process of the space-time attention module is:
[0025] The feature is input into the space-time attention module, and the operation process of the time encoder is represented by the following formula:
[0026]
[0027] wherein, represents an output of a previous time encoder, represents an output of a current time encoder, The time attention module is composed of 5 layers of time encoders, and therefore the output of the time attention module is ;
[0028] The operation process of the space encoder is represented by the following formula:
[0029]
[0030] wherein, represents a maximum value pooling, represents a convolutional operation layer with a convolution kernel size of 1*1, represents an output of a previous space encoder, represents the output of the current spatial encoder, , a 5-layer spatial encoder is used to form a spatial attention module, and the output of the spatial attention module is , so the output of the spatiotemporal attention module is .
[0031] Furthermore, the channel attention module adopts the SE module, the output result of the spatiotemporal attention module is used as the input of the SE module, and the SE module outputs 3D posture features.
[0032] Furthermore, the fusing of the 2D posture feature and the 3D posture feature to obtain the fused posture feature includes:
[0033] The 2D posture features and 3D posture features are input into the posture feature fusion network to obtain the final fused posture features. The posture feature fusion network uses the corresponding positions in the 2D posture features and 3D posture features as the input of the two branches. The two branches are processed by the third linear layer to obtain the query, key and value used in the multi-head self-attention operation, and complete their respective local attention calculations. After the local attention calculation is completed, the query calculated by the branch where the 3D posture features are located and the key and value combination calculated by the branch where the 2D posture features are located are used to perform cross-attention calculation to obtain global attention; after completing the global attention calculation, the two branches respectively map the attention calculation results through the fourth linear layer and add them to the input features. After the feature dimensions of the two branches are adjusted through the fifth linear layer, the features are connected, and finally the cross-attention feature information that can represent the global and local is obtained. As fused posture features.
[0034] Furthermore, S3 includes:
[0035] Each layer of the fused posture features is respectively subjected to temporal behavior recognition by a temporal decoupling adaptive graph convolutional network to obtain the corresponding temporal features; the temporal decoupling adaptive graph convolutional network is composed of m sequentially cascaded adaptive graph convolutional layers and a subsequent decoupling network; the operation process of the adaptive graph convolutional layer is as follows
[0036]
[0037] in, is the matrix multiplication operator symbol, Indicates the The output features of adaptive graph convolution, Indicates the The output features of adaptive graph convolution, , represents a temporal graph convolutional network, Represents the softmax activation function operation, representing the physical topology connection of human body, representing the data-driven joint topology graph; finally output after m adaptive graph convolution layers ;
[0038] Subsequently input to the decoupling network, and the operation process of the decoupling network is as follows:
[0039]
[0040] wherein, representing the classification network, representing the feature connection operation, representing the tensor absolute value difference calculation, representing the feature of the odd frame branch in the decoupling network, representing the feature of the even frame branch in the decoupling network, representing the output behavior recognition result; the sequence composed of all behavior recognition results of the video is input into the risk evaluation network to obtain the risk detection result.
[0041] The application also provides a power grid operator safety risk detection device, comprising:
[0042] a video acquisition module for acquiring video data in different work scenes of the operator;
[0043] a posture estimation module for extracting 2D posture features of each frame and 3D posture features of each frame from the input video segment, and then fusing the 2D posture features and the 3D posture features to obtain fused posture features;
[0044] a time sequence behavior recognition module for dividing the obtained fused posture features into layers, wherein k is an integer greater than or equal to 0, each layer of the fused posture features is subjected to time sequence behavior recognition to obtain corresponding time sequence features for recognizing behaviors in the current video, then adjacent two divided time sequence features are fused to obtain longer time sequence features for behavior recognition, and the process is repeated times to obtain time sequence features of the whole video, the time sequence features are subjected to behavior recognition, a sequence composed of all behavior recognition results of the video is input into a risk evaluation network to obtain a risk detection result, and the risk evaluation network is composed of three fully connected networks in series.
[0045] Further, the video acquisition module is further used for:
[0046] acquiring operator videos to construct a data set , which is divided into a training set and a test set The video in the dataset is divided into video segments and the specific risk type is labeled, and the training set is represented as wherein the training set contains video segments, represents the video segment of the training set, and the type label of the corresponding video segment is wherein , represents the number of new video segments divided from each video, represents the type label of the video segment of the training set, represents the type label of the first new video segment of the 0th layer divided from the video segment, represents the type label of the first new video segment of the cth layer divided from the video segment, represents the type label of the new video segment of the cth layer divided from the video segment; the test set is divided in the same way, represented as , and the type label is wherein , the test set contains video segments, represents the video segment of the test set, represents the type label of the video segment of the test set, represents the type label of the first new video segment of the 0th layer divided from the video segment of the test set, represents the type label of the first new video segment of the cth layer divided from the video segment of the test set, represents the type label of the new video segment of the cth layer divided from the video segment of the test set.
[0047] Further, the extraction method of the 2D pose feature is:
[0048] Each frame in the input video is input into a 2D pose estimation network to output a 2D pose feature, and the working process of the 2D pose estimation network is to embed each frame in the input video into a block embedding layer to obtain a block embedding result , and then input into an encoder composed of multiple sequentially cascaded converters to obtain an encoding feature The operation of each transformer is expressed as follows
[0049]
[0050] wherein, represents the intermediate variable of the th transformer, represents the output of the th transformer, represents the layer normalization operation, represents the multi-head self-attention operation, represents a combined network comprising a first linear layer, a nonlinear activation function , Dropout and a second linear layer; and then the encoded features are input into the decoder to obtain 2D pose features , which are expressed as follows:
[0051]
[0052] wherein, represents a bilinear interpolation operation, represents a convolutional operation layer with a convolution kernel size of 3*3, linear transformation operation.
[0053] Further, the extraction manner of the 3D pose features is as follows:
[0054] The 2D pose features are taken as inputs of a 3D pose estimation network, and 3D pose features are output, the 3D pose estimation network being a sequentially connected feature extraction module, a space-time attention module and a channel attention module, the space-time attention module being formed by connecting a plurality of time encoders and a plurality of space encoders in series.
[0055] Further, the processing process of the feature extraction module is as follows:
[0056] Another block embedding layer is adopted to obtain a block embedding result , and then convolution, a nonlinear activation function and Dropout processing are performed to combine position encoding information to obtain features for 3D pose estimation, the calculation process being as follows:
[0057]
[0058] wherein, represents a batch normalization operation, represents a Dropout operation, represents a position embedding operation.
[0059] Further, the working process of the space-time attention module is as follows:
[0060] characteristics The operation process of the input spatio-temporal attention module and the time encoder is represented by the following formula:
[0061]
[0062] wherein, represents the output of the previous time encoder, represents the output of the current time encoder, The time attention module is composed of 5 layers of time encoders, and the output of the time attention module is ;
[0063] The operation process of the spatial encoder is represented by the following formula:
[0064]
[0065] wherein, represents the maximum pooling, represents a convolution operation layer with a convolution kernel size of 1*1, represents the output of the previous spatial encoder, represents the output of the current spatial encoder, The spatial attention module is composed of 5 layers of spatial encoders, and the output of the spatial attention module is , and the output of the spatio-temporal attention module is .
[0066] Further, the channel attention module adopts an SE module, the output of the spatio-temporal attention module is taken as the input of the SE module, and the SE module outputs a 3D pose feature.
[0067] Further, the 2D pose feature and the 3D pose feature are fused to obtain a fused pose feature, which comprises:
[0068] The 2D pose feature and the 3D pose feature are input into a pose feature fusion network to obtain a final fusion pose feature, the pose feature fusion network uses positions corresponding to each other in the 2D pose feature and the 3D pose feature as inputs of two branches, the two branches respectively complete local attention calculation through a third linear layer for processing to obtain queries, keys and values used in multi-head self-attention operation, after the local attention calculation, cross-attention calculation is performed on the combination of the queries calculated by the branch where the 3D pose feature is located and the keys and values calculated by the branch where the 2D pose feature is located to obtain global attention; after the global attention calculation is completed, the two branches respectively map the attention calculation results through a fourth linear layer and add the attention calculation results to the input features, and then the two branches are adjusted in feature dimension through a fifth linear layer, then feature connection is performed, and finally cross-attention feature information capable of representing the global and the local is obtained as the fusion pose feature.
[0069] Further, the time sequence behavior recognition module is further configured to:
[0070] Each layer of the fusion pose feature is subjected to time sequence behavior recognition through a time sequence decoupling adaptive graph convolution network to obtain a corresponding time sequence feature; the time sequence decoupling adaptive graph convolution network is composed of m sequentially cascaded adaptive graph convolution layers and a subsequent decoupling network; the operation process of the adaptive graph convolution layer is as follows
[0071]
[0072] wherein, is a matrix multiplication operator, denotes the output feature of the mth adaptive graph convolution, denotes the output feature of the mth adaptive graph convolution, , denotes a time sequence graph convolution network, denotes a softmax activation function operation, denotes a human physical topology connection, denotes a data-driven joint topology graph; the final output of the m adaptive graph convolution layers is . Subsequently, the output is input into the decoupling network, and the operation process of the decoupling network is as follows:
[0073]
[0074]
[0075] wherein, denotes a classification network, denotes a feature connection operation, representing a tensor absolute difference calculation, representing features of the odd frame branch in the decoupled network, representing features of the even frame branch in the decoupled network, representing an output behavior recognition result; a sequence composed of all behavior recognition results of the video segment is input into a risk evaluation network to obtain a risk detection result.
[0076] The present application has the following advantages:
[0077] (1) The 2D pose feature and the 3D pose feature of the target are extracted and fused, so as to avoid pose estimation errors caused by occlusion, similar actions and the like, improve the recognition accuracy, and obtain a longer time sequence feature by fusing adjacent two divided time sequence features to perform behavior recognition, and repeat the process to obtain behavior recognition results of different duration, so as to avoid that a too long video segment easily ignores some behaviors with a short duration, and a too short video segment makes the model unable to capture the time sequence feature of the entire behavior, thereby accurately identifying the behavior type, ensuring the safety risk detection of the power grid operation site, reducing the probability of accidents, and protecting the personal safety of the operating personnel.
[0078] (2) The fusion network of the 2D pose feature and the 3D pose feature is used to extract accurate human pose features, the time sequence decoupling adaptive graph convolution network is used to extract parallax feature enhancement information to improve the information richness, multi-time scale time sequence behavior recognition is realized, and finally the safety risk detection of the power grid operating personnel is realized.
[0079] (3) The present application designs a 2D pose feature and 3D pose feature fusion network module, which is connected with the 2D pose feature and the 3D pose feature aligned in space by using a cross-attention operation on the basis of using a self-attention network module, extracts more stable pose features in the presence of occlusion, deformation and other interference, avoids pose estimation errors caused by occlusion, similarity actions and the like, and improves the recognition accuracy.
[0080] (4) The present application designs a time sequence decoupling adaptive graph convolution network, which increases a decoupling network module of time sequence features on the basis of the adaptive graph convolution network, uses 2D pose and 3D fusion features, extracts parallax features between continuous frames and adds them to the original pose features, improves the description ability of the pose features to the detailed information, improves the capture ability of the model to the subtle action differences, and solves the problem that the representation ability of the existing fusion features to the detailed information is limited.
[0081] (5) The application designs a time sequence behavior recognition network structure based on posture estimation features, and can integrate posture feature sequences of different time scales to extract time multi-scale time sequence behavior features, represent the diversity and complexity of time sequence behaviors, use multi-time scale behavior recognition result sequences to mine hidden safety risks, ensure safety risk detection of power grid operation sites, reduce the probability of accidents, protect the personal safety of operation personnel, and improve the refinement and accuracy of safety risk detection of power grid operation. BRIEF DESCRIPTION OF DRAWINGS
[0082] Figure 1 A flowchart of a power grid operator safety risk detection method disclosed by the embodiment of the application;
[0083] FIG. 2(a) is a schematic diagram of the overall architecture of a 2D posture feature estimation network in a power grid operator safety risk detection method disclosed by the embodiment of the application, FIG. 2(b) is a schematic diagram of the converter structure of an encoder in the 2D posture feature estimation network, and FIG. 2(c) is a schematic diagram of the predictor structure of a decoder in the 2D posture feature estimation network;
[0084] FIG. 3(a) is a schematic diagram of the overall architecture of a 3D posture feature estimation network in a power grid operator safety risk detection method disclosed by the embodiment of the application, FIG. 3(b) is a schematic diagram of the structure of a feature extraction module in the 3D posture feature estimation network, FIG. 3(c) is a schematic diagram of the structure of a space-time attention module in the 3D posture feature estimation network, and FIG. 3(d) is a schematic diagram of the structure of a channel attention module in the 3D posture feature estimation network;
[0085] Figure 4 FIG. 4 is a schematic diagram of a 2D and 3D posture feature fusion network structure in a power grid operator safety risk detection method disclosed by the embodiment of the application;
[0086] FIG. 5(a) is a schematic diagram of the overall architecture of a time sequence behavior recognition network in a power grid operator safety risk detection method disclosed by the embodiment of the application, FIG. 5(b) is a schematic diagram of the structure of a time sequence decoupling adaptive graph convolution network in the time sequence behavior recognition network, and FIG. 5(c) is a schematic diagram of the adaptive graph convolution structure of the time sequence decoupling adaptive graph convolution network in the time sequence behavior recognition network. DETAILED DESCRIPTION
[0087] To make the objectives, technical solutions and advantages of the embodiments of the application clearer, the technical solutions in the embodiments of the application will be described below in a clear and complete manner with reference to the embodiments of the application. Obviously, the described embodiments are some, but not all, of the embodiments of the application. Based on the embodiments in the application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the application.
[0088] Example 1
[0089] Embodiment 1 of the present invention provides a method for detecting safety risks of power grid workers, which can be mainly divided into two parts: posture estimation and temporal behavior recognition. First, in the training phase of the model, a training data set and a test data set are constructed by collecting visible light videos, and then the training data set is used to train the posture estimation network and the temporal behavior recognition network model; in the testing phase, the test data set data is input into the previously trained network model according to the segmented video segments to verify the model effect; in the implementation phase, the real-time collected video data is input into the network model according to the video segments divided by fixed time steps to obtain the behavior recognition results to detect the current safety risks of power grid workers. Figure 1 As shown, the method includes the following steps:
[0090] S1. Video data collection of operators in different working scenes; the specific process is as follows:
[0091] Collect videos of power grid workers to build a dataset , including different safety risk behaviors in different work scenarios, divided into training sets according to 7:3 and test set , divide the videos in the dataset into video segments according to fixed time length and mark the specific risk type. The length ratio can be set according to the actual situation. Here, the video segmentation is performed in a way that the long-duration video length contains 3 short-duration videos. The training set is represented as ,in, Indicates that the training set contains Video clips, Represents the training set video clips, and the corresponding video clip types are marked as ,in , Indicates the number of new video segments each video is divided into, Represents the training set Type annotation of video clips, Indicates the The type label of the first new video segment of layer 0 divided by the video segments, Indicates the The type label of the first new video segment of the cth layer divided by the video segments, Indicates the The cth layer is divided into The type of new video clips is annotated; the test set is divided in the same way and is represented as , the type is marked as wherein , represents that the test set contains video clips, represents the first video clip of the test set, represents the first video clip of the test set, represents the type label of the first new video clip of the 0th layer divided from the first video clip of the test set, represents the type label of the first new video clip of the cth layer divided from the first video clip of the test set, represents the type label of the first new video clip of the cth layer divided from the first video clip of the test set, represents the type label of the first new video clip of the cth layer divided from the first video clip of the test set, represents the type label of the first new video clip of the cth layer divided from the first video clip of the test set, represents the type label of the first new video clip of the cth layer divided from the first video clip of the test set, represents the type label of the first new video clip of the cth layer divided from the first video clip of the test set, the number of video division layers of the test set is consistent with that of the training set, and the safety risk probability of each video is labeled according to the specification, wherein the safety risk is set as 0~1 from low to high, and the labeling standard is determined by manual judgment. S2, 2D pose features and 3D pose features of each frame of the input video clip are extracted, and the 2D pose features and the 3D pose features are fused to obtain fused pose features; the specific process is as follows:
[0092] This step mainly designs a 2D-3D spatial attention feature fusion network to fuse 2D human body pose and 3D human body pose features, extracts target 2D pose features and 3D pose depth features through feature mapping, and avoids pose estimation errors caused by occlusion, similarity actions and the like.
[0093] Fig. 2(a) is a schematic diagram of the overall architecture of the 2D pose feature estimation network, Fig. 2(b) is a schematic diagram of the converter structure of the encoder in the 2D pose feature estimation network, and Fig. 2(c) is a schematic diagram of the predictor structure of the decoder in the 2D pose feature estimation network, the 2D pose estimation network adopts a vision transformer (Vision Transformer) network, first embeds a block embedding layer for each frame of input video to obtain a block embedding result
[0094] , then inputs the block embedding result into an encoding layer composed of multiple transformers (Transformer layers) to obtain an encoding feature , and the present application adopts 5 layers of transformers to compose the encoding layer, wherein the operation of each layer is represented by the following formula:
[0095]
[0096] wherein, represents the first video clip of the test set, intermediate variable of the converter, output of the first converter layer, layer normalization layer, multi-head self-attention layer, representing a combined network module comprising a linear layer, a nonlinear activation function (ReLU), Dropout, and a linear layer. The encoded features input into the decoder to obtain 2D pose features , which are represented as follows:
[0097]
[0098] wherein, represents a bilinear interpolation operation, represents a convolution operation layer with a convolution kernel size of 3*3, represents a linear transformation operation, which is implemented through a linear layer, and finally obtains 2D pose features of each frame of the input video, which are represented as position information of human key points in each frame, and have a size of , represents the number of frames of the input video, represents the number of feature channels of , represents the number of joints in the features.
[0099] Figure 3(a) is a schematic diagram of the overall architecture of the 3D pose feature estimation network, Figure 3(b) is a schematic diagram of the structure of the feature extraction module in the 3D pose feature estimation network, Figure 3(c) is a schematic diagram of the structure of the spatial-temporal attention module in the 3D pose feature estimation network, and Figure 3(d) is a schematic diagram of the structure of the channel attention module in the 3D pose feature estimation network. The heat map in the 2D pose feature network is taken as the input of the 3D pose estimation network, another block embedding layer is used to obtain block embedding results , and then through convolution, nonlinear activation function (ReLU), Dropout, and other processing, combined with position encoding information, features for 3D pose estimation are obtained, and the calculation process is as follows:
[0100]
[0101] wherein, represents a batch normalization operation, represents a Dropout operation, which is a commonly used regularization technique in deep learning, and the specific operation is to randomly set the output of part of the neurons to 0 during the training process, represents a position embedding operation, which uses a relative position encoding method to represent the position relationship between different joints, and the specific calculation method is as follows:
[0102]
[0103] wherein, represents the current calculated embedding block, , represents the total number of blocks into which the block embedding operation divides the input, and the sum of the relative position deviations between the current block and other blocks with existing node information, represents the first embedding block. According to the above calculation principle, the calculation method of
[0104]
[0105] is as follows input into the space-time attention module and the channel attention module. The space-time attention module is composed of a plurality of time encoder modules and a plurality of space encoder modules connected in series. The time encoder is composed of layer normalization, multi-head attention operation, linear layer, nonlinear activation function, and Dropout, and its operation process can be represented by the following formula:
[0106]
[0107] wherein, represents the output of the previous time encoder, represents the output of the current time encoder, The present application adopts 5 layers of time encoders to form the time attention module, so the output of the time attention module is .
[0108] The space encoder is composed of layer normalization, multi-head attention operation, 1*1 convolution layer, nonlinear activation function, Dropout, and maximum value pooling layer, and its operation process can be represented by the following formula:
[0109]
[0110] wherein, represents the maximum value pooling, represents the convolution operation layer with a convolution kernel size of 1*1, represents the output of the previous space encoder, represents the output of the current space encoder, The present application adopts 5 layers of space encoders to form the space attention module, so the output of the space attention module is .
[0111] The channel attention module can reduce the feature weight of the interference object and enhance the feature weight of the human body, thereby improving the accuracy of 3D pose estimation. This paper adopts the SE module (Squeeze-and-Excitation Attention Networks) as the channel attention module. Its network structure is shown in Figure 3 (d). It consists of an adaptive pooling layer, a linear layer, a nonlinear activation function (ReLU), and a Sigmoid activation function. The specific operation process is as follows:
[0112]
[0113] in, The obtained 3D posture features are represented as the 3D position information of the key points of the human body in each frame, and its size is , Indicates the number of input video frames, express The number of feature channels, Indicates the number of joints in the feature. express function, represents adaptive pooling, In the channel attention module, the input feature tensor and the output feature tensor have the same channel feature size, which can maintain the spatiotemporal information integrity of the pose features.
[0114] After obtaining the 2D and 3D posture features of the input video, the two are input into the 2D and 3D posture feature fusion network to obtain the final posture estimation features as the input of the subsequent temporal behavior recognition network. The 2D and 3D posture feature fusion network structure is as follows: Figure 4 As shown in Figure 2, the network realizes feature mapping between 2D pose and 3D pose, realizes global and local feature modeling of pose information, enhances the weight of effective pose features, and improves the model's ability to capture features of complex, subtle and occluded poses. The fusion network uses 2D pose features and 3D pose features The corresponding positions in the 3D pose feature branch are used as input, and the query, key, and value used in the multi-head self-attention operation are obtained through linear layer processing to complete the respective attention calculations. At the same time, the two use the cross-attention operation to calculate the global attention using the query calculated by the 3D pose feature branch and the key and value calculated by the 2D pose feature branch. The cross-attention operation is calculated as follows:
[0115]
[0116] in, Represents the global attention calculation result, represents the query calculated from the 3D pose features, represents the key calculated by the 2D pose feature, represents the value calculated by the 2D pose feature.
[0117] After the calculation of global attention and local attention is completed, the attention calculation result is mapped through a linear layer, added to the input feature, and then adjusted in feature dimension through a linear layer mapping, connected, and finally the cross-attention feature information capable of representing the global and local is obtained , the size of which is , represents the number of input video frames, represents the number of feature channels, represents the number of joints in the feature.
[0118] S3, the obtained fusion pose feature is divided into layers, is an integer greater than or equal to 0, each layer of the fusion pose feature respectively performs time sequence behavior recognition to obtain corresponding time sequence features for the recognition of behaviors within the current video, then two adjacent divided time sequence features are fused to obtain longer time sequence features for behavior recognition, and the process is repeated times to obtain the time sequence features of the entire video, and the behavior recognition is performed. The sequence composed of all behavior recognition results of the video is input into a risk evaluation network to obtain a risk detection result. The risk evaluation network is composed of three layers of fully connected networks in series; the specific process is as follows:
[0119] After the pose feature of the video is extracted, it is input into a time sequence behavior recognition network to obtain a behavior recognition result, and the sequence of behavior recognition results is input into a risk evaluation network to obtain a risk detection result of the current behavior. The specific implementation process is shown in FIG. 5.
[0120] First, the cross-attention feature information is divided into layers, is the number of new video segments divided from each video in the data set, and then the time sequence features obtained by a time sequence decoupling adaptive graph convolution network are used for the recognition of behaviors within the current video. In addition, two divided features are fused to obtain longer time sequence features for behavior recognition, and the process is repeated times to obtain the time sequence behavior features of the entire video, and the behavior recognition result is obtained through classification. All behavior recognition results of the video are input into a risk evaluation network to obtain a risk detection result.
[0121] As a further improved scheme, Fig. 5 (a) is a schematic diagram of the overall architecture of the time sequence behavior recognition network, Fig. 5 (b) is a schematic diagram of the time sequence decoupling adaptive graph convolution network structure in the time sequence behavior recognition network, and Fig. 5 (c) is a schematic diagram of the adaptive graph convolution structure of the time sequence decoupling adaptive graph convolution network in the time sequence behavior recognition network. It is composed of a group of adaptive graph convolution layers and subsequent decoupling networks. The present application uses 10 adaptive graph convolution layers in the time sequence decoupling adaptive graph convolution network. In the adaptive graph convolution, 3 different 1*1 convolution kernel operations are first performed to obtain feature tensors of different sizes, two of which are multiplied and subjected to softmax processing to obtain a data-related joint topology graph, which is then connected with the human physical topology and the data-driven topology graph are added to obtain a topology feature graph, wherein the human physical topology connection represents the physical connection strength between the joints of the human body, and is a joint topology graph given by a person, and the data-driven joint topology graph is optimized during the training process to dynamically represent the connection strength of the joints, that is, the joint topology graph B is the result of adjusting the weight of the human physical topology connection A by the network, but the joint topology graph B may be connected with joints that should not be connected due to deformation, occlusion, etc. during the optimization process, so the human physical topology connection A and the joint topology graph B are considered together during calculation to improve accuracy. Multiply the topology feature graph by another convolution result to obtain a time sequence feature, and add the input feature to reduce the spatiotemporal information decay, and then input it to the time sequence graph convolution (TGC) module to obtain the output time sequence action feature. The time sequence graph convolution is a commonly used network module, and will not be introduced here. Therefore, the operation process is as follows:
[0122]
[0123] wherein, is a matrix multiplication operator, represents the output feature of the current adaptive graph convolution, represents the output feature of the previous adaptive graph convolution, , represents the time sequence graph convolution network, represents the softmax activation function operation, represents the human physical topology connection, represents the data-driven joint topology graph, and finally outputs .
[0124] Subsequently, Input to the decoupling network, the network is composed of two branches, one through the maximum pooling of the original features are time dimension sampling, the other branch will be extracted according to the features of the features, the corresponding subtraction to obtain the disparity feature enhancement information rich degree, and then connected to the sampling features can describe the behavior details of the time sequence characteristics. After linear layer mapping for behavior recognition, the operation process is as follows:
[0125]
[0126] wherein, indicates a classification network, the present application adopts a softmax layer for classification, indicates a feature connection operation, indicates the absolute value difference calculation of tensor, indicates the output behavior category, a plurality of behavior recognition results are input into the risk assessment network to perform risk detection, and the final risk probability is output.
[0127] The risk assessment network is mainly composed of three fully connected networks, the first two layers perform linear mapping and channel adjustment of the features, and finally output the safety risk level assessment result of the current video, which displays the safety risk existing in the current operation in the form of probability, which can be expressed as follows:
[0128]
[0129] wherein, indicates the input behavior recognition result sequence, which is composed of a plurality of connections, indicates function, the output is . The greater, the higher the safety risk existing in the current video.
[0130] Through the above technical scheme, the present application adopts the method of improving 2D pose estimation to 3D pose estimation, designs a 2D-3D space attention feature fusion network to fuse 2D human pose and 3D human pose features, extracts the 2D pose features and 3D pose depth features of the target through feature mapping, avoids pose estimation errors caused by occlusion and similar actions, etc.; adopts a spatio-temporal convolution network to associate the time sequence relationship of key points, extracts time sequence behavior features for time sequence behavior recognition. This method can solve the action confusion problem existing in the process of human pose estimation, and combine the time sequence pose features for behavior recognition of different duration, to ensure the safety risk detection of power grid operation site, reduce the probability of accidents, and protect the personal safety of operating personnel.
[0131] Embodiment 2
[0132] Based on Example 1, Example 2 of the present invention further provides a power grid operator safety risk detection device, including:
[0133] Video acquisition module, used for collecting video data of operators in different working scenes;
[0134] The posture estimation module is used to extract the 2D posture features and 3D posture features of each frame of the input video clip, and then fuse the 2D posture features and 3D posture features to obtain the fused posture features;
[0135] The temporal behavior recognition module is used to divide the obtained fusion posture features into layer, is an integer greater than or equal to 0. Each layer of the fused posture feature is used to perform temporal behavior recognition to obtain the corresponding temporal features for behavior recognition in the current video. Then, the two adjacent divided temporal features are fused to obtain a longer temporal feature for behavior recognition. This process is repeated. After that, the temporal features of the entire video are obtained, and behavior recognition is performed. The sequence composed of all behavior recognition results of the video is input into the risk assessment network to obtain the risk detection result. The risk assessment network is composed of three layers of fully connected networks in series.
[0136] Specifically, the video acquisition module is further used to:
[0137] Collect videos of workers to build a dataset , divided into training set and test set , the videos in the dataset are divided into video clips and labeled with specific risk types. The training set is represented as ,in, Indicates that the training set contains Video clips, Represents the training set video clips, and the corresponding video clip types are marked as ,in , Indicates the number of new video segments each video is divided into, Represents the training set Type annotation of video clips, Indicates the The type label of the first new video segment of layer 0 divided by the video segments, Indicates the The type label of the first new video segment of the cth layer divided by the video segments, Indicates the The cth layer is divided into The type of new video clips is annotated; the test set is divided in the same way and is represented as , the type is marked as ,in , Indicates that the test set contains Video clips, The test set Video clips, The test set Type annotation of video clips, The test set The type label of the first new video segment of layer 0 divided by the video segments, The test set The type label of the first new video segment of the cth layer divided by the video segments, The test set The cth layer is divided into Type label of each new video clip.
[0138] Specifically, the 2D posture feature is extracted as follows:
[0139] Each frame in the input video is input into the 2D pose estimation network, and the 2D pose features are output. The working process of the 2D pose estimation network is to embed each frame in the input video into the block embedding layer to obtain the block embedding result. , then Input multiple sequentially cascaded converters to form an encoder to obtain the encoded features , the operation of each converter is expressed as follows
[0140]
[0141] in, Indicates the The intermediate variable of the converter, Indicates the The output of the converter, Representation layer normalization operation, represents a multi-head self-attention operation, Indicates that it contains the first linear layer and nonlinear activation function , Dropout and the second linear layer; then encode the features Input decoder to get 2D posture features , expressed as the following formula:
[0142]
[0143] in, represents the bilinear interpolation operation, Indicates a convolution operation layer with a convolution kernel size of 3*3. Linear transformation operation.
[0144] Specifically, the 3D posture feature is extracted as follows:
[0145] The 2D pose features are used as the input of the 3D pose estimation network, and the 3D pose features are output. The 3D pose estimation network is composed of a sequentially connected feature extraction module, a spatiotemporal attention module, and a channel attention module. The spatiotemporal attention module is composed of multiple time encoders and multiple spatial encoders in series.
[0146] More specifically, the processing process of the feature extraction module is as follows:
[0147] Use another block embedding layer to get the block embedding result , and then through convolution, nonlinear activation function and Dropout processing, combined with position encoding information to obtain features for 3D pose estimation , the calculation process is as follows:
[0148]
[0149] in, represents the batch normalization operation, represents the Dropout operation, Represents a positional embedding operation.
[0150] More specifically, the working process of the spatiotemporal attention module is as follows:
[0151] The features Enter the spatiotemporal attention module, and the operation process of the temporal encoder is expressed as follows:
[0152]
[0153] in, represents the output of the previous time encoder, Represents the output of the current time encoder, , a 5-layer temporal encoder is used to form a temporal attention module, so the output of the temporal attention module is ;
[0154] The operation process of the spatial encoder is expressed as follows:
[0155]
[0156] in, represents maximum pooling, Represents a convolution operation layer with a convolution kernel size of 1*1. represents the output of the previous spatial encoder, represents the output of the current spatial encoder, , a 5-layer spatial encoder is used to form a spatial attention module, and the output of the spatial attention module is , so the output of the spatiotemporal attention module is .
[0157] More specifically, the channel attention module adopts the SE module, the output result of the spatiotemporal attention module is used as the input of the SE module, and the SE module outputs 3D posture features.
[0158] More specifically, fusing the 2D posture feature and the 3D posture feature to obtain the fused posture feature includes:
[0159] The 2D posture features and 3D posture features are input into the posture feature fusion network to obtain the final fused posture features. The posture feature fusion network uses the corresponding positions in the 2D posture features and 3D posture features as the input of the two branches. The two branches are processed by the third linear layer to obtain the query, key and value used in the multi-head self-attention operation, and complete their respective local attention calculations. After the local attention calculation is completed, the query calculated by the branch where the 3D posture features are located and the key and value combination calculated by the branch where the 2D posture features are located are used to perform cross-attention calculation to obtain global attention; after completing the global attention calculation, the two branches respectively map the attention calculation results through the fourth linear layer and add them to the input features. After the feature dimensions of the two branches are adjusted through the fifth linear layer, the features are connected, and finally the cross-attention feature information that can represent the global and local is obtained. As fused posture features.
[0160] More specifically, the temporal behavior recognition module is further used to:
[0161] Each layer of the fused posture features is respectively subjected to temporal behavior recognition by a temporal decoupling adaptive graph convolutional network to obtain the corresponding temporal features; the temporal decoupling adaptive graph convolutional network is composed of m sequentially cascaded adaptive graph convolutional layers and a subsequent decoupling network; the operation process of the adaptive graph convolutional layer is as follows
[0162]
[0163] in, is the matrix multiplication operator symbol, Indicates the The output features of adaptive graph convolution, Indicates the The output features of adaptive graph convolution, , denotes a time series graph convolution network, denotes a softmax activation function operation, denotes a human physical topology connection, denotes a data-driven joint topology graph; finally output after m adaptive graph convolution layers
[0164] Subsequently is input into a decoupling network, and the decoupling network operation process is as follows:
[0165]
[0166] wherein, denotes a classification network, denotes a feature connection operation, denotes a tensor absolute value difference calculation, denotes a feature of an odd frame branch in the decoupling network, denotes a feature of an even frame branch in the decoupling network, denotes an output behavior recognition result; a sequence composed of all behavior recognition results of the video segment is input into a risk evaluation network to obtain a risk detection result.
[0167] The above embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, it should be understood by those skilled in the art that the technical solutions recorded in the foregoing embodiments can still be modified, or some technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for detecting safety risks of power grid operators, characterized in that: include: S1. Video data collection of operators in different working scenes; S2, extracting the 2D posture features and 3D posture features of each frame of the input video clip, and then fusing the 2D posture features and 3D posture features to obtain a fused posture feature; S3, divide the obtained fusion posture features into layer, is an integer greater than or equal to 0, Indicates the number of new video segments divided from each video. Each layer of the fused posture feature is used for temporal behavior recognition to obtain the corresponding temporal feature for behavior recognition in the current video. Then, the two adjacent temporal features are fused to obtain a longer temporal feature for behavior recognition. The process of fusing the two adjacent temporal features to obtain a longer temporal feature for behavior recognition is repeated. After that, the temporal features of the entire video are obtained, and behavior recognition is performed. The sequence composed of all behavior recognition results of the entire video is input into the risk assessment network to obtain the risk detection result. The risk assessment network is composed of three layers of fully connected networks in series.
2. A method for detecting safety risks of power grid operators according to claim 1, characterized in that: S1 includes: Collect videos of workers to build a dataset , divided into training set and test set , the videos in the dataset are divided into video clips and labeled with specific risk types. The training set is represented as ,in, Indicates that the training set contains Video clips, Represents the training set video clips, and the corresponding video clip types are marked as ,in , Represents the training set Type annotation of video clips, Indicates the The type label of the first new video segment of layer 0 divided by the video segments, Indicates the The type label of the first new video segment of the cth layer divided by the video segments, Indicates the The cth layer is divided into The type of new video clips is annotated; the test set is divided in the same way and is represented as , the type is marked as ,in , Indicates that the test set contains Video clips, The test set Video clips, The test set Type annotation of video clips, The test set The type label of the first new video segment of layer 0 divided by the video segments, The test set The type label of the first new video segment of the cth layer divided by the video segments, The test set The cth layer is divided into Type label of each new video clip.
3. A method for detecting safety risks of power grid operators according to claim 1, characterized in that: The 2D posture feature is extracted as follows: Each frame in the input video is input into the 2D pose estimation network, and the 2D pose features are output. The working process of the 2D pose estimation network is to embed each frame in the input video into the block embedding layer to obtain the block embedding result. , then Input multiple sequentially cascaded converters to form an encoder to obtain the encoded features , the operation of each converter is expressed as follows in, Indicates the The intermediate variable of the converter, Indicates the The output of the converter, Representation layer normalization operation, represents a multi-head self-attention operation, Indicates that it contains the first linear layer and nonlinear activation function , Dropout and the second linear layer; then encode the features Input decoder to get 2D posture features , expressed as the following formula: in, represents the bilinear interpolation operation, Indicates a convolution operation layer with a convolution kernel size of 3*3. Linear transformation operation.
4. A method for detecting safety risks of power grid operators according to claim 1, characterized in that: The 3D posture feature is extracted as follows: The 2D pose features are used as the input of the 3D pose estimation network, and the 3D pose features are output. The 3D pose estimation network is composed of a sequentially connected feature extraction module, a spatiotemporal attention module, and a channel attention module. The spatiotemporal attention module is composed of multiple time encoders and multiple spatial encoders in series.
5. A method for detecting safety risks of power grid operators according to claim 4, characterized in that: The processing process of the feature extraction module is as follows: Use another block embedding layer to get the block embedding result , and then through convolution, nonlinear activation function and Dropout processing, combined with position encoding information to obtain features for 3D pose estimation , the calculation process is as follows: in, represents the batch normalization operation, represents the Dropout operation, Represents a positional embedding operation.
6. A method for detecting safety risks of power grid operators according to claim 5, characterized in that: The working process of the spatiotemporal attention module is as follows: The features Enter the spatiotemporal attention module, and the operation process of the temporal encoder is expressed as follows: in, represents the output of the previous time encoder, Represents the output of the current time encoder, , a 5-layer temporal encoder is used to form a temporal attention module, so the output of the temporal attention module is ; The operation process of the spatial encoder is expressed as follows: in, represents maximum pooling, Represents a convolution operation layer with a convolution kernel size of 1*1. represents the output of the previous spatial encoder, represents the output of the current spatial encoder, , a 5-layer spatial encoder is used to form a spatial attention module, and the output of the spatial attention module is , so the output of the spatiotemporal attention module is .
7. A method for detecting safety risks of power grid operators according to claim 6, characterized in that: The channel attention module adopts the SE module, the output result of the spatiotemporal attention module is used as the input of the SE module, and the SE module outputs 3D posture features.
8. A method for detecting safety risks of power grid operators according to claim 7, characterized in that: The fusing of the 2D posture feature and the 3D posture feature to obtain the fused posture feature includes: The 2D posture features and 3D posture features are input into the posture feature fusion network to obtain the final fused posture features. The posture feature fusion network uses the corresponding positions in the 2D posture features and 3D posture features as the input of the two branches. The two branches are processed by the third linear layer to obtain the query, key and value used in the multi-head self-attention operation, and complete their respective local attention calculations. After the local attention calculation is completed, the query calculated by the branch where the 3D posture features are located and the key and value combination calculated by the branch where the 2D posture features are located are used to perform cross-attention calculation to obtain global attention; after completing the global attention calculation, the two branches respectively map the attention calculation results through the fourth linear layer and add them to the input features. After the feature dimensions of the two branches are adjusted through the fifth linear layer, the features are connected, and finally the cross-attention feature information that can represent the global and local is obtained. As fused posture features.
9. A method for detecting safety risks of power grid workers according to claim 8, characterized in that S3 include: Each layer of the fused posture features is respectively subjected to temporal behavior recognition by a temporal decoupling adaptive graph convolutional network to obtain the corresponding temporal features; the temporal decoupling adaptive graph convolutional network is composed of m sequentially cascaded adaptive graph convolutional layers and a subsequent decoupling network; the operation process of the adaptive graph convolutional layer is as follows in, is the matrix multiplication operator symbol, Indicates the The output features of adaptive graph convolution, Indicates the The output features of adaptive graph convolution, , represents a temporal graph convolutional network, Represents the softmax activation function operation, Represents the physical topological connection of the human body, Represents the data-driven joint topology; after m adaptive graph convolution layers, the final output ; Then it will Input to the decoupling network, the decoupling network operation process is as follows: in, represents the classification network, represents the feature connection operation, Indicates the calculation of the absolute value difference of the tensor, represents the characteristics of odd-numbered frame branches in the decoupled network, represents the characteristics of the even-numbered frame branches in the decoupled network, Represents the output behavior recognition result; the sequence composed of all behavior recognition results of this video is input into the risk assessment network to obtain the risk detection result.
10. A device for detecting safety risks of power grid workers, characterized in that: include: Video acquisition module, used for collecting video data of operators in different working scenes; The posture estimation module is used to extract the 2D posture features and 3D posture features of each frame of the input video clip, and then fuse the 2D posture features and 3D posture features to obtain the fused posture features; The temporal behavior recognition module is used to divide the obtained fusion posture features into layer, is an integer greater than or equal to 0, Indicates the number of new video segments divided from each video. Each layer of the fused posture feature is used for temporal behavior recognition to obtain the corresponding temporal feature for behavior recognition in the current video. Then, the two adjacent temporal features are fused to obtain a longer temporal feature for behavior recognition. The process of fusing the two adjacent temporal features to obtain a longer temporal feature for behavior recognition is repeated. After that, the temporal features of the entire video are obtained, and behavior recognition is performed. The sequence composed of all behavior recognition results of the entire video is input into the risk assessment network to obtain the risk detection result. The risk assessment network is composed of three layers of fully connected networks in series.
Citation Information
Patent Citations
Video three-dimensional human body posture estimation method and system based on multistage supervision graph convolution
CN114694261A
Human body abnormal behavior recognition method under transformer substation video monitoring based on posture estimation
CN116912930A
Multi-view fusion three-dimensional point cloud real-time semantic segmentation method and device and vehicle
CN119048764A
Human skeleton action recognition method for electric power field operation
CN119152571A