A behavioral trajectory prediction system and method for distribution station operation and maintenance personnel
Through the behavioral trajectory prediction system of multimodal fusion characteristics and attention mechanism, the problem of behavior identification and trajectory prediction of operation and maintenance personnel in distribution station building is solved, and the accurate identification and prediction of operation and maintenance personnel behavior is achieved, safety hazards are reduced, and the safe and stable operation of distribution station building is ensured.
Patent Information
- Application Number
- CN202311351969.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-18
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2043-10-18
AI Technical Summary
The prior art is difficult to accurately identify and predict the behavioral trajectory of power station operation and maintenance personnel, resulting in frequent safety hazards.
A behavioral trajectory prediction system for power station operation and maintenance personnel is adopted, including a data processing module, a behavior recognition module and a trajectory prediction module. The multimodal fusion feature, attention mechanism and graph convolution neural network are used to identify and predict human behaviors and tracks, and combined with the long and short-term memory network of the visual attention network for behavior prediction.
It realizes accurate identification and trajectory prediction of the behavior of operation and maintenance personnel, reduces the probability of safety hazards, and ensures the safe and stable operation and efficient operation and maintenance of distribution stations.
Smart Images

Figure CN117197723B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of distribution network operation, and in particular to a system and method for predicting the behavior trajectory of distribution station operation and maintenance personnel. Background Art
[0002] As the core link in power distribution, the safe and reliable operation of distribution stations directly impacts electricity consumption, the safety of life and property, and the well-being of consumers. In recent years, the rapid development of distributed resources such as photovoltaics, energy storage, and new energy electric vehicles, coupled with an increase in extreme weather events like storms, has led to frequent safety issues related to personnel, equipment, and the environment in distribution station operations and maintenance, requiring urgent resolution.
[0003] Safety management in distribution substations has always been a key issue. Preventing safety hazards to equipment, the environment, and O&M personnel is crucial for ensuring safe and stable operation. Key safety hazards that impact the normal operation of distribution substations include equipment corrosion, cracks, fireworks, accumulated water, foreign object intrusion, and improper attire and illegal operation by O&M personnel. Human behavior recognition and trajectory prediction can help identify safety hazards faced by O&M personnel, but these are challenging to do, as they are difficult to accurately identify human behavior and predict its trajectory. Summary of the Invention
[0004] The purpose of the present invention is to provide a system and method for predicting the behavior trajectory of distribution station operation and maintenance personnel, so as to solve the technical problem that the existing technology is difficult to accurately perform behavior recognition and trajectory prediction on operation and maintenance personnel.
[0005] The purpose of the present invention can be achieved through the following technical solutions:
[0006] Solution 1: A behavioral trajectory prediction system for substation operation and maintenance personnel, including:
[0007] Communication-connected data processing module, behavior recognition module, and trajectory prediction module;
[0008] The data processing module obtains video data of different modes of the power distribution station, performs feature extraction, modality alignment and feature fusion on the video data to obtain multimodal fusion features, where the modality represents the type of the video data;
[0009] The behavior recognition module extracts the human behavior semantic information from the multimodal fusion features, and performs human behavior recognition on the operation and maintenance personnel of the power distribution station based on the human behavior semantic information to obtain a behavior recognition result; the behavior recognition module is a graph convolutional neural network based on the attention mechanism;
[0010] The trajectory prediction module predicts the behavior trajectory of the operation and maintenance personnel according to the behavior recognition result to obtain the behavior trajectory prediction result of the operation and maintenance personnel; the trajectory prediction module is a long short-term memory network based on a visual attention network.
[0011] Optionally, the data processing module includes:
[0012] a first data processing unit, a second data processing unit, and a feature fusion unit;
[0013] Wherein, the first data processing unit extracts features from the audio data and the two-dimensional image in the video data to obtain a first video feature;
[0014] The second data processing unit performs position encoding, modality labeling, and feature dimensionality reduction on the point cloud data and the three-dimensional image in the video data to obtain a second video feature;
[0015] The feature fusion unit performs feature fusion on the first video feature and the second video feature to obtain a multimodal fusion feature.
[0016] Optionally, the first data processing unit includes:
[0017] Communication-connected policy neural networks and improved lightweight object detection networks;
[0018] Wherein, the strategy neural network obtains the audio data and two-dimensional image data in the video data and performs feature extraction;
[0019] The lightweight target detection network performs target detection, assigns adaptive attention weights to detected operation and maintenance personnel, and performs feature enhancement to obtain a first video feature; the lightweight target detection network is improved based on YOLOv5.
[0020] Optionally, the trajectory prediction module includes:
[0021] The communication connection of the attention estimation unit, encoder LSTM, interaction network and decoder LSTM, LSTM stands for long short-term memory network;
[0022] The attention estimation unit is used to estimate the estimated attention weight of each operation and maintenance personnel and perform field of view constraints;
[0023] The encoder LSTM processes the position sequence of the operation and maintenance personnel over time to obtain corresponding motion pattern features;
[0024] The interactive network estimates the social interaction characteristics of the operation and maintenance personnel based on the motion pattern characteristics and the social background characteristics; the social background characteristics are obtained based on the relative displacement of all the operation and maintenance personnel;
[0025] The decoder LSTM predicts the behavior trajectory of the operation and maintenance personnel based on the social interaction features.
[0026] Optionally, the attention estimation unit includes:
[0027] attentional networks and visual field filters that communicate connections;
[0028] The attention network assigns attention weights to the operator based on the operator's position and current speed relative to a specific person.
[0029] The field of view filter adjusts the attention weight according to the field of view of the real world to obtain a corresponding estimated attention weight.
[0030] Solution 2: A method for predicting the behavior trajectory of distribution station operation and maintenance personnel, including:
[0031] Obtain video data of different modalities of the power distribution station, and use a preset data processing module to perform feature extraction, modality alignment, and feature fusion on the video data to obtain multimodal fusion features, where the modality represents the type of the video data;
[0032] A preset behavior recognition module is used to extract human behavior semantic information from the multimodal fusion features, and human behavior recognition is performed on the operation and maintenance personnel of the power distribution station based on the human behavior semantic information to obtain a behavior recognition result; the behavior recognition module is a graph convolutional neural network based on the attention mechanism;
[0033] A preset trajectory prediction module is used to predict the behavior trajectory of the operation and maintenance personnel according to the behavior recognition result to obtain the behavior trajectory prediction result of the operation and maintenance personnel; the trajectory prediction module is a long short-term memory network based on a visual attention network.
[0034] Optionally, the data processing module includes a first data processing unit, a second data processing unit, and a feature fusion unit, and the using of the preset data processing module to perform feature extraction, modality alignment, and feature fusion on the video data to obtain multimodal fusion features includes:
[0035] Using the first data processing unit to extract features from the audio data and the two-dimensional image in the video data to obtain a first video feature;
[0036] Using the second data processing unit to perform position encoding, modality labeling, and feature dimensionality reduction on the point cloud data and the three-dimensional image in the video data to obtain a second video feature;
[0037] The feature fusion unit is used to perform feature fusion on the first video feature and the second video feature to obtain a multimodal fusion feature.
[0038] Optionally, the first data processing unit includes a strategy neural network and an improved lightweight object detection network that are communicatively connected, and the first data processing unit is used to extract features from the audio data and the two-dimensional image in the video data to obtain the first video feature, including:
[0039] Acquire audio data and a two-dimensional image from the video data, and extract features from the audio data and the two-dimensional image using the strategy neural network;
[0040] The lightweight target detection network is used to perform target detection, assign adaptive attention weights to the detected operation and maintenance personnel, and perform feature enhancement to obtain the first video feature; the lightweight target detection network is improved based on YOLOv5.
[0041] Solution 3: An electronic device comprising: a processor and a memory;
[0042] The memory stores a computer program, and the processor implements the steps of solution one when executing the computer program.
[0043] Solution 4: A computer-readable storage medium having a computer program stored thereon, which implements the steps of Solution 1 when executed by a processor.
[0044] The present invention provides a system and method for predicting the behavior trajectory of distribution station operation and maintenance personnel, wherein the system includes: a data processing module, a behavior recognition module and a trajectory prediction module that are communicatively connected; wherein the data processing module obtains video data of different modes of the distribution station, performs feature extraction, modality alignment and feature fusion on the video data to obtain multimodal fusion features, and the modality represents the type of the video data; the behavior recognition module extracts human behavior semantic information from the multimodal fusion features, and performs human behavior recognition on the operation and maintenance personnel of the distribution station according to the human behavior semantic information to obtain a behavior recognition result; the behavior recognition module is a graph convolutional neural network based on an attention mechanism; the trajectory prediction module predicts the behavior trajectory of the operation and maintenance personnel according to the behavior recognition result to obtain a behavior trajectory prediction result of the operation and maintenance personnel; and the trajectory prediction module is a long short-term memory network based on a visual attention network.
[0045] Based on the above technical solution, the beneficial effects brought about by the present invention are:
[0046] The data processing module is used to obtain video data of different modes of the distribution station, and feature extraction, modality alignment and feature fusion are performed to obtain multimodal fusion features. The video data of different modes can be fused according to the differences between heterogeneous and homogeneous modes in the specific scene of the distribution station, so as to understand and analyze the behavioral semantic information of the operation and maintenance personnel through video data; the behavior recognition module is used to extract the semantic information of human behavior in the multimodal fusion features, and at the same time, the attention mechanism is introduced to constrain the field of vision of the operation and maintenance personnel. The actions and behaviors of the people in the video are identified through semantic understanding, and then it is judged whether the people have accidents, dangerous behaviors and other hidden dangers; finally, the trajectory prediction module is used to predict the behavioral trajectory of the operation and maintenance personnel. It can automatically and efficiently predict the action trajectory of key operation and maintenance personnel, prevent and reduce the probability of safety hazards in advance, realize efficient operation and maintenance and safety management of the distribution station, and ensure the safe and stable operation of the power system.
[0047] The present invention can effectively identify the behavior of distribution station operation and maintenance personnel and predict their movement trajectories. By establishing a distribution station status detection and personnel trajectory prediction system that takes personnel, equipment and environment into consideration, the probability of safety hazards occurring to distribution station operation and maintenance personnel is reduced, which is of great significance for achieving efficient operation and maintenance of distribution stations and unmanned operation. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 A schematic structural diagram of an embodiment of the system of the present invention;
[0049] Figure 2 This is a structural diagram of a data processing module in an embodiment of the system of the present invention;
[0050] Figure 3 Schematic diagram of the structure of a lightweight target detection network in an embodiment of the system of the present invention;
[0051] Figure 4 Schematic diagram of an adaptive attention module of a lightweight object detection network in an embodiment of the system of the present invention;
[0052] Figure 5 Schematic diagram of the process of increasing the receptive field and the number of parameters for target detection tasks in an embodiment of the system of the present invention;
[0053] Figure 6 This is a schematic diagram of the structure of the trajectory prediction module in the system embodiment of the present invention;
[0054] Figure 7 This is a schematic diagram of the structure of the encoder LSTM in the system embodiment of the present invention;
[0055] Figure 8 This is a schematic diagram of the structure of the decoder LSTM in the system embodiment of the present invention;
[0056] Figure 9 Schematic diagram of the structure of the attention estimation unit in the system embodiment of the present invention;
[0057] Figure 10 Schematic diagram of a process of an embodiment of the method of the present invention;
[0058] Figure 11 Schematic diagram of the overall framework of the method embodiment of the present invention. DETAILED DESCRIPTION
[0059] The embodiments of the present invention provide a system and method for predicting the behavior trajectory of power distribution station operation and maintenance personnel, so as to solve the technical problem that the existing technology is difficult to accurately perform behavior recognition and trajectory prediction on operation and maintenance personnel.
[0060] To facilitate understanding of the present invention, the present invention will be described more fully below with reference to the accompanying drawings. Preferred embodiments of the present invention are shown in the drawings. However, the present invention may be embodied in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and comprehensive disclosure of the present invention.
[0061] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this invention pertains. The terms used in this specification of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0062] In today's era of big data and the construction of "smart cities," effectively analyzing and understanding distribution station video data is crucial for innovative development and comprehensive safety management of smart power plants. Human behavior recognition and trajectory prediction are among the most representative video understanding tasks.
[0063] See also Figure 1 The present invention provides an embodiment of a behavior trajectory prediction system for power distribution station operation and maintenance personnel, comprising:
[0064] The data processing module 11, the behavior recognition module 22 and the trajectory prediction module 33 are communicatively connected;
[0065] The data processing module 11 obtains video data of different modes of the power distribution station, performs feature extraction, modality alignment and feature fusion on the video data to obtain multimodal fusion features, where the modality represents the type of the video data;
[0066] The behavior recognition module 22 extracts the human behavior semantic information from the multimodal fusion features, and performs human behavior recognition on the operation and maintenance personnel of the power distribution station according to the human behavior semantic information to obtain a behavior recognition result; the behavior recognition module is a graph convolutional neural network based on the attention mechanism;
[0067] The trajectory prediction module 33 predicts the behavior trajectory of the operation and maintenance personnel according to the behavior recognition result to obtain the behavior trajectory prediction result of the operation and maintenance personnel; the trajectory prediction module is a long short-term memory network based on a visual attention network.
[0068] In this embodiment of the present invention, modality represents the type of video data. Modal data refers to a specific type of data within the video data. Any representative type of data within the video data can be selected as modal data. Multimodality is the fusion of different types of video data. For example, if a user is sensitive to sound, audio data may be used; if the user is sensitive to nighttime, infrared image data may be used; and if the user is sensitive to space, point cloud data may be used. This representative data is based on the environment of the distribution station. The data processing module selects the most appropriate modal data from different segments of the video data as input data and performs fusion.
[0069] In one embodiment, the data processing module may include: a first data processing unit, a second data processing unit, and a feature fusion unit;
[0070] The first data processing unit extracts features from the audio data and the two-dimensional image in the video data to obtain first video features;
[0071] The second data processing unit performs position encoding, modality labeling, and feature dimensionality reduction on the point cloud data and the three-dimensional image in the video data to obtain a second video feature;
[0072] The feature fusion unit performs feature fusion on the first video feature and the second video feature to obtain a multimodal fusion feature.
[0073] Specifically, the data in the video data is divided into two parts according to the difference in modality: the part containing audio data and two-dimensional images is processed by the first data processing unit; the part containing point cloud data and three-dimensional images is processed by the second data processing unit, and the uninformative markers are dynamically detected through multimodal token fusion (TokenFusion).
[0074] It can be understood that the data processing module is an adaptive multimodal fusion structure.
[0075] In one embodiment, the first data processing unit may include: a policy neural network and an improved lightweight object detection network in communication connection;
[0076] Among them, the policy neural network obtains audio data and two-dimensional image data in the video data and performs feature extraction; the lightweight object detection network performs target detection, assigns adaptive attention weights to the detected operation and maintenance personnel, and performs feature enhancement to obtain the first video feature; the lightweight object detection network is improved based on YOLOv5.
[0077] The lightweight target detection network Nano-YOLOv5 is an improved operation based on YOLOv5, such as Figure 3 As shown in , it includes optimizing the detection head, designing a new lightweight FPN Neck network to replace the original PAN, and using SIOU loss to improve the convergence speed. The adaptive attention module is shown in Figure 4 As shown, the feature layer C5 size is h×w, and the context features of three sizes are obtained through the adaptive pooling layer (AdaptivePooling Layer), and then the channel dimensions are made the same through 1×1 convolution. The spatial attention mechanism merges the three context feature channels through the Concat layer, and then generates the corresponding spatial weights through the 1×1 convolution layer, ReLU layer, 3×3 convolution layer and Sigmoid layer in sequence. The generated weights and the features after merging the channels are subjected to Hadamard product operation, separated and added to the M5 feature layer, and the context features are aggregated into M6. Finally, multi-scale context information is obtained, which alleviates the information loss caused by the reduction in the number of channels. Specifically, the lightweight target detection network Nano-YOLOv5 optimizes the detection head based on YOLOv5, avoids suboptimal error back propagation by introducing an asymmetric multi-level channel compression decoupling head, and divides the network into three paths to complete the corresponding tasks, deepens the network path for the target scoring task, and uses 3 convolutions to increase the receptive field and number of parameters of the target detection task, as shown above. Figure 5 As shown in the figure, each feature map is learned using dilated convolutions, and receptive fields of varying sizes are obtained through three dilated convolutions of different sizes. Furthermore, a new lightweight FPN Neck network is designed to replace the original PAN. PAN is a method similar to FPN, differing primarily in the feature fusion method. PAN incorporates an additional bottom-up path compared to FPN, preserving more details but increasing computational effort. FPN Neck, on the other hand, adds an adaptive attention mechanism and feature enhancement to FPN. The front end reduces channels, minimizing contextual information loss, while the back end enhances feature representation, accelerating inference. Furthermore, the SIOU loss function is used to improve convergence speed. The SIOU loss function is a loss function for bounding box regression learning. Based on the IOU, the SIOU loss factor is redefined to take into account the angle between the vectors of the desired regressions, improving training speed and inference accuracy. Compared to the original algorithm, the mean average prediction accuracy (mAP) is improved by 2.8%, the number of parameters is reduced by 28.3%, and the computational effort is reduced by 20.2%, resulting in a 2-3 times faster performance.
[0078] See also Figure 2 The data of different modalities in the video data are divided into two parts. The first part of the data mainly includes two-dimensional image data such as RGB images, depth images, and skeleton images, as well as audio data. The RGB images, depth images, and skeleton images are differentially operated and input into the strategy neural network together with the audio data. After feature extraction and long short-term memory network, they enter the fully connected layer to obtain the corresponding feature map and input it into the lightweight target detection network. The lightweight target detection network is used to perform target detection, that is, to detect operation and maintenance personnel. The detected operation and maintenance personnel are given adaptive attention weights to obtain the first video feature and input it into the feature fusion unit. The second part of the data mainly includes point cloud data and three-dimensional images. The point cloud data is sampled to obtain sampled point cloud data, and the sampled point cloud data and three-dimensional images are input into the second data processing unit. The second data processing unit performs position encoding on the sampled point cloud data and three-dimensional images, and then inputs the multi-head attention Transformer block to obtain multimodal tags. Finally, the second video features are output after feature dimensionality reduction through a feedforward neural network. The features output by the two units, namely the first video features and the second video features, are input into the feature fusion unit. The feature fusion unit is used to fuse the first and second video features to obtain multimodal fusion features, and the multimodal fusion features are input into the behavior recognition module.
[0079] The behavior recognition module 22 extracts the human behavior semantic information from the multimodal fusion features, and performs human behavior recognition on the operation and maintenance personnel of the power distribution station based on the human behavior semantic information to obtain a behavior recognition result.
[0080] Specifically, human behavior can be represented using data modalities such as RGB, skeleton, depth, infrared, point cloud, event stream, and audio. In power distribution station scenarios, RGB, depth, and skeleton modalities are the most popular. Each modality has its own advantages and disadvantages. RGB data is easy to collect and contains rich appearance information, but is significantly affected by the environment. Depth data is insensitive to lighting, but lacks color and texture information, has distance limitations, and is easily affected by occlusions. Skeleton data provides a more accurate description of movement and is less sensitive to the environment, but lacks appearance and shape information and is relatively noisy.
[0081] It should be noted that the depth data in the video data, also called depth image or range image, refers to an image that uses the distance (depth) from the image collector to each point in the specific scene of the distribution station as the pixel value. It directly reflects the geometric shape of the visible surface of the scene.
[0082] Behavior recognition refers to the process of observing and analyzing the behavior of people or objects to determine their identity, intentions, or status. Currently, there are two main types of behavior recognition methods: unimodal methods based on a single data source, and multimodal methods based on multiple data sources. While unimodal methods often present challenges in behavior recognition tasks, multimodal methods can combine the strengths of data from each modality to achieve more accurate behavior recognition results. However, implementation is more challenging.
[0083] Human behavior recognition can help identify safety hazards for operations and maintenance personnel, but different data modalities have their own advantages and disadvantages in distribution substation scenarios. Existing techniques for tracking human trajectories in crowd sizes fail to consider the importance of different neighbors, resulting in inaccurate trajectory predictions.
[0084] Trajectory prediction involves predicting the future movement paths of individuals based on observed human behavior data. Considering the importance of different neighbors in trajectory prediction is crucial, especially in scenarios with varying crowd sizes. In this context, this invention draws on the diverse cognitive modes humans use to perceive the environment, enabling machines to process and associate multiple modal information for multimodal learning and fusion, and to semantically understand the behavior of maintenance personnel based on multimodal video data.
[0085] The embodiment of the present invention intelligently predicts the behavioral trajectories of operation and maintenance personnel based on the video data of the distribution station room. It understands and analyzes the behavioral semantic information of the human body through video data, recognizes human movements, judges human behavior, and predicts future trajectories. On this basis, it can be applied to carry out safety hazard identification tasks, which is of great significance for realizing the prediction and identification of personnel safety hazards based on video data.
[0086] The trajectory prediction module 33 predicts the behavior trajectory of the operation and maintenance personnel according to the behavior recognition result to obtain the behavior trajectory prediction result of the operation and maintenance personnel; the trajectory prediction module is a long short-term memory network based on a visual attention network.
[0087] Taking into account the limited computing power and space of mobile devices, a solution to the high-density crowd problem is proposed. First, a graph structure is used to represent the crowd state. Second, the gaze data of human operators performing bird's-eye view navigation tasks are used to learn an attention network. The network assigns different weights to different individuals in the crowd according to the importance of attention measurement, and uses the field of view constraints of each individual to constrain the weights. Finally, variational inference is used to simulate the randomness of trajectories. Based on this, an embodiment of the present invention proposes a graph neural network ALVGCN with field of view constraints guided by human attention for trajectory prediction. The goal of this model is to generate possible future behavior trajectories of specific / all personnel in the distribution station scene. The overall network model for trajectory prediction is as follows: Figure 3 shown.
[0088] It should be noted that the trajectory prediction module in this embodiment is a graph neural network ALVGCN with field of view constraints guided by human attention. The graph neural network ALVGCN serves as a GCN-based codec and includes two long short-term memory networks LSTM, namely, an encoder LSTM and a decoder LSTM.
[0089] In one embodiment, the trajectory prediction module may include: a communicatively connected attention estimation unit, an encoder LSTM, an interaction network, and a decoder LSTM, where LSTM stands for Long Short-Term Memory Network;
[0090] Among them, the attention estimation unit is used to estimate the estimated attention weight of each operation and maintenance personnel and perform field of view constraints;
[0091] The encoder LSTM processes the position sequence of the operation and maintenance personnel over time to obtain the corresponding motion pattern features;
[0092] The interaction network estimates the social interaction characteristics of the operation and maintenance personnel based on the movement pattern characteristics and social background characteristics; the social background characteristics are obtained based on the relative displacement of all operation and maintenance personnel;
[0093] The decoder LSTM predicts the behavior trajectory of operation and maintenance personnel based on social interaction characteristics.
[0094] In one embodiment, the attention estimation unit may include: an attention network and a visual field filter in communication connection;
[0095] Among them, the attention network assigns attention weights to the operation and maintenance personnel based on their position and current speed relative to specific personnel; the field of view filter adjusts the attention weights according to the field of view of the real world to obtain corresponding estimated attention weights.
[0096] See also Figure 3 , the trajectory prediction module contains a GCN-based variant encoder-decoder backbone (encoder LSTM and decoder LSTM) for sequence-to-sequence trajectory prediction. For each pedestrian, an attention network is used to assign attention to neighboring pedestrians based on their position relative to a specific pedestrian, such as pedestrian i, and their speed; then, a field of view filter is used to adjust the attention according to the field of view limitations of the real world to obtain the corresponding attention weights. The resulting attention weights are applied to the attention set and the adjacency matrix of the modulated GCN. Sequence-to-sequence prediction is achieved by two LSTMs, namely the encoder LSTM and the decoder LSTM. In Figure 4 and Figure 5 In the figure, the mapping from data input to data output in the encoder LSTM and decoder LSTM can be clearly shown.
[0097] See also Figure 3 , Figure 3 The trajectory prediction process of maintenance personnel i is shown in Figure 1. The trajectory prediction module contains a GCN-based variant encoder-decoder backbone for sequence-to-sequence trajectory prediction. The trajectory of personnel i is defined as shown in formula (1), which represents the trajectory of maintenance personnel i at t b ~t e The position sequence within the timestamp is as follows:
[0098]
[0099] The position change within the k~k-1 timestamp is shown in formula (2):
[0100]
[0101] Time step t obs The relative displacement from person j to person i is shown in formula (3):
[0102]
[0103] Time step t=1,...t obs The input trajectory of all people in the scene is Indicates; t = t obs +1, ..., t obs The true future trajectory of +T time steps is used Indicates; N is the total number of people, and the relative displacement of all people to the person of interest is also used As input to better capture the social background characteristics of personnel.
[0104] The individual trajectory is represented by a fixed-length vector, and a single-layer MLP (FC) is first applied to each relative position. Then, an LSTM is used to process the sequence over time, as shown in Equation (4):
[0105]
[0106] Among them, W mot is the weight of a single-layer MLP, MLP mot Calculate the weighted sum, is the position change of person i within the timestamp, is the hidden layer factor of the encoder LSTM, W en represents the weight of the encoder LSTM, the feature vector e i is the hidden state vector of the encoder LSTM on the time axis.
[0107] The FC layer first accepts the relative displacement of all people, and then passes the attention sharing module MLPcont Calculate the weighted sum of embeddings between different groups of people to obtain the social background characteristics p of the people i , as shown in formula (5), where W cont represents the weight of a single-layer MLP, a i represents the attention weight assigned to each person.
[0108]
[0109] The motion mode feature i and social background characteristics i As the input feature s of operation and maintenance personnel i i A two-layer GCN (interaction network) is used to integrate all the operation and maintenance personnel in the crowd [c1, ..., c N ] input features, and obtain N social interaction features v1,...v N , as shown in formula (6):
[0110]
[0111] Each person corresponds to a node in the graph, and the adjacency matrix A si The i-th row of W contains the attention vector of person i to each person in the crowd, and the subscript si represents social interaction. si Represents the weight matrix of the two-layer GCN. The use of two-layer GCN can capture more complex interactions than pairwise interactions.
[0112] The social interaction features computed by the GCN layer are estimated using the mean and variance of the trajectory feature distribution of a pair of MLPs, as shown in Equation (7):
[0113] μ z ,∑ z =MLP mean (v i ;W mean ), MLP var (v i ;W var ); (7)
[0114] Among them, W mean and W var represents the weight matrix, W enc 、W de 、W dec is the weight parameter, from z to N(μ z ,∑ z ) distribution to extract a trajectory feature z used to predict the trajectory. Figure 5 A decoder LSTM connected in a feedback configuration is used to obs +1, ..., tobs +T], the embedding of trajectory feature z before entering the decoder LSTM and the predicted position at time t-1 Connected. Hidden features of decoder LSTM is input to the decoding MLP, which dec Generate trajectory predictions Combining formulas (4) and (5), as shown in formula (8):
[0115] For each individual, the attention network allocates attention to neighboring individuals based on their relative position and current velocity relative to person i. A field of view filter then adjusts attention based on real-world field of view constraints. The resulting attention weights are used to aggregate attention and regulate the GCN's adjacency matrix. Sequence-to-sequence prediction is implemented using two LSTMs. Figure 4 and Figure 5 Unfolding the LSTM clearly shows the mapping from input to output.
[0116] See also Figure 6 ,In order to further optimize the performance in the motion ,prediction task, a network estimation attention structure is introduced to ,learn the attention weights.
[0117] In the prior art, attention is directly learned by minimizing the attention estimation error defined based on gaze data. In the embodiments of the present invention, motion prediction error and attention estimation error are combined for learning.
[0118] As mentioned above, a two-layer GCN is used to estimate the attention weight of each pedestrian. The predicted attention weight is used to summarize the output of the first GCN layer. Using the mixed output features, a fully connected single-layer neural network is used to predict the movement of people. The input of the attention network is the position of people in the crowd relative to person i and the time t obs The instantaneous speed, is the number of people in the crowd at time t obs Social interaction in the x-direction is shown in (9):
[0119]
[0120] Let j∈{1,...,N}, embed these features into a high-dimensional space of a two-layer MLP to obtain high-dimensional features b i , as shown in formula (10):
[0121]
[0122] Embedded data b i is fed into a two-layer GCN, which produces two outputs at each node, with the true attention weight qi and estimated attention weight a i As shown in formula (11):
[0123]
[0124] W att and W gcn is the weight, set the adjacency matrix A att , person i is the central connection point, the rows are normalized, and the sum is 1. The output dimension of each node in the first layer of graph convolution layer is 64, and the output dimension of each node in the second layer is 2.
[0125] Define the weighted minimization loss function of the attention network. Formula (12) calculates the deviation between the motion prediction error and the true gaze, a i is the estimated attention weight of person i to other nodes, and β is the weight coefficient in the combination.
[0126]
[0127] By obs The gaze points collected in the 0.2 second time window of the point are used to calculate the true attention weight a gt (σ). The Gaussian mixture density at the personnel position of the Gaussian mixture model is normalized to the attention weight, and the KL divergence of the two probability distributions is calculated using x g represents the gaze data position, as shown in formula (13):
[0128]
[0129] In order to calculate the first term in Equation (12), the next frame motion of person i is An estimate was made, Generated It comes from a single FC layer, which takes the embedding after the first graph convolutional layer as input and multiplies it by the estimated attention weight, as shown in Equation (14):
[0130]
[0131] Because the gaze data used to estimate true attention is collected from human operators performing robot navigation tasks in a crowd, based on a top-down observation of the environment, the operators can therefore obtain a different perspective than people in the real world, whose field of view is limited to the area in front of them. To compensate for this mismatch, the embodiment of the present invention incorporates a field of view constraint into the estimated attention weights, assuming that each person only pays attention to people within a certain visual angle in the front direction of person i. It has been verified that a 120° field of view works best.
[0132] The present invention first performs modal alignment and network fusion on multimodal video data through a data processing module. Modality refers to a certain type of data in a video, and multimodality is the study of the fusion of different types of data. Representative types of data in a video can be selected as modal data. For example, if you are sensitive to sound, you focus on audio data; at night, you focus on infrared image data; if you are sensitive to space, you focus on point cloud data. This representativeness is reflected according to the environment. The most appropriate modality in different clips is selected as input and fused. The model is divided into two parts according to the difference in modality. The part of the learning data containing audio information is composed of a policy network and a lightweight detection network Nano-YOLOv5. The part of the learning data containing point cloud information is dynamically detected by multimodal token fusion (TokenFusion). Nano-YOLOv5 is an improved operation based on YOLOv5, including optimizing the detection head, designing a new lightweight FPN Neck network to replace the original PAN, and using SIOU loss to improve the convergence speed. Then, a neural network with field of view constraint and processing graph data guided by human attention is proposed to predict human body trajectories.
[0133] In order to further optimize the performance in motion prediction tasks, a network estimation attention structure is introduced to learn the attention weights; finally, by changing the input form, covering functions such as target detection, trajectory prediction, and behavior recognition, and after training and testing, it is proved that the method proposed in the embodiment of the present invention can better model interactions and generate more effective and accurate predicted trajectories.
[0134] The behavior trajectory prediction system for distribution station operation and maintenance personnel provided by the embodiment of the present invention uses a data processing module to obtain video data of different modes of the distribution station, and performs feature extraction, modal alignment and feature fusion to obtain multimodal fusion features. It can fuse video data of different modalities according to the differences between heterogeneous and homogeneous modalities in the specific scene of the distribution station, so as to understand and analyze the behavioral semantic information of the operation and maintenance personnel through video data; use a behavior recognition module to extract human behavior semantic information in the multimodal fusion features, and at the same time introduce an attention mechanism to constrain the field of vision of the operation and maintenance personnel, and identify the actions and behaviors of the people in the video through semantic understanding, and then judge whether the people have accidents, dangerous behaviors and other hidden dangers; finally, use the trajectory prediction module to predict the behavior trajectory of the operation and maintenance personnel, which can automatically and efficiently predict the action trajectory of key operation and maintenance personnel, prevent and reduce the probability of safety hazards in advance, realize efficient operation and maintenance and safety management of the distribution station, and ensure the safe and stable operation of the power system.
[0135] The embodiments of the present invention can effectively identify the behavior of distribution station operation and maintenance personnel and predict the movement trajectories of the operation and maintenance personnel. By establishing a distribution station status detection and personnel trajectory prediction system that takes personnel, equipment and environment into consideration, the probability of safety hazards occurring to distribution station operation and maintenance personnel is reduced, which is of great significance for achieving efficient operation and maintenance of distribution stations and unmanned operation.
[0136] See also Figure 7 The present invention provides an embodiment of a method for predicting the behavior trajectory of a power distribution station operation and maintenance personnel, comprising:
[0137] S100: Obtain video data of different modalities of a power distribution station, and use a preset data processing module to perform feature extraction, modality alignment, and feature fusion on the video data to obtain multimodal fusion features, where the modality represents the type of the video data;
[0138] S200: extracting human behavior semantic information from the multimodal fusion features using a preset behavior recognition module, and performing human behavior recognition on the operation and maintenance personnel of the power distribution station based on the human behavior semantic information to obtain a behavior recognition result; the behavior recognition module is a graph convolutional neural network based on an attention mechanism;
[0139] S300: Utilizing a preset trajectory prediction module to predict the behavior trajectory of the operation and maintenance personnel according to the behavior recognition result, to obtain a behavior trajectory prediction result of the operation and maintenance personnel; the trajectory prediction module is a long short-term memory network based on a visual attention network.
[0140] The embodiment of the present invention performs semantic understanding of human behavior based on video data, and then predicts the behavioral trajectories of distribution station operation and maintenance personnel. Through semantic understanding, the behavior and trajectory of personnel in the video are identified and predicted, and then it is determined whether the personnel have accidents, dangerous behaviors and other hidden dangers. The trajectory of key personnel is predicted to prevent and reduce the probability of safety hazards in advance.
[0141] The proposed method can effectively predict human behavior and trajectory, reducing the probability of potential safety hazards. This is crucial for establishing a comprehensive, integrated system for monitoring and predicting human trajectories in power distribution stations, ensuring efficient, unmanned operation and maintenance. The model is constructed by fusing different data modalities based on the differences between heterogeneous and homogeneous modalities in specific power distribution station scenarios. An attention mechanism is introduced to constrain pedestrian fields of view, and trajectory prediction is performed based on a graph neural network.
[0142] The embodiment of the present invention provides a method for predicting the behavioral trajectory of distribution station operation and maintenance personnel. The method uses a data processing module to obtain video data of different modalities of the distribution station, and performs feature extraction, modal alignment and feature fusion to obtain multimodal fusion features. The method can fuse video data of different modalities according to the differences between heterogeneous and homogeneous modalities in the specific scene of the distribution station, so as to understand and analyze the behavioral semantic information of the operation and maintenance personnel through video data; the method uses a behavior recognition module to extract the semantic information of human behavior in the multimodal fusion features, and at the same time introduces an attention mechanism to constrain the field of vision of the operation and maintenance personnel, and recognizes the actions and behaviors of the personnel in the video through semantic understanding, and then judges whether the personnel have accidents, dangerous behaviors and other hidden dangers; finally, the trajectory prediction module is used to predict the behavioral trajectory of the operation and maintenance personnel, which can automatically and efficiently predict the action trajectory of key operation and maintenance personnel, prevent and reduce the probability of safety hazards in advance, realize efficient operation and maintenance and safety management of the distribution station, and ensure the safe and stable operation of the power system.
[0143] The embodiments of the present invention can effectively identify the behavior of distribution station operation and maintenance personnel and predict the movement trajectories of the operation and maintenance personnel. By establishing a distribution station status detection and personnel trajectory prediction system that takes personnel, equipment and environment into consideration, the probability of safety hazards occurring to distribution station operation and maintenance personnel is reduced, which is of great significance for achieving efficient operation and maintenance of distribution stations and unmanned operation.
[0144] See also Figure 8 This method uses video data from distribution station rooms to intelligently predict the trajectory of maintenance personnel. By analyzing the semantic information of human behavior through video understanding, it can identify human movements, judge human behavior, and predict future trajectories. The overall framework consists of three parts: a data processing module (acquisition input module), a behavior recognition module, and a trajectory prediction module (function output module).
[0145] The acquisition input module, primarily composed of a modality alignment module and a network fusion module, aligns the RGB data of different modalities to compensate for static spatial information and eliminate differences between them. The aligned modality data is then fed into the network fusion module, where feature learning and fusion are performed simultaneously. The fused multimodal features are then fed into the behavior recognition module.
[0146] The behavior recognition module uses a lightweight, high-precision target detection algorithm to extract as much rich and in-depth semantic information about human behavior as possible, resulting in more accurate human movements and trajectories. The trajectory prediction module, by inputting data of the corresponding modality, can cover functions such as target detection, behavior recognition, and trajectory prediction. The trajectory prediction module can modify the form of input data and perform related operations to obtain corresponding functions. For example, when only image data is input, it can detect people in the image and identify their behavior and movements. When inputting video data of a certain duration (e.g., 20 seconds), it can predict the movement trajectory of people in the video.
[0147] The subject guides a "virtual pedestrian" to try to walk through a crowd from top to bottom according to the scene. The crowd is constructed from real pedestrian trajectories used to train the pedestrian trajectory prediction network. When performing this task, the subject's gaze is detected, and the gaze data is used only to train the attention network model.
[0148] The leave-one-out method is used to evaluate trajectory prediction and attention networks, taking the trajectory of 6 time steps as the observation object and evaluating the trajectory prediction of the next 10 time steps. The performance is measured by the average displacement error (ADE) and the final displacement error (FDE).
[0149] By calculating the mean absolute error of motion prediction on 5 different test datasets, the learning of attention improved the prediction accuracy by an average of 6.8%. Quantitative evaluation was also conducted with some advanced and excellent baseline models (S-GAN, CoMoGCN, STGAT, NMMP, etc.). ALVGCN achieved the smallest average ADE / FDE, with an average improvement of 11.1% in ADE and 13.6% in FDE. Finally, several ablation experiments were conducted to verify the effectiveness of using GCN and attention mechanisms. By comparing the AGCN, GCN and ALVGCN, VGCN models, the attention embedded in human gaze has a certain improvement on ADE and FDE. In summary, through these experiments, the ALVGCN model models better interactions and generates behavior trajectories more effectively and accurately.
[0150] The present invention also provides an electronic device, comprising: a processor and a memory;
[0151] The memory stores a computer program, and the processor implements the steps of the method when executing the computer program.
[0152] The present invention also provides a computer-readable storage medium having a computer program stored thereon, and the computer program implements the steps of the method when executed by a processor.
[0153] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0154] In the embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.
[0155] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0156] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0157] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0158] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A behavior trajectory prediction system for power distribution station operation and maintenance personnel, characterized in that: include: Communication-connected data processing module, behavior recognition module, and trajectory prediction module; The data processing module obtains video data of different modes of the power distribution station, performs feature extraction, modality alignment and feature fusion on the video data to obtain multimodal fusion features, where the modality represents the type of the video data; The behavior recognition module extracts the human behavior semantic information from the multimodal fusion features, and performs human behavior recognition on the operation and maintenance personnel of the power distribution station based on the human behavior semantic information to obtain a behavior recognition result; the behavior recognition module is a graph convolutional neural network based on the attention mechanism; The trajectory prediction module predicts the behavior trajectory of the operation and maintenance personnel according to the behavior recognition result to obtain the behavior trajectory prediction result of the operation and maintenance personnel; the trajectory prediction module is a long short-term memory network based on the visual attention network; The trajectory prediction module includes: an attention estimation unit connected in communication, an encoder LSTM, an interaction network and a decoder LSTM, where LSTM stands for long short-term memory network; The attention estimation unit is used to estimate the estimated attention weight of each operation and maintenance personnel and perform field of view constraints; The encoder LSTM processes the position sequence of the operation and maintenance personnel over time to obtain corresponding motion pattern features; The interactive network estimates the social interaction characteristics of the operation and maintenance personnel based on the motion pattern characteristics and the social background characteristics; the social background characteristics are obtained based on the relative displacement of all the operation and maintenance personnel; The decoder LSTM predicts the behavior trajectory of the operation and maintenance personnel based on the social interaction characteristics; The attention estimation unit includes: an attention network and a visual field filter in communication connection; The attention network assigns attention weights to the operator based on the operator's position and current speed relative to a specific person. The field of view filter adjusts the attention weight according to the field of view of the real world to obtain a corresponding estimated attention weight.
2. The behavior trajectory prediction system for power distribution station operation and maintenance personnel according to claim 1 is characterized in that: The data processing module includes: a first data processing unit, a second data processing unit, and a feature fusion unit; Wherein, the first data processing unit extracts features from the audio data and the two-dimensional image in the video data to obtain a first video feature; The second data processing unit performs position encoding, modality labeling, and feature dimensionality reduction on the point cloud data and the three-dimensional image in the video data to obtain a second video feature; The feature fusion unit performs feature fusion on the first video feature and the second video feature to obtain a multimodal fusion feature.
3. The behavior trajectory prediction system for power distribution station operation and maintenance personnel according to claim 2 is characterized in that: The first data processing unit includes: Communication-connected policy neural networks and improved lightweight object detection networks; Wherein, the strategy neural network obtains the audio data and two-dimensional image data in the video data and performs feature extraction; The lightweight target detection network performs target detection, assigns adaptive attention weights to detected operation and maintenance personnel, and performs feature enhancement to obtain a first video feature; the lightweight target detection network is improved based on YOLOv5.
4. A method for predicting the behavior trajectory of power distribution station operation and maintenance personnel, characterized in that: include: Obtain video data of different modalities of the power distribution station, and use a preset data processing module to perform feature extraction, modality alignment, and feature fusion on the video data to obtain multimodal fusion features, where the modality represents the type of the video data; A preset behavior recognition module is used to extract human behavior semantic information from the multimodal fusion features, and human behavior recognition is performed on the operation and maintenance personnel of the power distribution station based on the human behavior semantic information to obtain a behavior recognition result; the behavior recognition module is a graph convolutional neural network based on the attention mechanism; Using a preset trajectory prediction module to predict the behavior trajectory of the operation and maintenance personnel according to the behavior recognition result, to obtain a behavior trajectory prediction result of the operation and maintenance personnel; The trajectory prediction module is a long short-term memory network based on the visual attention network; The trajectory prediction module includes: an attention estimation unit connected in communication, an encoder LSTM, an interaction network and a decoder LSTM, where LSTM stands for long short-term memory network; wherein, the attention estimation unit is used to estimate the estimated attention weight of each operation and maintenance personnel and perform field of view constraints; Using the encoder LSTM to process the position sequence of the operation and maintenance personnel over time to obtain corresponding motion pattern features; estimating social interaction characteristics of the operation and maintenance personnel based on the motion pattern characteristics and social background characteristics using the interaction network; the social background characteristics are obtained based on the relative displacement of all the operation and maintenance personnel; Using the decoder LSTM to predict the behavior trajectory of the operation and maintenance personnel based on the social interaction characteristics; The attention estimation unit includes: an attention network and a visual field filter in communication connection; wherein, the attention network is used to assign an attention weight to the operation and maintenance personnel according to the position and current speed of the operation and maintenance personnel relative to the specific personnel; The field of view filter is used to adjust the attention weight according to the field of view of the real world to obtain a corresponding estimated attention weight.
5. The method for predicting the behavior trajectory of power distribution station operation and maintenance personnel according to claim 4 is characterized in that: The data processing module includes a first data processing unit, a second data processing unit and a feature fusion unit. The method of using the preset data processing module to perform feature extraction, modality alignment and feature fusion on the video data to obtain multimodal fusion features includes: Using the first data processing unit to extract features from the audio data and the two-dimensional image in the video data to obtain a first video feature; Using the second data processing unit to perform position encoding, modality labeling, and feature dimensionality reduction on the point cloud data and the three-dimensional image in the video data to obtain a second video feature; The feature fusion unit is used to perform feature fusion on the first video feature and the second video feature to obtain a multimodal fusion feature.
6. The method for predicting the behavior trajectory of power distribution station operation and maintenance personnel according to claim 5 is characterized in that: The first data processing unit includes a communication-connected strategy neural network and an improved lightweight target detection network. The first data processing unit is used to extract features from the audio data and the two-dimensional image in the video data to obtain the first video features, including: Acquire audio data and a two-dimensional image from the video data, and extract features from the audio data and the two-dimensional image using the strategy neural network; The lightweight target detection network is used to perform target detection, assign adaptive attention weights to the detected operation and maintenance personnel, and perform feature enhancement to obtain the first video feature; the lightweight target detection network is improved based on YOLOv5.
7. An electronic device, characterized in that: include: processor and memory; The memory stores a computer program, and the processor implements the steps of the method according to any one of claims 4 to 6 when executing the computer program.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 4 to 6 are implemented.
Citation Information
Patent Citations
Abnormal behavior identification method and system based on feature object and human body key point, and medium
CN115482502A