Operation inspection personnel behavior identification method and system of extra-high voltage transformer substation, electronic equipment and medium

By combining the temporal video state space model and the video attitude transformer model, the problems of insufficient processing of temporal video redundancy information and spatiotemporal correlation modeling in UHV substations are solved, enabling accurate and real-time identification of the behavior of operation and maintenance personnel, and improving the identification accuracy and processing speed.

CN121768065APending Publication Date: 2026-03-31MAINTENANCE BRANCH COMPANY STATE GRID ZHEJIANG ELECTRIC POWER
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-27
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively handle redundant information in time-series videos in UHV substations, and insufficient spatiotemporal correlation modeling makes it difficult to balance the accuracy of feature extraction and computational efficiency in recognizing the behavior of operation and maintenance personnel.

Method used

Employing a temporal video state-space model and a video pose transformer model, and through a marker pruning and restoration mechanism, we achieve low-redundancy temporal feature extraction and 3D behavioral feature recognition, including temporal feature compression and reconstruction operations, and combine them with an edge computing platform for real-time processing.

Benefits of technology

It improved the accuracy of operation and maintenance personnel behavior recognition to 96.3%, and the processing speed reached 2.8 times that of traditional methods, thereby improving the automation level and reliability of substation safety monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121768065A_ABST
    Figure CN121768065A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of safety production monitoring of extra-high voltage substations, in particular to an operation inspection personnel behavior identification method and system of an extra-high voltage substation, electronic equipment and a medium. The method comprises the following steps: acquiring time sequence video data from operation and maintenance personnel of the extra-high voltage transformer substation; extracting low-redundancy time sequence features from the time sequence video data through the time sequence video state space model; inputting the low-redundancy time sequence features into a pre-configured video attitude converter model, executing hierarchical feature compression and reconstruction operation of a mark pruning and recovery mechanism through the video attitude converter model, and extracting three-dimensional behavior features from the low-redundancy time sequence features; and according to the three-dimensional behavior characteristics, identifying to obtain the behavior category of the operation inspection personnel. In this way, the technical problem that feature extraction precision and calculation efficiency are difficult to consider at the same time in the current operation and maintenance personnel behavior recognition technology is solved, and the automation level and reliability of substation safety monitoring are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of safety monitoring technology for ultra-high voltage substations, and in particular to a method, system, electronic equipment, and medium for identifying the behavior of operation and maintenance personnel in ultra-high voltage substations. Background Technology

[0002] With the continuous and rapid development of the power industry, the construction scale of ultra-high voltage (UHV) substations is constantly expanding, making the actual need for monitoring the work behavior and safety risk assessment of power operation and maintenance personnel increasingly urgent. UHV substations are high-risk operating environments with complex work processes and numerous potential risk points. The ability to promptly and accurately capture non-standard behaviors of operation and maintenance personnel and quickly identify potential safety hazards directly affects the stable operation of the power system and the safety of personnel. Therefore, efficient monitoring of the behavior of operation and maintenance personnel has become a key task in the power system safety assurance system.

[0003] In the field of behavior recognition and safety early warning, existing technologies have made some progress. Behavior recognition technology based on video surveillance, due to its intuitiveness and real-time capability, has been widely applied in monitoring the behavior of operation and maintenance personnel in UHV substations. Current mainstream time-series video analysis methods mostly rely on various deep learning technologies to extract behavioral features through video data processing and combine 3D pose estimation technology to improve the accuracy of behavior recognition. However, the operating environment of UHV substations is complex and dynamic, and time-series video data often has characteristics such as high information redundancy and complex spatiotemporal correlations. Existing technologies still have significant limitations in addressing these issues.

[0004] Prior art document 1 (application publication number CN116631050A) discloses a user behavior recognition method and system for intelligent video conferencing. It extracts temporal and spatial features by constructing a spatiotemporal dual-branch network to improve the accuracy and efficiency of user behavior recognition in conferencing scenarios. However, this method is primarily designed for the specific scenario of video conferencing. Its feature extraction logic and optimization objectives are more suited to the conventional behavior recognition of participants, failing to fully consider the high redundancy characteristics of time-series videos in the complex environment of UHV substations and the dynamic spatiotemporal correlation requirements of operation and maintenance personnel's work behaviors. Specifically, existing technologies generally suffer from two prominent problems: firstly, they fail to effectively handle redundant information in time-series videos, especially invalid data in long time-series videos, significantly increasing the computational burden and reducing processing efficiency; secondly, their ability to model spatiotemporal correlations is insufficient, making it difficult to accurately capture the dynamic changes in the work behaviors of operation and maintenance personnel, resulting in limited accuracy in behavior feature extraction. These problems overlap, making it difficult for existing technologies to balance the accuracy and computational efficiency of behavior recognition in practical applications at UHV substations, and failing to meet the needs of rapid and accurate safety monitoring in high-risk operation scenarios.

[0005] Therefore, existing video-based UHV substation operation and maintenance personnel behavior recognition technology faces technical problems such as poor handling of redundant information and insufficient spatiotemporal correlation modeling, which in turn makes it difficult to balance feature extraction accuracy and computational efficiency. Summary of the Invention

[0006] To address the aforementioned shortcomings or deficiencies, this invention provides a method, system, electronic device, and medium for recognizing the behavior of operation and maintenance personnel in ultra-high voltage substations, which can solve the technical problem of the difficulty in balancing feature extraction accuracy and computational efficiency in current operation and maintenance personnel behavior recognition technologies.

[0007] This invention provides a method for recognizing the behavior of operation and maintenance personnel in ultra-high voltage substations, comprising: Acquire time-series video data from UHV substation operation and maintenance personnel.

[0008] Low-redundancy temporal features are extracted from temporal video data using a pre-defined temporal video state space model.

[0009] Low-redundancy temporal features are input into a pre-configured video pose transformer model. The video pose transformer model performs hierarchical feature compression and reconstruction operations using a tag pruning and restoration mechanism to extract three-dimensional behavioral features from the low-redundancy temporal features.

[0010] The behavior categories of maintenance personnel are identified based on three-dimensional behavioral characteristics.

[0011] According to a second aspect, this invention provides a behavior recognition system for operation and maintenance personnel in ultra-high voltage substations, comprising: The video data acquisition module is used to acquire time-series video data from UHV substation operation and maintenance personnel.

[0012] The temporal feature extraction module is used to extract low-redundancy temporal features from temporal video data using a preset temporal video state space model.

[0013] The 3D feature extraction module is used to input low-redundancy temporal features into a pre-configured video pose transformer model. The video pose transformer model performs hierarchical feature compression and reconstruction operations using a tag pruning and restoration mechanism to extract 3D behavioral features from the low-redundancy temporal features.

[0014] The behavior category recognition module is used to identify the behavior category of maintenance personnel based on three-dimensional behavioral features.

[0015] According to a third aspect, the present invention provides an electronic device comprising: At least one processor; and The memory that is communicatively connected to the at least one processor; The memory stores instructions that can be executed by the at least one processor, which enables the at least one processor to execute the operation and maintenance personnel behavior recognition method for any UHV substation in the embodiments of the present invention.

[0016] According to another aspect of the present invention, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute the operation and maintenance personnel behavior recognition method of any UHV substation in the embodiments of the present invention.

[0017] The present invention provides a method for recognizing the behavior of operation and maintenance personnel in ultra-high voltage substations. The method includes: first, acquiring time-series video data from operation and maintenance personnel in ultra-high voltage substations, with a video frame rate of no less than 25 frames per second and a duration covering the entire operation process; next, extracting spatiotemporal features from the original video data using a preset time-series video state space model to eliminate visual redundancy and obtain low-redundancy time-series features; then, inputting the low-redundancy time-series features into a pre-configured video posture transformer model, which removes redundant feature segments through a label pruning mechanism and reconstructs key motion trajectories through a recovery mechanism, ultimately outputting a feature vector representing the three-dimensional motion posture of the human body; finally, based on the three-dimensional behavior feature vector, using a classifier to identify the behavior category of the operation and maintenance personnel (such as normal operation, violation, and dangerous action).

[0018] In this technical solution, the present invention addresses the problems of high data redundancy and difficulty in effective feature extraction in substation video monitoring, as described in the background art. It achieves a compact representation of the original video data through a temporal video state-space model, overcoming the shortcomings of traditional methods such as high processing latency and heavy storage burden due to large data volumes. Regarding the issues of strong spatiotemporal feature coupling and high modeling complexity in dynamic behavior recognition, the invention utilizes a labeling, pruning, and restoration mechanism in a video pose converter model to achieve accurate capture of key motion information and effective suppression of redundant noise. Finally, addressing the difficulties in extracting three-dimensional behavioral features and low recognition accuracy, the invention achieves accurate mapping from two-dimensional image sequences to three-dimensional motion features through hierarchical feature compression and reconstruction operations. Therefore, the technical solution of this invention solves the technical problem of balancing feature extraction accuracy and computational efficiency in current operation and maintenance personnel behavior recognition technologies, achieving accurate and real-time recognition of operation and maintenance personnel's operational behaviors, and improving the automation level and reliability of substation safety monitoring. Attached Figure Description

[0019] Figure 1 This is a schematic diagram of the extraction of behavioral features of power personnel according to the present invention; Figure 2 This is a schematic diagram of the overall framework of the Hot invention; Figure 3This is a flowchart of a method for recognizing the behavior of operation and maintenance personnel in an ultra-high voltage substation according to an embodiment of the present invention; Figure 4 This is an example diagram of Mamba blocks used for 1D and 2D sequences in one embodiment of the present invention. Figure 5 This is a schematic diagram of a VPT architecture according to an embodiment of the present invention; Figure 6 This is a comparison diagram of the VPT and HoT models according to an embodiment of the present invention; Figure 7 This is a structural block diagram of a behavior recognition system for operation and maintenance personnel in an ultra-high voltage substation according to an embodiment of the present invention; Figure 8 This is a block diagram of an electronic device used to implement embodiments of the present invention. Detailed Implementation

[0020] The following description, in conjunction with the accompanying drawings, illustrates exemplary embodiments of the present invention, including various details to aid understanding. These details should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the invention. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0021] During the development of this invention, the inventors, through extensive experiments and data analysis, revealed the intrinsic relationship between spatiotemporal redundancy in long-term video and the accuracy of 3D behavioral feature extraction: spatiotemporal redundancy not only increases the computational burden but also dilutes the salience of key behavioral features, making it difficult for traditional methods to balance accuracy and efficiency. Based on this relationship, the inventors innovatively proposed this technical solution, utilizing the linear modeling capability of the state-space model for spatiotemporal dynamics and the adaptive compression and reconstruction mechanism of the hourglass word segmenter, thereby achieving a significant improvement in computational efficiency while ensuring feature extraction accuracy, reflecting the core concept of this invention.

[0022] Specifically, such as Figure 1 As shown, the invention team discovered through comparative experiments that the traditional VPT (Video Pose Transformer) model has a redundant frame rate of 40% to 60% when processing long videos with more than 300 frames. This redundant information leads to an increase in 3D pose estimation error of about 15%. However, by using the selective scanning mechanism of the state-space model, key spatiotemporal segments can be effectively focused, reducing invalid calculations.

[0023] Furthermore, such as Figure 2As shown, clustering analysis using the TPC (Token Pruning Cluster) module demonstrates that retaining only 20% of representative tags with high semantic diversity is sufficient to preserve over 95% of the effective behavioral feature information. This quantitative relationship provides a theoretical basis for designing efficient tag pruning strategies. Experimental data confirms that the proposed solution, in the actual scenario of UHV substations, improves the behavior recognition accuracy to 96.3% while achieving a processing speed 2.8 times faster than traditional methods. This breakthrough performance improvement is based on a profound understanding of the intrinsic relationship between spatiotemporal redundancy and feature saliency, as well as an innovative technical architecture design.

[0024] Therefore, this invention provides a method for recognizing the behavior of operation and maintenance personnel in ultra-high voltage substations, based on the first aspect. This method can be applied to an intelligent safety monitoring system for ultra-high voltage substations (hereinafter referred to as the "system"). The system can be deployed locally or run on an edge computing platform or server cluster in a cloud-based collaborative manner to achieve real-time recognition and safety warning of the behavior of operation and maintenance personnel.

[0025] Specifically, the system's physical equipment includes, but is not limited to, high-definition surveillance cameras, video analytics servers, data storage units, and network transmission equipment. These devices must possess capabilities for video data acquisition, real-time stream processing, deep learning model inference, and high-speed data communication to support efficient analysis of large-scale time-series video, ensuring the system can accurately identify the behavior of maintenance personnel and issue timely warning signals in complex environments.

[0026] like Figure 3 As shown, the method may include: Step S110: Obtain time-series video data from UHV substation operation and maintenance personnel.

[0027] Temporal video data refers to a continuous sequence of images captured by the system through video acquisition equipment and arranged in chronological order. It is used to record the dynamic behavior of operation and maintenance personnel in the working environment of UHV substations. Temporal video data typically includes a spatial dimension (image pixels) and a temporal dimension (frame sequence), and can fully capture the spatiotemporal evolution characteristics of behavior.

[0028] Specifically, the system collects raw video streams by monitoring equipment deployed in key operating areas of UHV substations (such as high-voltage equipment areas and inspection routes), and preprocesses the video streams, including frame rate standardization, resolution adjustment, and data format unification, in order to eliminate environmental interference and ensure data consistency.

[0029] For example, the system uses an electromagnetic interference-resistant industrial camera (such as a high-resolution camera). The industrial camera captures video at a sampling rate of 25 frames per second (fps), and the video sequence length can be set from several seconds to several minutes according to monitoring needs. Data is transmitted via a gigabit Ethernet interface, and an integrated temperature control system (operating temperature range of -40°C to 85°C) ensures stable operation in environments with strong electromagnetic fields and temperature variations. The system adds timestamps and location tags to the acquired video data, forming a structured time-series video dataset for subsequent model processing. If the video brightness is below 50 lux (unit of illuminance) or motion blur exists, the system automatically triggers a resampling mechanism to ensure data quality.

[0030] Step S120: Extract low-redundancy temporal features from temporal video data using a preset temporal video state space model.

[0031] Among them, the temporal video state-space model refers to the model based on selective state-space modeling. The architecture (a selective state-space model specifically designed for video understanding, which achieves dynamic spatiotemporal background modeling through a linear complexity method, effectively handling spatiotemporal redundancy and correlation issues in time-series videos) is used to dynamically model the spatiotemporal background of video sequences, eliminating redundant information and enhancing spatiotemporal correlation with linear complexity. Low-redundancy temporal features refer to the key spatiotemporal information representations retained after model processing, with reduced dimensionality and the removal of duplicate or irrelevant frames, facilitating efficient subsequent processing.

[0032] Specifically, the model maps the video frame sequence to the state space through the state projection module to generate an initial state sequence; then, it dynamically focuses on the state sequence through the selective scanning module to capture long-term spatiotemporal dependencies; finally, it outputs the optimized feature sequence through the spatiotemporal correlation modeling module.

[0033] For example, for a segment containing 300 frames, resolution The video input is pixels, and the temporal video state space model transforms it into a dimension. The model employs low-redundancy feature sequences, achieving a redundant frame removal rate of over 40%, while retaining over 95% of key behavioral features (such as arm swing amplitude and torso tilt angle). Running on edge computing devices (such as NVIDIA Jetson modules, an embedded AI platform for edge computing with low power consumption and high-performance AI inference capabilities, suitable for real-time video analysis tasks in industrial environments), the inference latency is less than 100 milliseconds, meeting real-time requirements.

[0034] Step S130: Input the low-redundancy temporal features into the pre-configured video pose transformer model, and perform hierarchical feature compression and reconstruction operations through the tag pruning and restoration mechanism of the video pose transformer model to extract three-dimensional behavioral features from the low-redundancy temporal features.

[0035] The video pose transformer model refers to the VPT model based on the Transformer architecture, specifically designed for regressing 3D human pose from video sequences. The token pruning and recovery mechanism refers to the core operations within the HoT framework (an efficient 3D human pose estimation framework based on an hourglass tokenizer structure, optimizing the computational efficiency of the video pose transformer model through token pruning and recovery mechanisms). This includes the TPC module and the Token Recovering Attention (TRA) module, optimizing computational efficiency through compression followed by reconstruction. The hierarchical feature compression and reconstruction operation involves multi-stage processing of the input tokenized sequence: first, dynamically selecting representative tokens to reduce the sequence length, and then restoring it to the original resolution to ensure accuracy. 3D behavioral features refer to the quantitative data describing the motion trajectory of the operator's joints in 3D space, typically represented as a sequence of joint coordinates.

[0036] Specifically, the model first performs cluster analysis on low-redundancy temporal features through a label pruning clustering module, and selects labels with high semantic diversity to generate compressed sequences; then, it reconstructs the complete sequence based on attention weights through a label recovery attention module; finally, it maps the labels to three-dimensional pose coordinates through a regression head module.

[0037] For example, the input low-redundancy feature sequence contains 200 labels. The label pruning and clustering module selects 50 representative labels (cluster centers) using K-means clustering, achieving a compression rate of 75%. The label recovery attention module restores the sequence to 200 labels using a multi-head attention mechanism and outputs a 3D coordinate sequence of 17 human joints (such as shoulder, elbow, and wrist), with a coordinate accuracy error of less than 5 millimeters. The entire processing is completed on a GPU (Graphics Processing Unit) server, reducing computation time by 40% compared to the traditional VPT model.

[0038] Step S140: Identify the behavior category of the maintenance personnel based on the three-dimensional behavioral features.

[0039] Among them, behavior category refers to the predefined classification of the operation and maintenance personnel's work actions, such as normal operation (e.g., walking, inspection), abnormal behavior (e.g., falling, unauthorized climbing), or risky action (e.g., approaching high-voltage equipment). The identification process is based on the spatiotemporal pattern analysis of three-dimensional behavioral features, which is mapped to discrete category labels through a classification model.

[0040] Specifically, the system inputs 3D behavioral features into the behavior classification module (such as a neural network based on fully connected layers and the Softmax function). First, it performs spatiotemporal aggregation on the features (e.g., using global average pooling) to generate behavioral feature vectors. Then, it calculates the class probability distribution through a classifier and selects the class with the highest probability as the output. The Softmax function is a normalized exponential function used for multi-class classification tasks. It transforms the raw scores output by the neural network into probability distributions for each class through exponential transformation and probability mapping, thereby supporting the model in making interpretable classification decisions.

[0041] For example, for a 3D pose sequence containing 10 seconds of behavior (30 frames per second, 300 frames in total), the behavior classification module extracts a feature vector with a dimension of 256. The module then outputs the probabilities of five behavior categories (e.g., "normal walking," "abnormal running," "dangerous climbing," "safe gesture," and "fall") using the Softmax function, setting a probability threshold of 0.7. If the probability of "dangerous climbing" exceeds the threshold, a real-time warning signal is triggered. The system supports dynamically updating the behavior category library to adapt to the needs of different work scenarios.

[0042] Therefore, according to the above implementation method, the system first acquires time-series video data from UHV substation operation and maintenance personnel, with a video frame rate of no less than 25 frames per second and a duration covering the entire operation process; then, it extracts spatiotemporal features from the original video data through a preset time-series video state space model to eliminate visual redundancy and obtain low-redundancy time-series features; then, it inputs the low-redundancy time-series features into a pre-configured video posture converter model, which removes redundant feature segments through a label pruning mechanism and reconstructs key motion trajectories through a recovery mechanism, finally outputting a feature vector representing the three-dimensional motion posture of the human body; finally, based on the three-dimensional behavior feature vector, a classifier is used to identify the behavior category of the operation and maintenance personnel (such as normal operation, violation behavior, dangerous action).

[0043] Specifically, in this implementation, addressing the issues of high data redundancy and difficulty in effective feature extraction in substation video monitoring as described in the background technology, a time-series video state-space model is used to achieve a compact representation of the original video data, solving the shortcomings of traditional methods such as high processing latency and heavy storage burden due to large data volume. Regarding the issues of strong spatiotemporal feature coupling and high modeling complexity in dynamic behavior recognition, a labeling, pruning, and restoration mechanism of the video pose converter model is used to achieve accurate capture of key motion information and effective suppression of redundant noise. Addressing the difficulties in extracting three-dimensional behavioral features and low recognition accuracy, hierarchical feature compression and reconstruction operations are used to achieve accurate mapping from two-dimensional image sequences to three-dimensional motion features. Therefore, this implementation solves the technical problem of balancing feature extraction accuracy and computational efficiency in current operation and maintenance personnel behavior recognition technologies, achieving accurate and real-time recognition of operation and maintenance personnel's operational behaviors, and improving the automation level and reliability of substation safety monitoring.

[0044] In some embodiments, the temporal video state space model includes a selective state space modeling module and a spatiotemporal correlation modeling module; low-redundancy temporal features are extracted from temporal video data using a preset temporal video state space model, including: The selective state-space modeling module performs state-space transformation on the temporal video data to generate intermediate spatiotemporal feature sequences.

[0045] State-space transformation processing refers to the operation of mapping temporal video data to a state-space representation, and describing the dynamic evolution of the video sequence through state-space equations; intermediate spatiotemporal feature sequence refers to the feature vector sequence containing spatiotemporal information obtained after transformation, whose dimensionality is significantly reduced compared to the original video data, making it easier for subsequent processing.

[0046] Specifically, the selective state-space modeling module is based on the Video Mamba architecture (a selective state-space model specifically designed for video understanding tasks, which effectively solves the computational efficiency and feature correlation problems in long video sequence processing through a linear complexity dynamic spatiotemporal modeling mechanism). It transforms the input video frame sequence into a state sequence through linear complexity state-space modeling. The module first projects the video data onto the state space through a state projection operation to generate an initial state representation; then, it dynamically focuses on key spatiotemporal regions through a selective scanning mechanism to optimize feature extraction efficiency.

[0047] For example, for a 10-second segment with a resolution The video sequence consists of pixels sampled at 25 frames per second, for a total of 250 frames. The selective state-space modeling module converts each frame into a feature vector, which can be expressed using state-space equations (e.g., differential equations): ,in (A is the state vector, and A and B are parameter matrices used to describe the dynamics of the video sequence.) The generation dimension is... The intermediate spatiotemporal feature sequence, in which The spatial dimension is represented by 200, the time frame number is 256, and the feature channel number is 256. The module runs on an edge computing device with a processing latency of less than 50 milliseconds.

[0048] In some embodiments, examples of Mamba blocks used for 1D and 2D sequences are as follows: Figure 4 As shown, the Chinese translation of the Mamba block is State Space Model Block (or, based on its functional characteristics, Dynamic State Space Block). This term specifically refers to the basic computational unit based on Selective State Space Modeling, whose core is to dynamically model sequential data through state space equations, possessing linear computational complexity and the ability to capture long-range dependencies. Furthermore, as a core component of the VideoMamba architecture, the Mamba block achieves efficient processing of video sequences through a selective scanning mechanism. For example, when processing time-series video data of UHV substation maintenance personnel, the Mamba block can perform dynamic feature extraction on a 256-frame video sequence, with a single forward propagation taking only 3.2 milliseconds, nearly three times more computationally efficient than traditional Transformer modules. Figure 4 (a) The unidirectional Mamba block shown maps the input sequence to the latent space through a linear projection layer. After the sequence is transformed by the state space model (SSM), the output features are generated through gating mechanism and convolution operation, which is suitable for 1D sequence processing such as time series signals. Figure 4 (b) The bidirectional Mamba block captures contextual dependencies from both directions through a bidirectional scanning mechanism (marked "flip" to indicate sequence flipping), and then outputs the results after feature fusion. This is particularly suitable for modeling 2D spatiotemporal data such as video frame sequences. This hierarchical architecture allows the present invention to flexibly adapt to monitoring data of different dimensions in UHV substations. For example, 1D Mamba blocks can be used for sensor time-series signal analysis (such as vibration frequency sequences), and 2D bidirectional Mamba blocks can be used for behavioral feature extraction of video frame sequences. This achieves millisecond-level real-time processing on edge devices while maintaining more than 90% feature reconstruction accuracy.

[0049] The spatiotemporal correlation modeling module captures long-term spatiotemporal dependencies in intermediate spatiotemporal feature sequences and outputs low-redundancy time series features.

[0050] Among them, long-term spatiotemporal dependency refers to the continuous association of actions between long-distance frames in a video sequence, as well as the coordinated changes between human body joints in the spatial dimension; low-redundancy temporal features refer to the optimized feature representation after redundancy removal, which retains key behavioral information while reducing computational burden.

[0051] Specifically, the spatiotemporal correlation modeling module analyzes long-term dependencies in feature sequences through attention mechanisms (such as multi-head self-attention) or spatiotemporal convolutional networks. The module first decomposes the intermediate spatiotemporal feature sequences into spatiotemporal dimensions to capture temporal and spatial correlations; then, it enhances key features through weighted fusion and outputs a compressed feature sequence.

[0052] For example, the dimension of the input intermediate spatiotemporal feature sequence is... After processing by the spatiotemporal correlation modeling module, using a Transformer structure with 8 attention heads and 512 hidden layer dimensions, long-term dependencies are captured, and the output dimension is... It exhibits low-redundancy temporal characteristics, with a redundant frame removal rate of approximately 50% and a key behavioral feature retention rate exceeding 95%. The module runs on a GPU server, reducing processing time by 40% compared to traditional methods.

[0053] Therefore, according to the above implementation method, the system can efficiently eliminate video redundancy, enhance spatiotemporal correlation, provide optimized input for behavior recognition, and improve processing speed and accuracy.

[0054] In some embodiments, the temporal video state space model further includes a state projection module and a selective scanning module; the selective state space modeling module performs state space transformation processing on the temporal video data to generate an intermediate spatiotemporal feature sequence, including: The state projection module performs state space projection processing on the time-series video data to generate an initial state sequence; the state space projection processing includes a linear transformation operation that converts the video frame sequence into a state vector.

[0055] Among them, the state projection module is the component in the temporal video state space model responsible for mapping the input video data to the state space representation. It realizes the conversion from high-dimensional video frames to low-dimensional state vectors through linear transformation, so as to reduce computational complexity and retain key spatiotemporal information. State space projection processing is a feature extraction method based on matrix operations, which converts each video frame or frame sequence into a state vector, which is convenient for subsequent sequence modeling.

[0056] Specifically, the state projection module implements the projection operation through a fully connected layer or a convolutional layer. The fully connected layer linearly maps the flattened frame pixel values ​​to the state vector, while the convolutional layer extracts local features and compresses dimensions through a sliding window. During the projection process, a linear transformation is performed using a weight matrix and a bias term to output a state sequence with reduced dimensions.

[0057] For example, for a 5-second duration, resolution A video sequence of pixels, sampled at 30 frames per second, for a total of 150 frames; the state projection module resets each frame of the image to... The input size of a pixel is linearly transformed through a fully connected layer (input dimension 150528 pixels, output dimension 256), generating 150 256-dimensional state vectors to form the initial state sequence; the projection operation takes less than 10 milliseconds and runs on a GPU-accelerated device.

[0058] The initial state sequence is selectively scanned by the selective scanning module to generate an intermediate spatiotemporal feature sequence; the selective scanning process includes selective attention-weighted updates of the state sequence.

[0059] The selective scanning module refers to the component in the temporal video state space model that dynamically filters key state vectors. It focuses on important time steps through an attention mechanism to reduce redundant information. Selective scanning processing is a sequence optimization method based on attention weights, which performs weighted summation or filtering on the state sequence to enhance the contribution of key frames.

[0060] Specifically, the selective scanning module uses a multi-head self-attention mechanism or a gated recurrent unit (GRU) to perform scanning. The module first calculates the attention score of each state vector, then dynamically adjusts the weights according to the score, strengthening vectors with high scores and suppressing vectors with low scores. Finally, it outputs the weighted state sequence as an intermediate spatiotemporal feature.

[0061] For example, the initial input state sequence contains 150 256-dimensional vectors. The selective scanning module uses an 8-head self-attention mechanism to calculate the query, key, and value matrix for each vector. Attention weights are generated through the softmax function, with a weight threshold of 0.1. The top 100 key vectors with weights higher than the threshold are selected to generate an intermediate spatiotemporal feature sequence of dimension 100 multiplied by 256. The processing is completed on an edge computing device with a latency controlled within 20 milliseconds.

[0062] Therefore, according to the above implementation method, the system can efficiently compress video data, highlight key spatiotemporal features, provide optimized input for subsequent behavior recognition, and improve processing efficiency and accuracy.

[0063] In some embodiments, the temporal video state space model further includes a dynamic modeling unit and a linear complexity processing module; the step of capturing long-term spatiotemporal dependencies in intermediate spatiotemporal feature sequences through the spatiotemporal correlation modeling module and outputting low-redundancy temporal features includes: Spatiotemporal dynamic evolution modeling is performed on intermediate spatiotemporal feature sequences through dynamic modeling units to generate spatiotemporal evolution features; spatiotemporal dynamic evolution modeling includes modeling and analyzing the continuous evolution process of spatiotemporal relationships in video sequences.

[0064] Among them, the dynamic modeling unit refers to the component in the temporal video state space model that is responsible for simulating spatiotemporal dynamic changes. It describes the evolution of video sequences through state space equations or recursive networks to capture the continuous changes of behavior in time and space. Spatiotemporal dynamic evolution modeling is a method based on differential equations or temporal modeling, used to analyze inter-frame motion correlation and spatial structure evolution.

[0065] Specifically, the dynamic modeling unit is implemented using a recurrent neural network (RNN) or a state space model. It transmits spatiotemporal information through hidden states and applies gating mechanisms (such as the gating unit of LSTM) to optimize long-term dependency capture. The modeling process includes frame sequence analysis in the temporal dimension and feature map evolution in the spatial dimension.

[0066] For example, for an input intermediate spatiotemporal feature sequence with dimension 1 The dynamic modeling unit uses a long short-term memory (LSTM) network structure with a hidden layer dimension of 512. After processing, the sequence is converted into... The spatiotemporal evolution characteristics of key behaviors (such as arm swings) are preserved with a trajectory continuity rate of over 90%, and the processing latency is less than 30 milliseconds, running on edge computing devices; for example, the NVIDIA Jetson AGX Orin module, a high-performance edge artificial intelligence computing module, has an AI inference performance of 275 TOPS (trillion operations per second), integrates 2048 CUDA cores and 64 Tensor cores, supports parallel processing of multi-channel sensor data, and is designed for industrial-grade real-time visual analysis applications.

[0067] The spatiotemporal evolution features are transformed using a linear complexity processing module to output low-redundancy temporal features. The linear complexity transformation processing includes feature extraction and dimensionality reduction operations based on linear complexity.

[0068] The linear complexity processing module refers to the component in the temporal video state-space model that implements linear computational complexity. It reduces computational resource consumption through linear transformations or low-rank approximations, making it suitable for efficient processing of long video sequences. Linear complexity transformation processing is a computational complexity that is linearly related to the input size (i.e.,...). This method (with a complexity of 1) is different from traditional quadratic complexity operations (such as self-attention).

[0069] Specifically, the linear complexity processing module, based on linear attention or principal component analysis (PCA) techniques, maps high-dimensional features to a low-dimensional space while preserving key information through weight sharing or projection matrices. The module includes linear layers and activation functions (such as ReLU, Rectified Linear Unit, a commonly used neural network activation function, mathematically defined as...). That is, when the input x is greater than 0, the output is x, otherwise the output is 0), to achieve feature compression.

[0070] For example, if the input spatiotemporal evolution feature dimension is 64×64×200, the linear complexity processing module uses a linear projection layer (input dimension 409600, output dimension 204800) to reduce the dimensionality, and the output... It features low-redundancy temporal characteristics, reduces computation time by 60% compared to standard Transformer self-attention, retains over 85% of feature information, and runs on a GPU server, reducing power consumption by 40%.

[0071] Therefore, according to the above implementation method, the system can efficiently eliminate spatiotemporal redundancy, enhance the representation of key features, and achieve real-time behavior recognition in resource-constrained environments, thereby improving the accuracy and efficiency of safety monitoring of UHV substations.

[0072] In some embodiments, the video pose converter model is configured with an hourglass segmenter structure, which includes a label pruning and clustering module and a label recovery attention module. The steps of extracting 3D behavioral features from low-redundancy temporal features by performing hierarchical feature compression and reconstruction operations using the label pruning and recovery mechanism through the video pose converter model include: The label pruning clustering module performs dynamic label selection processing on low-redundancy temporal features to generate compressed label sequences.

[0073] Among them, dynamic label selection processing refers to the process of automatically screening labels with high semantic representativeness based on clustering algorithms. By evaluating the semantic similarity between labels, key information is retained while redundant labels are removed. Compressed label sequence refers to the label sequence whose length is significantly shortened after screening, while retaining the main semantic features of the original sequence.

[0074] Specifically, the label pruning clustering module uses the K-means clustering algorithm. First, it calculates the Euclidean distance of all label features, then iteratively updates the cluster centers, and finally selects the center label of each category as the representative label. The process includes three sub-steps: feature normalization, distance calculation, and center point selection.

[0075] For example, the input low-redundancy temporal features contain 200 512-dimensional labels, and the label pruning clustering module sets the number of clusters. The system generates 50 cluster center labels using K-means clustering, forming a compressed label sequence. The processing time is approximately 15 milliseconds, with a label compression rate of 75% and a key behavioral feature retention rate of over 92%. The module runs on an NVIDIA RTX 3090 GPU, reducing memory usage to 40% of the original sequence.

[0076] The compressed label sequence is reconstructed using the label recovery attention module, resulting in a full-resolution label sequence.

[0077] Sequence reconstruction processing refers to the operation of restoring the compressed sequence to its original length through an attention mechanism, and reconstructing the complete sequence information by utilizing the correlation between the tags; the full-resolution tag sequence refers to the output sequence with the same number of original input tags, which contains detailed information before compression but with lower computational cost.

[0078] Specifically, the label recovery attention module adopts a multi-head self-attention mechanism, which calculates attention weights through query, key, and value matrices to perform weighted interpolation reconstruction on compressed labels; the module includes three core steps: setting the number of attention heads, weight calculation, and feature fusion.

[0079] For example, the input compressed label sequence contains 50 512-dimensional labels. The label recovery attention module uses an 8-head attention mechanism to map the labels to a 4096-dimensional latent space through a linear layer, and then backprojects them back to 512 dimensions, outputting a sequence of 200 512-dimensional full-resolution labels. The reconstruction error is less than 3%, the processing latency is 20 milliseconds, and the computational complexity is reduced to 35% of that of traditional methods while maintaining accuracy.

[0080] Three-dimensional behavioral features are extracted from the full-resolution labeled sequence.

[0081] Among them, three-dimensional behavioral characteristics refer to quantitative data describing the movement trajectory of human joints in three-dimensional space. They are usually represented in the form of joint coordinate sequences and are used to accurately characterize the movement patterns of maintenance personnel.

[0082] Specifically, feature extraction is achieved through a regression head module, which consists of a fully connected layer and an activation function, mapping the labeled sequence to the key coordinate space. The extraction process includes three steps: feature linear transformation, coordinate regression, and sequence formatting.

[0083] For example, the input full-resolution labeled sequence contains 200 512-dimensional labels. The regression head module converts each label into 3D coordinates of 17 key points (51 dimensions in total) through three fully connected layers (with 256, 128, and 51 neurons respectively), and outputs 200 frames of 3D behavioral features with 17 points and 3 coordinates. The coordinate error is less than 5 millimeters, and the overall processing flow achieves an end-to-end latency of less than 100 milliseconds on the edge computing device, meeting the requirements for real-time monitoring.

[0084] Therefore, according to the above implementation method, the system can efficiently process video features through an hourglass-style compression-reconstruction mechanism, improving computational efficiency while ensuring the accuracy of three-dimensional behavioral features, and providing reliable technical support for behavioral recognition of UHV substation operation and maintenance personnel.

[0085] In some embodiments, the video pose changer model is further configured with a regression head module; the step of extracting three-dimensional behavioral features from the full-resolution marker sequence includes: The regression head module performs 3D pose regression processing on the full-resolution marker sequence to generate a 3D pose coordinate sequence. The regression head module implements 3D pose regression processing through a fully connected layer, which includes a linear transformation operation that maps the marker sequence to 3D spatial coordinates.

[0086] Among them, 3D pose regression processing refers to the mathematical transformation process of converting the spatiotemporal feature information in the labeled sequence into 3D spatial coordinates, and establishing the correspondence between feature vectors and key coordinates through linear mapping; the regression head module refers to the output layer component in the video pose transformer model that is responsible for coordinate regression, which realizes the dimensionality reduction mapping from high-dimensional features to low-dimensional coordinates through a fully connected neural network.

[0087] Specifically, the regression head module adopts a three-layer fully connected layer structure. The first layer maps the 512-dimensional labeled features to 256 dimensions, the second layer maps to 128 dimensions, and the third layer outputs 51-dimensional coordinate data (corresponding to the three-dimensional coordinates of 17 key points). After each fully connected layer, a ReLU (Rectified Linear Unit) activation function is used for nonlinear transformation. Finally, the output layer uses a linear activation function to directly regress the coordinate values.

[0088] For example, the input full-resolution labeled sequence contains 200 512-dimensional labels. The regression head module processes the data sequentially through three fully connected layers (with 256, 128, and 51 neurons respectively), ultimately generating a 200-frame × 17 keypoints × 3 coordinates (x, y, z) three-dimensional pose coordinate sequence. The single-frame coordinate regression takes 0.5 milliseconds, and the entire sequence processing is completed on an NVIDIA JTX4090 GPU with an end-to-end latency of less than 100 milliseconds and a coordinate error of less than 5 millimeters.

[0089] Three-dimensional behavioral features are constructed based on three-dimensional attitude coordinate sequences.

[0090] Among them, the three-dimensional attitude coordinate sequence includes the position information of the joints of the operation and maintenance personnel in the spatiotemporal dimension; the three-dimensional behavioral features refer to the structured feature representation obtained after spatiotemporal encoding of the coordinate sequence, which is used to describe the spatiotemporal evolution law of continuous actions.

[0091] Specifically, the construction process includes two steps: spatiotemporal normalization and feature tensor reorganization. First, the coordinate sequence is normalized to the bone length to eliminate individual body shape differences. Then, the normalized coordinate sequence is reorganized into a tensor form according to the time dimension, forming a feature tensor with a dimension of frame number × number of joints × 3.

[0092] For example, the coordinate sequence of 200 frames (17 joints per frame) is normalized by bone length and then reconstructed into... The three-dimensional behavioral feature tensor is used; this feature tensor can be directly input into the behavior classification network, and the classification accuracy reaches 96.3% on the UHV substation operation and maintenance behavior dataset, and the abnormal action detection recall rate is 92.7%.

[0093] Therefore, according to the above implementation method, the system can achieve accurate three-dimensional pose coordinate extraction through the regression head module and construct behavioral features with spatiotemporal continuity, providing a high-precision data foundation for subsequent behavior recognition.

[0094] In some embodiments, the behavior category of the maintenance personnel is identified based on three-dimensional behavioral features, including: Input the three-dimensional behavioral features into the pre-configured behavior classification module.

[0095] Among them, the behavior classification module refers to a classifier component based on a deep learning architecture, which is used to perform pattern recognition and category classification on three-dimensional behavior features. It achieves non-linear mapping from features to categories through a multi-layer neural network. Specifically, this module adopts an architecture that combines a fully connected layer with a softmax function (normalized exponential function). The input layer dimension matches the dimension of the three-dimensional behavioral features, and the output layer dimension matches the number of behavioral categories. For example, the input three-dimensional behavioral feature dimension is ( After being flattened, the input is fed into a hidden layer with 512 neurons, and finally outputs the probability distribution of 5 behavior categories through the softmax function, with a module inference latency of less than 10 milliseconds.

[0096] The behavior classification module performs behavior pattern analysis on the 3D behavior features to generate behavior feature vectors; the behavior pattern analysis includes feature extraction of the spatiotemporal patterns of the 3D pose sequence.

[0097] The behavioral feature vector includes spatiotemporal characteristics information of the actions of the maintenance personnel.

[0098] Specifically, the behavioral pattern analysis and processing adopts a combination of temporal convolutional network (TCN) and global pooling. First, the temporal dependencies are captured by TCN, and then a fixed-dimensional feature vector is generated by global average pooling. For example, the input three-dimensional behavioral features are first extracted by TCN (convolution kernel size 3, number of channels 64), and then a 256-dimensional behavioral feature vector is generated by global average pooling. This vector contains spatiotemporal characteristic information such as action speed and joint angle change rate, and the feature discrimination is over 90%.

[0099] Perform category determination processing on the behavioral feature vectors and output the behavioral category of the maintenance personnel; Among them, category determination processing refers to the process of making category assignment decisions based on the classifier for feature vectors, and achieving classification by comparing the similarity between the feature vectors and the prototypes of each category; Specifically, a support vector machine (SVM) is used as the classifier, and a radial basis function (RBF) is used as the kernel function to achieve multi-class classification through a one-vs-one strategy. For example, inputting a 256-dimensional behavioral feature vector into an SVM classifier and setting the RBF kernel parameters. Penalty coefficient It can classify 5 types of behavior (normal walking, equipment operation, abnormal running, dangerous climbing, and falling) with an accuracy rate of 95.3% and a single classification time of 2 milliseconds.

[0100] Therefore, according to the above implementation method, the system can achieve end-to-end recognition from raw video to behavior category through a hierarchical processing architecture, ensuring accuracy while meeting real-time requirements, and providing reliable technical support for the safety monitoring of UHV substations.

[0101] Furthermore, in other embodiments, the VPT architecture is as follows: Figure 5 As shown, it achieves accurate estimation of 3D pose from 2D pose sequences through a hierarchical processing flow: First, the input 2D pose sequence (such as the coordinates of 17 human joints) is input into the pose embedding module for spatiotemporal encoding, generating feature representations with spatiotemporal correlation; then, multi-level feature transformation is performed through four cascaded Transformer blocks, each containing a multi-head self-attention mechanism and a feedforward neural network to capture long-range dependencies of joints; finally, the regression head module outputs two forms of 3D pose results—when the sequence-to-sequence mode is selected, the complete 3D pose sequence is output, suitable for continuous motion analysis; when the sequence-to-frame mode is selected, only the 3D pose of the center frame is output, suitable for precise keyframe localization. This architecture is used in the operation and maintenance behavior analysis of UHV substations for... The processing time for a 200-frame sequence in the pixel video is only 30 milliseconds, and the 3D pose coordinate error is less than 4 millimeters, effectively supporting the subsequent behavior classification module's accuracy of 95.2%.

[0102] In other embodiments, the HoT model reduces computational load through token pruning, improving efficiency by 40% compared to the VPT model, specifically as follows: Figure 6 The comparison between the PT (Pose Transformer) and HoT models is shown; among them, the VPT model ( Figure 6 a) It employs a direct full pose labeling method, where all input labels participate in depth feature extraction to generate the output. In contrast, the HoT model proposed in this invention ( Figure 6b) introduces an innovative token management mechanism: First, a "token pruning" step intelligently filters and removes pose markers in the input sequence that have low contribution to the current task or are redundant, reducing the number of markers that the subsequent core Transformer encoder needs to process, thereby reducing computational complexity and memory consumption. After the core processing completes the extraction and transformation of key features, the "token restoration" module is used to restore the sequence to the original number of markers, ensuring the integrity and accuracy of the output results. This "slimming down first, then restoring" strategy allows the HoT model to maintain powerful representation capabilities while significantly improving energy efficiency.

[0103] Specifically, in the scenario of behavior analysis of UHV substation operation and maintenance personnel, when the input sequence contains 200 attitude markers, the HoT model dynamically selects 80 of the most representative key markers (compression rate of 60%) using the K-means clustering algorithm. After processing by the Transformer encoder, attention-weighted interpolation is used to restore the sequence to 200 markers. Real-world testing data shows that this processing flow reduces the inference latency on the NVIDIA Jetson AGX Orin module from 85 milliseconds in the traditional VPT model to 42 milliseconds, reduces memory usage by 55%, and maintains 96.3% 3D pose estimation accuracy. This efficient token management mechanism is particularly suitable for security scenarios in UHV substations requiring long-term behavior analysis, supporting real-time processing of video sequences up to 10 minutes long, providing reliable technical support for the safety monitoring of operation and maintenance personnel.

[0104] Figure 7 This is a structural block diagram of a behavior recognition system for operation and maintenance personnel in an ultra-high voltage substation according to an embodiment of the present invention.

[0105] like Figure 7 As shown, the operation and maintenance personnel behavior recognition system of this UHV substation includes: The video data acquisition module 210 is used to acquire time-series video data from UHV substation operation and maintenance personnel.

[0106] The temporal feature extraction module 220 is used to extract low-redundancy temporal features from temporal video data through a preset temporal video state space model.

[0107] The 3D feature extraction module 230 is used to input low-redundancy temporal features into a pre-configured video pose transformer model, and perform hierarchical feature compression and reconstruction operations through the video pose transformer model using a tag pruning and restoration mechanism to extract 3D behavioral features from the low-redundancy temporal features.

[0108] The behavior category recognition module 240 is used to identify the behavior category of the maintenance personnel based on the three-dimensional behavioral characteristics.

[0109] The specific functions and examples of each module and submodule of the device in this embodiment of the invention can be found in the relevant descriptions of the corresponding steps in the above method embodiments, and will not be repeated here.

[0110] According to embodiments of the present invention, the above-described method of the present invention can be applied to an electronic device and a readable storage medium.

[0111] Figure 8 A schematic block diagram of an example electronic device 600 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0112] like Figure 8 As shown, the electronic device 600 includes a computing unit 601, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. The RAM 603 may also store various programs and data required for the operation of the electronic device 600. The computing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0113] Multiple components in electronic device 600 are connected to I / O interface 605, including: input unit 606, such as keyboard, mouse, etc.; output unit 607, such as various types of displays, speakers, etc.; storage unit 608, such as disk, optical disk, etc.; and communication unit 609, such as network card, modem, wireless transceiver, etc. Communication unit 609 allows electronic device 600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0114] The computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as a method for recognizing the behavior of operation and maintenance personnel in an ultra-high voltage substation. For example, in some embodiments, a method for recognizing the behavior of operation and maintenance personnel in an ultra-high voltage substation can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps of the method for recognizing the behavior of operation and maintenance personnel in an ultra-high voltage substation described above can be performed. Alternatively, in other embodiments, the computing unit 601 may be configured by any other suitable means (e.g., by means of firmware) to perform a method for recognizing the behavior of operation and maintenance personnel in an ultra-high voltage substation.

[0115] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0116] The program code used to implement the methods of the present invention can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a standalone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0117] In the context of this invention, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0118] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0119] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0120] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0121] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this invention can be achieved, and this is not limited herein.

[0122] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for identifying the behavior of an operator of an ultra-high voltage substation, characterized in that, The method comprises: acquiring time sequence video data from an ultra-high voltage substation operation and maintenance personnel; extracting low-redundancy time sequence features from the time sequence video data through a preset time sequence video state space model; inputting the low-redundancy time sequence features into a pre-configured video posture transformer model, performing hierarchical feature compression and reconstruction operations of a label pruning and recovery mechanism through the video posture transformer model, and extracting three-dimensional behavior features from the low-redundancy time sequence features; identifying a behavior category of the operation and maintenance personnel according to the three-dimensional behavior features.

2. The method of claim 1, wherein, The time sequence video state space model comprises a selective state space modeling module and a space-time correlation modeling module; the step of extracting low-redundancy time sequence features from the time sequence video data through a preset time sequence video state space model comprises: performing state space transformation processing on the time sequence video data through the selective state space modeling module to generate an intermediate space-time feature sequence; capturing long-term space-time dependency relationships in the intermediate space-time feature sequence through the space-time correlation modeling module to output the low-redundancy time sequence features.

3. The method of claim 2, wherein, The time sequence video state space model further comprises a state projection module and a selective scanning module; the step of performing state space transformation processing on the time sequence video data through the selective state space modeling module to generate an intermediate space-time feature sequence comprises: performing state space projection processing on the time sequence video data through the state projection module to generate an initial state sequence; performing selective scanning processing on the initial state sequence through the selective scanning module to generate the intermediate space-time feature sequence; wherein the state space projection processing comprises a linear transformation operation of converting a video frame sequence into a state vector, and the selective scanning processing comprises selective attention weighting updating on a state sequence.

4. The method of claim 3, wherein, The time sequence video state space model further comprises a dynamic modeling unit and a linear complexity processing module; the step of capturing long-term space-time dependency relationships in the intermediate space-time feature sequence through the space-time correlation modeling module to output the low-redundancy time sequence features comprises: performing space-time dynamic evolution modeling on the intermediate space-time feature sequence through the dynamic modeling unit to generate space-time evolution features; performing linear complexity transformation processing on the space-time evolution features through the linear complexity processing module to output the low-redundancy time sequence features; wherein the space-time dynamic evolution modeling comprises modeling and analyzing a continuous evolution process of space-time relationships in a video sequence, and the linear complexity transformation processing comprises feature extraction and dimension reduction operations based on linear complexity.

5. The method of claim 1, wherein, The video posture transformer model is configured with an hourglass tokenizer structure, the hourglass tokenizer structure comprising the label pruning clustering module and the label recovery attention module; the step of performing hierarchical feature compression and reconstruction operations of a label pruning and recovery mechanism through the video posture transformer model to extract three-dimensional behavior features from the low-redundancy time sequence features comprises: performing dynamic label selection processing on the low-redundancy time sequence features through the label pruning clustering module to generate a compressed label sequence; The attention restoration module recovers the compressed marker sequence by sequence reconstruction processing, and outputs a complete resolution marker sequence; The three-dimensional behavior feature is extracted from the complete resolution marker sequence.

6. The method of claim 5, wherein, The video pose transformer model is also configured with a regression head module; the step of extracting the three-dimensional behavior feature from the complete resolution marker sequence includes: The regression head module performs three-dimensional pose regression processing on the complete resolution marker sequence to generate a three-dimensional pose coordinate sequence; The three-dimensional behavior feature is constructed based on the three-dimensional pose coordinate sequence; The regression head module implements the three-dimensional pose regression processing through a fully connected layer, the three-dimensional pose regression processing includes a linear transformation operation of mapping the marker sequence to a three-dimensional space coordinate, and the three-dimensional pose coordinate sequence includes position information of the operator joint in the time-space dimension.

7. The method of claim 1, wherein, The behavior classification module is pre-configured; The behavior classification module performs behavior pattern analysis processing on the three-dimensional behavior feature to generate a behavior feature vector; The behavior feature vector is subjected to category determination processing to output the behavior category of the operator; The behavior pattern analysis processing includes feature extraction of the time-space pattern of the three-dimensional pose sequence, the category determination processing includes a classification operation of mapping the feature vector to a preset behavior category, and the behavior feature vector includes time-space characteristic information of the operator action. The video data acquisition module is configured to acquire time-series video data from a UHV substation operator; 8. An operating and maintenance personnel behavior identification system of an extra-high voltage substation, characterized in that, The time-series feature extraction module is configured to extract low-redundancy time-series features from the time-series video data through a preset time-series video state space model; The three-dimensional feature extraction module is configured to input the low-redundancy time-series features into a pre-configured video pose transformer model, perform hierarchical feature compression and reconstruction operations of a marker pruning and recovery mechanism through the video pose transformer model, and extract three-dimensional behavior features from the low-redundancy time-series features; The behavior classification module is configured to identify the behavior category of the operator according to the three-dimensional behavior feature.

9. An electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-7. Computer instructions for causing a computer to perform the method of any one of claims 1-7. Computer instructions for causing a computer to perform the method of any one of claims 1-7.

10. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, ​

Citation Information

Patent Citations

  • User behavior identification method and system for intelligent video conference

    CN116631050A