Operation maintenance personnel behavior analysis method and system of extra-high voltage transformer substation, electronic equipment and medium
By employing feature alignment and attention fusion methods, combined with a temporal video state space model and a video pose transformer model, the heterogeneity problem of multimodal data in UHV substations was solved, enabling efficient and accurate identification of operation and maintenance personnel behavior and precise early warning of safety risks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-28
- Publication Date
- 2026-03-31
AI Technical Summary
Existing technologies in UHV substations suffer from insufficient effectiveness in multimodal data fusion, limited accuracy in behavior recognition, and low levels of real-time and intelligent safety early warning, making it difficult to meet the needs for efficient monitoring of operation and maintenance personnel behavior and accurate early warning of safety risks in high-risk operation scenarios.
By acquiring multimodal data, performing feature alignment and attention fusion, and utilizing temporal video state space models and video pose transformer models to extract temporal three-dimensional behavioral features, efficient identification of operation and maintenance personnel behavior can be achieved.
It improves the accuracy, robustness, and practicality of operation and maintenance personnel behavior analysis, enabling precise and efficient identification of operation and maintenance personnel behavior, reducing computational load, and improving identification accuracy and processing speed.
Smart Images

Figure CN121765302A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of safety monitoring technology for ultra-high voltage substations, and in particular to a method, system, electronic equipment, and medium for analyzing the behavior of operation and maintenance personnel in ultra-high voltage substations. Background Technology
[0002] With the rapid development of the power industry and the continuous expansion of the power grid, as of January 2024, there were over 55,000 substations of 35 kV and above in China, placing higher demands on the operational safety and behavior monitoring of power operation and maintenance personnel. Ultra-high voltage (UHV) substation operation and maintenance is characterized by multi-person collaboration, frequent high-altitude cross-operations, high personal danger, and tight work schedules. The safety of operation and maintenance work is directly related to the stable operation of the power grid. Therefore, building an efficient system for analyzing the behavior of operation and maintenance personnel and for safety early warning has become a core task for ensuring the safety of the power system.
[0003] Intelligent technologies have been gradually applied to the power operation and maintenance field. The combination of intelligent monitoring systems and artificial intelligence algorithms provides technical support for real-time monitoring of work sites and identification of risk factors, effectively reducing reliance on manpower. Currently, related technologies are mainly developing around three core directions: multimodal data processing, personnel behavior recognition, and safety early warning. On the one hand, there are intelligent monitoring technologies already applied in the power industry, which collect video, environmental, and other data through various sensing devices and combine them with traditional computer vision algorithms to achieve behavior monitoring. Some solutions attempt to integrate multi-source data for risk assessment. On the other hand, there is the application of multimodal analysis technologies from fields such as security and industrial inspection, combined with relevant deep learning models, to dynamically identify and make simple predictions of personnel actions, and adapt them to substation application scenarios.
[0004] The prior art (application publication number: CN120219903A) discloses a video surveillance early warning method and system based on multimodal behavior patterns. It extracts multimodal data from video streams, aligns and fuses it, constructs a behavior pattern library and scene behavior baseline, updates the pattern library through online learning, and adjusts the anomaly detection threshold to achieve early warning functionality. However, the technical solution in this prior art mainly focuses on multimodal data processing related to video streams. Its data fusion method and behavior analysis logic emphasize short-term behavior recognition in general monitoring scenarios, failing to fully adapt to the complex scenario characteristics of UHV substation operation and maintenance, and not effectively solving key issues such as multimodal data heterogeneity, long-term behavioral feature capture, and intelligent early warning mechanisms.
[0005] Existing technologies still have significant limitations in practical applications. First, the efficiency of comprehensive utilization of multi-source and multi-modal data is not high. Differences in the format and semantics of different types of data make it difficult to fully explore complementary information between data, and cannot comprehensively reflect the relationship between the behavior of operation and maintenance personnel and the status of the environment and equipment. Second, the accuracy of behavior recognition is insufficient, making it difficult to effectively cope with complex scenarios such as personnel obstruction, changes in lighting, and dense equipment in UHV substations. The ability to capture dynamic behavioral characteristics during long-term operations is limited, resulting in insufficient reliability of recognition results. Finally, the real-time performance and intelligence level of safety warnings are low. There is a lack of a comprehensive evaluation mechanism that integrates multi-dimensional information and historical data, which easily leads to problems such as delayed warnings, high false alarm rates, or weak targeting, making it difficult to predict potential dangers in advance.
[0006] In summary, existing technologies for analyzing the behavior of UHV substation operation and maintenance personnel and providing safety early warnings face a combination of technical challenges, including insufficient effectiveness of multimodal data fusion, limited accuracy of behavior recognition, and low levels of real-time performance and intelligence in safety early warnings. These technologies are insufficient to meet the actual needs of efficient monitoring of operation and maintenance personnel behavior and accurate early warning of safety risks in high-risk operation scenarios at UHV substations. Summary of the Invention
[0007] To address the aforementioned shortcomings or deficiencies, this invention provides a method, system, electronic device, and medium for analyzing the behavior of operation and maintenance personnel in ultra-high voltage substations. This solves the technical problem that existing technologies cannot meet the requirements for efficient monitoring of operation and maintenance personnel behavior and accurate early warning of safety risks in high-risk operation scenarios of ultra-high voltage substations.
[0008] This invention provides a method for analyzing the behavior of operation and maintenance personnel in ultra-high voltage substations, including: Acquire multimodal data in the operation and maintenance scenarios of ultra-high voltage substations.
[0009] The multimodal data is sequentially processed by feature alignment and attention fusion to generate fused multimodal features.
[0010] The fused multimodal features are input into a pre-configured temporal video state space model and a video pose transformer model. Selective state space modeling and long-term spatiotemporal dependency capture are performed through the temporal video state space model, and hierarchical feature compression and reconstruction operations are performed through the video pose transformer model to extract temporal 3D behavioral features from the fused multimodal features.
[0011] The behavior categories of operation and maintenance personnel are identified based on the temporal three-dimensional behavioral characteristics.
[0012] According to a second aspect, this invention provides a behavior analysis system for operation and maintenance personnel in ultra-high voltage substations, comprising: The multimodal data acquisition module is used to acquire multimodal data in the operation and maintenance scenarios of UHV substations.
[0013] The multimodal feature generation module is used to sequentially perform feature alignment and attention fusion processing on multimodal data to generate fused multimodal features.
[0014] The 3D behavioral feature analysis module is used to input fused multimodal features into a pre-configured temporal video state space model and a video pose transformer model. It performs selective state space modeling and long-term spatiotemporal dependency capture through the temporal video state space model, and performs hierarchical feature compression and reconstruction operations through the video pose transformer model to extract temporal 3D behavioral features from the fused multimodal features.
[0015] The personnel behavior type identification module is used to identify the behavior category of operation and maintenance personnel based on time-series three-dimensional behavioral characteristics.
[0016] According to a third aspect, the present invention provides an electronic device comprising: At least one processor; and The memory that is communicatively connected to the at least one processor; The memory stores instructions that can be executed by the at least one processor, which enables the at least one processor to execute the operation and maintenance personnel behavior analysis method for any UHV substation in the embodiments of the present invention.
[0017] According to another aspect of the present invention, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute the operation and maintenance personnel behavior analysis method of any UHV substation in the embodiments of the present invention.
[0018] This invention provides a method for analyzing the behavior of operation and maintenance personnel in ultra-high voltage (UHV) substations. This method is achieved through four core steps: multimodal data acquisition, feature alignment and fusion, temporal three-dimensional behavioral feature extraction, and behavior recognition. Multimodal data comprehensively characterizes the operation and maintenance scenario; feature alignment and fusion address the integration of multi-source heterogeneous data; and a temporal video state space model and a video attitude transformer model extract spatiotemporally consistent behavioral features from the fused features. The method includes: first, simultaneously acquiring multimodal data from the operation and maintenance scenario using multiple acquisition devices deployed in the UHV substation; then, sequentially performing feature alignment and attention fusion processing on the multimodal data to generate unified fused multimodal features; next, inputting the fused multimodal features into a pre-configured temporal video state space model and a video attitude transformer model; performing selective state space modeling and long-term spatiotemporal dependency capture through the temporal video state space model, and performing hierarchical feature compression and reconstruction operations through the video attitude transformer model to extract temporal three-dimensional behavioral features; finally, identifying the behavioral category of the operation and maintenance personnel based on the temporal three-dimensional behavioral features.
[0019] In this technical solution, the present invention addresses the multimodal data heterogeneity problem described in the background technology by unifying the feature dimensions and sampling rates of different modal data through feature alignment processing, thus overcoming the shortcomings of traditional methods in effectively aligning multi-source data. Regarding the problem of redundant information interference in long-term video, the invention utilizes a selective state-space modeling mechanism in the temporal video state-space model to filter key spatiotemporal features and capture long-term dependencies. Furthermore, addressing the issue of low recognition accuracy due to insufficient extraction of 3D behavioral features, the invention employs hierarchical feature compression and reconstruction operations in the video pose converter model to achieve effective mapping from 2D image sequences to 3D pose sequences. Therefore, the technical solution of this invention solves the technical problem of existing technologies failing to meet the requirements of efficient monitoring of operation and maintenance personnel behavior and accurate early warning of safety risks in high-risk operation scenarios of UHV substations, achieving accurate and efficient identification of operation and maintenance personnel behavior and improving the accuracy, robustness, and practicality of behavior analysis. Attached Figure Description
[0020] Figure 1 This is a block diagram of the overall implementation scheme of the present invention; Figure 2 This is a structural diagram of the feature alignment sampling module of the present invention; Figure 3 This is a diagram illustrating the automated detection and early warning system for safety hazards according to the present invention. Figure 4 This is a flowchart of a method for analyzing the behavior of operation and maintenance personnel in an ultra-high voltage substation according to an embodiment of the present invention; Figure 5 This is a structural diagram of the attention mechanism feature fusion module according to an embodiment of the present invention; Figure 6 This is a schematic diagram illustrating the operation steps of intra-frame static attention and inter-frame dynamic attention according to an embodiment of the present invention. Figure 7 This is a schematic diagram of the structure of a behavior monitoring model with multi-dimensional feature fusion according to an embodiment of the present invention; Figure 8 This is a schematic diagram of the operation steps of a multi-level state data monitoring model based on LSTM fusion according to an embodiment of the present invention; Figure 9 This is a schematic diagram of the structure of an ultra-high voltage substation operation and maintenance personnel behavior analysis system according to an embodiment of the present invention; Figure 10 This is a block diagram of an electronic device used to implement embodiments of the present invention. Detailed Implementation
[0021] The following description, in conjunction with the accompanying drawings, illustrates exemplary embodiments of the present invention, including various details to aid understanding. These details should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the invention. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0022] During the development of this invention, the inventors, through extensive experiments and data analysis, revealed the intrinsic relationship between multimodal data redundancy and feature dilution: direct fusion of multi-source heterogeneous data not only increases the computational burden but also dilutes the representational strength of key behavioral features, making it difficult for traditional methods to balance computational efficiency and recognition accuracy. Based on this relationship, the inventors innovatively proposed this technical solution, utilizing a synergistic mechanism of feature alignment and attention fusion. By selectively modeling the state space of a temporal video state space model to filter spatiotemporal redundant information, and combining hierarchical feature compression and reconstruction operations of a video pose converter model to enhance the extraction of key behavioral features, this solution achieves a 40% reduction in computational load while increasing behavior recognition accuracy to over 90%, embodying the core concept of "trading feature quality for computational efficiency."
[0023] Specifically, such as Figure 1 As shown, the invention team discovered through comparative experiments that traditional multimodal fusion methods suffer from issues such as feature channel mismatch and inconsistent sampling rates when dealing with complex scenarios in UHV substations. These technical deficiencies lead to the loss of cross-modal semantic relevance. However, the Feature Alignment Module (FCU) proposed in this invention can effectively unify the feature dimensions of multi-source data and reduce information loss during heterogeneous data fusion.
[0024] Furthermore, such as Figure 2 As shown, the attention weight distribution analysis demonstrates that dynamic enhancement of key modal features can be achieved using only a cross-attention mechanism. This adaptive weighting relationship between modalities provides a theoretical basis for designing an attention-based feature fusion (AFF) module. Experimental data confirms that the proposed solution improves the accuracy of behavior analysis to over 90% in real-world scenarios of UHV substations, while achieving a processing speed 2.5 times faster than traditional methods. This breakthrough performance improvement is based on a deep understanding of the heterogeneous nature of multimodal data and an innovative "alignment-fusion-extraction" technical architecture design.
[0025] Therefore, this invention provides a method for analyzing the behavior of operation and maintenance personnel in ultra-high voltage substations, based on the first aspect. This method can be applied to an intelligent operation and maintenance safety management system for ultra-high voltage substations (hereinafter referred to as the "system"). The system can be deployed locally or run on the substation's on-site industrial control computer and edge computing nodes in a cloud-edge collaborative manner to complete the entire process analysis from multimodal data acquisition to behavior recognition and early warning.
[0026] Specifically, the system's physical equipment includes, but is not limited to, multimodal acquisition devices, edge computing units, and central processing servers. These devices must possess high-precision data acquisition, real-time processing capabilities, and electromagnetic interference resistance to support multi-dimensional behavioral analysis in complex environments and ensure the system can operate stably under harsh conditions such as strong electromagnetic interference and large temperature and humidity variations in UHV substations.
[0027] In other embodiments, such as Figure 3 As shown, the system hardware configuration comprises three layers: the acquisition layer is equipped with explosion-proof high-definition cameras, depth vision sensors, directional microphone arrays, and multi-parameter environmental sensors; the edge layer deploys an industrial-grade edge computing gateway, providing data preprocessing and local analysis capabilities; and the platform layer employs a highly available server cluster to run multimodal fusion analysis algorithms. This layered architecture ensures both the real-time nature of data acquisition and the efficient execution of complex algorithms.
[0028] In specific application scenarios, when maintenance personnel are working in different areas of a substation (as shown in the attached diagram, personnel are inspecting or climbing equipment areas), the multimodal sensors at the acquisition layer can simultaneously acquire panoramic data of the work site: explosion-proof high-definition cameras acquire 1080-pixel resolution video streams at a rate of 25 frames per second; depth vision sensors generate point cloud data in real time to construct a 3D spatial model; directional microphone arrays capture sound wave signals from specific directions; and multi-parameter environmental sensors continuously monitor environmental indicators such as temperature, humidity, and sulfur hexafluoride gas concentration. This multi-source data is transmitted to the edge layer computing gateway via industrial Ethernet. The gateway's built-in time synchronization module (using the IEEE 1588 precise time protocol) controls the timestamp deviation of the heterogeneous data to within milliseconds and completes data preprocessing through a lightweight feature extraction algorithm.
[0029] After receiving feature data processed by the edge layer, the server cluster at the platform layer uses digital twin technology to construct a virtual model that fully corresponds to the physical substation. Through multimodal fusion analysis algorithms, closed-loop management of behavior recognition is achieved: when the system detects personnel entering a live hazardous area, it can complete the entire response process from data acquisition to early warning push within 200 milliseconds; when it identifies non-standard operating behaviors (such as failure to maintain a safe distance or failure to use insulated tools), the system automatically generates a risk assessment report and pushes it to the safety management platform. This layered and collaborative architecture design enables the system to maintain 99.9% communication reliability even in the complex electromagnetic environment of the substation, providing comprehensive technical support for the safety of operation and maintenance work.
[0030] like Figure 4 As shown, the method may include: Step S110: Obtain multimodal data under the operation and maintenance scenario of UHV substation.
[0031] Multimodal data refers to a complementary heterogeneous data set that is synchronously acquired through different types of acquisition devices, including video data, audio data, sensor data, location data, and environmental data, which is used to comprehensively characterize the diverse features of operation and maintenance scenarios.
[0032] Specifically, the system can synchronously acquire multi-source data through a cluster of acquisition devices deployed in key areas of the substation: video acquisition devices record the visual behavior information of operation and maintenance personnel, audio acquisition devices capture environmental sound characteristics, sensor acquisition devices monitor the physiological state of personnel, positioning devices track spatial location trajectories, and environmental monitoring devices collect operating parameters such as temperature and humidity.
[0033] For example, the system employs an explosion-proof infrared camera (2560×1440 pixels resolution) to capture video streams at 30 frames per second, a directional microphone array to capture audio at a sampling rate of 48 kHz, an inertial measurement unit (IMU) to output nine-axis motion data at a frequency of 100 Hz, an ultra-wideband positioning base station (accuracy ±10 cm) to update personnel coordinates in real time, and a temperature and humidity sensor (range -40 degrees Celsius to 85 degrees Celsius) to report environmental parameters every minute. All devices use a time synchronization protocol (such as IEEE 1588) to control timestamp deviations within 50 milliseconds, and aggregate data to edge computing nodes via an industrial switch.
[0034] Step S120: Perform feature alignment and attention fusion processing on the multimodal data in sequence to generate fused multimodal features.
[0035] Feature alignment processing refers to the operation of unifying the feature dimensions and sampling rates of data from different modalities through the Feature Alignment Module (FAM), while attention fusion processing refers to the operation of cross-modal feature weighted fusion through the Attention-based Feature Fusion Module (AFFM).
[0036] Specifically, the feature alignment module first unifies the number of feature channels of each modality to 512 dimensions through a Full Connection Layer (FC), and then normalizes the feature size to 64×64 through sampling operations (upsampling uses bilinear interpolation, downsampling uses average pooling). The attention fusion module then adopts a cross-attention mechanism, using video features as queries and sensor features as key-value pairs to calculate the intermodal association weights.
[0037] For example, for a 5-second multimodal sequence, the feature alignment module upscales the audio features (original 128 dimensions) to 512 dimensions and reduces the video features (original 1024 dimensions) to 512 dimensions, generating a unified sequence with a time step of 150 frames after sampling. The attention fusion module calculates modal weights through an 8-head attention mechanism, assigning a weight of 0.8 or higher to the sensor features of key action frames (such as the moment of tool operation) and a weight of 0.3 or lower to static scenes, ultimately generating a 512-dimensional fused feature vector.
[0038] Step S130: Input the fused multimodal features into the pre-configured temporal video state space model and video pose transformer model, perform selective state space modeling and long-term spatiotemporal dependency capture through the temporal video state space model, and perform hierarchical feature compression and reconstruction operations through the video pose transformer model to extract temporal three-dimensional behavioral features from the fused multimodal features.
[0039] Among them, temporal three-dimensional behavioral characteristics refer to human motion representation data that includes time, space and joint dimensions, used to describe the continuous motion trajectory of the joints of maintenance personnel in three-dimensional space.
[0040] Specifically, the temporal video state space model filters key video frames through a selective scanning mechanism, retaining less than 40% of high-information frames; the video pose converter model first compresses the feature sequence to one-quarter of its original length through an hourglass word segmenter structure, and then restores the full resolution through attention reconstruction.
[0041] For example, inputting 10 seconds of fused features (300 frames), the temporal video state space model outputs 112 frames of low-redundancy features. The video pose transformer model further compresses these features to 28 representative markers, which are then reconstructed to generate a 300-frame sequence of 17 key points with a coordinate accuracy error of ±3 cm. The entire process takes 1.2 seconds on an NVIDIA RTX 3080 graphics card (10 gigabytes of video memory), which is 2.8 times faster than traditional methods.
[0042] Step S140: Identify the behavior category of the operation and maintenance personnel based on the temporal three-dimensional behavioral features.
[0043] Among them, the behavior category refers to the classification of operational behaviors predefined in the substation safety regulations, including three categories and 12 subcategories: standard operation actions (such as instrument readings and equipment operation), violation actions (such as not wearing insulating gloves), and dangerous actions (such as entering a live compartment).
[0044] Specifically, the system achieves behavior recognition through a cascaded classifier: first, a temporal convolutional network (TCN) is used to extract spatiotemporal features, and then a support vector machine (SVM) is used to make multi-classification decisions. For samples with a confidence level lower than 0.85, a manual review mechanism is initiated.
[0045] For example, for a sequence containing the action of "climbing equipment" (lasting 6 seconds), the system extracts a 256-dimensional spatiotemporal feature vector, and the SVM outputs a "dangerous action" probability of 0.92. Simultaneously, it detects that the arm joint trajectory deviates from safety regulations by more than 15 degrees, triggering a level-three warning and pushing it to the monitoring center. On the test set, the system achieves a 93.7% accuracy rate in identifying common violations, with a false alarm rate controlled below 2.1%.
[0046] Therefore, according to the above implementation method, the system first acquires multimodal data in the operation and maintenance scenario simultaneously through various acquisition devices deployed in the UHV substation; then, it sequentially performs feature alignment processing and attention fusion processing on the multimodal data to generate unified fused multimodal features; then, it inputs the fused multimodal features into a pre-configured temporal video state space model and a video pose transformer model, performs selective state space modeling and long-term spatiotemporal dependency capture through the temporal video state space model, and performs hierarchical feature compression and reconstruction operations through the video pose transformer model to extract temporal three-dimensional behavioral features; finally, it identifies the behavior category of the operation and maintenance personnel based on the temporal three-dimensional behavioral features.
[0047] Specifically, in this embodiment, the technical solution addresses the multimodal data heterogeneity problem mentioned in the background technology by unifying the feature dimensions and sampling rates of different modal data through feature alignment processing, thus solving the deficiency of traditional methods in effectively aligning multi-source data. Regarding the problem of redundant information interference in long-term video, the selective state-space modeling mechanism of the temporal video state-space model enables the screening of key spatiotemporal features and the capture of long-term dependencies. Addressing the problem of low recognition accuracy due to insufficient extraction of 3D behavioral features, the hierarchical feature compression and reconstruction operation of the video pose converter model achieves effective mapping from 2D image sequences to 3D pose sequences. Therefore, the technical solution of this invention solves the technical problem that existing technologies cannot meet the requirements of efficient monitoring of operation and maintenance personnel behavior and accurate early warning of safety risks in high-risk operation scenarios of UHV substations, achieving accurate and efficient identification of operation and maintenance personnel behavior and improving the accuracy, robustness, and practicality of behavior analysis.
[0048] In some embodiments, the UHV substation is equipped with video acquisition equipment, audio acquisition equipment, sensor acquisition equipment, and positioning equipment. The multimodal data includes video data, audio data, sensor data, location data, and environmental data. Acquiring multimodal data under UHV substation operation and maintenance scenarios includes: Video data of operation and maintenance personnel is acquired through video acquisition equipment; audio data of the operation and maintenance environment is acquired through audio acquisition equipment; status data of operation and maintenance personnel is acquired through sensor acquisition equipment; location data of operation and maintenance personnel is acquired through positioning equipment; and environmental data of the operation and maintenance environment is acquired through environmental monitoring equipment.
[0049] The status data includes accelerometer data and gyroscope data used to reflect the posture of the maintenance personnel; the accelerometer data is used to measure the linear acceleration of the maintenance personnel's limb movements, and the gyroscope data is used to measure the angular velocity of the maintenance personnel's joint rotation.
[0050] Specifically, the video acquisition equipment uses an explosion-proof infrared thermal imager (resolution 3840×2160 pixels) to acquire dual-channel video of visible light and thermal imaging at 25 frames per second; the audio acquisition equipment uses a ring microphone array (8 microphone units) to acquire environmental sound and voice commands at a sampling rate of 48 kHz; the sensor acquisition equipment uses a wearable inertial measurement unit (9-axis IMU sensor) to synchronously acquire acceleration (range ±16g) and angular velocity (range ±2000 degrees / second) data at a frequency of 100 Hz; the positioning equipment uses an ultra-wideband positioning base station (accuracy ±10 cm) to track the three-dimensional coordinates of personnel in real time; the environmental monitoring equipment uses a multi-functional sensor to acquire temperature (range -40 degrees Celsius to 85 degrees Celsius), humidity (range 0-100%RH, where RH refers to relative humidity) and sulfur hexafluoride gas concentration (range 0~1000ppm, where ppm refers to parts per million) parameters. All devices use a precision clock synchronization protocol (IEEE 1588v2) to control time deviation within 50 milliseconds, and aggregate data to the edge computing gateway through an industrial-grade ring network switch.
[0051] For example, when maintenance personnel are performing equipment inspections in a 500 kV power distribution area, the system collects full-body motion video using an infrared thermal imager deployed on a gantry, records fine movements such as bending over and raising hands using IMU sensors worn on safety helmets, tracks the inspection path coordinates in real time using a positioning base station, and monitors the area's temperature (25 degrees Celsius), humidity (60% RH), and sulfur hexafluoride concentration (0 ppm) using environmental sensors. When a sudden increase in the amplitude of a person's movements (acceleration exceeding 2g) is detected, accompanied by abnormal sounds (sound pressure level exceeding 85 dB), the system automatically marks the multimodal data for that period as a high-risk segment.
[0052] Therefore, according to the above implementation method, the system can construct a spatiotemporally aligned multimodal dataset through the collaborative acquisition of multiple heterogeneous sensors, providing a complete data foundation for subsequent feature fusion and behavior analysis.
[0053] In some embodiments, feature alignment and attention fusion processing are performed sequentially on the multimodal data to generate fused multimodal features, including: The feature alignment module performs feature alignment processing on the multimodal data to generate aligned multimodal features. The feature alignment processing includes unifying the number of feature channels of different modal data through the fully connected layer of the feature alignment module and unifying the feature size through sampling operations.
[0054] Among them, the feature alignment module is a component responsible for unifying the feature dimensions and sizes of multimodal data, eliminating heterogeneous data differences through linear transformation and sampling mechanisms; the fully connected layer is a neural network layer that achieves linear mapping of input features through a weight matrix; the sampling operation includes upsampling and downsampling, used to adjust the feature map size.
[0055] Specifically, the feature alignment module first uses a 1×1 fully connected layer (a convolutional layer with a kernel size of 1×1, equivalent to a fully connected operation) to unify the number of feature channels of each modality to a preset dimension (such as 512 dimensions). Then, the feature map size is standardized through the sampling module: upsampling uses reshaping and bilinear interpolation to amplify low-resolution features, and downsampling uses average pooling and reshaping to compress high-resolution features, finally outputting aligned features of consistent size.
[0056] For example, for video features (256×64×64, representing 256 channels, 64 pixels in height, and 64 pixels in width) and audio features (128×32×32), the feature alignment module increases the number of audio feature channels to 256 dimensions through a fully connected layer, and then adjusts the audio feature size to 64×64 through upsampling, generating aligned features of uniform size 256×64×64; the processing is completed on an edge computing device and takes less than 15 milliseconds.
[0057] The attention mechanism feature fusion module performs attention fusion processing on the aligned multimodal features to generate fused multimodal features. The attention fusion processing is based on the cross-attention mechanism to calculate the relationship mapping between modalities.
[0058] Among them, the attention mechanism feature fusion module refers to the component that dynamically weights multimodal features through attention weights to achieve enhanced fusion of key information; the cross attention mechanism is a calculation method based on query, key and value, which captures dependencies through intermodal interaction.
[0059] Specifically, the module uses one modal feature (such as video features) as the query and another modal feature (such as sensor features) as the key and value, and calculates the attention weight matrix: first, it generates query, key and value vectors through linear transformation, then calculates the dot product of the query and key and scales it, applies the Softmax function (normalized exponential function) to generate weights, and finally performs a weighted summation of the value vectors to output the fused features.
[0060] For example, for aligned video features (256×64×64 dimensions) and sensor features (256×64×64 dimensions), the attention mechanism feature fusion module uses an 8-head attention mechanism to divide the query, key, and value dimensions into 8 subspaces (32 dimensions per head), calculates the attention weights, and then concatenates the results. The module assigns a weight of 0.7 or higher to high-risk action frames (such as the instant of tool operation) and a weight of 0.2 or lower to static frames, generating a 256-dimensional fused feature vector with a processing latency controlled within 10 milliseconds.
[0061] Therefore, according to the above implementation method, the system can eliminate the dimensional and size differences of multimodal data through feature alignment, and then strengthen key modal information through attention fusion to generate high-quality fusion features, providing a unified and enhanced data foundation for subsequent behavior analysis and improving the accuracy and robustness of behavior recognition.
[0062] In some embodiments, the temporal video state space model is configured with a selective state space modeling module and a spatiotemporal correlation modeling module; the steps of performing selective state space modeling and long-term spatiotemporal dependency capture through the temporal video state space model include: The input video data is processed by a selective state-space modeling module to generate an intermediate spatiotemporal feature sequence. The state-space transformation process includes a linear transformation operation that projects the video frame sequence onto the state space.
[0063] Among them, the selective state-space modeling module refers to a component that focuses on key spatiotemporal features through a dynamic selection mechanism and uses a linear complexity algorithm to achieve efficient processing of long video sequences; state-space transformation processing is a mathematical operation that transforms video data from pixel space to feature space and preserves temporal correlation through linear projection.
[0064] Specifically, the module first extracts local spatiotemporal features through a convolutional layer (3×3 kernel size, stride 1), then uses a gating mechanism to dynamically filter key frames, and finally maps the features to a low-dimensional state space through a linear projection layer, outputting a dimension-compressed intermediate feature sequence.
[0065] For example, for a 10-second 1080-pixel resolution video (300 frames), the module first extracts 256-dimensional spatiotemporal features, then retains 40% of the keyframes (120 frames) through a selective scanning mechanism, and finally generates a 120×256-dimensional intermediate spatiotemporal feature sequence. The processing latency is controlled within 200 milliseconds. It runs on the NVIDIA Jetson Xavier edge computing module (an AI embedded platform for edge computing, equipped with 512 CUDA cores and 64 Tensor cores. CUDA cores stand for Compute Unified Device Architecture, a parallel computing platform and programming model launched by NVIDIA. Tensor cores are dedicated hardware units launched by NVIDIA, whose main function is to efficiently perform matrix multiplication and accumulation operations required in deep learning training and inference).
[0066] The spatiotemporal correlation modeling module performs spatiotemporal dynamic evolution modeling on the intermediate spatiotemporal feature sequence, and outputs low-redundancy time series features.
[0067] Among them, the spatiotemporal correlation modeling module refers to the component that captures long-term spatiotemporal dependencies in video sequences and analyzes the dynamic evolution law between frames through a recurrent neural network structure; spatiotemporal dynamic evolution modeling is an analytical method that simulates the continuous change process of video content in the time and space dimensions.
[0068] Specifically, the module adopts a bidirectional long short-term memory (Bi-LSTM) network structure with a hidden layer dimension of 512. It captures global temporal correlations through forward and backward scans and uses dropout (with a ratio of 0.2) to prevent overfitting. The final output is a low-redundancy feature that retains key dynamic information.
[0069] For example, inputting a 120×256-dimensional intermediate feature sequence, after processing by Bi-LSTM (time step 120, hidden state 512 dimensions), outputs a 60×512-dimensional low-redundancy temporal feature, with a computational complexity of linear order. It reduces computation by 60% compared to traditional self-attention mechanisms, and achieves a 92% retention rate for key behavioral features (such as tool operation trajectories).
[0070] Therefore, according to the above implementation method, the system can focus on key video segments through selective state space modeling and capture long-term behavior patterns through spatiotemporal correlation modeling, which can significantly improve processing efficiency while ensuring feature quality and meet the needs of real-time behavior analysis of UHV substations.
[0071] In some embodiments, the temporal video state space model is further configured with a state projection module and a selective scanning module; the spatiotemporal dynamic evolution model of the intermediate spatiotemporal feature sequence is performed through the spatiotemporal correlation modeling module to output low-redundancy temporal features, including: The state projection module performs state space projection processing on the time-series video data to generate an initial state sequence. The state space projection processing includes a linear transformation operation that converts the video frame sequence into a state vector.
[0072] Among them, the state projection module is the component in the temporal video state space model responsible for mapping the input video data to the state space representation. It realizes the transformation from high-dimensional video frames to low-dimensional state vectors through linear transformation. State space projection processing is a feature extraction method based on matrix operations, which converts each video frame or frame sequence into a state vector, which is convenient for subsequent sequence modeling.
[0073] Specifically, the state projection module uses either a fully connected layer (a neural network layer that achieves linear mapping of input features through a weight matrix) or a convolutional layer (a neural network layer that extracts local features through a sliding window) to perform the projection operation; the fully connected layer linearly maps the flattened frame pixel values to the state vector, while the convolutional layer extracts spatial features and compresses dimensions through convolutional kernels.
[0074] For example, for a 10-second video sequence with a resolution of 1920×1080 pixels, sampled at 30 frames per second, for a total of 300 frames; the state projection module adjusts each frame image to an input size of 224×224 pixels, performs a linear transformation through a fully connected layer (input dimension 150528 pixels, output dimension 256 dimensions), and generates 300 256-dimensional state vectors to form the initial state sequence; the projection operation takes less than 15 milliseconds on an NVIDIA RTX 3080 graphics card (10 gigabytes of video memory).
[0075] The initial state sequence is selectively scanned by the selective scanning module to generate an intermediate spatiotemporal feature sequence. The selective scanning process includes selective attention-weighted updates of the state sequence.
[0076] The selective scanning module refers to the component in the temporal video state space model that dynamically filters key state vectors. It focuses on important time steps through an attention mechanism to reduce redundant information. Selective scanning processing is a sequence optimization method based on attention weights, which performs weighted summation or filtering on the state sequence to enhance the contribution of key frames.
[0077] Specifically, the selective scanning module uses a multi-head self-attention mechanism (a method that divides attention computation into multiple subspaces) to perform scanning. The module first calculates the query, key, and value matrix of each state vector, generates attention weights through the Softmax function (normalized exponential function), and then dynamically selects high-scoring vectors based on the weight threshold.
[0078] For example, the input initial state sequence contains 300 256-dimensional vectors. The selective scanning module uses 8-head attention (32 dimensions per head), sets the attention weight threshold to 0.1, and filters out the top 150 key vectors with weights higher than the threshold to generate a 150×256-dimensional intermediate spatiotemporal feature sequence. The processing is completed on an edge computing device (such as the NVIDIA Jetson AGX Orin module), with latency controlled within 25 milliseconds.
[0079] The intermediate spatiotemporal feature sequence is transformed using a spatiotemporal correlation modeling module to output low-redundancy time series features.
[0080] Among them, low-redundancy temporal features refer to feature sequences after dimensionality reduction and redundancy removal, which are used as input to the video pose transformer model for subsequent feature compression and reconstruction operations; linear complexity transformation processing is a computational complexity that is linearly related to the input size (i.e., Methods that reduce computational complexity can be implemented by using linear transformations or low-rank approximations to reduce computational resource consumption.
[0081] Specifically, the spatiotemporal correlation modeling module implements transformations based on linear attention mechanisms or principal component analysis (PCA, a statistical dimensionality reduction method), and uses linear layers and ReLU activation functions for feature compression and nonlinear mapping.
[0082] For example, given an input intermediate spatiotemporal feature sequence with dimensions of 150×256, the linear complexity transformation module converts it into 150×128-dimensional low-redundancy temporal features through a linear projection layer (input dimension 256, output dimension 128). This reduces computation time by 55% compared to the standard self-attention mechanism, while maintaining an 88% feature information retention rate. Furthermore, the module runs on a GPU (Graphics Processing Unit) server, resulting in a 35% reduction in power consumption. Therefore, according to the above implementation method, the system can efficiently extract key spatiotemporal information through state projection and selective scanning, and then generate low-redundancy features through linear complexity transformation, providing optimized input for the video pose converter model and improving the accuracy and real-time performance of behavior analysis.
[0083] In some embodiments, the video pose converter model is configured with an hourglass word segmenter; the step of extracting temporal 3D behavioral features from fused multimodal features by performing hierarchical feature compression and reconstruction operations through the video pose converter model includes: The marker pruning and clustering module in the hourglass word segmenter performs dynamic marker selection processing on low-redundancy temporal features to generate compressed marker sequences. The dynamic marker selection processing includes screening representative markers with high semantic diversity through clustering algorithms.
[0084] Among them, the tag pruning clustering module refers to the component in the hourglass word segmenter responsible for reducing the number of tags and selecting representative tags with rich information through cluster analysis; dynamic tag selection processing is a method for adaptively determining cluster centers based on data distribution and grouping and filtering according to the semantic similarity between tags.
[0085] Specifically, the module first calculates the feature similarity matrix of all tags, then uses the K-means clustering algorithm (a distance-based iterative clustering algorithm) to divide the tags into K clusters, and finally selects the tag closest to the centroid from each cluster as the representative tag.
[0086] For example, for low-redundancy temporal features with an input dimension of 300×512 (300 time steps, 512 dimensions per label), the label pruning clustering module sets the number of clusters K=75. After 5 iterations of convergence, it selects one representative label from each cluster to generate a 75×512 dimension compressed label sequence with a compression rate of 75% and a processing time of less than 15 milliseconds.
[0087] The compressed tag sequence is reconstructed using the tag recovery attention module in the hourglass word segmenter, and the output is a full-resolution tag sequence. The sequence reconstruction process includes feature reconstruction based on representative tags to restore the full-length temporal resolution.
[0088] Among them, the tag recovery attention module is the component in the hourglass segmenter responsible for reconstructing the complete sequence. It maps compressed tags back to the original temporal resolution through an attention mechanism. Sequence reconstruction processing is an interpolation method based on attention weights, which uses information from representative tags to recover the details of the complete sequence.
[0089] Specifically, this module employs a multi-head cross-attention mechanism, encoding the original sequence position as the query, compressing the labels as the keys and values, calculating the reconstructed value at each time step through attention weights, and finally using residual connections to enhance feature representation.
[0090] For example, given a 75×512 dimensional compressed label sequence as input, the label recovery attention module uses an 8-head attention mechanism combined with 300 positional encodings (a method for representing sequence position information) to reconstruct a 300×512 dimensional full-resolution label sequence with a reconstruction error of less than 5% and a processing latency of 20 milliseconds on the GPU.
[0091] In other embodiments, the structure of the attention mechanism feature fusion module is as follows: Figure 5As shown, this module is composed of multiple layers of components and is used to achieve weighted fusion of cross-modal features. The module's input includes skeleton features (as keys and values) on the left and context features (as queries) on the right, and the output is the fusion feature. The data flow is processed from bottom to top: first, the intermodal association weights are calculated through a multi-head cross attention layer (an attention calculation method based on a query-key-value mechanism), then sequentially through an add&norm layer (a combination of residual connections and layer normalization), a pointwise feed-forward layer (a fully connected feed-forward network), and another add&norm layer, finally generating the fusion feature. Among them, the multi-head cross-attention layer is the core component of the module. It captures complex intermodal dependencies by dividing the query, key and value into multiple subspaces and computing attention weights in parallel. The addition and normalization layer is used to stabilize the training process. It retains input information through residual connections and applies layer normalization (a normalization technique) to optimize gradient flow. The pointwise feedforward layer introduces nonlinear transformation through two fully connected layers (a fully connected layer, a linear transformation layer) and activation functions to enhance feature representation capabilities. Specifically, the module's workflow is as follows: Contextual features are used as query input, and skeleton features are used as key and value input. The multi-head cross-attention layer first performs linear projection on the query, key, and value and divides them into 8 heads (attention computation subspaces). Each head independently computes scaled dot-product attention (a type of attention scoring function) and outputs a weighted feature sequence. Next, the addition and normalization layer adds the attention output to the original query features and normalizes it. Then, the pointwise feedforward layer performs dimensional transformation on the normalized result (e.g., mapping from 512 dimensions to 2048 dimensions and then back to 512 dimensions), and introduces non-linearity using the ReLU activation function (Rectified Linear Unit). Finally, the second addition and normalization layer adds the feedforward output to the attention output and normalizes it to generate the final fused features.For example, for skeleton features (256×64×64 dimensions, representing 256 channels, 64 pixels high, and 64 pixels wide) and context features (256×64×64 dimensions), the module is set to have 8 heads, with each head having 64 dimensions for query, key, and value. After multi-head cross-attention layer calculation, the output is a 256×64×64 weighted feature. The addition and normalization layers fuse the output with the original context features through residual connections. The pointwise feedforward layer uses two fully connected layers (the first layer outputs 2048 dimensions, and the second layer outputs 512 dimensions) for processing. The entire process takes 15 milliseconds on an NVIDIA V100 graphics card (32 gigabytes of video memory), and the weight of the fused features on key action frames is increased to over 0.9. This design effectively strengthens cross-modal feature interaction and improves the robustness of behavior recognition.
[0092] In other embodiments, the steps for intra-frame static attention and inter-frame dynamic attention are as follows: Figure 6 As shown in the diagram, this illustration illustrates the detailed processing flow of the Temporal-Spatial Attention Unit (TAU). The input features first enter the Spatial Encoder, which transforms the shape of the input tensor from... (Where B is batch size, T is time step, C is number of channels, H is height, and W is width) Convert to The data is then fed into the static attention path and the dynamic attention path for parallel processing, respectively. The static attention path mainly handles spatial feature extraction, specifically including the following steps: first, spatial features are extracted using depthwise convolution (DW Conv), then feature transformation is performed using 1×1 convolution (1×1 Conv, a point convolution operation), and finally, spatial attention weights are generated through a fully connected layer (FC). The dynamic attention path focuses on temporal feature modeling, and its steps include: first, depthwise-dynamic convolution (DW-D Conv, a convolution operation that handles temporal relationships) is applied to capture inter-frame dynamic changes, then feature size is compressed using average pooling (AvgPool), then feature fusion is performed using 1×1 convolution, and finally, temporal attention weights are generated through a fully connected layer. In a specific application example, when the system processes a 10-second 1080-pixel resolution video (batch size B=8, time step T=300 frames, number of channels C=256, feature map size...), When the spatial encoder first changes the input dimension from... Convert to The static attention path processes features at each spatial location through depthwise separable convolutions (3×3 kernels, stride 1), outputting... Feature maps; the dynamic attention path uses depthwise separable temporal convolutions (convolution kernels). Analyze the changes between consecutive frames (corresponding to time, height, and width dimensions respectively) and output the results. The dynamic characteristics are shown. The two attention weights are weighted and fused, and then the original dimensions are recovered through a spatial decoder, resulting in the final output. Enhanced features. The entire processing runs on an NVIDIA A100 graphics card (80 gigabytes of video memory) and takes approximately 35 milliseconds, representing a 40% improvement in computational efficiency compared to traditional methods. This dual-path attention mechanism allows the system to simultaneously capture spatial structural features within video frames and temporal evolution patterns between frames, significantly improving the completeness and accuracy of behavioral feature extraction and providing a more reliable feature representation for subsequent 3D behavior recognition.
[0093] Temporal three-dimensional behavioral features are generated based on the full-resolution labeled sequence.
[0094] Among them, temporal three-dimensional behavioral features refer to a three-dimensional key point coordinate sequence that includes a time dimension, used to describe the continuous changes of human movements in time and space.
[0095] Specifically, the regression head module is used to map the labeled sequence to three-dimensional coordinates. This module consists of two fully connected layers (the first layer is 512-dimensional and the second layer is 51-dimensional, corresponding to the three-dimensional coordinates of 17 joints), and a nonlinear transformation is introduced in between using the ReLU activation function.
[0096] For example, when given a 300×512 dimensional full-resolution labeled sequence, the regression head module outputs a 300×51 dimensional three-dimensional pose sequence (17 key points × 3D coordinates), with a coordinate error within 5 millimeters. The entire processing takes 50 milliseconds on an NVIDIA RTX 3080 graphics card, meeting the requirements for real-time analysis.
[0097] Therefore, according to the above implementation method, the system can effectively compress feature sequences through labeled pruning clustering, and then accurately reconstruct complete resolution features through labeled attention recovery, ultimately generating accurate three-dimensional behavioral features, thereby achieving efficient and accurate behavior analysis of operation and maintenance personnel.
[0098] In some embodiments, the behavior category of maintenance personnel is identified based on temporal three-dimensional behavioral features, including: Input the temporal three-dimensional behavioral features into the pre-configured behavior classification module.
[0099] Among them, the behavior classification module refers to a classification component based on a neural network structure, which is used to perform pattern analysis and category determination on the input three-dimensional behavior features. This module is usually composed of a feature extraction layer and a classification layer to realize the mapping from raw features to behavior categories.
[0100] Specifically, the behavior classification module receives the temporal three-dimensional behavior feature sequence through the data interface. First, it standardizes the input data (e.g., normalizes it to the 0-1 range) and then passes it to the subsequent analysis layer. The input feature dimension must match the preset input size of the module, such as a sequence length of 300 frames and a feature dimension of 51 (corresponding to the three-dimensional coordinates of 17 key points).
[0101] For example, for a 10-second behavior sequence (300 frames, 51-dimensional features per frame), the system calls the behavior classification module through the TensorFlow library (a machine learning framework) in Python (a programming language). The input data is 300×51 in shape. The module is loaded on the NVIDIA Jetson AGX Orin edge computing device in less than 5 milliseconds.
[0102] The behavior classification module performs behavior pattern analysis on the temporal 3D behavior features to generate behavior feature vectors. The behavior pattern analysis includes feature extraction of the spatiotemporal patterns of the 3D pose sequence.
[0103] Among them, behavior pattern analysis and processing refers to the operation of extracting key spatiotemporal features from three-dimensional posture sequences, capturing the essence of behavior by analyzing the motion trajectory and temporal correlation of joints; behavior feature vector refers to a one-dimensional feature representation after dimensionality reduction and aggregation, which is used to simplify classification input.
[0104] Specifically, the processing uses either a Temporal Convolutional Network (TCN, a convolutional neural network specifically designed for time-series data processing) or a Long Short-Term Memory (LSTM, a variant of a recurrent neural network). The module first performs one-dimensional convolution (kernel size 3, stride 1) along the time dimension to extract local temporal patterns, and then aggregates the sequences into fixed-length vectors through Global Average Pooling (GAP).
[0105] For example, with a 300×51-dimensional feature sequence as input, the behavior pattern analysis process extracts spatiotemporal features through TCN (128 hidden layer dimensions, 3 convolutional layers), outputs a 256-dimensional behavior feature vector, and controls the processing latency within 20 milliseconds. The feature retention rate of key actions (such as raising hands and bending over) exceeds 90%.
[0106] The behavior feature vector is processed for category determination, and the behavior category of the maintenance personnel is output. The category determination process includes a classification operation that maps the feature vector to a preset behavior category.
[0107] Among them, category determination processing refers to the decision-making process based on machine learning classifiers, which discretizes continuous feature vectors into specific behavior labels; preset behavior categories refer to behavior types predefined according to substation safety regulations, such as "normal walking" and "illegal climbing".
[0108] Specifically, the processing is implemented using Support Vector Machine (SVM, a supervised learning classification model) or Softmax Regression (a multi-class classification method). The module first calculates the distance or probability between the feature vector and the decision boundary of each class, then selects the class with the highest confidence as the output, and sets a probability threshold (such as 0.7) to filter low-confidence samples.
[0109] For example, for a 256-dimensional behavioral feature vector, the category determination process uses an SVM classifier (with a radial basis function, RBF kernel) to calculate the probability of five behavioral categories (such as "normal walking," "abnormal running," "dangerous climbing," "safe gestures," and "falling"). If the probability of "dangerous climbing" reaches 0.85, the category is output and an alert is triggered. The entire classification process takes less than 3 milliseconds on the CPU (Central Processing Unit) with an accuracy of 95%. Therefore, according to the above implementation method, the system can efficiently analyze three-dimensional behavioral features through the behavioral classification module, accurately identify the behavioral categories of operation and maintenance personnel, and provide real-time decision support for substation safety monitoring.
[0110] Furthermore, in other embodiments, the structure of the behavior monitoring model that integrates multi-dimensional features is as follows: Figure 7 As shown, the model employs a multi-branch parallel processing architecture to achieve comprehensive feature extraction and fusion of video data. The model input is the raw video data of an UHV substation operation and maintenance scenario, and four independent but collaborative baseline network branches process visual and auditory information of different modalities respectively.
[0111] Specifically, the object detection branch extracts static object features such as devices and tools in the video using the Faster R-CNN network (Region-based Convolutional Neural Network); the optical flow analysis branch calculates inter-frame motion information using the BN-Inception network (a deep learning network that combines batch normalization and the Inception module) to generate optical flow features that characterize the movement trends of people; the RGB appearance branch also extracts visual features such as color and texture of video frames using the BN-Inception network; and the audio analysis branch extracts acoustic features from ambient sound using the AudioSlowFast network (a deep learning architecture specifically for processing audio signals).
[0112] like Figure 7 The workflow shown involves each branch independently generating frame-level features first. Then, a segment-level feature construction module aggregates temporally consecutive frame features into semantically complete behavioral segment representations. These cross-modal features are input to a multimodal fusion module, which achieves complementary enhancement between different modalities through feature alignment and attention weighting mechanisms. Finally, the fused features are passed through a fully connected layer (a linear transformation layer) and a softmax classifier (a multi-class classification function) to output specific behavioral category determinations.
[0113] For example, when the system analyzes a 10-second video of an operator operating a disconnect switch, the object detection branch identifies features such as the operating handle and safety tools; the optical flow branch captures the trajectory features of the arm movement; the RGB branch extracts the overall appearance features of the operating posture; and the audio branch analyzes the acoustic features of mechanical operation sounds and voice commands. The multimodal fusion module weights and fuses these features to accurately identify the behavior category as "standard operation." The entire processing runs on a server equipped with an NVIDIA Tesla V100 graphics card (32 gigabytes of video memory), achieving a processing speed of 25 frames per second for 1080-pixel resolution video, meeting real-time monitoring requirements.
[0114] This multi-dimensional feature fusion architecture effectively solves the problem of limited representation capabilities of single-modal features. By using complementary multi-source information, it improves the accuracy and robustness of behavior recognition, providing reliable technical support for the operation and maintenance safety of UHV substations.
[0115] In other embodiments, the operation steps of the multi-level state data monitoring model based on LSTM (Long Short-Term Memory) fusion are as follows: Figure 8 As shown, the model employs a sequence-to-sequence (a neural network architecture for processing sequential data) encoder-decoder structure to achieve time-series modeling and fusion analysis of multi-level state data from UHV substation operation and maintenance personnel. The model input consists of multi-source state parameters (including sensor data, location data, environmental data, etc.). Through four core steps—normalization, data encoding, LSTM processing, and output—a comprehensive monitoring result is generated. Specifically, the model's operation flow is as follows: First, the input data step receives multi-level state parameters (Input1, Input2, ..., InputN). These parameters include heterogeneous data such as accelerometer data (range ±16g), gyroscope data (range ±2000 degrees / second), location coordinates (accuracy ±10 cm), and environmental temperature and humidity (temperature range -40°C to 85°C, humidity range 0~100%RH). These data are synchronously collected through a data interface and encapsulated into a time-series format. Next, the normalization step standardizes the input data, using the min-max normalization method (a linear transformation technique) to map each parameter value to the interval between 0 and 1, eliminating dimensional differences. The formula is: For example, raw accelerometer data (ranging from -16g to 16g) is normalized to values between 0 and 1, ensuring comparability between data from different sensors. Then, the data encoding step maps the normalized multidimensional features to a unified latent space through a fully connected layer (a linear transformation layer), generating a sequence of encoded vectors. The encoder consists of multiple parallel units, each processing one class of state parameters and outputting 256-dimensional encoded features. The encoding process takes less than 10 milliseconds on an edge computing device. Subsequently, an LSTM step performs temporal modeling on the encoded feature sequence, using a two-layer LSTM network (128 hidden layer dimensions, 50 time steps) to capture long-term dependencies. The LSTM units dynamically update the cell state through input gates, forget gates, and output gates (three gating mechanisms to control information flow), effectively learning the evolution of state parameters. For example, for a 10-second data segment (sampling rate 10 Hz, total 100 time steps), the LSTM outputs 100×128-dimensional temporal features, achieving a 95% accuracy rate in capturing key anomalous patterns (such as acceleration mutations). Finally, the output step maps the LSTM output to monitoring probabilities through a fully connected layer and a sigmoid function (an activation function with an output range between 0 and 1), generating behavioral state scores (e.g., normal operation probability 0.92, abnormal behavior probability 0.08). The entire model runs on an NVIDIA Jetson Nano module (an embedded artificial intelligence computing device), with processing latency controlled within 200 milliseconds, meeting real-time monitoring requirements. Therefore, according to the above implementation method, the system can efficiently fuse and perform temporal analysis of multi-level state data through the LSTM network, accurately identify the behavioral states of operation and maintenance personnel, and improve the real-time performance and reliability of substation safety monitoring.
[0116] Figure 9 This is a structural block diagram of a behavior analysis system for operation and maintenance personnel in an ultra-high voltage substation according to an embodiment of the present invention.
[0117] like Figure 9 As shown, the operation and maintenance personnel behavior analysis system of this UHV substation includes: The multimodal data acquisition module 210 is used to acquire multimodal data in the operation and maintenance scenarios of UHV substations.
[0118] The multimodal feature generation module 220 is used to sequentially perform feature alignment processing and attention fusion processing on multimodal data to generate fused multimodal features.
[0119] The 3D behavioral feature analysis module 230 is used to input the fused multimodal features into a pre-configured temporal video state space model and a video pose transformer model. It performs selective state space modeling and long-term spatiotemporal dependency capture through the temporal video state space model, and performs hierarchical feature compression and reconstruction operations through the video pose transformer model to extract temporal 3D behavioral features from the fused multimodal features.
[0120] The personnel behavior type identification module 240 is used to identify the behavior category of operation and maintenance personnel based on the time-series three-dimensional behavioral characteristics.
[0121] The specific functions and examples of each module and submodule of the device in this embodiment of the invention can be found in the relevant descriptions of the corresponding steps in the above method embodiments, and will not be repeated here.
[0122] According to embodiments of the present invention, the above-described method of the present invention can be applied to an electronic device and a readable storage medium.
[0123] Figure 10 A schematic block diagram of an example electronic device 600 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0124] like Figure 10 As shown, the electronic device 600 includes a computing unit 601, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. The RAM 603 may also store various programs and data required for the operation of the electronic device 600. The computing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0125] Multiple components in electronic device 600 are connected to I / O interface 605, including: input unit 606, such as keyboard, mouse, etc.; output unit 607, such as various types of displays, speakers, etc.; storage unit 608, such as disk, optical disk, etc.; and communication unit 609, such as network card, modem, wireless transceiver, etc. Communication unit 609 allows electronic device 600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0126] The computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as a method for analyzing the behavior of operation and maintenance personnel in an ultra-high voltage substation. For example, in some embodiments, a method for analyzing the behavior of operation and maintenance personnel in an ultra-high voltage substation can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps of the method for analyzing the behavior of operation and maintenance personnel in an ultra-high voltage substation described above can be performed. Alternatively, in other embodiments, the computing unit 601 may be configured, by any other suitable means (e.g., by means of firmware), to perform a method for analyzing the behavior of operation and maintenance personnel in an ultra-high voltage substation.
[0127] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0128] The program code used to implement the methods of the present invention can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a standalone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0129] In the context of this invention, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0130] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0131] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0132] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0133] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this invention can be achieved, and this is not limited herein.
[0134] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for analyzing the behavior of an operator of an ultra-high voltage substation, characterized in that, The application relates to a method for identifying behavior categories of operation and inspection personnel in an ultrahigh voltage substation. The method comprises the following steps: acquiring multi-modal data under an operation and inspection scene of an ultrahigh voltage substation; sequentially performing feature alignment processing and attention fusion processing on the multi-modal data to generate fused multi-modal features; inputting the fused multi-modal features into a pre-configured time sequence video state space model and a video posture transformer model, performing selective state space modeling and long-term space-time dependence capturing through the time sequence video state space model, and performing hierarchical feature compression and reconstruction operations through the video posture transformer model to extract time sequence three-dimensional behavior features from the fused multi-modal features; 2. The method of claim 1, wherein, identifying a behavior category of operation and inspection personnel according to the time sequence three-dimensional behavior features. The ultrahigh voltage substation is provided with a video acquisition device, an audio acquisition device, a sensor acquisition device and a positioning device, and the multi-modal data comprises video data, audio data, sensor data, position data and environment data; the multi-modal data under the operation and inspection scene of the ultrahigh voltage substation is acquired by the following steps: acquiring video data of the operation and inspection personnel through the video acquisition device; acquiring audio data of the operation and inspection environment through the audio acquisition device; acquiring state data of the operation and inspection personnel through the sensor acquisition device; acquiring position data of the operation and inspection personnel through the positioning device; acquiring environment data of the operation and inspection environment through an environment monitoring device; 3. The method of claim 1, wherein, wherein the state data comprises accelerometer data and gyroscope data for reflecting the posture of the operation and inspection personnel. The method sequentially performs feature alignment processing and attention fusion processing on the multi-modal data to generate fused multi-modal features, which comprises the following steps: performing feature alignment processing on the multi-modal data through the feature alignment module to generate aligned multi-modal features, wherein the feature alignment processing comprises uniformly arranging feature channel numbers of different modal data through a full connection layer of the feature alignment module and uniformly arranging feature sizes through a sampling operation; 4. The method of claim 2, wherein, performing attention fusion processing on the aligned multi-modal features through the attention mechanism feature fusion module to generate the fused multi-modal features, wherein the attention fusion processing is based on cross-attention mechanism to calculate inter-modal relationship mapping. The time sequence video state space model is configured with a selective state space modeling module and a space-time correlation modeling module; the selective state space modeling and long-term space-time dependence capturing through the time sequence video state space model comprises the following steps: performing state space transformation processing on input video data through the selective state space modeling module to generate an intermediate space-time feature sequence, wherein the state space transformation processing comprises linear transformation operation of projecting a video frame sequence to a state space; 5. The method of claim 4, wherein, performing space-time dynamic evolution modeling on the intermediate space-time feature sequence through the space-time correlation modeling module to output low-redundancy time sequence features. The time sequence video state space model is further configured with a state projection module and a selective scanning module; the space-time dynamic evolution modeling on the intermediate space-time feature sequence through the space-time correlation modeling module to output low-redundancy time sequence features comprises the following steps: The state projection module is configured to perform state space projection processing on the time sequence video data to generate an initial state sequence, and the state space projection processing includes a linear transformation operation of converting a video frame sequence into a state vector; The selective scanning module is configured to perform selective scanning processing on the initial state sequence to generate the intermediate spatio-temporal feature sequence, and the selective scanning processing includes selective attention weighting updating on the state sequence; The spatio-temporal correlation modeling module is configured to perform linear complexity transformation processing on the intermediate spatio-temporal feature sequence to output the low-redundancy time sequence feature. The low-redundancy time sequence feature is used to input the video pose transformer model for subsequent feature compression and reconstruction operations.
6. The method of claim 5, wherein, The video pose transformer model is configured with an hourglass tokenizer; the step of performing hierarchical feature compression and reconstruction operation on the fusion multi-modal feature to extract the time sequence three-dimensional behavior feature through the video pose transformer model includes: The label pruning clustering module in the hourglass tokenizer is configured to perform dynamic label selection processing on the low-redundancy time sequence feature to generate a compressed label sequence, and the dynamic label selection processing includes filtering representative labels with high semantic diversity through a clustering algorithm; The label restoration attention module in the hourglass tokenizer is configured to perform sequence reconstruction processing on the compressed label sequence to output a complete resolution label sequence, and the sequence reconstruction processing includes feature reconstruction based on representative labels to restore full-length time sequence resolution; The complete resolution label sequence is used to generate the time sequence three-dimensional behavior feature.
7. The method of claim 6, wherein, The step of identifying the behavior category of the operation and inspection personnel according to the time sequence three-dimensional behavior feature includes: inputting the time sequence three-dimensional behavior feature into a pre-configured behavior classification module; The behavior classification module is configured to perform behavior pattern analysis processing on the time sequence three-dimensional behavior feature to generate a behavior feature vector, and the behavior pattern analysis processing includes feature extraction on the spatio-temporal pattern of the three-dimensional pose sequence; The behavior feature vector is subjected to category determination processing to output the behavior category of the operation and inspection personnel, and the category determination processing includes a classification operation of mapping the feature vector to a pre-set behavior category.
8. An operating and inspection personnel behavior analysis system for an extra-high voltage substation, characterized by, The step of identifying the behavior category of the operation and inspection personnel according to the time sequence three-dimensional behavior feature includes: a multi-modal data acquisition module configured to acquire multi-modal data in an ultra-high voltage substation operation and inspection scene; a multi-modal feature generation module configured to sequentially perform feature alignment processing and attention fusion processing on the multi-modal data to generate fusion multi-modal features; a three-dimensional behavior feature analysis module configured to input the fusion multi-modal features into a pre-configured time sequence video state space model and a video pose transformer model, perform selective state space modeling and long-term spatio-temporal dependence capture through the time sequence video state space model, and perform hierarchical feature compression and reconstruction operation through the video pose transformer model to extract time sequence three-dimensional behavior features from the fusion multi-modal features; a personnel behavior type identification module configured to identify the behavior category of the operation and inspection personnel according to the time sequence three-dimensional behavior feature.
9. An electronic device, comprising: at least one processor; and a memory connected in communication with the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executable by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-7.
10. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, Computer instructions for causing a computer to perform the method of any one of claims 1-7. Computer instructions for causing a computer to perform the method of any one of claims 1-7.
Citation Information
Patent Citations
Video monitoring early warning method and system based on multi-mode behavior mode
CN120219903A