Cable external damage prevention human behavior identification method and system based on space-time diagram convolution
Through the cable anti-outbreaking human behavior recognition method based on spatiotemporal graph convolution, the multi-spectral camera and sensors acquire data, build a skeletal motion topology map, and use an improved spatiotemporal graph convolution network and a cross-modal adversarial discriminator to solve the problems of low recognition accuracy and long response time in complex scenarios, real-time protection of high-precision and low false alarm rates is achieved.
Patent Information
- Application Number
- CN202510699398.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-28
AI Technical Summary
The existing cable anti-outbreak monitoring system has difficulty in detecting human bodies in complex urban market scenarios, insufficient representation of the space-time continuity of disrupting actions, low recognition accuracy, high misjudgment rate, and cloud computing architecture results in too long response time, which cannot meet the real-time protection needs.
Using a cable anti-outbreak human behavior recognition method based on spatiotemporal graph convolution, images and electrical data are acquired through multispectral cameras and sensors, bone motion topology maps containing dynamic associations of time series are constructed, and a modified spatiotemporal graph convolution network and cross-modal adversarial discriminator are used to output behavior recognition probability and risk levels to achieve hierarchical response.
It significantly improves the accuracy of identifying high-risk behaviors of cable breaking, reduces the false alarm rate, achieves rapid response, and meets the real-time protection needs of urban power grids.
Smart Images

Figure CN120217169A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of cable monitoring and protection, and particularly relates to a method and system for identifying human behaviors of preventing external damage to cables based on spatio-temporal graph convolution. Background Art
[0002] With the deep implementation of the new urbanization strategy, the underground cable network has become the core artery of urban energy supply, and its safe operation is directly related to the stability of the urban lifeline system. However, against the backdrop of the rapid development of municipal construction, the accidents of external force damage to buried cables show a trend of high frequency and concealment. This situation poses more stringent requirements for cable anti-external damage technology, and there is an urgent need to build intelligent protection devices and systems with accurate identification methods and real-time early warning capabilities.
[0003] In recent years, academia and industry have carried out a large number of exploratory studies in the field of cable anti-external damage, forming a research pattern with multiple technical paths running in parallel. Existing methods construct a protection system by integrating technical means such as computer vision and Internet of Things sensing. However, practical applications show that the current technology is still insufficient in adapting to complex construction scenarios.
[0004] Currently, cable anti-external damage monitoring systems based on RGB videos generally use two-dimensional convolutional neural networks for behavior recognition, which essentially model action features through two-dimensional pixel space. Such methods have three problems: in complex urban scenarios, interference factors such as object occlusion and sudden light changes easily lead to difficulties in human body detection and destroy the spatio-temporal continuity representation of action; most existing methods do not consider the specific spatio-temporal patterns of cable damage behaviors, have low recognition accuracy for specific cable damage behaviors, and misjudge the behavior risk level; existing systems mostly adopt a cloud centralized computing architecture, and the video stream transmission and processing delay result in a response time exceeding the time threshold for handling high-risk events. Existing systems have a high false alarm rate for typical external damage behaviors such as excavation and metal knocking, and it is difficult to meet the real-time protection requirements of urban power grids. Summary of the Invention
[0005] The purpose of the present invention is to provide a method and system for identifying human behaviors of preventing external damage to cables based on spatio-temporal graph convolution to overcome the deficiencies of the prior art.
[0006] The purpose of the present invention can be achieved through the following technical solutions: A method for identifying human behaviors of preventing external damage to cables based on spatio-temporal graph convolution, the steps are as follows: S1: Obtain image data and electrical data through a multispectral camera and sensors respectively; S2: Preprocess the image data and construct a skeletal motion topology graph including time series dynamic associations; S3: Use the improved spatio-temporal graph convolutional network (ST-GCN) to jointly model the multi-scale global spatio-temporal features of actions through multi-layer spatio-temporal convolutions and output the behavior recognition probability; the improved spatio-temporal graph convolutional network includes an improved spatio-temporal graph convolutional layer and a channel-spatio-temporal attention module, and the design of the improved spatio-temporal graph convolutional layer includes a spatial graph partitioning strategy, time dimension modeling, and dynamic adjacency matrix fusion; S4: Use the cross-modal adversarial discriminator to output the fused behavior recognition probability, and judge the risk level and formulate a hierarchical response strategy; the cross-modal adversarial discriminator includes: Generator G, which is used to map the image features and electrical features to a common embedding space to generate joint features: ; where, is the joint feature, is the image feature, and the image feature is extracted from the image data; is the electrical feature, and the electrical feature is extracted from the electrical data; Discriminator : Judge whether the feature pair comes from the true data distribution, and the loss function is: ; where, is the discriminator loss, is the mathematical expectation, is the forged electrical feature with added Gaussian noise.
[0007] Further preferably, the preprocessing of the image data in step S2 refers to: fusing the three-channel data into a three-channel fused feature map : ; where, are the weight coefficients of the visible light image , the thermal infrared image , and the degree of polarization respectively, and are dynamically adjusted according to the illumination intensity.
[0008] Further preferably, in step S2, a skeletal motion topology map including time series dynamic associations is constructed through a three-dimensional human pose estimation algorithm.
[0009] Further preferably, in step S2, a pre-trained three-dimensional human pose estimation algorithm is used to extract the three-dimensional coordinates of 17 joints from the three-channel fused feature map generated by the preprocessing to construct a spatio-temporal topology map; a loss function is designed as a constraint on the continuity of the skeletal trajectory to punish the change amplitude of the joint positions in adjacent frames; Calculate the joint Euler angles according to the three-dimensional coordinates of the joints , and compare with the preset biomechanical limits based on human anatomy and kinematics standards 、 for comparison, is the minimum allowable angle of the current degree of freedom, is the maximum allowable angle of the current degree of freedom; Define the physical loss function : ; where t is the index of the frame and i is the joint number; Constrain the joint Euler angles within the biomechanical limits through backpropagation; Design a second-order smoothing loss to achieve motion continuity constraints.
[0010] Further preferably, the improved spatio-temporal graph convolutional layer: In the spatial dimension, divide the neighborhood nodes of each joint into three categories through a spatial graph partitioning strategy: Root node: the current joint itself; Centripetal nodes: adjacent joints pointing to the root node; Centrifugal nodes: adjacent joints far from the root node; Assign an independent learnable weight matrix to each partition to capture local features of different motion patterns and obtain a spatial adjacency matrix; Temporal dimension modeling means: in the temporal dimension, adopt a one-dimensional temporal convolution kernel to perform temporal modeling on the motion trajectories of the same joint in consecutive frames to capture the temporal changes of actions; Set learnable weight coefficients to fuse the spatial adjacency matrix and the temporal adjacency matrix into a dynamic adjacency matrix; use the inverse square root of the degree matrix as the normalization factor to symmetrically normalize the dynamic adjacency matrix.
[0011] Further preferably, the channel-spatio-temporal attention module calculates the spatial attention weight , the temporal attention weight and the channel attention weight , and achieve feature adaptive enhancement through element-wise multiplication of the spatial-temporal-channel of the triple attention weights, and then perform weighted fusion with the original feature to form multi-scale global spatio-temporal features sensitive to disruptive behaviors; ; Among them, is the multi-scale global spatio-temporal feature.
[0012] Further preferably, the cross-modal adversarial discriminator adopts alternating training, and in each round of iteration, update the discriminator 3 times first and then update the generator 1 time; The first discriminator update inputs real data pairs , calculate the discriminator output , calculate the real loss , backpropagate to update the discriminator parameters ; The second discriminator update inputs the generated forged electrical features and the forged joint features , is the added Gaussian noise, calculate the forged discriminant loss , backpropagate to update the discriminator parameters ; The third discriminator update mixes real and forged electrical features, calculate the total loss , update the discriminator parameters to maximize the total loss ; then, freeze the discriminator parameters , generate the forged joint features , calculate the generator loss , backpropagate to update the generator parameters .
[0013] In step S4, the process of outputting the fusion behavior recognition probability and realizing the hierarchical response of external break with the electrical data at the current moment is as follows: Define the cable status anomaly detection function to quantify the sudden change in current exceeding 30% or the insulation resistance being lower than the safety threshold of the electrical anomaly event: ; Among them, is the current current, is the rated current; is the insulation resistance; When the improved spatio-temporal graph convolutional network outputs the behavior recognition probability , activate the electrical parameter verification, c represents the behavior category; the electrical parameter acquisition unit monitors the sudden change in cable current and the insulation resistance in real time; the electrical anomaly confidence is defined as: ; Establish the spatio-temporal alignment constraint condition : ; Among them, is the start time when the improved spatio-temporal graph convolutional network detects a dangerous behavior, is the electrical anomaly trigger time, s represents seconds; Judge the risk level: ; Trigger responses according to risk level classification.
[0014] Further preferably, an edge-side incremental learning module under the federated learning cooperation framework is constructed to update the parameters of the three-dimensional human pose estimation algorithm and the improved spatio-temporal graph convolutional network layer by layer.
[0015] The present invention also provides a cable external damage prevention human behavior recognition system based on spatio-temporal graph convolution, including a multispectral camera, an electrical parameter acquisition unit, an edge computing terminal, a spatio-temporal graph convolution inference module, a fusion analysis module, and a hierarchical response device; The multispectral camera is connected to the edge computing terminal, and the multispectral camera is used to collect image data; The electrical parameter acquisition unit collects the surface current fluctuation and insulation resistance value of the cable in real time; The spatio-temporal graph convolution inference module uses an improved spatio-temporal graph convolutional network to jointly model the multi-scale global spatio-temporal features of actions through multi-layer spatio-temporal convolution and outputs the behavior recognition probability; The edge computing terminal integrates an edge-side incremental learning module and a three-dimensional human pose estimation algorithm. The three-dimensional human pose estimation algorithm constructs a skeletal motion topology graph including time series dynamic association; the edge-side incremental learning module optimizes the parameters of the improved spatio-temporal graph convolutional network and the three-dimensional human pose estimation algorithm; The fusion analysis module uses a cross-modal adversarial discriminator to output the fusion behavior recognition probability and judge the risk level; The hierarchical response device specifies a hierarchical response strategy according to the risk level.
[0016] The present invention has the following advantages: High-precision behavior recognition: By using an improved spatio-temporal graph convolutional network and jointly modeling through multi-layer spatio-temporal convolution, it comprehensively captures the multi-scale global spatio-temporal features of actions. With the designs such as the spatial graph partitioning strategy, time dimension modeling, and dynamic adjacency matrix fusion, as well as the channel-spatio-temporal attention module to enhance features, the recognition accuracy of high-risk cable external damage behaviors such as excavation and knocking is significantly improved, overcoming the problem of low behavior recognition accuracy of traditional two-dimensional convolutional neural networks in complex scenarios.
[0017] Strong anti-interference ability: The multispectral camera adopts a three-light fusion imaging technology, where visible light, thermal infrared, and polarized light cooperate to take pictures and automatically switch the imaging mode according to the lighting conditions. It can accurately capture the human body contour and behaviors in different environments, effectively coping with interference factors such as object occlusion and sudden lighting changes in complex urban scenarios, and ensuring the stability of behavior recognition.
[0018] Fast response and low false alarm rate: Construct a visual-electrical two-way verification mechanism, and combine a cross-modal adversarial discriminator to fuse image and electrical features. Through spatio-temporal alignment constraints and risk level determination, the full-link calculation is completed within 200 ms. When the risk is high, the power supply can be quickly cut off, and when the risk is low, an early warning can be given in a timely manner, greatly reducing the false alarm rate and meeting the real-time protection requirements of the urban power grid.
[0019] Good adaptability and scalability: Construct an edge-side incremental learning module under the federated learning cooperation framework, and hierarchically update the parameters of the three-dimensional human pose estimation algorithm and the improved spatio-temporal graph convolutional network. It can adapt to the changes in behavior patterns brought by new construction equipment, continuously optimize the recognition ability, and improve the practicality and scalability of the system. Description of the Drawings
[0020] Figure 1 It is a flowchart of a cable external damage prevention human behavior recognition method based on spatio-temporal graph convolution; Figure 2 It is a structural diagram of an improved spatio-temporal graph convolutional network; Figure 3 It is a flowchart of a visual-electrical two-way verification mechanism and hierarchical alarm response. Detailed Embodiments
[0021] The following will describe in detail the specific embodiments of the present application with reference to the accompanying drawings. According to these detailed descriptions, those skilled in the art can clearly understand the present application. Without departing from the principle of the present application, the features in different embodiments can be combined to obtain new embodiments, or some features in some embodiments can be replaced to obtain other preferred embodiments.
[0022] As Figure 1 shown, the cable external damage prevention human behavior recognition method based on spatio-temporal graph convolution is as follows: S1: Obtain image data and electrical data through a multi-spectral camera and a sensor respectively; S2: Preprocess the image data and construct a skeletal motion topology graph including time series dynamic associations; S3: Use an improved spatio-temporal graph convolutional network (ST-GCN) to jointly model the multi-scale global spatio-temporal features of actions through multi-layer spatio-temporal convolution and output the behavior recognition probability; S4: Use a cross-modal adversarial discriminator to output the fused behavior recognition probability, and judge the risk level and formulate a hierarchical response strategy.
[0023] In step S1, image data is acquired by a multispectral camera. To improve the robustness of video behavior feature detection, the multispectral camera collects video data through visible light and thermal infrared dual-mode imaging technology, and the image sampling frequency is 30 fps. The video data is transmitted to the edge computing terminal through the PoE power supply interface. The multispectral camera includes: The polarized light channel uses a polarization sensor to capture the metal reflection characteristics through four polarizers in the directions of 0°, 45°, 90°, and 135°, and calculates the degree of polarization : ; In the formula respectively represent the light intensity values captured by the polarizers in the directions of 0°, 45°, 90°, and 135°.
[0024] Visible light channel and thermal infrared channel: Synchronously collect visible light images and thermal infrared images to solve interferences such as backlight and haze.
[0025] In step S2, the preprocessing of the image data refers to: fusing the three-channel data into a three-channel fusion feature map : ; In the formula, is the visible light image, is the thermal infrared image, are respectively the visible light image , the thermal infrared image , the degree of polarization 's weight coefficients, which are dynamically adjusted according to the light intensity (Lux value) to improve the video behavior feature detection rate in strong light scenarios.
[0026] In step S2, a skeletal motion topology map including time-series dynamic associations is constructed through a three-dimensional human pose estimation algorithm, as Figure 2 shown, and the specific steps are as follows: S2-1: Extract three-dimensional skeletal data and construct a spatio-temporal topology map: Use a pre-trained three-dimensional human pose estimation algorithm to extract the three-dimensional coordinates of 17 joints from the three-channel fusion feature map generated by preprocessing, and construct a spatio-temporal topology map ; among which the vertex set , represents the three-dimensional coordinates of the rd joint in the th frame, T is the duration; the edge set , includes spatial edges , connects adjacent joints based on human anatomical structure, represents the three-dimensional coordinates of the rd joint in the th frame; and temporal edges , associate the motion trajectories of the same joint in three consecutive frames . Denote the three-dimensional coordinates of the -th joint in the -th frame; the dynamic adjacency matrix contains the spatial connection of joints and the temporal motion association, where is the spatial adjacency matrix, encoding the natural connection relationship of human joints; is the temporal adjacency matrix, encoding the temporal continuity of joint motion; is the learnable weight coefficient; S2-2: Design a loss function as the continuity constraint of the bone trajectory, punishing the change amplitude of joint positions in adjacent frames: Traditional methods only smooth joint coordinates, which may produce postures that violate human kinematics, such as the reverse bending of the elbow joint. To suppress the sudden change of the bone trajectory caused by occlusion or blind spots, add physical constraints through the inverse kinematics model: First, calculate the joint Euler angles according to the three-dimensional coordinates of the joints , and compare them with the biomechanical limits , preset based on human anatomy and kinematics standards is the minimum allowable angle of the current degree of freedom, is the maximum allowable angle of the current degree of freedom.
[0027] Define the physical loss function : ; where t is the index of the frame and i is the joint number; Constrain the joint Euler angles within the biomechanical limits through backpropagation.
[0028] To solve the problem of bone jitter caused by video jitter, design a second-order smoothing loss to achieve the motion continuity constraint: ; In the formula, denotes the three-dimensional coordinates of the -th joint in the -th frame, denotes the three-dimensional coordinates of the -th joint in the
[0029] The improved spatio-temporal graph convolutional network (ST-GCN) is used to jointly model the spatio-temporal features of actions through multi-layer spatio-temporal convolutions. Referring to Figure 2 , the specific implementation steps are as follows: S3-1: The design of the improved spatio-temporal graph convolutional layer includes a spatial graph partitioning strategy, temporal dimension modeling, and dynamic adjacency matrix fusion; In the spatial dimension, the neighborhood nodes of each joint are divided into three categories through the spatial graph partitioning strategy: Root node: The current joint itself; Centripetal node: The adjacent joint pointing to the root node; Centrifugal node: The adjacent joint far from the root node; Each partition is assigned an independent learnable weight matrix to capture the local features of different motion patterns.
[0030] Temporal dimension modeling means that in the temporal dimension, a one-dimensional temporal convolution kernel with a size of is used to perform temporal modeling on the motion trajectories of the same joint in consecutive frames to capture the temporal changes of actions: ; Among them, is the weight of the k-th temporal convolution kernel, is the feature of the layer, is the feature input of the layer at time step t + k, is the bias term of the frame.
[0031] Set learnable weight coefficients to fuse the spatial adjacency matrix and the temporal adjacency matrix: ; Among them is the spatial adjacency matrix, is the temporal adjacency matrix, are the learnable weight coefficients of the spatial adjacency matrix and the temporal adjacency matrix respectively, controlling the strength of spatial and temporal correlations respectively.
[0032] Use the inverse square root of the degree matrix as the normalization factor to perform symmetric normalization on the dynamic adjacency matrix to ensure the numerical stability of the features during the graph convolution process and avoid gradient explosion or disappearance.
[0033] S3-2: Use a channel-spatio-temporal attention module to achieve feature enhancement: In the spatial dimension, first construct dynamic attention weights based on the spatio-temporal correlation of joint motion trajectories: Spatial attention performs concatenation mapping on the current joint feature and its neighborhood maximum response feature through two learnable parameter matrices respectively, realizes feature extraction and non-linear transformation, generates spatial attention weights, and generates spatial weights through the Sigmoid function , which is used to strengthen the motion features of key joints; The expression is: ; In the formula represents the spatial attention weight, and represent two learnable parameter matrices of spatial attention, ‖ represents concatenation, is the set of neighborhood nodes of the th joint, represents the feature of the th joint, represents the feature of the set of neighborhood nodes of the th joint, represents the maximum pooling operation, is the Sigmoid function.
[0034] Temporal attention models the temporal dependence relationship through the LSTM network and calculates the frame-level temporal attention weight to highlight the key frames of the action. The expression is: ; Among them, is the learnable parameter matrix of temporal attention, is the temporal feature sequence of the th joint within the time window from 1 to T. LSTM is the long short-term memory network; T is the duration; Channel attention captures the statistical characteristics of the feature channels through global average pooling and generates the channel attention weight to suppress the noise feature dimensions. The expression is: ; Among them, is the learnable parameter matrix of channel attention, is the original feature; The process of feature fusion is as follows: Feature adaptive enhancement is achieved through element-wise multiplication of the spatial-time-channel of the triple attention weights, and then weighted fusion is performed with the original feature to form multi-scale global spatio-temporal features sensitive to disruptive behaviors; ; Among them, is the multi-scale global spatio-temporal feature.
[0035] This mechanism enables the model to have high feature focusing accuracy for small but crucial action patterns in complex backgrounds.
[0036] S3-3: Adopt the design of spatio-temporal feature classification layer to process multi-scale global spatio-temporal features After global average pooling, input to the fully connected layer, and output the behavior recognition probability through the Softmax function: ; In the formula, is the behavior recognition probability, is the classification weight matrix, represents the dimension, C is the number of behavior categories, and D is the feature dimension; is the bias term.
[0037] In step S4, to achieve the joint determination of behavior risk and abnormal electrical parameters, build a multi-modal adversarial training framework, combine the transient analysis of electrical parameters carried out by the lightweight current sensor and the insulation resistance detection module, and establish a two-way verification mechanism for the visual behavior recognition - electrical linkage model. Its mathematical modeling includes the following core elements: Refer to Figure 3 , to solve the modal differences between image data and electrical data, design a cross-modal adversarial discriminator: Generator : Map the image features and electrical features to a common embedding space to generate joint features: ; In the formula, is the joint feature, is the image feature, which is extracted from image data and comes from the ST-GCN network; is the electrical feature, which is extracted from electrical data and comes from the sensor; Discriminator : Judge whether the feature pair comes from the true data distribution. The loss function is: ; In the formula, is the discriminator loss, is the mathematical expectation, that is, the calculation of the expected value of all samples in the data distribution; is the forged electrical feature with added Gaussian noise, forcing the generator to learn noise invariance. Adopt alternating training, update the discriminator 3 times first and then update the generator 1 time in each round of iteration to avoid mode collapse.
[0038] In each round of iteration, update the discriminator 3 times first: The first discriminator update inputs the true data pair , calculate the discriminator output , calculate the real loss , update the discriminator parameters through backpropagation ; The second discriminator update inputs the forged electrical features and the forged joint features , , calculate the forged discriminant loss for the added Gaussian noise , update the discriminator parameters through backpropagation ; The third discriminator update mixes the real and forged electrical features and calculates the total loss , update to maximize ; In each iteration, the generator is also updated once: Freeze the discriminator parameters , generate the forged joint features , calculate the generator loss , update the generator parameters through backpropagation .
[0039] Constraining the covariance matrix of the image features and the electrical features to approach zero and retaining the unique information of each modality has a significant supplementary effect on the detection of cable external damage, especially when the insulation layer is slightly damaged and the electrical signal is weak.
[0040] First, define the cable status anomaly detection function to quantify the sudden change in current exceeding 30% or the insulation resistance being lower than the safety threshold for electrical anomaly events: ; where is the current at present, is the rated current; is the insulation resistance; When the output behavior recognition probability of the improved spatio-temporal graph convolutional network (ST-GCN) is and the electrical parameter verification is activated by the system. c represents the behavior category. The electrical parameter acquisition unit monitors the sudden change in cable current and the insulation resistance in real time. The electrical anomaly confidence is defined as: Furthermore, establish the spatio-temporal alignment constraint condition , forcing the start time when the improved spatio-temporal graph convolutional network detects a dangerous behavior Synchronize within a time window of ±0.5 seconds to avoid misassociating non-causal events; ; In the formula, s represents seconds.
[0041] Determine the risk level: ; Trigger responses according to the risk level classification.
[0042] Through strict spatio-temporal correlation and multi-modal feature fusion, the false alarm rate of traditional single-modal systems is greatly reduced, and at the same time, the full-link calculation is completed within 200 ms to meet the real-time protection requirements of urban power grids.
[0043] Implement a hierarchical response mechanism for different risk levels: for low-risk events, the system will activate the on-site sound and light alarm device for early warning, and at the same time upload the alarm record to the central monitoring platform; when a high-risk situation is detected, the hierarchical alarm system will trigger the intelligent circuit breaker power cut function within 200 ms, and synchronously transmit the fault location data to the repair terminal in real time.
[0044] Furthermore, to adapt to new construction equipment, such as hydraulic shears and laser cutters, and effectively identify different types of damage, an edge-side incremental learning module under the federated learning cooperation framework is constructed to update the parameters of the three-dimensional human pose estimation algorithm and the improved spatio-temporal graph convolutional network layer by layer: Frozen layer: The backbone network parameters of HRNet (three-dimensional human pose estimation algorithm) are fixed, retaining the general feature extraction ability to avoid catastrophic forgetting; Fine-tunable layer: Only update the graph convolution weights of the last two layers of the improved spatio-temporal graph convolutional network (ST-GCN), and the update formula is: ; In the formula, is the gradient, is the learning rate, is the elastic coefficient, is the newly generated data set, is the learnable parameter matrix of time attention, is the federated learning global model parameter to prevent the local model from deviating too far.
[0045] Build a federated learning cooperation framework, aggregate differential privacy parameters, and each edge node Based on local data Calculate the gradient For local training; Add Laplace noise to protect privacy: ; In the formula, is an edge node 's gradient, is the gradient of the edge node after adding Laplace noise 's gradient, is a Laplace noise generator, is the sensitivity, which can take the value of 0.1, is the privacy budget, satisfying differential privacy.
[0046] A cable anti-external damage human behavior recognition system based on spatio-temporal graph convolution includes a multispectral camera, an electrical parameter acquisition unit, an edge computing terminal, a spatio-temporal graph convolution inference module, a fusion analysis module, and a hierarchical response device; The multispectral camera adopts triple-light fusion imaging technology, supports collaborative shooting of visible light, thermal infrared, and polarized light, enables visible light high-definition imaging when the daytime light is sufficient, and automatically switches to the thermal imaging mode under low visibility conditions such as night or haze. It enhances the contour extraction ability through the infrared radiation characteristics emitted by the human body. Its image sampling frequency is 30fps, and it is powered by PoE and connected to the edge computing terminal.
[0047] The electrical parameter acquisition unit collects the surface current fluctuation and insulation resistance value of the cable in real time; The improved spatio-temporal graph convolutional network (ST-GCN) jointly models the multi-scale global spatio-temporal features of actions through multi-layer spatio-temporal convolution and outputs the behavior recognition probability.
[0048] The edge computing terminal adopts an embedded intelligent chip, integrates an edge-side incremental learning module and a three-dimensional human pose estimation algorithm, can parse the three-dimensional spatial coordinates of 17 key bone nodes of the human body in real time through the depth visual data obtained by the multispectral camera, and constructs a bone motion topology graph containing time-series dynamic associations; this bone motion topology graph is double-modeled through the connection relationship of joint points in the spatial dimension and the continuity of actions in the time dimension, providing a structured feature expression for subsequent behavior recognition; The fusion analysis module solves the modal differences between videos and electrical signals by building a multi-modal adversarial training framework, jointly analyzes multi-modal data, collects the surface current fluctuation and insulation resistance value of the cable in real time through the current sensor and the insulation resistance detection module, and combines the edge computing node for transient analysis of electrical parameters; when the video behavior recognition detects a cable damage behavior, it automatically activates the spatio-temporal alignment verification of the cable status monitoring unit, and realizes double verification through the matching degree calculation of the electrical anomaly time window of ±0.5s and the dangerous action video segment, and controls the system response delay within 200ms.
[0049] The hierarchical response device includes an acoustic-optic warning unit, a remote alarm interface, and an intelligent circuit breaker controller, and adopts a hierarchical response strategy: Low risk: If dangerous actions (such as excavation, metal knocking) are identified, start the on-site strong light flashing warning and synchronously upload the alarm information to the monitoring center; High risk: When video analysis confirms the existence of cable damage behavior by personnel and the electrical abnormal data exceeds the safety threshold, automatically cut off the power supply of the corresponding section and push the positioning information to the emergency repair terminal.
[0050] The overall system architecture collects video data through a multi-spectral camera. The edge computing terminal extracts three-dimensional skeleton data and constructs a spatio-temporal topology. The spatio-temporal graph convolutional inference module completes the analysis and classification of behavior characteristics. Combining with the two-way verification mechanism of the electrical parameter acquisition unit, finally trigger the hierarchical response device to respond.
[0051] The above only expresses the preferred embodiments of the present invention and does not limit the present invention in other forms. Any person skilled in the art may use the disclosed content to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention still belong to the protection scope of the technical solution of the present invention.
Claims
1. A method for identifying human behaviors of preventing external damage to cables based on spatio-temporal graph convolution, characterized in that the steps As follows: S1: Obtain image data and electrical data through a multispectral camera and sensors respectively; S2: Preprocess the image data and construct a skeletal motion topology map containing time-series dynamic associations; S3: Use an improved spatio-temporal graph convolutional network to jointly model the multi-scale global spatio-temporal features of actions through multi-layer spatio-temporal convolutions and output the behavior recognition probability; The improved spatio-temporal graph convolutional network includes an improved spatio-temporal graph convolutional layer and a channel-spatio-temporal attention module. The design of the improved spatio-temporal graph convolutional layer includes a spatial graph partitioning strategy, time dimension modeling, and dynamic adjacency matrix fusion; S4: Use a cross-modal adversarial discriminator to output the fused behavior recognition probability, and judge the risk level and formulate a hierarchical response strategy; The cross-modal adversarial discriminator includes: A generator G for mapping image features and electrical features to a common embedding space to generate joint features: ; In the formula, is the combined feature, is the image feature, and the image feature is extracted from the image data; is the electrical feature, and the electrical feature is extracted from the electrical data; Discriminator : Determine whether the feature pair comes from the true data distribution, and the loss function is: ; In the formula, is the discriminator loss, is the mathematical expectation, is the forged electrical feature with added Gaussian noise.
2. The cable anti-external-breakage human behavior recognition method according to claim 1, wherein The preprocessing of the image data in step S2 means: fusing the three-channel data into a three-channel fused feature map : ; In the formula, are respectively the visible light image , the thermal infrared image , and the polarization degree 's weight coefficients, which are dynamically adjusted according to the illumination intensity.
3. The cable anti-external damage human behavior recognition method according to claim 1, characterized in that, In step S2, construct a skeletal motion topology map containing time-series dynamic associations through a three-dimensional human pose estimation algorithm.
4. The cable anti-external-break human behavior recognition method according to claim 3, characterized in that, In step S2, use a pre-trained three-dimensional human pose estimation algorithm to extract the three-dimensional coordinates of 17 joints from the three-channel fusion feature map generated by preprocessing and construct a spatio-temporal topology map; Design a loss function as a constraint on the continuity of the skeletal trajectory to punish the change amplitude of the joint positions in adjacent frames; Calculate the Euler angles of the joints based on their three-dimensional coordinates and compare them with the biomechanical limits , preset according to human anatomy and kinematic standards where is the minimum allowable angle of the current degree of freedom and is the maximum allowable angle of the current degree of freedom; Define the physical loss function : ; where t is the index of the frame and i is the joint number; Constrain the joint Euler angles within the biomechanical limit through backpropagation; Design second-order smoothing loss Implement motion continuity constraints.
5. The cable anti-external-breakage human behavior recognition method according to claim 1, characterized in that, The improved spatio-temporal graph convolutional layer: In the spatial dimension, divide the neighborhood nodes of each joint into three categories through a spatial graph partitioning strategy: Root node: The current joint itself; Centripetal node: The adjacent joint pointing to the root node; Centrifugal node: The adjacent joint far from the root node; Assign an independent learnable weight matrix to each partition to capture the local features of different motion patterns and obtain a spatial adjacency matrix; Time dimension modeling means: In the time dimension, use a one-dimensional time convolution kernel to perform temporal modeling on the motion trajectories of the same joint in consecutive frames to capture the temporal changes of the action; Set a learnable weight coefficient to fuse the spatial adjacency matrix and the temporal adjacency matrix into a dynamic adjacency matrix; Use the inverse square root of the degree matrix as a normalization factor to symmetrically normalize the dynamic adjacency matrix.
6. The cable anti-external damage human behavior recognition method according to claim 1, characterized in that The channel-spatiotemporal attention module calculates the spatial attention weights respectively , the temporal attention weights , and the channel attention weights , and realizes the feature adaptive enhancement through the element-wise multiplication of the spatial-temporal-channel of the triple attention weights, and then performs weighted fusion with the original features to form multi-scale global spatiotemporal features; ; Among them, is the multi-scale global spatio-temporal feature.
7. The cable anti-external damage human behavior recognition method according to claim 1, characterized in that The cross-modal adversarial discriminator uses alternating training, updating the discriminator 3 times first and then the generator 1 time in each round of iteration; The first discriminator update input real data pair , calculate the discriminator output , calculate the real loss , backpropagate to update the discriminator parameters ; Forged electrical features generated by the second discriminator update input and forged joint features , For the added Gaussian noise, calculate the forged discriminant loss , and update the discriminator parameters by backpropagation ; The third discriminator updates the mixed real and forged electrical features and calculates the total loss , and updates the discriminator parameters to maximize the total loss ; Then, freeze the discriminator parameters , generate forged joint features , calculate the generator loss , and update the generator parameters by backpropagation .
8. The cable anti-external damage human behavior recognition method according to claim 1, characterized in that In step S4, the process of outputting the fused behavior recognition probability and realizing a hierarchical response to external breakdown with the electrical data at the current moment is as follows: Define a cable status anomaly detection function to quantify the sudden change in current Exceeding 30% or the insulation resistance being lower than the safety threshold of electrical anomaly events: ; Among them, is the current at present, is the rated current; is the insulation resistance; When the improved spatio-temporal graph convolutional network outputs the behavior recognition probability activate the electrical parameter verification, where c represents the behavior category; the electrical parameter acquisition unit monitors the sudden change in cable current and the insulation resistance in real time; the electrical anomaly confidence is defined as: ; Establish spatio-temporal alignment constraint conditions : ; Among them, is the starting time of the dangerous behavior detected by the improved spatio-temporal graph convolutional network, is the electrical anomaly trigger time, and s represents seconds; Judge the risk level: , Trigger a response according to the risk level classification.
9. The cable anti-external damage human behavior recognition method according to claim 1, characterized in that Construct an edge-side incremental learning module under the federated learning collaboration framework to update the parameters of the three-dimensional human pose estimation algorithm and the improved spatio-temporal graph convolutional network layer by layer.
10. A cable anti-external damage human behavior recognition system based on spatio-temporal graph convolution, characterized in that It includes a multispectral camera, an electrical parameter acquisition unit, an edge computing terminal, a spatio-temporal graph convolution inference module, a fusion analysis module, and a hierarchical response device; The multispectral camera is connected to the edge computing terminal, and the multispectral camera is used to collect image data; The electrical parameter acquisition unit collects the surface current fluctuation and insulation resistance value of the cable in real time; The spatio-temporal graph convolution inference module uses an improved spatio-temporal graph convolutional network to jointly model the multi-scale global spatio-temporal features of actions through multi-layer spatio-temporal convolutions and outputs the behavior recognition probability; The edge computing terminal integrates an edge-side incremental learning module and a three-dimensional human pose estimation algorithm, and the three-dimensional human pose estimation algorithm constructs a skeletal motion topology map containing time-series dynamic associations; The edge-side incremental learning module optimizes the parameters of the improved spatio-temporal graph convolutional network and the three-dimensional human pose estimation algorithm; The fusion analysis module uses a cross-modal adversarial discriminator to output the fusion behavior recognition probability and judge the risk level; The hierarchical response device specifies a hierarchical response strategy according to the risk level.
Citation Information
Patent Citations
Human behavior recognition method and electronic equipment
CN110472612A
Electric power system field operation action risk identification method based on graph convolution
CN112200030A
Cross-modal identification technology based on adversarial network
CN112711670A
Complex supply and transmission mechanism fault diagnosis method based on sparse self-encoding auxiliary classification generative adversarial network
CN114676733A
Multi-modal anomaly detection method based on graph attention network and time convolution network
CN116701992A
Cited By
Microgrid digital twin modeling method
CN120764117A
Overhead line threat grading evaluation method and system based on dynamic trajectory tracking
CN120851626A
Iron tower safety operation monitoring method and system based on intelligent AI identification and medium
CN121581620A
Action recognition method and system based on human skeleton structure
CN121861729A