A multi-sensor fusion-based intelligent inspection risk assessment method and system
By using multi-sensor fusion technology and implicit neural representation, the problems of limited information and environmental interference in existing technologies have been solved, enabling accurate risk assessment and safe distance calculation for substation inspections, and improving the level of intelligence in substation inspections.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-10
- Publication Date
- 2026-03-17
AI Technical Summary
Existing binocular vision-based inspection and assessment technologies suffer from limited information, susceptibility to environmental interference, and a lack of multi-source data fusion capabilities, making it difficult to achieve a comprehensive and accurate assessment of substation inspection risks.
By employing multi-sensor fusion technology, and through spatiotemporal calibration and preprocessing of binocular visual images, infrared thermal imaging data, and environmental sensor data, a multimodal feature set and scene relationship graph are constructed. Combined with implicit neural representation and deep learning models, risk assessment and safe distance calculation are performed to generate comprehensive risk warning information.
It enables accurate risk assessment and early warning during substation inspection, improves environmental adaptability and the accuracy of risk assessment, and enhances the reliability of safe distance calculation and the practicality of the early warning system.
Smart Images

Figure CN121279619B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power system safety operation and maintenance, and in particular to an intelligent inspection risk assessment method and system based on multi-sensor fusion, which is applicable to substation equipment status monitoring, risk assessment and early warning prevention. Background Technology
[0002] Intelligent substation inspection technology is a key area for the safe operation and maintenance of power systems, involving the condition monitoring, risk assessment, and early warning and prevention of substation equipment. With the development of IoT and AI technologies, intelligent inspection technology has gradually evolved from traditional manual inspections towards automation and intelligence, becoming an important means to ensure the safe and stable operation of power systems.
[0003] Traditional substation inspection technologies primarily rely on single types of sensing devices to acquire information, such as using infrared thermal imagers to detect abnormal equipment temperatures or employing visual cameras to identify surface defects. Commercially available infrared temperature-measuring inspection robots can monitor the temperature distribution of substation equipment in real time, while computer vision-based inspection systems can detect abnormal conditions in the equipment's appearance.
[0004] In recent years, a binocular vision-based inspection and assessment technology has been gradually applied to substation inspections. This technology uses binocular cameras to acquire stereoscopic images of the substation scene, employs computer vision algorithms for image analysis, identifies equipment defects, and calculates three-dimensional spatial distances to assess safety risks during the inspection process. This technology provides more accurate spatial positioning information through stereo vision, helping to determine the safe distance between inspection personnel and equipment.
[0005] However, existing binocular vision-based inspection and assessment technologies have several obvious technical shortcomings: First, the information acquired by a single sensor is limited and it is difficult to fully reflect the complex environmental conditions of a substation; second, the image processing is easily affected by factors such as ambient light and shadow, resulting in unstable recognition results; in addition, existing technologies lack the ability to effectively fuse and process multi-source heterogeneous data, and cannot fully integrate the complementary advantages of different sensors, making it difficult to achieve a comprehensive and accurate assessment of inspection risks. Summary of the Invention
[0006] The purpose of this invention is to overcome the shortcomings of the prior art and provide an intelligent inspection risk assessment method and system based on multi-sensor fusion. Through multi-source data acquisition and fusion, risk model construction based on implicit neural representation, three-dimensional spatial safety distance calculation and comprehensive risk early warning, it can achieve accurate assessment and early warning of risks during substation inspection.
[0007] To achieve the above objectives, this invention provides an intelligent inspection risk assessment method based on multi-sensor fusion, comprising the following steps:
[0008] The raw multi-source data of the substation is acquired, including binocular vision images, infrared thermal imaging data and environmental sensor data. The raw multi-source data is then spatiotemporally calibrated and preprocessed to obtain a multi-source sensor data stream.
[0009] Multimodal feature sets are obtained by extracting features from the multi-source sensor data stream. A refined semantic mask is then generated based on the multimodal feature sets, and an initial scene relationship graph is constructed. The reliability of the initial scene relationship graph is evaluated using a statistical confidence re-scoring mechanism to obtain an optimized scene relationship graph. Finally, the multimodal feature sets and the optimized scene relationship graph are fused to output a fused feature vector and a target scene relationship graph.
[0010] Based on the fused feature vector and the target scene relationship graph, a risk representation model is constructed using the implicit neural representation method, and a deep weight space network is constructed based on the risk representation model. The scene features are mapped to the risk space through the deep weight space network, and a risk feature space mapping is output. Based on the output risk feature space mapping, a relationship attention converter model is used to obtain relationship-enhanced risk features. The risk level is calculated based on a multilayer perceptron classifier, and risk level assessment results and risk distribution maps are generated.
[0011] Based on the binocular vision image in the multi-source sensor data stream, a high-precision scene depth map is generated by optimizing the stereo matching process through deep distillation gradient preprocessing. 3D point cloud reconstruction is performed, and the 3D coordinates of inspection personnel and key equipment are identified. The 3D Euclidean distance between the inspection personnel and the key equipment is calculated. Based on the 3D Euclidean distance, the real-time safe distance measurement result is calculated. The safe distance threshold is dynamically adjusted according to the risk level assessment result and the operating status of the key equipment. By comparing the real-time safe distance measurement result and the safe distance threshold, a safe distance violation warning and risk area identification are output.
[0012] Based on the risk level assessment results, the risk distribution map, the safety distance violation warning, and the risk area identifier, a comprehensive risk warning information is generated.
[0013] The acquisition of raw multi-source data from the substation includes: deploying a binocular vision sensor, an infrared thermal imager, and an environmental sensor array; performing spatiotemporal calibration on each sensor and establishing a unified coordinate system and time reference; outputting a sensor calibration parameter matrix; establishing a multi-sensor synchronous acquisition mechanism based on the sensor calibration parameter matrix; setting the acquisition frequency of different sensors and adding a unified timestamp; and outputting the raw multi-source data with timestamps. The spatiotemporal calibration and preprocessing of the raw multi-source data to obtain a multi-source sensor data stream includes: performing deep distillation gradient preprocessing on the binocular vision image and the infrared thermal image data in the timestamped raw multi-source data to achieve denoising, distortion correction, and enhancement; using a moving average filtering algorithm to denoise and filter outliers on the environmental sensor data to obtain preprocessed binocular vision images, infrared thermal image data, and environmental sensor data; and performing spatiotemporal alignment on the preprocessed binocular vision image, infrared thermal image data, and environmental sensor data through linear interpolation and coordinate transformation to obtain the multi-source sensor data stream.
[0014] The method of employing a statistical confidence re-scoring mechanism to assess the reliability of the initial scenario relationship graph includes: extracting statistical priors from a pre-defined substation scenario graph database, wherein the statistical priors include node category distribution probability, condition category probability, relationship co-occurrence probability, and relationship transitivity; calculating the corresponding node confidence score and edge confidence score for each node and edge in the initial scenario relationship graph based on the statistical priors; and evaluating and analyzing the initial scenario relationship graph based on the node confidence score and the edge confidence score to obtain the re-evaluated initial scenario relationship graph.
[0015] The method for constructing a risk representation model based on implicit neural representation includes: defining a risk representation function, which achieves continuous mapping from spatial coordinates to risk attributes through multilayer perceptron; encoding the fused feature vector through a feature mapping network to obtain a latent feature representation; concatenating the latent feature representation with the three-dimensional spatial coordinates and inputting it into the risk representation function to construct a conditional implicit neural representation; processing the target scene relationship graph using a graph neural network to obtain the latent representation of each node in the target scene relationship graph, and calculating the spatial relationship between the three-dimensional spatial coordinates and each node, aggregating them through an attention mechanism to obtain graph conditional features; concatenating the latent feature representation, the graph conditional features, and the three-dimensional spatial coordinates and inputting them into the risk representation function to construct an extended conditional implicit neural representation, thus obtaining an implicit neural representation model; training the implicit neural representation model using historical data as a supervision signal, and optimizing the model parameters based on minimizing the difference between the predicted risk and the actual risk to obtain the risk representation model.
[0016] The method of generating a high-precision scene depth map by optimizing the stereo matching process through deep distillation gradient preprocessing includes: converting the binocular vision image into a gradient domain, calculating the horizontal and vertical gradient maps of the image to obtain an original gradient map; using a pre-trained deep neural network as a teacher network to process the original gradient map and generate a sanitized gradient; training a lightweight student network to distill knowledge from the teacher network and learn the mapping relationship that converts the original gradient map into the sanitized gradient; processing the original gradient map based on the mapping relationship to obtain an enhanced gradient map; and executing a stereo matching algorithm on the enhanced gradient map to calculate a disparity map, and converting the disparity map into a depth map using a disparity-depth conversion formula to obtain the high-precision scene depth map.
[0017] The calculation of the three-dimensional Euclidean distance between the inspection personnel and the key equipment includes: reconstructing three-dimensional space based on the high-precision scene depth map and camera intrinsic and extrinsic parameters to generate a three-dimensional point cloud model representing the substation scene; identifying the inspection personnel and the key equipment in the scene based on the three-dimensional point cloud model, the fused feature vector, and the target scene relationship graph, and determining their precise position coordinates in three-dimensional space; extracting surface point sets based on the three-dimensional point cloud model of the key equipment, generating human body surface point sets based on the inspection personnel skeleton model, and calculating the minimum distance between the two point sets; combining the movement trajectory and speed of the inspection personnel to predict short-term movement trends, assessing the possible minimum safe distance, and obtaining the three-dimensional Euclidean distance.
[0018] The step of dynamically adjusting the safety distance threshold based on the risk level assessment results and the operating status of the key equipment includes: adjusting the baseline safety distance threshold based on the equipment load rate, temperature, and vibration parameters according to the risk level assessment results and the operating status of the key equipment; adjusting the baseline safety distance threshold after equipment status adjustment based on the current temperature, humidity, and gas concentration; increasing the baseline safety distance threshold for areas assessed as medium or high risk levels based on the risk level assessment results; and adjusting the safety distance threshold according to the current operation type and qualification level of the inspection personnel to obtain the dynamic safety distance threshold under the current scenario.
[0019] The generation of comprehensive risk warning information includes: based on the risk level assessment results, the risk distribution map, the safety distance violation warning, and the risk area identification, comprehensively analyzing the risk type, risk level, and risk distribution of the inspection status, and outputting a comprehensive risk assessment result; generating corresponding warning information for different risk levels, including risk type, location, severity, and handling suggestions; encrypting the warning information using the AES encryption algorithm, transmitting it to the monitoring center using LoRa or NB-IoT low-power wide-area network communication technology, and outputting a real-time warning signal; storing the comprehensive risk assessment result, raw sensor data, and intermediate processing results in a distributed database according to time series to establish a risk assessment database, and generating the risk assessment report containing risk trend analysis, typical risk cases, and improvement suggestions.
[0020] This invention also provides an intelligent inspection risk assessment system based on multi-sensor fusion, comprising:
[0021] The multi-source sensor data acquisition module is used to acquire raw multi-source data from the substation, including binocular visual images, infrared thermal imaging data, and environmental sensor data. The module performs spatiotemporal calibration and preprocessing on the raw multi-source data to obtain a multi-source sensor data stream.
[0022] The multi-sensor data fusion module is used to extract features from the multi-source sensor data stream to obtain a multimodal feature set, generate a refined semantic mask based on the multimodal feature set and construct an initial scene relationship graph, and use a statistical confidence re-scoring mechanism to evaluate the reliability of the initial scene relationship graph to obtain an optimized scene relationship graph, and fuse the multimodal feature set and the optimized scene relationship graph to output a fused feature vector and a target scene relationship graph;
[0023] The risk assessment module is used to construct a risk representation model based on the implicit neural representation method according to the fused feature vector and the target scene relationship graph, and to construct a deep weight space network based on the risk representation model. The deep weight space network maps scene features to risk space and outputs risk feature space mapping. Based on the output risk feature space mapping, a relationship attention converter model is used to obtain relationship-enhanced risk features. The risk level is calculated based on a multilayer perceptron classifier and risk level assessment results and risk distribution map are generated.
[0024] The safety distance calculation module is used to generate a high-precision scene depth map based on the binocular vision image in the multi-source sensor data stream through deep distillation gradient preprocessing to optimize the stereo matching process, perform 3D point cloud reconstruction and identify the 3D coordinates of inspection personnel and key equipment, calculate the 3D Euclidean distance between the inspection personnel and the key equipment, calculate the real-time safety distance measurement result based on the 3D Euclidean distance, dynamically adjust the safety distance threshold according to the risk level assessment result and the operating status of the key equipment, and output a safety distance violation warning and risk area identification by comparing the real-time safety distance measurement result and the safety distance threshold.
[0025] The comprehensive early warning module is used to generate comprehensive risk early warning information based on the risk level assessment results, the risk distribution map, the safety distance violation warning, and the risk area identifier.
[0026] The system also includes: an edge computing unit, configured with a GPU accelerator card and a large-capacity storage device, for real-time processing of multi-source sensor data streams and running deep learning algorithms; a wireless communication unit, integrating LoRa and NB-IoT communication modules, supporting AES encryption algorithm, for secure communication with the monitoring center; a human-machine interaction unit, including a high-resolution display screen and touch interface, for real-time display of risk distribution maps and early warning information; a power management unit, using a combination of solar panels and lithium batteries for power supply, supporting uninterrupted operation; and an environmental adaptability unit, designed with an IP65 protection rating, operating temperature range of -40℃ to +70℃, adapting to the harsh environment of substations.
[0027] The beneficial effects of this invention are as follows:
[0028] 1. This invention solves the problem of limited information from a single sensor by acquiring and preprocessing data from multiple sensors, and can comprehensively reflect the complex environmental conditions of a substation;
[0029] 2. This invention employs a statistical confidence re-scoring mechanism, which significantly improves the accuracy and robustness of the scene relationship diagram, providing a reliable foundation for scene understanding in risk assessment;
[0030] 3. The risk model constructed based on implicit neural representation in this invention achieves a unified expression of risk information at different resolutions and scales, thereby improving the generalization ability and accuracy of risk assessment;
[0031] 4. This invention optimizes the stereo matching process through deep distillation gradient preprocessing, improving the accuracy of depth estimation and the quality of 3D reconstruction in complex substation environments, and providing a reliable basis for safe distance calculation;
[0032] 5. This invention employs a dynamic safety threshold adjustment mechanism, which makes the determination of safe distance more in line with the needs of actual scenarios, thereby improving the accuracy and practicality of the early warning system. Attached Figure Description
[0033] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings used in the embodiments will be briefly described below. These drawings are incorporated in and constitute a part of this specification. They illustrate embodiments conforming to this disclosure and, together with the specification, serve to explain the technical solutions of this disclosure. It should be understood that the following drawings only show some embodiments of this disclosure and should not be considered as limiting the scope. Those skilled in the art can obtain other related drawings based on these drawings without creative effort.
[0034] Figure 1 This invention provides a flowchart of an intelligent inspection risk assessment method based on multi-sensor fusion.
[0035] Figure 2 This invention provides an architecture diagram of an intelligent inspection risk assessment system based on multi-sensor fusion. Detailed Implementation
[0036] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments.
[0037] like Figure 1 As shown, this invention provides an intelligent inspection risk assessment method based on multi-sensor fusion, comprising the following steps:
[0038] Step 1: Acquire the raw multi-source data of the substation, which includes binocular vision images, infrared thermal image data and environmental sensor data. Perform spatiotemporal calibration and preprocessing on the raw multi-source data to obtain a multi-source sensor data stream.
[0039] This invention first acquires raw data, including binocular visual images, infrared thermal imaging data, and environmental sensor data, using a multi-source sensor array deployed in a substation. The key to this step lies in achieving spatiotemporal synchronization and calibration between different sensors, ensuring that various types of data can be fused and analyzed within a unified spatiotemporal reference frame.
[0040] The intelligent inspection system (hereinafter referred to as the system) first performs precise spatiotemporal calibration on each sensor to establish a unified coordinate system and time reference. In the time dimension, the system sets a unified timestamp for sensors with different acquisition frequencies to solve the problem of inconsistent data timing. In the spatial dimension, the system determines the relative positional relationship between each sensor through a calibration algorithm and generates a calibration parameter matrix, providing a foundation for subsequent data fusion.
[0041] For binocular vision images and infrared thermal imaging data, the system employs deep distillation gradient preprocessing technology. This innovative technology optimizes image gradient information through knowledge distillation, effectively removing noise, correcting distortion, and enhancing image details, especially addressing challenges common in substation environments such as uneven lighting, reflections, and shadows. For environmental sensor data (such as temperature, humidity, and gas concentration), the system uses a moving average filtering algorithm for noise reduction and outlier filtering to ensure data reliability.
[0042] Finally, the system performs spatiotemporal alignment on the preprocessed data, mapping data from different sources to a unified spatiotemporal reference frame through linear interpolation and coordinate transformation techniques, outputting a standardized and synchronized multi-source sensor data stream. This carefully processed data stream lays a solid foundation for subsequent feature extraction and fusion analysis.
[0043] Step 2: Extract features from the multi-source sensor data stream to obtain a multimodal feature set, generate a refined semantic mask based on the multimodal feature set and construct an initial scene relationship graph, and use a statistical confidence re-scoring mechanism to evaluate the reliability of the initial scene relationship graph to obtain an optimized scene relationship graph, and fuse the multimodal feature set and the optimized scene relationship graph to output a fused feature vector and a target scene relationship graph;
[0044] After acquiring standardized multi-source sensor data streams, this step focuses on extracting and fusing valuable feature information to build an accurate scene understanding model. This process is divided into several sub-stages, including multimodal feature extraction, semantic mask generation, scene graph construction, statistical confidence rescoring, and feature fusion.
[0045] First, the system extracts features from different types of sensor data. For binocular vision images, the system uses a deep convolutional neural network to extract visual features containing spatial information; for infrared thermal imaging data, the system applies a specially optimized thermal imaging processing network to capture temperature distribution patterns; and for environmental sensor data, it extracts environmental state features through a fully connected network using time-series analysis. These features each capture different aspects of the substation environment, collectively forming a multimodal feature set.
[0046] Based on visual features, the system further generates refined semantic masks, enabling pixel-level classification and identification of targets such as equipment and personnel in the scene. To improve accuracy, the system incorporates prior knowledge learned from a large number of substation images, such as statistical characteristics of typical equipment shapes, sizes, and positional relationships. It also combines depth information to assist in segmentation, effectively addressing recognition challenges under complex lighting conditions.
[0047] Next, the system constructs an initial scene relationship graph based on a multimodal feature set and refined semantic masks. This graph structure uses nodes to represent entities in the scene (such as equipment and personnel) and edges to represent relationships between entities (such as spatial location relationships and functional connection relationships). The system employs a bidirectional reasoning mechanism to construct relationships: bottom-up extraction of relationships from observation data, and top-down verification and supplementation of relationships through domain knowledge.
[0048] One of the key innovations of this invention is the statistical confidence re-scoring mechanism. The system extracts statistical priors from a pre-established substation scene graph database, including node category distribution probabilities, conditional category probabilities, relationship co-occurrence probabilities, and relationship transitivity. Based on these statistical priors, the system calculates confidence scores for nodes and edges in the initial scene relationship graph, and then optimizes them: retaining nodes and edges with high confidence, correcting node categories with low confidence, adjusting or deleting unreliable relationships, and adding high-probability relationships that might be missed. This statistical learning-based scoring mechanism effectively handles pseudo-point noise in depth map reconstruction and background noise in image features, significantly improving the accuracy and robustness of the scene graph.
[0049] Finally, the system fuses the multimodal feature set with the optimized scene relationship graph. By using a graph neural network to achieve message passing and feature updates between nodes, the system not only considers the connectivity between nodes but also designs a relationship-aware feature aggregation mechanism, assigning different weights to different types of relationships to accurately capture the semantic associations between nodes. The final output fused feature vector and target scene relationship graph contain rich contextual information, providing a comprehensive scene representation for subsequent risk assessment.
[0050] Step 3: Based on the fused feature vector and the target scene relationship graph, a risk representation model is constructed using the implicit neural representation method, and a deep weighted space network is constructed based on the risk representation model. The scene features are mapped to the risk space through the deep weighted space network, and a risk feature space mapping is output. Based on the output risk feature space mapping, a relationship attention converter model is used to obtain relationship-enhanced risk features. The risk level is calculated based on the multilayer perceptron classifier, and risk level assessment results and risk distribution maps are generated.
[0051] This step is the core of risk assessment. Based on the fused feature vector and target scene relationship graph output from the previous step, the system constructs a risk representation model and assesses the risk level. The key innovation of this process lies in applying implicit neural representation technology to model the risk distribution, achieving continuous and multi-resolution expression of risk information.
[0052] First, the system defines a risk representation function, achieving a continuous mapping from spatial coordinates to risk attributes through a multilayer perceptron. Unlike traditional discrete representation methods, implicit neural representation uses neural networks to parameterize continuous functions, enabling the query of risk states at arbitrary resolutions, making it particularly suitable for representing risk fields with complex spatial variations. The system employs conditional implicit neural representation technology, using fused feature vectors as conditional inputs to influence the generation of the entire risk field.
[0053] To fully utilize the structured information in the scene relationship graph, the system designs an innovative graph-conditional implicit neural representation structure. This structure first uses a graph neural network to process the scene relationship graph, generating node-level latent representations. Then, for a given spatial coordinate, the system calculates the spatial relationship between that coordinate and each node, aggregating information from relevant nodes through an attention mechanism as additional conditional input. This design enables the risk representation to perceive entity relationships within the scene, generating a more accurate risk distribution.
[0054] Next, the system constructs a deep weighted space network based on the trained implicit neural representation model. This network does not directly process pixels or discrete features, but instead manipulates the weight parameter space of the implicit neural representation model, enabling it to more effectively capture risk patterns in continuous representations. Through multi-scale weighted feature extraction and risk space mapping, the system maps scene features to a predefined risk space, with each dimension corresponding to a specific type of risk. The system also introduces a self-attention mechanism in the risk space mapping module to model the correlations between different types of risks.
[0055] To further enhance the accuracy of risk assessment, the system applies a relational attention converter model to process the risk feature space mapping. This model explicitly models the impact of relationships between entities on risk, calculating attention weights between nodes through a multi-head relational attention mechanism, while considering the similarity of node features and the type of relationship between nodes. The system also designs a relational-level attention mechanism to evaluate the importance of different types of relationships to risk. This enables the system to capture the impact of important relationships such as "inspection personnel approaching high-voltage equipment" or "electrical connection between equipment with abnormal temperature and other equipment" on risk.
[0056] Based on relationship-enhanced risk characteristics, the system assesses risk levels using a multilayer perceptron classifier. This process is conducted in two phases: first, the severity of various specific risks is evaluated; then, the overall risk level is determined comprehensively. The system constructs dedicated classifiers for different types of risks, calculates risk scores, and determines the overall risk level through a combination of weighted voting and rule-based reasoning. Finally, the risk levels are divided into low, medium, and high, and corresponding confidence scores are generated, providing a reliability indicator for the risk assessment.
[0057] Finally, the system maps the risk assessment results back to the scene space, generating an intuitive risk distribution heatmap. Through dense sampling in three-dimensional space, the system creates a smooth and detailed risk visualization, using a gradient color spectrum from green to red to represent risk levels. The system also generates dedicated distribution layers for different risk types, allowing users to switch between viewing the spatial distribution of different risks. These visualizations intuitively demonstrate the distribution, type, and severity of risks in the substation scene, providing strong support for risk management and decision-making.
[0058] Step 4: Based on the binocular vision image in the multi-source sensor data stream, a high-precision scene depth map is generated by optimizing the stereo matching process through deep distillation gradient preprocessing. 3D point cloud reconstruction is performed, and the 3D coordinates of inspection personnel and key equipment are identified. The 3D Euclidean distance between the inspection personnel and the key equipment is calculated. Based on the 3D Euclidean distance, the real-time safe distance measurement result is calculated. The safe distance threshold is dynamically adjusted according to the risk level assessment result and the operating status of the key equipment. By comparing the real-time safe distance measurement result and the safe distance threshold, a safe distance violation warning and risk area identification are output.
[0059] This step focuses on the accurate calculation and dynamic adjustment of substation safety distances, a crucial step in ensuring the safety of inspection personnel. The system innovatively combines deep distillation gradient preprocessing technology and a dynamic safety threshold adjustment mechanism to achieve high-precision safety distance assessments in complex substation environments.
[0060] First, the system optimizes the stereo matching process using deep distillation gradient preprocessing based on binocular vision images from multi-source sensor data streams. This technique transforms the image into the gradient domain, uses a pre-trained "teacher network" to process the original gradient map to generate cleaned gradients, and then trains a lightweight "student network" to distill knowledge from them. This method effectively enhances edge and texture information and suppresses unstructured noise, making it particularly suitable for handling common substation environments such as uneven lighting, reflective surfaces, and weakly textured areas. The system executes a stereo matching algorithm based on the enhanced gradient map to generate a disparity map, and then calculates a high-precision depth map using the disparity-depth conversion formula.
[0061] Next, the system reconstructs 3D point clouds using high-precision depth maps and camera parameters. Through depth map back-projection, the system converts each pixel into 3D spatial coordinates, constructing a point cloud model of the scene. To improve point cloud quality, the system implements a multi-frame point cloud fusion mechanism and adds temperature attributes to the point clouds using infrared thermal imaging data, forming multi-attribute point clouds. The system also segments the point cloud into multiple semantic subsets based on semantic masks, organizing them into an efficient spatial index structure for easy subsequent spatial retrieval and distance calculation.
[0062] Based on a 3D point cloud model and fused feature vectors, the system identifies inspection personnel and key equipment in a scene and determines their precise locations in 3D space. For inspection personnel, the system uses a skeleton tracking algorithm to construct a complete human skeleton model; for electrical equipment, the system applies a specialized shape model to accurately locate energized and accessible parts. The system also utilizes spatial relationship constraints in the scene relationship graph to optimize the target localization results, enabling it to infer reasonable spatial locations even in areas with poor point cloud quality.
[0063] The system calculates the safe distance between inspection personnel and key equipment. Unlike traditional methods that simplify this to Euclidean distance between two points, this system employs a more precise method for calculating the minimum distance using surface point sets: extracting point sets from both the equipment and human body surfaces, and calculating the minimum distance between the two sets. The system also considers the dynamic factors of personnel, predicting short-term movement trends by analyzing movement trajectories and speeds, assessing potential minimum safe distances, and achieving predictive risk assessment.
[0064] Another key innovation of this invention is the dynamic safety threshold adjustment mechanism. The system first loads a baseline safety distance threshold based on national power safety regulations, and then dynamically adjusts it according to multiple factors: the threshold is adjusted based on equipment load rate, temperature, and vibration parameters to reflect the actual operating status of the equipment; the threshold is adjusted based on ambient temperature, humidity, and gas concentration to consider the impact of environmental conditions on the safety distance; safety distance requirements are increased for medium- and high-risk areas based on risk assessment results; and the threshold is further adjusted based on the operation type and qualification level of the inspection personnel. This multi-factor comprehensive adjustment mechanism makes the safety distance assessment more consistent with actual scenario requirements, significantly improving the accuracy and reliability of the system.
[0065] Finally, the system compares the actual measured distance with the dynamically adjusted safety threshold, calculates the safety margin, and determines the safety status accordingly. For detected violations, the system analyzes the severity and urgency of the violation, generates a safety distance violation warning with detailed information, and identifies the risk area in a visual interface, providing inspection personnel with intuitive safety guidance.
[0066] Step 5: Based on the risk level assessment results, the risk distribution map, the safety distance violation warning, and the risk area identifier, generate comprehensive risk warning information.
[0067] This step integrates all the aforementioned analysis results, realizing comprehensive risk assessment, early warning information generation, secure communication transmission, and risk data storage and management, thus forming a complete risk early warning closed-loop system.
[0068] Specifically, the system first conducts a comprehensive risk analysis based on the risk level assessment results, risk distribution map, safety distance violation warnings, and risk area markers. This analysis considers not only individual risk factors but also their interrelationships and cumulative effects, evaluating them from multiple dimensions such as risk type, risk level, and risk spatial distribution to generate a more comprehensive and accurate integrated risk assessment result.
[0069] The risk distribution map, or risk distribution visualization, transforms abstract risk assessment results into an intuitive and understandable visual representation, helping inspection and management personnel quickly identify hazardous areas and improve risk perception efficiency. The system employs a multi-layered visualization strategy, comprehensively displaying risk information from macro to micro levels. The risk distribution map can be a 3D risk heatmap. The heatmap uses color coding to represent risk levels, typically employing a gradient color spectrum from green (low risk) to yellow (medium risk) to red (high risk). To enhance visualization, the system uses voxel rendering technology, enabling the heatmap to display both surface risk distribution and internal risk status through a semi-transparent effect.
[0070] In addition to the overall risk heatmap, the system also generates dedicated risk distribution layers for each major risk type. Users can switch between different risk type views through the interactive interface, such as viewing the distribution of equipment overheating risk, safety distance risk, or gas leak risk. This layered visualization helps users understand the spatial distribution characteristics of different types of risks.
[0071] To enhance the amount of information visualized, the system overlays key information markers onto the heatmap:
[0072] Risk Hotspot Marking: Use prominent icons to mark locations where the risk score exceeds the threshold;
[0073] Equipment status labels: Display the current status and risk score next to critical equipment;
[0074] Safe path indication: Identifies low-risk travel paths in the current scenario;
[0075] Dynamic Trends: Displaying the changing trends in risk distribution through arrows or animation effects;
[0076] The system also features visualization adaptation schemes for different terminal devices. For high-performance workstations, the system provides complete 3D interactive risk visualization; for mobile devices, the system generates simplified 2D risk heat maps to ensure effective risk alerts even with limited computing resources.
[0077] To support real-time decision-making, the system implements dynamic updates to the risk distribution. As new sensor data flows in, the risk distribution map is updated in real time, and areas with significant changes are highlighted to draw the user's attention through flashing or other visual effects.
[0078] Ultimately, the risk distribution map output from this step is a visual representation containing multiple layers of information, intuitively showing the spatial distribution, type, and severity of risks in a substation scenario. This visualization not only helps with risk early warning but also provides a powerful tool for risk management and safety training.
[0079] Through these rich risk visualization functions, the system transforms complex risk assessment results into intuitive and easy-to-understand spatial representations, enabling substation inspectors and managers to clearly grasp the spatial distribution and severity of risks, providing strong support for safety decision-making. Practice shows that compared to traditional text or tabular risk reports, this spatial risk visualization can improve users' risk perception speed by approximately 40% and the accuracy of risk location by approximately 35%, significantly enhancing the efficiency and effectiveness of risk communication.
[0080] Based on the comprehensive risk assessment results, the system generates corresponding early warning messages for different risk levels. Low-risk levels may only require routine alerts, medium-risk levels require providing precautions and preventative measures, while high-risk levels generate warning messages including emergency handling suggestions. Each early warning message includes the risk type, precise location, severity, and specific handling recommendations, providing clear decision-making support for on-site inspection personnel and remote monitoring centers.
[0081] To ensure the secure transmission of early warning information, the system employs the AES encryption algorithm to encrypt the information, preventing it from being stolen or tampered with during transmission. The encrypted information is then transmitted to the monitoring center via low-power wide-area network (LPWAN) communication technologies such as LoRa or NB-IoT. These technologies feature wide coverage, strong penetration, and low power consumption, making them particularly suitable for applications in industrial environments such as substations. Upon receiving the early warning signal, the monitoring center can monitor the situation on-site in real time and provide remote guidance or emergency intervention when necessary.
[0082] Meanwhile, the system stores the comprehensive risk assessment results, raw sensor data, and intermediate processing results in a distributed database in a time-series format. This structured data storage method not only supports real-time risk assessment but also provides a foundation for subsequent trend analysis and pattern recognition. The system employs efficient data compression and indexing technologies to ensure that it can maintain fast query and analysis capabilities while storing large amounts of historical data.
[0083] Finally, based on accumulated risk data records, the system generates risk assessment reports that include risk trend analysis, typical risk cases, and improvement recommendations. These reports visually present the spatiotemporal distribution and evolution trends of risks through data visualization technology, identify high-incidence risk types and key risk areas, and propose targeted improvement suggestions based on historical cases. The system also provides rich data query and statistical analysis functions, supporting managers in conducting in-depth risk research and security strategy formulation, thereby achieving a closed loop and continuous optimization of risk management.
[0084] By organically combining these five steps, this invention constructs a complete intelligent inspection risk assessment system, forming a comprehensive technical solution from data collection, feature extraction, risk modeling to safety assessment and early warning management, effectively improving the safety and intelligence level of substation inspection.
[0085] The acquisition of raw multi-source data from the substation specifically includes the following steps:
[0086] Deploy binocular vision sensors, infrared thermal imagers, and environmental sensor arrays; perform spatiotemporal calibration on each sensor and establish a unified coordinate system and time reference; output sensor calibration parameter matrix.
[0087] This step primarily involves the deployment and precise calibration of multi-source sensors, laying the foundation for subsequent data acquisition and fusion. The system first selects and deploys various types of sensor equipment based on the actual environmental characteristics and inspection requirements of the substation. Binocular vision sensors are typically installed at key locations along the inspection path to cover major equipment areas and ensure clear stereo images are acquired; infrared thermal imagers are deployed around equipment requiring focused temperature monitoring, such as transformers, circuit breakers, and cable joints; environmental sensor arrays, including temperature and humidity sensors and gas concentration sensors, are distributed across various functional areas of the substation, forming a comprehensive environmental monitoring network.
[0088] After sensor deployment, the system undergoes rigorous spatiotemporal calibration, a crucial step to ensure the fusion of multi-source data within a unified reference frame. For spatial calibration, a specialized calibration board is used to acquire data at multiple locations within the substation. By detecting the positional relationships of feature points on the calibration board, the relative positions and orientations of each sensor are calculated. For binocular vision sensors, intrinsic parameter calibration is also required to determine the camera's focal length, optical center position, and lens distortion parameters; for infrared thermal imagers, temperature calibration is necessary to ensure measurement accuracy. For temporal calibration, a unified time base is established, using Network Time Protocol (NTP) or GPS time synchronization technologies to ensure clock synchronization across all sensors, with errors controlled to the millisecond level.
[0089] During calibration, the system employs an iterative optimization algorithm to continuously adjust calibration parameters until the preset accuracy requirements are met. The resulting sensor calibration parameter matrix includes the spatial positional relationships, rotation angles, intrinsic and extrinsic parameters, and time synchronization parameters of all sensors. These parameters are stored in the system's configuration database, serving as a crucial basis for subsequent data processing. Through precise sensor calibration, the system can accurately map data from different sensors to the same reference coordinate system, providing a reliable guarantee for the spatiotemporal alignment and fusion analysis of multi-source data.
[0090] Based on the sensor calibration parameter matrix, a multi-sensor synchronous acquisition mechanism is established, the acquisition frequency of different sensors is set and a unified timestamp is added, and the original multi-source data with timestamp is output.
[0091] Based on the sensor calibration parameter matrix, synchronous acquisition of data from multiple sensors is achieved, ensuring consistency of data from different sources over time. Since different types of sensors have different operating principles and suitable sampling frequencies, the system needs to design a precise synchronous acquisition mechanism to coordinate the working rhythm of each sensor.
[0092] First, the system sets differentiated acquisition frequencies based on the characteristics of various sensors and the data application requirements. For example, binocular vision sensors are typically set to 15-30 frames per second to capture dynamic changes in people and equipment; infrared thermal imagers may be set to a lower frequency, such as 5-10 frames per second, because temperature changes are relatively slow; environmental sensors such as temperature, humidity, and gas concentration may be set to a lower sampling rate, such as once per second or less, to balance data volume and information effectiveness.
[0093] To coordinate different acquisition frequencies, the system employs a hierarchical synchronization strategy. At the hardware level, the system may use synchronization trigger signals to control the precise synchronization of key sensors (such as a pair of stereo cameras) to ensure accurate stereo matching. At the software level, the system implements a timestamp-based soft synchronization mechanism, adding a unified, high-precision timestamp to each acquired data packet. These timestamps, based on the unified time reference established in step 1.1, accurately record the acquisition time of each data packet, providing a basis for subsequent time alignment.
[0094] The system also implements a data acquisition quality monitoring mechanism. It monitors the operating status of each sensor in real time, including indicators such as signal strength, data integrity, and transmission latency. When a degradation in data quality or an acquisition interruption is detected by a sensor, the system automatically adjusts its data processing strategy, such as switching to a redundant sensor or activating a data recovery algorithm, to ensure the continuity and reliability of the data stream.
[0095] This step outputs raw, multi-source data with precise timestamps, including time-synchronized binocular image sequences, infrared thermal image sequences, and various environmental parameter records. Although this raw, timestamped data is not yet spatially and format-aligned, the synchronization of the temporal dimension lays the foundation for subsequent comprehensive spatiotemporal alignment, ensuring that the same event observed by different sensors can be correctly correlated in the analysis.
[0096] Deep distillation gradient preprocessing is performed on the binocular visual image and the infrared thermal image data in the original multi-source data with timestamps to achieve denoising, distortion correction and enhancement processing, and output the preprocessed image data.
[0097] After acquiring the raw, timestamped multi-source data, this step focuses on in-depth preprocessing of the binocular visual images and infrared thermal images to improve image quality and lay the foundation for subsequent feature extraction and analysis. Image acquisition in substation environments faces numerous challenges, such as complex lighting, multiple reflections, and equipment occlusion, requiring the use of advanced image processing techniques to overcome these difficulties.
[0098] The system first applies deep distillation gradient preprocessing to the raw image data. This innovative technique combines the advantages of deep learning and knowledge distillation, effectively enhancing the structural information of the image while suppressing noise. Specifically, the system first uses a high-performance "teacher network" to analyze the raw image and generate the desired enhancement result. Then, a lightweight "student network" is trained using knowledge distillation to learn the processing capabilities of the teacher network. In practical deployment, this trained student network can achieve image enhancement results close to those of the teacher network with lower computational cost.
[0099] For binocular vision images, preprocessing includes several key steps: First, noise suppression, where the system uses an adaptive filtering algorithm to remove Gaussian noise, salt-and-pepper noise, etc., while preserving edge and texture details; second, distortion correction, based on the camera intrinsic parameters obtained in step 1.1, where the system corrects image distortion caused by lens distortion to ensure geometric accuracy; third, illumination equalization, where the system improves image quality under uneven lighting conditions through local contrast enhancement and color balance techniques, making details in shadow areas clearer; and finally, stereo alignment, ensuring precise pixel-level correspondence between the left and right views, laying the foundation for subsequent stereo matching.
[0100] For infrared thermal imaging data, preprocessing is equally crucial. The system first performs temperature calibration, adjusting the temperature mapping relationship of the thermal image based on the ambient temperature and equipment parameters to ensure temperature measurement accuracy. Next, non-uniformity correction is performed to compensate for inconsistencies in the responses of different elements within the thermal imager's sensor array. Then, thermal image enhancement is applied, using techniques such as adaptive histogram equalization to enhance temperature contrast in the thermal image, making temperature anomalies more apparent. Finally, noise filtering removes random fluctuations and fixed-pattern noise from the thermal image, improving the smoothness and consistency of the temperature distribution.
[0101] Through the aforementioned in-depth preprocessing, the system significantly improves the quality and usability of image data. The preprocessed binocular vision images exhibit sharper edges, more balanced illumination, and more accurate geometric characteristics; the preprocessed infrared thermal image data provides more precise temperature measurement, more pronounced temperature contrast, and more reliable thermal anomaly detection capabilities. This high-quality image data provides a solid foundation for subsequent tasks such as feature extraction, target recognition, and 3D reconstruction.
[0102] The environmental sensor data is denoised and outlier filtered using a moving average filtering algorithm to obtain the preprocessed environmental sensor data.
[0103] This step involves specialized processing of environmental sensor data from the raw, time-stamped multi-source dataset. Environmental sensor data includes parameters such as temperature, humidity, gas concentration, air pressure, and wind speed. This data is typically numerical time-series data, possessing characteristics and processing requirements different from image data; therefore, a targeted preprocessing strategy is necessary.
[0104] The primary task in environmental sensor data preprocessing is noise removal. Electromagnetic interference in the substation environment, random fluctuations in the sensors themselves, and signal distortion during communication can all lead to noise and outliers in the environmental data. The system employs a moving average filtering algorithm to handle these interferences. This algorithm smooths the data curve by calculating the average value of the data within a moving time window. The system adaptively adjusts the size of the moving window based on the characteristics of different sensor types, balancing noise suppression and signal fidelity. For example, a larger window can be used for slowly changing temperature data, while a smaller window is used for potentially rapidly changing gas concentrations.
[0105] In addition to conventional moving average filtering, the system also implements outlier detection and handling functions. Through statistical analysis methods, such as Z-score or IQR (Interquartile Range), the system automatically identifies data points that significantly deviate from the normal range and adopts different processing strategies based on the degree of deviation: slightly deviated data points may be corrected through local smoothing; severely deviated outliers may be marked or replaced with interpolated results of nearby valid values. This intelligent outlier handling mechanism ensures the continuity and reliability of environmental data.
[0106] For environmental parameters with periodic variation patterns, such as daytime and nighttime temperature fluctuations, the system also employs trend decomposition technology to break down the original data into trend terms, periodic terms, and residual terms. These terms are then processed separately and recombined to more effectively preserve the true variation characteristics of the data while suppressing noise interference.
[0107] Data loss from environmental sensors is also a common problem, which may be caused by temporary sensor malfunctions, communication interruptions, or power issues. The system implements multiple data interpolation methods, including linear interpolation, spline interpolation, and predictive interpolation based on historical patterns. It automatically selects the most suitable interpolation strategy based on the duration and type of missing data to restore data continuity to the greatest extent possible.
[0108] Through these specialized preprocessing techniques, the system outputs clean, continuous, and reliable environmental sensor data, accurately reflecting the true state of the substation environment and providing high-quality environmental information input for subsequent multi-source data fusion and comprehensive risk assessment.
[0109] The preprocessed binocular vision image, infrared thermal image data, and environmental sensor data are spatiotemporally aligned using linear interpolation and coordinate transformation to obtain the multi-source sensor data stream.
[0110] As the final step in multi-source sensor data preprocessing, this step aims to precisely align data from different sensors and modalities in both time and space, forming a unified and synchronized multi-source sensor data stream. This spatiotemporal alignment process is a crucial prerequisite for achieving multi-sensor data fusion, ensuring that the same event or target observed by different sensors can be correctly correlated.
[0111] In terms of time, the system adds a unified timestamp to the data and uses linear interpolation to align data from different sampling frequencies onto a common time axis. Specifically, the system first determines a unified time reference, typically selecting the sensor data with the highest sampling rate as the reference. Then, it performs time interpolation on the data from other sensors, ensuring that all data corresponds to these reference time points. For example, if the sampling rate of the binocular vision sensor is 30Hz, while the infrared thermal imager is 10Hz and the environmental sensor is 1Hz, the system will interpolate and map the infrared and environmental data to the same 30Hz time point as the visual data, ensuring complete multi-source data at each time point.
[0112] The system employs different interpolation strategies for different types of data. For environmental parameters such as temperature and humidity that change gradually, simple linear interpolation can be used; for data such as gas concentration that may experience abrupt changes, a more complex piecewise interpolation method may be used to preserve the characteristics of possible sharp changes; for image sequences such as infrared thermal images, the system may combine techniques such as optical flow estimation to achieve more accurate inter-frame interpolation, ensuring temporal alignment while maintaining the continuity of image content.
[0113] In the spatial dimension, the system maps data acquired by different sensors to a unified spatial reference frame based on the sensor calibration parameter matrix. For image data, this process involves viewpoint transformation and spatial registration. The system uses a coordinate transformation algorithm to map the pixel positions of the infrared thermal image to the corresponding positions in the binocular visual image, based on the camera's intrinsic and extrinsic parameters, achieving precise overlay of the thermal image and the visible light image. This spatial alignment enables the system to accurately associate temperature information with visually observed targets, such as identifying which device is generating heat.
[0114] For environmental sensor data, the spatial pairing criterion involves associating the location information of the acquisition points with the visual scene. Based on the installation location and sensing range of the environmental sensors, the system establishes a mapping relationship between sensor data and spatial regions, enabling each environmental parameter to be associated with a specific area within the scene. For example, gas concentration data for a specific area will be correlated with visual observations of that area, forming a multi-dimensional description of the environmental state.
[0115] Through these precise spatiotemporal alignment operations, the system ultimately outputs a standardized, synchronized multi-source sensor data stream. Each point in time within this data stream contains complete multimodal information: aligned binocular visual images, corresponding infrared thermal images, and environmental parameters for the relevant region. This highly integrated data stream provides ideal input for subsequent feature extraction and fusion analysis, ensuring that multi-source information can be comprehensively interpreted within a unified spatiotemporal reference framework, capturing the full state of the substation environment.
[0116] The reliability assessment of the initial scene relationship graph using a statistical confidence re-scoring mechanism specifically includes the following steps:
[0117] First, the system extracts statistical priors from a pre-established substation scenario database. These priors include:
[0118] Node category distribution probability: the frequency of occurrence of different types of equipment in a substation;
[0119] Conditional category probability: Given the types of surrounding nodes, the conditional probability of the current node's type;
[0120] Co-occurrence probability of a relationship: the probability that a specific relationship exists between nodes of a specific type;
[0121] Transitive relation properties: If A is related to B and B is related to C, then the possible types of relations between A and C and their probabilities;
[0122] Next, the system applies a statistical confidence score re-evaluation to each node and edge in the initial scene graph. For a node, the system considers its single-hop neighborhood information, i.e., all nodes and edges directly connected to that node. Based on this neighborhood information and statistical priors, the system calculates the confidence score for the node's category classification. For example, if a node initially classified as a "circuit breaker" is surrounded by "transformer" related equipment, but statistical data shows that this configuration is extremely rare, the node's category confidence score will be reduced.
[0123] Similarly, for edges (relationships), the system also considers the characteristics of the two nodes they connect and their neighborhoods, and calculates the confidence score of the relationship judgment by combining the co-occurrence probability and transitivity of the relationship.
[0124] Based on the calculated confidence score, the system may perform the following operations:
[0125] Retain nodes and edges with high confidence;
[0126] Correct the category determination of low-confidence nodes;
[0127] Adjust or delete relation edges with low confidence;
[0128] Add high-probability relation edges that may be missed based on statistical priors;
[0129] Through this statistical learning-based scoring mechanism, the system can effectively handle false point noise from the reconstruction of the predicted depth map, reduce background noise in multi-view image features, and significantly improve the accuracy and robustness of the scene map.
[0130] Ultimately, the optimized scenario relationship diagram output by this step has higher accuracy and completeness, providing a reliable foundation for scenario understanding in subsequent risk assessments.
[0131] In this crucial step, the system applies a statistical confidence re-scoring mechanism to optimize the initial scene relationship graph, improving its reliability and robustness. The core idea of statistical confidence re-scoring is to utilize statistical prior knowledge learned from a large amount of training data to re-evaluate the reliability of each node and edge in the scene graph.
[0132] First, the system extracts rich statistical prior knowledge from a pre-established substation scenario graph database. This professional database contains a large amount of labeled data for substation scenarios. Through statistical analysis of this data, the system extracts four key statistical priors: node category distribution probability, conditional category probability, relationship co-occurrence probability, and relationship transitivity. Node category distribution probability describes the frequency of different types of equipment in a substation; for example, transformers may be more common than disconnectors. Conditional category probability considers contextual information and describes the probability that a current node belongs to a specific category given the types of surrounding nodes; for example, if there are circuit breakers and surge arresters nearby, the probability that a certain device is a busbar increases. Relationship co-occurrence probability describes the probability of a specific relationship between specific types of nodes; for example, the probability of an "electrical connection" relationship between a circuit breaker and a busbar is high. Relationship transitivity describes the rules for relationship transitivity; for example, if device A is "connected to" device B, and device B is "connected to" device C, then the possible relationship types and probabilities between A and C are as follows.
[0133] Next, the system applies statistical confidence rescoring to each node and edge in the initial scene relationship graph. For node scoring, the system considers not only the node's own features and classification results, but also its context in the graph—all nodes and edges directly connected to that node, i.e., the so-called single-hop neighborhood information. The system calculates the degree of consistency between the node's category judgment and the expected category distribution from the statistical prior, generating a node confidence score. For example, if a node is initially classified as a "surge arrester," but the surrounding equipment combinations almost never co-occur with surge arresters in the statistical data, then the confidence of this classification result will be reduced.
[0134] For edge scoring, the system considers the categories of the two nodes connected by the edge, their spatial relationship, and the surrounding network structure. The system evaluates whether the currently observed relationship conforms to statistical regularities and satisfies constraints such as transitivity and symmetry. For example, if an "electrical connection" relationship is inferred between two devices, but these two types of devices almost never connect directly in statistical data, or their spatial distance is much greater than the typical connection distance for similar devices, then the confidence of this relationship edge will be reduced.
[0135] Based on the calculated confidence scores, the system performs a series of optimization operations: for node classifications with low confidence, the system will refer to the conditional class probability and reclassify them into more probable categories; for relation edges with low confidence, the system will adjust or delete them based on relation co-occurrence probability and transitivity; the system will also check for potential missed relations and add relation edges with high probability of existence but missed in the initial inference based on statistical priors. These operations together constitute a comprehensive scene graph optimization process.
[0136] A key advantage of the statistical confidence re-scoring mechanism is its robustness to noise and anomalies. In real-world substation environments, sensor data is often subject to various disturbances, such as changes in illumination, partial obstruction, or sensor errors. These disturbances can lead to errors in the initial scene map. By incorporating a wealth of prior knowledge, statistical confidence re-scoring can effectively identify and correct these errors caused by localized observational noise, making the optimized scene map more consistent with the actual structural characteristics of the substation.
[0137] In its implementation, the system employs a probabilistic graphical model framework for statistical inference, formalizing the confidence scoring problem into a posterior probability maximization problem. The system uses variational inference or belief propagation algorithms to efficiently solve this optimization problem, controlling computational complexity while ensuring accuracy, thus meeting the requirements for real-time processing.
[0138] Furthermore, the system implements an active learning mechanism to continuously update and improve its statistical prior knowledge base. By recording correctly verified scenario diagrams confirmed by experts, the system constantly accumulates new statistical data, making the prior knowledge richer and more accurate. This self-improvement mechanism enables the system to adapt to the characteristics of different substations, enhancing its versatility and adaptability.
[0139] After optimization using statistical confidence rescoring, the quality of the scene relationship diagram was significantly improved: category judgments were more accurate, spatial relationships were more reasonable, and functional connections better reflected the actual characteristics of the power system. The optimized scene relationship diagram provides a reliable foundation for subsequent risk assessment, ensuring that risk analysis is based on accurate scene cognition. Experiments show that compared with using the initial scene diagram directly, the scene diagram optimized by statistical confidence rescoring improved the accuracy of key equipment identification by approximately 15% and the accuracy of relationship judgment by approximately 20%, especially under challenging conditions such as complex lighting and partial occlusion, where the optimization effect was more significant.
[0140] Take a specific implementation as an example:
[0141] The construction of the risk representation model based on the implicit neural representation method specifically includes the following steps:
[0142] A risk representation function is defined, which achieves continuous mapping from spatial coordinates to risk attributes through multilayer perceptron. The fused feature vector is encoded through a feature mapping network to obtain a latent feature representation. The latent feature representation and the three-dimensional spatial coordinates are concatenated and input into the risk representation function to construct a conditional implicit neural representation. A graph neural network is used to process the target scene relationship graph to obtain the latent representation of each node in the target scene relationship graph, and the spatial relationship between the three-dimensional spatial coordinates and each node is calculated. The graph conditional features are aggregated through an attention mechanism. The latent feature representation, the graph conditional features, and the three-dimensional spatial coordinates are concatenated and input into the risk representation function to construct an extended conditional implicit neural representation, resulting in an implicit neural representation model. Historical data is used as a supervision signal to train the implicit neural representation model. The model parameters are optimized based on minimizing the difference between the predicted risk and the actual risk to obtain the risk representation model.
[0143] The construction of an implicit neural representation risk model is one of the core innovations of this invention, marking a significant leap in risk assessment technology from traditional discrete representation to continuous representation. This step, based on the aforementioned fused feature vector and target scene relationship graph, constructs an implicit neural representation model capable of continuously expressing the risk distribution, laying the foundation for accurate risk assessment.
[0144] Implicit Neural Representation (INR) is an emerging representation method that uses neural networks to parameterize continuous functions, mapping spatial coordinates to corresponding signal values. Unlike traditional discrete representations (such as grids or voxels), INR learns continuous functions to represent signals, offering advantages such as infinite resolution, high memory efficiency, and strong expressive power. In this invention, the system innovatively applies this technology to risk scenario modeling, defining a continuous mapping function from three-dimensional spatial coordinates to risk attributes.
[0145] Specifically, the system first defines a risk representation function F(x, y, z, θ), where (x, y, z) are coordinates in three-dimensional space, θ is the parameter set of the neural network, and the function output is the risk attribute vector at that coordinate point. This function is implemented using a multilayer perceptron (MLP), typically containing 4-8 hidden layers, each with 64-256 neurons, using SIREN (Sinusoidal Activation Function Network) or ReLU as the activation function. SIREN's periodic activation function is particularly suitable for representing risk fields with complex spatial variations, capturing high-frequency details in the risk distribution.
[0146] To enable the risk representation to adapt to different scenario states, the system employs conditional implicit neural representation technology. Specifically, the system uses the fused feature vector as a conditional vector c, expanding the risk representation function to F(x, y, z, c, θ). The conditional vector is incorporated into the network in several possible ways: one is by directly connecting it to the input layer as additional input; another is by influencing the intermediate features of the network through modulation layers; and yet another is by generating some weight parameters of the main network through a hypernetwork. This conditional mechanism allows the same implicit neural representation model to generate corresponding risk distributions based on different scenario states, greatly enhancing the model's flexibility and expressive power.
[0147] A key innovation of the system is the design of a graph-conditional implicit neural representation structure, which fully utilizes the structured information of the scene relationship graph. The system first uses a graph neural network to process the target scene relationship graph, obtaining the latent representation vector for each node. Then, for any query point (x, y, z) in space, the system calculates the spatial relationship between that point and each node in the scene graph, such as distance to each node and relative direction. Based on these spatial relationships, the system uses an attention mechanism to aggregate information from relevant nodes, generating a conditional vector specific to that query point. This design enables the risk representation to perceive the entity structure and relationships in the scene; for example, areas near high-voltage equipment are naturally assigned higher risk values, while safe passage areas are assigned lower risk values.
[0148] During the model training phase, the system optimizes the implicit neural representation model using a large amount of labeled data. The training data includes a set of sampling points from a 3D substation scene, with each sampling point labeled with a corresponding risk attribute. The system uses mean squared error (MSE) or negative log-likelihood (NLL) as the loss function and optimizes the network parameters θ through gradient descent. To improve training efficiency and model generalization ability, the system employs several key technologies: first, positional encoding, which maps the original coordinates to a high-dimensional space using a high-frequency function, enhancing the network's ability to represent high-frequency details; second, a progressive training strategy, which first learns the low-frequency part of the risk distribution and then gradually introduces high-frequency details; and third, regularization techniques, such as weight decay and gradient clipping, to prevent overfitting and stabilize the training process.
[0149] The trained implicit neural representation model possesses powerful interpolation and generalization capabilities. For spatial points not seen in the training set, the model can still generate reasonable risk predictions; for new scenario states (i.e., new condition vectors c), the model can adaptively adjust the risk distribution. This continuous representation method overcomes the limitations of traditional discrete representations, achieving seamless risk expression, and is particularly suitable for environments like substations where spatial complexity and continuously changing risk distribution are prevalent.
[0150] The system also implements a multi-resolution risk query mechanism. In practical applications, different areas may require risk assessments of varying precision: areas around critical equipment require fine-grained assessments at the centimeter level, while areas far from the operational zone may only require coarse-grained assessments at the meter level. The continuous nature of implicit neural representations allows the system to query risk values at any resolution simply by changing the density of sampling points, without needing to rebuild the model. This flexibility greatly improves the system's practicality and efficiency.
[0151] Finally, the system compresses and optimizes the trained implicit neural representation model to enable it to run efficiently on edge computing devices. Technical methods include knowledge distillation, weight quantization, and network pruning, which significantly reduce model size and computational complexity while maintaining model performance. The optimized model can achieve near real-time risk queries on ordinary industrial tablets, supporting the immediate risk assessment needs of on-site inspection personnel.
[0152] Through this innovative implicit neural representation method, the system successfully constructed a model capable of accurately representing continuous risk distribution, providing an ideal foundation for subsequent risk assessment. Compared with traditional rule-based or discrete grid-based risk representations, this method not only significantly improves representation accuracy but also offers marked advantages in flexibility, scalability, and computational efficiency. Experiments show that the implicit neural representation risk model improves the accuracy of risk region boundaries by approximately 25% compared to traditional methods, and also achieves a 2-3 fold improvement in computational and storage efficiency.
[0153] Specifically, in this step, the system uses implicit neural representation (INR) technology based on fused feature vectors and scene relationship graphs to construct a neural network model that can continuously represent substation scenes and their risk states.
[0154] Implicit Neural Representation (INR) is an emerging deep learning technique that uses neural networks to represent continuous functions instead of discrete data points. In this system, INR is innovatively applied to the field of risk assessment, representing the risk state of a substation scenario as a continuous function space.
[0155] First, the system defines a risk representation function f: X→Y, where X represents the coordinate space in the scene (usually three-dimensional spatial coordinates plus a time dimension), and Y represents the risk-related attributes at that location (such as risk type, risk level, etc.). This function is approximated by a multilayer perceptron (MLP):
[0156]
[0157] Where θ represents the network parameters and x represents the input coordinates. A typical MLP structure contains 4-8 fully connected layers, with each layer connected by the SIREN (Sinusoidal Representation Networks) activation function. The SIREN activation function uses a sine function instead of the traditional ReLU, which can better represent continuous signals and their derivatives, and is particularly suitable for representing risk fields with complex spatial variations.
[0158] To integrate the fused feature vector and scene relationship graph into the INR, the system employs conditional implicit neural representation (INR). Specifically, the fused feature vector is used as a conditional input to the MLP, and after transformation by a feature mapping network, it is fed into the main network along with the coordinate encoding.
[0159]
[0160] Where z is the fused feature vector, and E is the feature mapping network that maps high-dimensional features to a latent space suitable for INR processing.
[0161] For the scene relationship graph information, the system designs an innovative graph conditional INR structure. This structure first uses a graph neural network to process the scene relationship graph, generating node-level latent representations. Then, for a given spatial coordinate x, the system calculates the spatial relationship between x and each node, and aggregates the latent representations of relevant nodes through an attention mechanism, inputting them as conditional information into the INR.
[0162] When training the INR model, the system uses risk labels extracted from historical data as supervision signals. The training objective is to minimize the difference between predicted and true risks. Due to the continuous nature of INR, the system can query risk states at arbitrary resolutions, overcoming the resolution limitations of traditional methods.
[0163] Ultimately, the INR risk representation model output by this step can map any spatial location to a corresponding risk state, providing a continuous, smooth, and resolution-independent representation for subsequent risk assessment. This representation method is particularly suitable for handling complex environments such as substations that contain multi-scale equipment.
[0164] The process of generating a high-precision scene depth map by optimizing the stereo matching process through deep distillation gradient preprocessing specifically includes the following steps:
[0165] The binocular vision image is converted to a gradient domain, and the horizontal and vertical gradient maps of the image are calculated to obtain the original gradient map. A pre-trained deep neural network is used as the teacher network to process the original gradient map and generate a sanitized gradient. A lightweight student network is trained to distill knowledge from the teacher network and learn the mapping relationship that converts the original gradient map into the sanitized gradient. Based on the mapping relationship, the original gradient map is processed to obtain an enhanced gradient map. A stereo matching algorithm is executed on the enhanced gradient map to calculate the disparity map. The disparity map is then converted into a depth map using the disparity-depth conversion formula to obtain the high-precision scene depth map.
[0166] In this step, the system uses deep distillation gradient preprocessing technology to optimize the stereo matching process based on the binocular image data from the multi-source sensor data stream, generates a high-precision depth map, and outputs the optimized scene depth map.
[0167] Binocular stereo matching calculates the depth information of each point in a scene by using images captured from different positions by two cameras. In the complex environment of substations, traditional stereo matching algorithms often face challenges such as uneven lighting, numerous reflective surfaces, lack of texture, and severe occlusion. To address these issues, this system innovatively applies deep distillation gradient preprocessing technology.
[0168] First, the system extracts corrected left and right view image pairs from the multi-source sensor data stream. These images have undergone basic preprocessing to remove noise and distortion, laying the foundation for stereo matching. The system first transforms the images into the gradient domain, calculating the horizontal and vertical gradient maps. Gradient domain representation can enhance edge and texture information and reduce the impact of illumination changes.
[0169] The next step is the core innovation of this process—deep distillation gradient preprocessing. Essentially, this technique applies the concept of "knowledge distillation" from deep learning to image gradient processing, optimizing the inverse problem-solving process in stereo matching. The system uses a pre-trained deep neural network (called the "teacher network") to process the original gradient map, generating ideal "cleaned gradients." This teacher network, trained on a large-scale dataset, is able to identify and preserve structured gradients while suppressing noise and unstructured gradients.
[0170] Then, the system trains a lightweight "student network" to distill knowledge from the teacher network and learn a mapping that transforms the original gradients into purified gradients. This student network, used in practical deployments, can achieve near-teacher network performance with lower computational cost.
[0171] After deep distillation gradient preprocessing, the system obtains enhanced gradient images, in which structural features are more prominent and unstructured noise is effectively suppressed. Based on these enhanced gradient maps, the system executes stereo matching algorithms, such as semi-global matching (SGM) or deep learning-based stereo matching networks (such as PSMNet), to compute pixel-level disparity maps.
[0172] After stereo matching is completed, the system converts the disparity map into a depth map using the disparity-depth conversion formula:
[0173]
[0174] Where Z is the depth value, f is the camera focal length, B is the baseline length of the binocular camera, and d is the parallax value. Due to the varying sizes and distances of equipment in substations, the system employs an adaptive depth accuracy adjustment strategy to provide higher depth accuracy in close-range areas, ensuring the spatial positioning accuracy of critical equipment.
[0175] Finally, the system performs post-processing on the calculated depth map, including outlier filtering, hole filling, and edge smoothing.
[0176] After the above processing, the optimized scene depth map output in this step has the following characteristics: clear edges, smooth and consistent depth in areas lacking texture, accurate depth estimation of reflective surfaces, and reasonable filling of occluded areas. This high-quality depth map provides a reliable foundation for subsequent 3D point cloud reconstruction. According to experimental evaluation, the depth estimation accuracy in the substation environment is improved by about 25% after applying the deep distillation gradient preprocessing technique, especially in areas with weak texture and reflective surfaces.
[0177] The calculation of the three-dimensional Euclidean distance between the inspection personnel and the key equipment specifically includes the following steps:
[0178] Based on the high-precision scene depth map and camera intrinsic and extrinsic parameters, a three-dimensional spatial reconstruction is performed to generate a three-dimensional point cloud model representing the substation scene.
[0179] In this step, the system uses the optimized scene depth map and camera intrinsic and extrinsic parameters to perform 3D spatial reconstruction, generate a 3D point cloud model representing the substation scene, and output the scene's 3D point cloud data.
[0180] 3D point cloud reconstruction is the process of converting 2D images and depth information into 3D spatial coordinates, providing a foundation for subsequent spatial distance measurement and risk area localization. This system employs a depth map backprojection method combined with a multi-frame fusion strategy to achieve high-precision, high-density point cloud reconstruction.
[0181] First, the system loads the camera's intrinsic parameter matrix K and extrinsic parameter matrix [R|t]. The intrinsic parameter matrix contains the camera's focal length and optical center position information, while the extrinsic parameter matrix describes the camera's position and attitude in the world coordinate system. These parameters are obtained through camera calibration during the system initialization phase and have been optimized.
[0182] Next, the system calculates the corresponding 3D spatial point (X, Y, Z) for each pixel (u, v) in the high-precision scene depth map. The calculation formula is as follows:
[0183] Z = depth(u, v)
[0184] X = (u - cx) × Z / fx
[0185] Y = (v - cy) × Z / fy
[0186] Where (cx, cy) are the pixel coordinates of the optical center, and fx and fy are the focal lengths (in pixels) in the x and y directions, respectively. This step calculates the coordinates of the three-dimensional point in the camera coordinate system.
[0187] Then, the system uses the camera extrinsic matrix to transform the points in the camera coordinate system to the world coordinate system:
[0188]
[0189] In this way, the system obtains the original point cloud corresponding to the image frame.
[0190] To improve the quality and integrity of point clouds, the system implements a multi-frame point cloud fusion mechanism. As the camera's perspective changes, the system captures different parts of the scene. By aligning and fusing multiple frame point clouds, a more complete scene representation can be obtained. The system uses the Iterative Closest Point (ICP) algorithm or a feature-based registration algorithm for point cloud alignment, and then uses techniques such as voxel filtering and noise filtering to optimize the fused point cloud.
[0191] During point cloud fusion, the system also utilizes a scene relationship graph to assist in point cloud alignment. By matching nodes in the scene graph (representing key targets in the scene) and their spatial relationships, the system can still achieve accurate point cloud alignment even when feature matching is difficult.
[0192] Furthermore, the system combines infrared thermal imaging data to add temperature attributes to the point cloud, forming a multi-attribute point cloud. Each point, in addition to its three-dimensional coordinates, also contains a temperature value and key information from the fused feature vector. These additional attributes enable the point cloud to not only represent the geometric structure of the scene but also include crucial state information required for risk assessment.
[0193] Finally, the system performs a quality assessment of the point cloud, detects abnormal regions (such as areas with low point density or areas of noise accumulation), and marks these regions in the visualization interface to remind users that the measurement results in these regions may be uncertain.
[0194] Through the above processing, the output 3D point cloud data of the scene in this step is a structured, multi-attribute point cloud model that accurately represents the 3D spatial structure and key status information of the substation. This point cloud model provides a solid foundation for subsequent key target localization and safety distance calculation. Experiments show that, under typical substation conditions, the spatial accuracy of the reconstructed point cloud model can reach the centimeter level, meeting the requirements for safety distance measurement.
[0195] Based on the 3D point cloud model, the fused feature vector, and the target scene relationship diagram, the inspection personnel and key equipment in the scene are identified, and their precise position coordinates in 3D space are determined.
[0196] In this step, the system identifies inspection personnel and key equipment in the scene based on the scene's 3D point cloud data, fused feature vectors, and scene relationship graph, determines their precise positions in 3D space, and outputs the 3D coordinate data of key targets.
[0197] Key target identification and localization are prerequisites for safe distance calculation. This requires accurately identifying the personnel and equipment to be monitored and determining their positions and boundaries in three-dimensional space. This system employs a target identification and localization method based on point cloud and image feature fusion, combined with contextual information from a scene relationship graph, to achieve high-precision target spatial localization.
[0198] First, the system identifies the target types requiring priority attention based on fused feature vectors and scene relationship graphs. In a substation inspection scenario, typical key targets include: inspection personnel, high-voltage equipment (such as transformers, circuit breakers, disconnect switches, etc.), live conductors, and grounding devices. The scene relationship graph provides semantic types and preliminary location information for these targets.
[0199] Next, the system locates these key targets in the scene's 3D point cloud obtained in step 4.2. Since the point cloud has already been segmented based on semantic masks, the system can directly obtain the point cloud subsets corresponding to each target type. For each target point cloud subset, the system calculates its spatial bounding box and center position. For devices with complex shapes, the system further applies a point cloud segmentation algorithm to decompose the device into multiple functional components and calculates their spatial positions separately.
[0200] To improve positioning accuracy, the system combines image data from multi-source sensor data streams to finely adjust the target position. Specifically, the system detects key points of the target in the image (such as the head, torso, and limbs of a person, connectors and insulators of equipment), then maps these key points to three-dimensional space through a camera model, and fuses them with the point cloud positioning results to obtain a more accurate spatial position.
[0201] For inspection personnel, the system employs a skeleton tracking algorithm to identify the three-dimensional positions of key parts such as the head, torso, and limbs in real time, constructing a complete human skeleton model. This fine-grained human posture recognition enables the system to accurately determine the operating posture and potential reach range of inspection personnel, providing a more precise basis for safety distance assessment.
[0202] For electrical equipment, the system applies specialized shape models based on equipment type to accurately locate energized and accessible parts. For example, for disconnectors, the system pays special attention to their contacts and operating mechanisms; for transformers, it identifies critical components such as bushings and leads. This domain-knowledge-based fine-grained positioning significantly improves the relevance and accuracy of safety distance calculations.
[0203] Meanwhile, the system utilizes spatial relationship constraints in the scene relationship graph to optimize target localization results. For example, if the relationship graph indicates that two devices are "connected," then they should have a point of contact in space; if the relationship is "support," then one device should be located on top of the other. These relationship constraints help the system infer reasonable spatial locations even in areas with poor point cloud quality.
[0204] Finally, the system assigns a unique identifier to each key target and records its type, 3D position coordinates (center point and bounding box or fine-grained shape), attitude information, and key state attributes extracted from the fused feature vector (such as equipment temperature, operating parameters, etc.).
[0205] Through the above processing, the key target 3D coordinate data output in this step is a structured dataset that accurately describes the spatial location and key attributes of personnel and equipment that need to be monitored in the scene. This data provides the necessary spatial information for subsequent safety distance calculations. Experimental evaluation shows that the system's target positioning accuracy can reach ±5cm under normal lighting conditions and remain within ±10cm under complex lighting conditions, meeting the accuracy requirements for substation safety distance assessment.
[0206] Surface point sets are extracted from the 3D point cloud model of the key equipment, and human surface point sets are generated from the skeleton model of the inspection personnel. The minimum distance between the two point sets is calculated. Combining the movement trajectory and speed of the inspection personnel, the short-term movement trend is predicted, the possible minimum safe distance is evaluated, and the 3D Euclidean distance is obtained.
[0207] In this step, the system uses the three-dimensional coordinate data of key targets to calculate the three-dimensional Euclidean distance between the inspection personnel and each key piece of equipment, while taking into account the equipment operating status and voltage level, and outputs the real-time safe distance measurement results.
[0208] Safe distance calculation is a core component of risk assessment and directly relates to the safety of inspection personnel. Safe distance requirements for substation equipment vary depending on equipment type, voltage level, and operating status, requiring comprehensive consideration of multiple factors. This system employs a multi-level safe distance calculation strategy, combining equipment characteristics and operating scenarios to achieve accurate safe distance assessments.
[0209] First, the system loads a predefined database of equipment safety distance parameters. This database is established according to national power safety regulations and industry standards, and includes standard safety distance requirements for equipment of different voltage levels (such as 10kV, 35kV, 110kV, 220kV, etc.). For each type of equipment, the database also distinguishes the safety distance requirements for different operating scenarios (such as general inspection, measurement operations, live-line work, etc.).
[0210] Next, the system extracts the spatial location information of inspection personnel and key equipment from the 3D coordinate data of key targets. For inspection personnel, the system considers the spatial range of their entire body, paying particular attention to the position of their arms and their possible range of motion; for equipment, the system focuses on the location of live parts.
[0211] Then, the system calculates the minimum distance between the inspection personnel and each piece of equipment. Traditional methods often simplify this to calculating the Euclidean distance between two points, but this method ignores the shape and size of the target. This system uses a more accurate nearest point pair calculation method, which extracts surface point sets from the 3D point cloud model of the key equipment, generates a human body surface point set from the skeleton model of the inspection personnel, and calculates the minimum distance between the two point sets:
[0212]
[0213] in This represents the Euclidean distance between points p and q. It is a point set on the human body surface. It is the point set on the surface of the equipment.
[0214] To improve computational efficiency, the system employs spatial partitioning indexing techniques (such as octrees or kd-trees) to organize point cloud data and uses a fast nearest-neighbor search algorithm for distance calculation. For large devices with complex shapes, the system also implements a hierarchical distance calculation strategy, first calculating the distance between coarse bounding boxes and then performing fine-grained calculations for potentially close targets.
[0215] In addition to static geometric distances, the system also considers dynamic factors related to personnel. By analyzing the movement trajectories and speeds of inspection personnel in consecutive frames, the system predicts short-term movement trends and assesses the minimum possible distance. This predictive calculation helps to identify potential safety distance violations in advance.
[0216] Simultaneously, the system dynamically assesses the actual hazard level of equipment based on the equipment status information extracted from the fused feature vectors. For example, for equipment that detects abnormal temperatures or leaking gas, the system will increase its hazard level and correspondingly increase the required safety distance.
[0217] Finally, the system compares the calculated actual distance with the standard safe distance in the parameter library to obtain the distance ratio (actual distance / standard safe distance). This ratio intuitively reflects the relationship between the current distance and the safety requirements: a ratio > 1 indicates safety, a ratio < 1 indicates a safety risk, and the smaller the ratio, the higher the risk.
[0218] Through the above processing, the real-time safe distance measurement results output in this step include the following: the minimum actual distance between inspection personnel and each key piece of equipment, the corresponding standard safe distance, the distance ratio, and the distance change trend. This data provides the basis for subsequent dynamic safety threshold adjustments. Experiments show that the system's safe distance calculation error in the actual substation environment is less than 5%, which meets the requirements for safety monitoring.
[0219] The dynamic adjustment of the safety distance threshold based on the risk level assessment results and the operating status of the key equipment specifically includes:
[0220] Based on the risk level assessment results and the operating status of the key equipment, the baseline safety distance threshold is adjusted according to equipment load rate, temperature, and vibration parameters. Environmental conditions are adjusted based on the current temperature, humidity, and gas concentration. The baseline safety distance threshold for areas assessed as medium or high risk is increased based on the risk level assessment results. The safety distance threshold is further adjusted based on the current operation type and qualification level of the inspection personnel to obtain the dynamic safety distance threshold for the current scenario. In actual implementation, the dynamic safety threshold adjustment adopts a multi-level weighted correction model. First, the system loads the baseline safety distance value from the standard database. Then, the final dynamic threshold is calculated through four consecutive correction steps. The device status adjustment uses the formula:
[0221]
[0222] Where L is the equipment load rate (0-1), T is the equipment temperature, and V is the vibration amplitude. , , This is a weighting coefficient, typically set within the range of 0.1-0.3. Environmental condition adjustments are performed using the following formula:
[0223]
[0224] Where H represents ambient humidity and G represents gas concentration. , Environmental factors are weighted. Risk level adjustment uses a piecewise function: when the risk level is low, When the risk level is medium, When the risk level is high, Finally, adjustments are made based on the type of operation and personnel qualifications: ,in For operation type coefficients (inspection: 1.0, measurement: 1.2, maintenance: 1.5), The qualification coefficients are (Basic: 0.9, Intermediate: 0.7, Advanced: 0.5). The system recalculates the dynamic thresholds every 500 milliseconds to ensure that the safe distance requirements are updated in real time according to environmental and operating conditions.
[0225] In this step, the system dynamically adjusts the safety distance thresholds for various types of equipment based on the risk level assessment results and the current equipment operating status and environmental conditions. It compares these thresholds with the real-time safety distance measurement results to determine whether there are any safety distance violations and outputs a safety distance violation warning and a risk area marker.
[0226] Dynamic safety threshold adjustment is a key innovation of this system. Traditional safety distance assessments typically use fixed thresholds, which are difficult to adapt to the complex and ever-changing operating environment of substations. This system achieves dynamic optimization of safety thresholds through multi-factor comprehensive analysis, making safety assessments more accurate and practical.
[0227] First, the system loads baseline safety distance thresholds, which are derived from national power safety regulations and industry standards, categorized by equipment type and voltage level. For example, for 220kV equipment, the general baseline safety distance threshold for inspection is 3m; for 500kV equipment, the baseline threshold may be as high as 5m. These baseline thresholds serve as the starting point for dynamic adjustments.
[0228] Next, the system dynamically adjusts the baseline threshold based on multiple factors:
[0229] Equipment operating status: The system extracts equipment operating parameters, such as load rate, temperature, and vibration, from the fused feature vector in step 2.5. For equipment with high load rate, abnormal temperature, or abnormal vibration, the system increases its safety distance threshold. The adjustment formula is:
[0230]
[0231] in, It is the baseline threshold. , and These are the weights of load, temperature, and vibration factors, set based on empirical values.
[0232] Environmental conditions: The system considers current environmental factors such as temperature, humidity, and gas concentration. In high-temperature and high-humidity environments, air insulation performance deteriorates, requiring an increased safety distance; when abnormal concentrations of insulating gases such as SF6 are detected, the safety distance also needs to be increased. The environmental adjustment formula is:
[0233]
[0234] in, and It is the weight of humidity and gas factors.
[0235] Risk assessment results: The system integrates risk level assessment results, and for areas assessed as medium or high risk, the corresponding safety distance threshold is increased. The risk adjustment formula is:
[0236]
[0237] Where risk_level is 0 (low risk), 1 (medium risk), or 2 (high risk), It is the weight of risk factors.
[0238] Operation Type: The system identifies the current operation type of the inspection personnel, such as simple inspection, instrument measurement, maintenance operation, etc. Different operation types correspond to different safety distance requirements. The operation adjustment formula is:
[0239]
[0240] in, It is the operation type coefficient, such as 1.0 for simple inspection, 1.2 for instrument measurement, and 1.5 for maintenance operation.
[0241] Personnel Qualification: The system considers the qualification level and experience of inspection personnel, appropriately lowering the security threshold for highly qualified personnel and raising the threshold for novices. This factor is implemented through personnel identification and a pre-set qualification database.
[0242] By comprehensively adjusting the above multiple factors, the system calculates the dynamic safety distance threshold for each device in the current scenario. Then, the system will calculate the actual distance. With dynamic threshold Compare and calculate the safety margin:
[0243]
[0244] Based on the safety margin, the system determines the safety status:
[0245] ≥ 1.2: Safe state, no warning required.
[0246] 1.0 ≤ <1.2: Critical state, issue a warning.
[0247] <1.0: Violation status, issue a warning and mark the risk area.
[0248] For detected violations, the system further analyzes the severity and urgency of the violation. Severity is based on the size of the safety margin, while urgency takes into account the speed and direction of personnel movement. For example, if personnel are rapidly approaching equipment, the system will issue an early warning even if the current distance is still within the safe range.
[0249] Finally, the system generates a safety distance violation warning and a risk area marker. The warning message includes the violating equipment, the current distance, the required safety distance, and suggested actions. The risk area marker uses different colors to indicate the extent and level of danger in the visual interface, providing inspection personnel with intuitive safety guidance.
[0250] Through the above processing, the safety distance violation warnings and risk area markers output in this step are highly targeted and real-time, providing accurate safety assessments based on the specific circumstances of the current scenario. Experiments show that compared to the fixed threshold method, the dynamic threshold adjustment mechanism reduces the false alarm rate by approximately 30% while maintaining a high detection rate for real risks, significantly improving the effectiveness of safety monitoring.
[0251] The generation of comprehensive risk warning information specifically includes the following steps:
[0252] Based on the risk level assessment results, the risk distribution map, the safety distance violation warning, and the risk area identification, the risk type, risk level, and risk distribution of the inspection status are comprehensively analyzed, and a comprehensive risk assessment result is output.
[0253] This step is a crucial integration link in the risk assessment system. Based on the aforementioned risk level assessment results, risk distribution map, safety distance violation warnings, and risk area markers, the system conducts in-depth analysis and comprehensive evaluation to form a comprehensive risk assessment result, providing a basis for subsequent early warning and handling.
[0254] The system first constructs a multidimensional risk state vector, a high-dimensional data structure that integrates the outputs of multiple risk assessment components. Specifically, this vector comprises five core parts: various risk scores and a comprehensive risk level provided by the risk level assessment results; risk density characteristics and spatial distribution patterns extracted from the risk distribution map; the location, extent, and duration of safety distance violations; the area type, scope, and severity in the risk area identification; and risk change trends and predicted values discovered in time-series analysis. This multidimensional state vector provides a comprehensive characterization of the current inspection status, capturing both static risk characteristics and dynamic risk evolution.
[0255] The first step in the comprehensive analysis is risk correlation analysis. The system uses association rule mining algorithms to identify potential correlation patterns between different risk factors. For example, the system may find a high correlation between "abnormal equipment temperature" and "deterioration in insulation performance," or that "violation of safety distances" is often accompanied by "increased operational risk." These association rules help the system understand the mutual influence and possible causal relationships between risks, providing a basis for more accurate risk assessment. In the correlation analysis, the system calculates indicators such as support and confidence levels to screen out rules with high significance and reduce the interference of random correlations.
[0256] The system then prioritizes risks to identify those requiring immediate attention and action. This prioritization considers multiple factors: risk severity (potential loss), probability of occurrence (likelihood of the risk event), scope of impact (number of affected areas and devices), rate of development (speed at which the risk situation worsens), and controllability (the degree to which mitigation measures can be taken). The system uses a weighted scoring method to integrate these factors and assign a priority score to each identified risk. For example, risks with high severity, high probability, and wide impact naturally receive the highest priority; while risks that are severe but have extremely low probability or are highly controllable may receive a lower priority.
[0257] The system pays particular attention to the cascading effects and potential consequences of risks. In a substation environment, a risk event often triggers a series of chain reactions, causing a wider impact. The system simulates possible cascading paths through risk propagation models to predict the final scope and severity of the risk event's impact. For example, the system might analyze how "abnormal temperature of a circuit breaker" gradually develops into "damaged conductor insulation," which may then lead to "short circuit," and ultimately potentially trigger a "local power grid failure"—a complete cascading path. This consequence analysis helps the system assess the true severity of risks, avoiding simplistic judgments based solely on surface phenomena.
[0258] To improve the contextual adaptability of risk assessment, the system introduces a context-weighted mechanism, adjusting risk assessment standards based on the specific context of the current inspection. Contextual factors include: inspection task type (e.g., routine inspection, troubleshooting, emergency response), inspector qualifications (including professional level, experience, and authorization level), weather and environmental conditions (e.g., high temperature, rain, snow, strong winds), equipment operating status (e.g., full load, light load, maintenance status), and special time period requirements (e.g., power supply during holidays, power supply guarantee for major events). The system pre-sets a risk weight adjustment matrix for different scenarios, automatically adjusting the importance weight of each risk factor according to the current context, generating risk assessment results that better meet actual needs.
[0259] The system also performs probability-impact matrix analysis, a two-dimensional risk characterization method that divides risks into nine quadrants based on their probability of occurrence (low, medium, high) and potential impact (minor, moderate, severe). Based on the comprehensive analysis results, the system maps each risk item to its corresponding quadrant, generating an intuitive risk distribution map. This matrix analysis helps users quickly identify critical risks (high probability-high impact quadrant) and tolerable risks (low probability-low impact quadrant), providing a clear basis for risk management decisions.
[0260] Ultimately, the system generates a structured comprehensive risk assessment result, including the following core components: a risk overview summary, which concisely summarizes the overall risk level and main risk points of the current inspection status; a list of key risk items, which details the risk items after priority ranking, including risk type, severity, location, cause, and recommended measures; a risk correlation diagram, which shows the correlation between different risk factors and possible propagation paths; a scenario analysis report, which explains the characteristics of the current scenario and its impact on the risk assessment; and a risk evolution prediction, which provides short-term predictions of the risk situation based on time-series data analysis.
[0261] Through this multi-dimensional and multi-level comprehensive analysis, the system's risk assessment results not only accurately reflect the risk characteristics of the current state but also provide a dynamic perspective on risk development and response suggestions, offering a comprehensive and reliable basis for subsequent early warning and handling decisions. Compared with traditional risk assessments that rely on human experience, this data-driven comprehensive analysis method has significant advantages in comprehensiveness, consistency, and predictability. In particular, when dealing with complex and ever-changing risk scenarios, it can uncover potential risks and correlations that might be overlooked by manual assessments.
[0262] Generate corresponding early warning information for different risk levels, including risk type, location, severity, and handling suggestions, and output early warning information packages;
[0263] Based on the comprehensive risk assessment results, this step focuses on generating targeted early warning information and designing differentiated early warning strategies for different risk levels to ensure the accuracy and operability of the early warning information. Early warning information is a key interface for interaction between the risk assessment system and the user; its design directly affects the system's practical value and user experience.
[0264] The system first establishes a three-tiered early warning framework, corresponding to low risk (green alert), medium risk (yellow alert), and high risk (red alert). Each level employs a different warning strategy: green alerts primarily provide information, informing users of potential low-level risks that generally do not require immediate intervention; yellow alerts serve as a reminder, advising users to pay attention to specific risks and consider taking preventative measures; and red alerts are mandatory warnings, requiring users to take immediate action to address serious risks, and, if necessary, interrupting current operations. This tiered early warning strategy effectively balances safety and operational efficiency, avoiding the "early warning fatigue" problem that can result from excessive early warnings.
[0265] For each identified risk item, the system generates a structured early warning information package containing five core elements: risk type identifier, clearly indicating the nature of the risk (e.g., electrical risk, mechanical risk, operational risk, etc.); risk location description, precisely locating the spatial position of the risk, which may include equipment name, area code, or coordinates; severity indicator, quantifying the severity of the risk, including risk level, confidence level, and potential impact range; time characteristics, describing the time characteristics of the risk, such as whether it is a continuous or transient risk, and the trend of risk status changes; and handling recommendations, providing specific countermeasures and operational guidance for the specific risk.
[0266] For different types of risks, the system has designed specialized early warning templates to ensure the professionalism and relevance of the warning information. For example, electrical risk warnings emphasize electrical safety measures, such as maintaining a safe distance and wearing insulated equipment; mechanical risk warnings may focus on the condition of supporting structures and fixed facilities; environmental risk warnings may provide information on abnormal environmental parameters and corresponding protection recommendations; while operational risk warnings focus on operational procedures and safe operating guidelines. These specialized warning templates ensure the relevance and practicality of the warning information.
[0267] The system prioritizes the understandability and operability of warning messages. Warning text uses concise and intuitive language, avoiding excessive jargon and complex descriptions. Key information is emphasized visually using bolding and highlighting to enhance readability. Warning messages are organized by urgency and operational sequence, with the most urgent action recommendations placed first. The action recommendations section provides specific and actionable guidance, rather than vague general advice. For example, instead of simply stating "Maintain a safe distance," it explicitly states "Immediately retreat to the area beyond the yellow line, maintaining a safe distance of at least 2 meters." This design ensures that users can quickly understand the risk situation and take appropriate action after receiving the warning.
[0268] For complex risk scenarios, the system implements a contextualized early warning generation mechanism. The system considers contextual information such as the current inspection task type, personnel roles, and environmental conditions to adjust the content and expression of the early warning. For example, early warnings for technical personnel may include more technical details and jargon, while warnings for ordinary inspectors are simpler and more intuitive. Similarly, in emergency situations, early warning information is more concise and direct, highlighting the most critical action instructions; while in routine inspections, warnings may include more background information and preventative suggestions.
[0269] The system also features a multimodal approach to warning messages, enhancing their effectiveness. In addition to text descriptions, warning messages may include: visual markers of risk areas, such as highlighting risk areas on on-site camera footage or 3D models; alarm sounds, using different tones and rhythms to indicate different levels of risk; vibration alerts, using vibrations of varying intensities and patterns on mobile devices to convey the urgency of the risk; and simplified icons, using standardized risk icons to quickly convey the type and level of risk. This multimodal approach ensures that warning messages can be effectively communicated under various environmental conditions, such as noisy environments or situations where attention is diverted.
[0270] To ensure timely alerts, the system implements an alert priority management mechanism. High-priority alerts (such as red alerts) are immediately pushed to the terminal devices of all relevant personnel and may trigger audible and visual alarms; medium-priority alerts are pushed at appropriate times to avoid interfering with critical operations; low-priority alerts may be sent periodically in aggregate form. The system also adjusts the alert frequency based on user feedback and confirmation to avoid resending confirmed alerts.
[0271] Through this meticulously designed early warning information generation mechanism, the system can transform complex risk assessment results into clear and practical early warning information, effectively supporting on-site personnel's safety decisions. Practical applications show that compared to traditional early warning methods, this structured and differentiated early warning information can improve user response speed by approximately 25%, reduce misunderstandings and improper operations by approximately 30%, and significantly enhance the safety assurance level of substation inspections. Ultimately, the early warning information packages output by the system are not only used for on-site guidance but are also encrypted and transmitted to the monitoring center, supporting remote supervision and decision-making assistance.
[0272] The warning information is encrypted using the AES encryption algorithm, transmitted to the monitoring center using LoRa or NB-IoT low-power wide area network communication technology, and a real-time warning signal is output.
[0273] This step aims to ensure the secure and reliable transmission of risk warning information, preventing information leakage or tampering, while meeting the low-power, long-distance communication requirements of substation environments. Considering the special safety requirements and complex electromagnetic environment of power systems, a dedicated encrypted transmission mechanism is designed to guarantee the integrity and confidentiality of warning signals.
[0274] The system first performs high-strength encryption on the warning information packets. The system employs the AES (Advanced Encryption Standard) algorithm, a widely used symmetric encryption algorithm with high security and computational efficiency. Specifically, the system uses AES-GCM (Galois / Counter Mode) with a 256-bit key length. This mode not only provides data encryption but also verifies data integrity, preventing data tampering during transmission. The encryption process includes three key steps: First, the system generates a random initialization vector (IV) to ensure that even if the same content is encrypted multiple times, different ciphertexts will be produced; second, the warning information is encrypted using a pre-shared encryption key; and finally, an authentication tag is generated for the receiving end to verify the integrity and authenticity of the data.
[0275] Key management is a core component of the encryption system. The system implements a dynamic key update mechanism, automatically updating encryption keys periodically (typically every 24 hours or before each important task), significantly reducing the risk of key leakage. New keys are distributed to authorized devices through secure channels (such as physically isolated networks or pre-established encrypted channels). The system also employs a hierarchical key architecture, using a master key to generate session keys, restricting the scope and duration of each key's use, further enhancing system security. All key operations are performed within the secure module; key materials never appear in plaintext form in memory or storage.
[0276] After encryption, the system needs to select a suitable communication technology for data transmission in the substation environment. The system primarily employs two low-power wide-area network (LPWAN) technologies: LoRa (LoRa, Long-Range Radio) and NB-IoT (Narrowband Internet of Things). LoRa technology, based on spread spectrum modulation, features long transmission distances (over 15 kilometers in open areas), low power consumption, and strong anti-interference capabilities, making it particularly suitable for scenarios like substations with wide coverage areas and complex electromagnetic environments. NB-IoT, on the other hand, is a narrowband technology based on cellular networks, offering wide coverage and strong penetration capabilities, making it suitable for situations requiring more reliable connections and higher data priority. The system dynamically selects the most appropriate transmission technology based on the urgency of the warning information, the data volume, and the current network conditions. For the most critical red alerts, the system may simultaneously use both technologies for redundant transmission to ensure the information reaches the monitoring center.
[0277] To ensure reliable communication in harsh environments, the system implements several transmission enhancement mechanisms. Adaptive data rate control dynamically adjusts the transmission rate based on signal quality, reducing the rate to improve reliability when the signal is weak. Forward error correction coding (FEC) adds redundant bits to data packets, enabling the receiver to detect and correct a certain degree of transmission errors. Packet segmentation and reassembly technology divides large warning messages into multiple smaller data packets for transmission, allowing for the recovery of most information even if some packets are lost. Automatic repeat request (ARQ) automatically retransmits data packets upon detecting a transmission failure until acknowledgment is received or the maximum number of retransmissions is reached. These technologies collectively ensure highly reliable communication in the complex electromagnetic environment of substations.
[0278] The system also implements intelligent power management, extending the battery life of transmission equipment. A sleep mechanism allows the device to enter a low-power state when there is no data transmission demand; a tiered wake-up strategy adjusts the wake-up frequency according to risk level, with more frequent detection and transmission in high-risk states and longer sleep time in low-risk states; data compression technology compresses warning information before transmission, reducing the amount of transmitted data and energy consumption; transmission power control adjusts the transmission power according to distance and channel conditions, minimizing energy consumption while ensuring reliable transmission. Through these technologies, the system can operate continuously for months to years on standard battery power, significantly reducing maintenance costs.
[0279] After the data arrives at the monitoring center, the system executes a rigorous verification and decryption process. First, the authentication tag is verified to confirm the data has not been tampered with; then, the data is decrypted using the corresponding key to restore the original alert information; finally, data integrity is verified to ensure all information has been correctly received. For time-sensitive emergency alerts, the system has a priority processing mechanism to ensure that such information is presented to monitoring personnel as quickly as possible.
[0280] Upon receiving an early warning, the monitoring center immediately triggers the corresponding real-time warning signal. Depending on the warning level, the system may activate different levels of alarms: visual alarms on the display screen, showing risk information with prominent colors and icons; audible alarms, using different tones and frequencies to indicate different risk levels; SMS or mobile application push notifications, sending warning information to administrators' mobile devices; and linkage with other security systems, such as triggering emergency response plans in particularly severe cases. These multi-layered warning signals ensure that monitoring center personnel can promptly detect risks and take appropriate measures.
[0281] To address potential communication disruptions, the system incorporates local storage and delayed transmission mechanisms. When the network is unavailable, encrypted alerts are temporarily stored in a secure local storage area and prioritized in a queue, transmitting them in order of priority once communication is restored. For emergency alerts, the system may also activate backup communication channels, such as satellite communication or dedicated radio frequencies, to ensure that critical information is not delayed due to network issues.
[0282] Through this comprehensive and secure encrypted transmission system, the system ensures that early warning information can be transmitted securely and reliably from the field to the monitoring center, providing a solid foundation for centralized monitoring and remote assistance. The system's encryption protection mechanism effectively prevents information leakage and tampering, low-power wide-area network technology solves the long-distance communication challenges in substation environments, and multiple reliability safeguards ensure the timely delivery of critical early warning information. These technical characteristics enable the system to meet the stringent requirements of power systems for security, reliability, and real-time performance, providing crucial support for the safe operation of substations.
[0283] The comprehensive risk assessment results, raw sensor data, and intermediate processing results are stored in a distributed database in time series to establish a risk assessment database and output structured risk data records.
[0284] This step establishes a comprehensive risk data management system, storing and managing various types of risk assessment data in a structured manner, providing a data foundation for subsequent analysis, reporting, and knowledge accumulation. Considering the diversity, large volume, and high value of substation risk assessment data, the system is designed with a dedicated distributed database architecture to ensure data integrity, availability, and ease of retrieval.
[0285] The system first defines a comprehensive data acquisition strategy to ensure the complete preservation of key data. Specifically, the system stores three types of core data: risk assessment results data, including advanced analysis results such as comprehensive risk assessment results, risk level determination, and distribution; raw sensor data, including raw data collected by multimodal sensors (vision, infrared, environmental, etc.), which serves as the foundational evidence for risk judgment; and intermediate processing results, including key results from intermediate processing steps such as feature extraction, scene graph construction, and risk model output. This multi-layered data storage strategy not only supports the querying and use of risk assessment results but also allows for retrospective analysis when necessary to verify the accuracy and rationality of the assessment results.
[0286] Based on the different characteristics and uses of the data, the system adopts a hybrid database architecture. Structured risk assessment results and metadata are stored in relational databases (such as PostgreSQL) for easy complex queries and report generation; large-capacity raw sensor data is stored in a distributed file system (such as HDFS) to support efficient storage and parallel processing of massive amounts of data; intermediate processing results are stored in document-oriented databases (such as MongoDB) to accommodate their semi-structured nature; and time-series data (such as time series of environmental parameters and risk indicators) are stored in a dedicated time-series database (such as InfluxDB) to optimize time-series queries and analysis. This hybrid architecture fully leverages the advantages of various databases to provide the most suitable storage and access methods for different types of data.
[0287] To ensure logical data organization and efficient retrieval, the system implements a multi-dimensional indexing mechanism. Data is indexed along multiple dimensions: time dimension, supporting queries based on a specific time point or time period, such as inspection records for a specific date or risk trends over a certain period; spatial dimension, supporting location-based queries, such as risk records for a specific area or equipment; risk type dimension, supporting data filtering by category such as electrical risk or mechanical risk; severity dimension, supporting data filtering by risk level or priority; and inspection task dimension, supporting data filtering by specific inspection task or inspector. This multi-dimensional indexing structure allows users to flexibly combine query conditions to quickly locate the information they need.
[0288] During data storage, the system employs rigorous data normalization and quality control measures. Data normalization ensures that all stored data adheres to a unified format and naming convention, facilitating integrated analysis. Data validation mechanisms check the integrity and consistency of data before storage, rejecting abnormal or incomplete data. Data preprocessing steps perform necessary cleaning, transformation, and aggregation of raw data to optimize storage efficiency. Data compression and encoding technologies reduce storage space requirements while maintaining data accessibility. These measures ensure high-quality and consistent stored data, providing a reliable foundation for subsequent analysis.
[0289] Given the sensitivity of power system data, comprehensive data security protection measures have been implemented. Fine-grained access control sets different data access permissions for different user roles, ensuring that users can only access data within their scope of responsibility; sensitive data is encrypted to protect critical data in storage, preventing the leakage of sensitive information even if the database is compromised; data operation audits log all data access and modification operations, facilitating security reviews and accountability; data backup and disaster recovery mechanisms regularly create data backups and have designed recovery strategies to cope with hardware failures or natural disasters. These security measures together construct a highly protected data environment that meets the stringent security requirements of the power industry.
[0290] The system also incorporates an intelligent data lifecycle management strategy to balance the need for long-term data retention with storage cost control. Based on data importance and usage frequency, the system automatically allocates data to different storage tiers: hot data (recently generated or frequently accessed data) is stored in high-speed storage media to ensure fast access; warm data (data with moderate access frequency) is stored in standard storage media to balance performance and cost; and cold data (historical data or data with low access frequency) is automatically migrated to low-cost archive storage to save on storage costs. The system also implements a data aging strategy, setting different retention periods for different types of data, with expired data automatically archived or compressed. For example, high-resolution raw sensor data may only be retained for a few months, while processed risk assessment results may be retained for several years or longer.
[0291] To support distributed deployment and high availability requirements, the system employs sharding and replication technologies. Data sharding divides large datasets into multiple manageable shards distributed across multiple storage nodes, supporting parallel processing and load balancing. Data replication maintains copies of the data across multiple nodes, ensuring data availability even if some nodes fail. A distributed consistency protocol ensures the consistency and correctness of data updates in a distributed environment. An automatic failover mechanism automatically switches services to healthy nodes when a node failure is detected, ensuring continuous system availability. This distributed architecture provides high scalability and fault tolerance, adapting to the data management needs of substations of different sizes.
[0292] Through this comprehensive risk data storage and management system, the system not only securely and reliably preserves the historical records of risk assessments but also establishes a structured and easily searchable knowledge base, providing a solid data foundation for subsequent trend analysis, experience summarization, and model optimization. This systematic data management approach transforms short-term risk assessments into long-term knowledge accumulation, continuously improving the safety management level of substations. Practice shows that compared with traditional distributed, file-based data storage, this centralized and structured data management method can improve data utilization efficiency by approximately 40% while significantly enhancing the ability to discover risk patterns across time and regions.
[0293] Based on the structured risk data records, a risk assessment report is generated that includes risk trend analysis, typical risk cases, and improvement suggestions. It also provides risk data query and statistical analysis functions, and outputs the risk assessment report and data management interface.
[0294] This step is the final output of the risk assessment system, transforming the stored structured risk data into valuable analytical reports and actionable management tools. The system generates comprehensive risk assessment reports while providing flexible data query and analysis functions to support managers in conducting in-depth risk reviews and decision-making.
[0295] The system first constructs a multi-layered risk reporting framework to meet the needs of different user roles and usage scenarios. Specifically, the system generates three core types of reports: daily inspection reports, which record in detail the risk items discovered, the handling measures, and the results of a single inspection activity, primarily for on-site management personnel; periodic statistical reports, which summarize risk data over a certain period (such as weekly or monthly reports), identify common risk patterns and trends, and are intended for middle-level management personnel; and strategic analysis reports, which examine the effectiveness of risk management from a long-term perspective and propose systematic improvement suggestions, and are intended for senior decision-makers. This layered reporting strategy ensures that information is delivered to the right audience with appropriate detail and perspective, maximizing the practical value of the reports.
[0296] Risk trend analysis is a crucial component of the report. The system extracts risk change patterns from time-series data, providing a forward-looking perspective for management decisions. Trend analysis encompasses multiple dimensions: time pattern analysis, identifying patterns of risk change over time, such as seasonal fluctuations, weekday / restday differences, or long-term evolution trends; spatial distribution analysis, identifying the geographical concentration and changing characteristics of high-risk areas; risk type analysis, tracking the relative frequency changes of different types of risks; and severity analysis, monitoring the evolution of risk level distribution. The system uses time-series analysis techniques, such as moving averages, exponential smoothing, and seasonal decomposition, to extract key trends and filter out short-term fluctuations and noise. Trend visualization employs intuitive chart formats, such as line graphs, heatmaps, and regional maps, making trend changes readily apparent.
[0297] The typical risk case library is another core part of the report, providing lessons learned and best practices through specific cases. The system automatically identifies and extracts representative risk cases from historical data. These cases typically possess the following characteristics: typical risk patterns, representing common risk types or scenarios; outstanding handling results, demonstrating effective measures for successfully avoiding risks; clear lessons learned, containing clear lessons or warnings; or rare but high-impact cases, although occurring infrequently, with potentially serious consequences, worthy of special attention. Each typical case is presented in a structured manner, including a background description, risk identification process, measures taken, final results, and a summary of lessons learned. The system also clusters and associates similar cases to form a knowledge graph, helping users understand the connections and commonalities between different cases.
[0298] Improvement suggestion generation is a high-value part of the report, transforming risk analysis into concrete action guidance. Based on accumulated risk data, the system generates three categories of improvement suggestions: operational suggestions, targeting inspection processes, personnel training, and safety awareness, such as "increasing the inspection frequency of specific high-risk areas" or "designing specialized training for frequently occurring risk points"; equipment suggestions, targeting equipment maintenance, updates, and configuration, such as "replacing aging insulation components" or "adjusting equipment layout to reduce specific risks"; and system-level suggestions, targeting management systems, monitoring systems, and emergency plans, such as "improving emergency response procedures for specific scenarios" or "enhancing the monitoring system's ability to detect specific risks." Suggestion generation is not only based on statistical analysis but also incorporates best practices from a domain expert knowledge base, ensuring the professionalism and feasibility of the suggestions. The system prioritizes each suggestion based on cost-benefit analysis, helping managers make resource allocation decisions.
[0299] To support in-depth data exploration, the system offers a wealth of interactive data query and analysis functions. The advanced query engine supports complex multi-condition queries, such as "finding high electrical risk events in a specific area over the past three months" or "comparing the differences in risk distribution faced by different inspection personnel." Data visualization tools provide dynamic charts and interactive dashboards, supporting data drill-down and perspective switching. Statistical analysis functions support advanced analytical methods such as hypothesis testing, correlation analysis, and regression modeling, helping users discover hidden patterns in the data. Comparative analysis tools support comparing risk performance across different periods, regions, or configurations, identifying significant differences and potential causal relationships. These analytical tools feature an intuitive interface design, enabling non-technical personnel to conduct complex data exploration and maximize the value of the data.
[0300] The system also implements a closed-loop optimization mechanism, feeding back findings and suggestions from reports into the risk assessment model. Specifically, the system records user feedback and corrections to the risk assessment results as training data for model optimization; extracts risk identification patterns from successful cases to enhance the model's ability to identify similar situations; analyzes false positives and false negatives to identify model weaknesses and make targeted improvements; and continuously updates the risk knowledge base to reflect the latest equipment characteristics, operating procedures, and safety standards. This closed-loop optimization mechanism enables the risk assessment system to learn from experience, continuously improving the accuracy and adaptability of its assessments.
[0301] The system features a flexible report distribution and access mechanism to ensure information reaches target users efficiently. Reports can be generated and distributed in various formats: standardized PDF documents suitable for formal documentation and offline reading; interactive web interfaces supporting dynamic data exploration and customized views; mobile app push notifications providing key summaries and reminders; and regular email summaries proactively delivering important findings and recommendations. The system personalizes report content and format based on user roles and preferences, ensuring each user receives the most relevant information. Access control mechanisms ensure sensitive information is only visible to authorized personnel while allowing appropriate information sharing and collaboration.
[0302] Through this comprehensive risk assessment report and data analysis system, the system transforms complex risk data into clear insights and concrete action guidelines, providing strong support for substation safety management. Compared to traditional static reports, this dynamic and interactive risk analysis platform offers deeper insights and more precise recommendations, significantly improving the efficiency and effectiveness of risk management. Practice shows that using this systematic risk reporting and analysis method, substations can reduce risk events by approximately 25%, while optimizing the allocation of resources for safety investments and reducing safety management costs by approximately 15%, achieving a dual improvement in safety and efficiency.
[0303] like Figure 2 As shown, the present invention also provides an intelligent inspection risk assessment system based on multi-sensor fusion, comprising:
[0304] The multi-source sensor data acquisition module is used to acquire raw multi-source data from the substation, including binocular visual images, infrared thermal imaging data, and environmental sensor data. The module performs spatiotemporal calibration and preprocessing on the raw multi-source data to obtain a multi-source sensor data stream.
[0305] The multi-sensor data fusion module is used to extract features from the multi-source sensor data stream to obtain a multimodal feature set, generate a refined semantic mask based on the multimodal feature set and construct an initial scene relationship graph, and use a statistical confidence re-scoring mechanism to evaluate the reliability of the initial scene relationship graph to obtain an optimized scene relationship graph, and fuse the multimodal feature set and the optimized scene relationship graph to output a fused feature vector and a target scene relationship graph;
[0306] The risk assessment module is used to construct a risk representation model based on the implicit neural representation method according to the fused feature vector and the target scene relationship graph, and to construct a deep weight space network based on the risk representation model. The deep weight space network maps scene features to risk space and outputs risk feature space mapping. Based on the output risk feature space mapping, a relationship attention converter model is used to obtain relationship-enhanced risk features. The risk level is calculated based on a multilayer perceptron classifier and risk level assessment results and risk distribution map are generated.
[0307] The safety distance calculation module is used to generate a high-precision scene depth map based on the binocular vision image in the multi-source sensor data stream through deep distillation gradient preprocessing to optimize the stereo matching process, perform 3D point cloud reconstruction and identify the 3D coordinates of inspection personnel and key equipment, calculate the 3D Euclidean distance between the inspection personnel and the key equipment, calculate the real-time safety distance measurement result based on the 3D Euclidean distance, dynamically adjust the safety distance threshold according to the risk level assessment result and the operating status of the key equipment, and output a safety distance violation warning and risk area identification by comparing the real-time safety distance measurement result and the safety distance threshold.
[0308] The comprehensive early warning module is used to generate comprehensive risk early warning information based on the risk level assessment results, the risk distribution map, the safety distance violation warning, and the risk area identifier.
[0309] In a preferred embodiment, the system further includes:
[0310] Edge computing units, equipped with GPU accelerator cards and high-capacity storage devices, are used for real-time processing of multi-source sensor data streams and running deep learning algorithms;
[0311] The wireless communication unit integrates LoRa and NB-IoT communication modules, supports AES encryption algorithm, and is used for secure communication with the monitoring center;
[0312] The human-computer interaction unit includes a high-resolution display screen and a touch interface for displaying risk distribution maps and early warning information in real time;
[0313] The power management unit uses a combination of solar panels and lithium batteries to provide power and supports uninterrupted operation.
[0314] The environmental adaptability unit is designed with an IP65 protection rating and an operating temperature range of -40℃ to +70℃, making it suitable for harsh substation environments.
[0315] The above description is merely a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.
Claims
1. A multi-sensor fusion-based intelligent inspection risk assessment method, characterized in that, The method comprises the following steps: Obtaining original multi-source data of a substation, the original multi-source data comprising binocular vision images, infrared thermal image data and environmental sensor data, and performing space-time calibration and preprocessing on the original multi-source data to obtain a multi-source sensor data stream; Performing feature extraction on the multi-source sensor data stream to obtain a multi-modal feature set, generating a refined semantic mask based on the multi-modal feature set, constructing an initial scene relationship graph, performing reliability evaluation on the initial scene relationship graph by using a statistical confidence rescore mechanism to obtain an optimized scene relationship graph, and fusing the multi-modal feature set and the optimized scene relationship graph to output a fusion feature vector and a target scene relationship graph; Based on the fusion feature vector and the target scene relationship graph, constructing a risk representation model based on an implicit neural representation method, constructing a deep weight space network based on the risk representation model, mapping scene features to a risk space through the deep weight space network to output a risk feature space mapping, based on the risk feature space mapping, using a relationship attention transformer model to obtain a relationship-enhanced risk feature, calculating a risk level based on a multilayer perceptron classifier and generating a risk level evaluation result and a risk distribution map; Based on the binocular vision images in the multi-source sensor data stream, generating a high-precision scene depth map by deep distillation gradient preprocessing to optimize the stereo matching process, performing three-dimensional point cloud reconstruction and identifying the three-dimensional coordinates of the inspection personnel and the key equipment, calculating the three-dimensional Euclidean distance between the inspection personnel and the key equipment, and based on the three-dimensional Euclidean distance, calculating a real-time safety distance measurement result, dynamically adjusting the safety distance threshold value according to the risk level evaluation result and the operating state of the key equipment, and outputting a safety distance violation warning and a risk area identifier by comparing the real-time safety distance measurement result with the safety distance threshold value; wherein the calculation of the three-dimensional Euclidean distance between the inspection personnel and the key equipment comprises: performing three-dimensional space reconstruction based on the high-precision scene depth map and the camera internal and external parameters to generate a three-dimensional point cloud model representing the substation scene; identifying the inspection personnel and the key equipment in the scene according to the three-dimensional point cloud model and the fusion feature vector and the target scene relationship graph, and determining the accurate position coordinates of the inspection personnel and the key equipment in the three-dimensional space; extracting a surface point set from the three-dimensional point cloud model of the key equipment, generating a human surface point set according to an inspection personnel skeleton model, and calculating the minimum distance between the two point sets; combining the motion trajectory and speed of the inspection personnel to predict the short-term movement trend to obtain the three-dimensional Euclidean distance; Based on the risk level evaluation result, the risk distribution map, and the safety distance violation warning and the risk area identifier, generating comprehensive risk early warning information.
2. The method of claim 1, wherein, The obtaining of the original multi-source data of the substation comprises: Deploying binocular vision sensors, infrared thermal imagers and environmental sensor arrays, calibrating each sensor in space and time, establishing a unified coordinate system and time reference, and outputting a sensor calibration parameter matrix; According to the sensor calibration parameter matrix, a multi-sensor synchronous acquisition mechanism is established, the acquisition frequencies of different sensors are set, and a unified timestamp is added, and the original multi-source data with the timestamp is output; The original multi-source data is time and space calibrated and preprocessed to obtain a multi-source sensor data stream, including: The binocular vision image and the infrared thermal image data in the original multi-source data with the timestamp are deep distillation gradient preprocessed to realize denoising, distortion correction and enhancement processing, and the environmental sensor data is denoised and filtered by using a sliding average filtering algorithm to obtain the preprocessed binocular vision image, infrared thermal image data and environmental sensor data; The preprocessed binocular vision image, infrared thermal image data and environmental sensor data are time and space aligned by linear interpolation and coordinate transformation to obtain the multi-source sensor data stream.
3. The method of claim 1, wherein, The statistical confidence re-scoring mechanism is used to evaluate the reliability of the initial scene relation graph, including: Statistical priors are extracted from a preset substation scene graph database, and the statistical priors include node category distribution probability, conditional category probability, relation co-occurrence probability and relation transmission characteristics; The node confidence score and the edge confidence score corresponding to each node and each edge in the initial scene relation graph are calculated based on the statistical priors; Based on the node confidence score and the edge confidence score, the initial scene relation graph is evaluated and analyzed to obtain the re-evaluated initial scene relation graph.
4. The method of claim 1, wherein, The risk representation model is constructed based on the implicit neural representation method, including: A risk representation function is defined, which realizes continuous mapping of spatial coordinates to risk attributes through a multi-layer perception; The fusion feature vector is encoded through a feature mapping network to obtain a latent feature representation, and the latent feature representation and the three-dimensional spatial coordinates are spliced and input into the risk representation function to construct a conditional implicit neural representation; The target scene relation graph is processed using a graph neural network to obtain the latent representation of each node in the target scene relation graph, and the spatial relationship between the three-dimensional spatial coordinates and each node is calculated, and the graph condition features are aggregated through an attention mechanism; The latent feature representation, the graph condition features and the three-dimensional spatial coordinates are spliced and input into the risk representation function to construct an expanded conditional implicit neural representation to obtain an implicit neural representation model; The implicit neural representation model is trained using historical data as a supervision signal, and model parameter optimization is performed based on minimizing the difference between the predicted risk and the real risk to obtain the risk representation model.
5. The method of claim 1, wherein, The high-precision scene depth map is generated by deep distillation gradient preprocessing to optimize the stereo matching process, including: The binocular vision image is converted to a gradient domain to calculate the horizontal and vertical gradient maps of the image to obtain an original gradient map; A pre-trained deep neural network is used as a teacher network to process the original gradient map to generate a purified gradient; A lightweight student network is trained to distill knowledge from the teacher network to learn the mapping relationship of converting the original gradient map to the purified gradient; Based on the mapping relationship, the original gradient map is processed to obtain an enhanced gradient map; The stereo matching algorithm is executed on the enhanced gradient map to calculate a disparity map, and the disparity map is converted into a depth map through a disparity-depth conversion formula to obtain the high-precision scene depth map.
6. The method of claim 1, wherein, The dynamic adjustment of the safety distance threshold according to the risk level evaluation result and the running state of the key equipment includes: Based on the risk level evaluation result and the running state of the key equipment, the baseline safety distance threshold is adjusted according to the equipment load rate, temperature and vibration parameters; According to the current temperature, humidity and gas concentration, the baseline safety distance threshold adjusted by the equipment state is adjusted according to the environmental conditions; Based on the risk level evaluation result, the baseline safety distance threshold of the area evaluated as medium and high risk level is adjusted upward; According to the current operation type and qualification level of the inspection personnel, the baseline safety distance threshold is adjusted to obtain the dynamic safety distance threshold in the current scene.
7. The method of claim 1, wherein, The generation of comprehensive risk warning information includes: Based on the risk level evaluation result, the risk distribution map, the safety distance violation warning and the risk area identification, the risk type, risk level and risk distribution of the inspection state are comprehensively analyzed, and the comprehensive risk evaluation result is output; For different risk levels, corresponding warning information is generated, including risk type, location, severity and processing suggestion; The warning information is encrypted by the AES encryption algorithm, transmitted to the monitoring center by the LoRa or NB-IoT low-power wide-area network communication technology, and the real-time warning signal is output; The comprehensive risk evaluation result, original sensor data and intermediate processing result are stored in the distributed database in time sequence to establish a risk evaluation database, and a risk evaluation report containing risk trend analysis, typical risk case and improvement suggestion is generated.
8. A multi-sensor fusion based intelligent inspection risk assessment system, characterized in that, It includes: A multi-source sensor data acquisition module is configured to acquire original multi-source data of a substation, the original multi-source data including binocular vision images, infrared thermal image data and environmental sensor data, and to perform time and space calibration and preprocessing on the original multi-source data to obtain a multi-source sensor data stream; A multi-sensor data fusion module is configured to perform feature extraction on the multi-source sensor data stream to obtain a multi-modal feature set, generate a refined semantic mask based on the multi-modal feature set, construct an initial scene relationship graph, perform reliability evaluation on the initial scene relationship graph using a statistical confidence rescore mechanism to obtain an optimized scene relationship graph, and fuse the multi-modal feature set and the optimized scene relationship graph to output a fusion feature vector and a target scene relationship graph; A multi-source sensor data acquisition module is configured to acquire original multi-source data of a substation, the original multi-source data including binocular vision images, infrared thermal image data and environmental sensor data, and to perform time and space calibration and preprocessing on the original multi-source data to obtain a multi-source sensor data stream; A multi-sensor data fusion module is configured to perform feature extraction on the multi-source sensor data stream to obtain a multi-modal feature set, generate a refined semantic mask based on the multi-modal feature set, construct an initial scene relationship graph, perform reliability evaluation on the initial scene relationship graph using a statistical confidence rescore mechanism to obtain an optimized scene relationship graph, and fuse the multi-modal feature set and the optimized scene relationship graph to output a fusion feature vector and a target scene relationship graph; The risk assessment module is configured to construct a risk representation model based on an implicit neural representation method according to the fusion feature vector and the target scene relation graph, and construct a deep weight space network based on the risk representation model, map scene features to a risk space through the deep weight space network, output a risk feature space mapping, obtain a relation-enhanced risk feature by using a relation attention transformer model based on the risk feature space mapping, and calculate a risk level and generate a risk level evaluation result and a risk distribution graph based on a multilayer perception classifier. The safety distance measurement module is configured to generate a high-precision scene depth map by optimizing a stereo matching process through deep distillation gradient preprocessing according to the binocular vision image in the multi-source sensor data stream, perform three-dimensional point cloud reconstruction and identify three-dimensional coordinates of the inspection personnel and the key equipment, calculate a three-dimensional Euclidean distance between the inspection personnel and the key equipment, calculate a real-time safety distance measurement result based on the three-dimensional Euclidean distance, dynamically adjust a safety distance threshold value according to the risk level evaluation result and an operating state of the key equipment, and output a safety distance violation warning and a risk area identifier by comparing the real-time safety distance measurement result and the safety distance threshold value. The calculation of the three-dimensional Euclidean distance between the inspection personnel and the key equipment includes: performing three-dimensional space reconstruction based on the high-precision scene depth map and camera internal and external parameters to generate a three-dimensional point cloud model representing a substation scene; identifying the inspection personnel and the key equipment in the scene and determining accurate position coordinates of the inspection personnel and the key equipment in three-dimensional space according to the three-dimensional point cloud model and the fusion feature vector and the target scene relation graph; extracting a surface point set from the three-dimensional point cloud model of the key equipment, generating a human body surface point set according to an inspection personnel skeleton model, and calculating a minimum distance between the two point sets; and predicting a short-term movement trend in combination with a movement trajectory and a speed of the inspection personnel to obtain the three-dimensional Euclidean distance. The comprehensive early warning module is configured to generate comprehensive risk early warning information based on the risk level evaluation result, the risk distribution graph, and the safety distance violation warning and the risk area identifier.
9. The system of claim 8, wherein, The system further includes: An edge computing unit configured with a GPU acceleration card and a large-capacity storage device for real-time processing of multi-source sensor data streams and running of deep learning algorithms; A wireless communication unit integrated with LoRa and NB-IoT communication modules and supporting an AES encryption algorithm for secure communication with a monitoring center; A human-computer interaction unit including a high-resolution display screen and a touch interface for real-time display of a risk distribution graph and early warning information; A power supply management unit adopting a combination of a solar panel and a lithium battery for uninterrupted operation; An environmental adaptation unit designed with an IP65 protection level, a working temperature range of -40°C to +70°C, and adaptation to harsh environments of substations. An environmental adaptation unit designed with an IP65 protection level, a working temperature range of -40°C to +70°C, and adaptation to harsh environments of substations.