A method and system for multimodal data fusion and analysis for intelligent substation inspection

By using a multi-sensor system and deep analysis technology, the problems of poor environmental adaptability and low spatiotemporal synchronization accuracy in intelligent substation inspection have been solved. High-precision multimodal data fusion across all scenarios has been achieved, improving the accuracy of equipment defect detection and construction risk prediction, and supporting real-time alarms and automatic inspection.

CN122133076APending Publication Date: 2026-06-02STATE GRID JIBEI ELECTRIC POWER COMPANY

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
STATE GRID JIBEI ELECTRIC POWER COMPANY
Filing Date
2026-04-01
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing intelligent substation inspection multimodal data fusion technologies suffer from poor environmental adaptability, low spatiotemporal synchronization accuracy, insufficient fusion depth, incomplete scenario coverage, and weak analysis robustness, failing to meet the high-precision inspection requirements in complex environments.

Method used

A quadruped inspection robot employing a multi-sensor system collects multimodal data. Spatiotemporal calibration is achieved through adaptive electromagnetic interference denoising, temperature-weighted enhancement, federated Kalman filtering, and joint calibration of QR codes and point clouds. A three-level fusion architecture of 'physical layer-feature layer-decision layer' is constructed. In-depth analysis is performed by combining a temporal attention CNN-LSTM network and a Substation-VLM model to generate real-time alarms and automatic inspection reports.

Benefits of technology

High-precision multimodal data fusion was achieved in complex environments, improving spatiotemporal synchronization accuracy and fusion depth, covering all scenarios of equipment, personnel and construction, improving the accuracy of equipment defect identification, response time to personnel violations and accuracy of construction risk prediction, and reducing labor costs and safety risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122133076A_ABST
    Figure CN122133076A_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for multimodal data fusion and analysis for intelligent substation inspection, relating to the field of intelligent operation and maintenance of power systems. The method includes: acquiring multimodal data such as lidar point clouds, dual-light images, and IMU attitude data; preprocessing data through adaptive electromagnetic interference denoising and federated Kalman filtering; constructing a three-level fusion architecture of "physical layer - feature layer - decision layer" to achieve spatiotemporal unification, geometric-semantic association, and semantic-rule reasoning; performing full-scene analysis based on models such as temporal attention CNN-LSTM; and outputting real-time alarms, 3D SLAM maps, and structured reports. This invention is adaptable to the special operating conditions of substations, has high spatiotemporal synchronization accuracy, covers the entire scenario of "equipment-personnel-construction," and significantly improves analysis accuracy and decision-making practicality, effectively supporting intelligent inspection and operation and maintenance decisions for substations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent operation and maintenance technology of power systems, specifically relating to a method and system for multimodal data fusion and analysis for intelligent inspection of substations. Background Technology

[0002] With the advancement of "dual-carbon" goals and the construction of new power systems, the inspection needs of substations, as core hubs of the power grid, have evolved from traditional "manual inspection + fixed equipment monitoring" to "fully autonomous, high-precision, and all-scenario" intelligent inspection during renovation and expansion projects. Multimodal data fusion technology, due to its ability to integrate complementary information from different sensors, has become a core means to improve the intelligence level of inspection. However, existing technologies have the following key shortcomings in substation applications: Poor environmental adaptability: unable to meet the reliability requirements of inspection in complex environments; insufficient spatiotemporal synchronization accuracy: difficult to achieve accurate correlation between geometric position and semantic information; limited fusion depth and scenario coverage: unable to take into account both personnel violation identification and construction risk prediction, making it difficult to support the full-process supervision of renovation and expansion projects; insufficient robustness of the analysis model: unable to meet the real-time alarm requirements.

[0003] Therefore, there is an urgent need for a multimodal data fusion and analysis method and system that is adapted to the special operating conditions of substations, has high-precision spatiotemporal synchronization, deep integration across the entire link, and full-scenario analysis capabilities, in order to overcome the shortcomings of existing technologies. Summary of the Invention

[0004] This invention aims to address the core problems existing in current substation intelligent inspection multimodal data fusion technologies, such as poor environmental adaptability, low spatiotemporal synchronization accuracy, insufficient fusion depth, incomplete scene coverage, and weak analysis robustness. This invention provides a method and system for multimodal data fusion and analysis for substation intelligent inspection, in order to solve the aforementioned technical defects.

[0005] In a first aspect, this invention proposes a method for multimodal data fusion and analysis for intelligent substation inspection, which includes the following steps: S1. Multimodal Data Acquisition: Based on a quadrupedal inspection robot equipped with a multi-sensor system, LiDAR point cloud data is collected in substation inspection scenarios. Dual-light camera image data Beidou positioning data Temperature collected by environmental sensors ,humidity and electromagnetic interference intensity The dual-light camera image data Including visible light images With infrared thermal imaging ; S2. Multimodal data preprocessing: This includes denoising enhancement and spatiotemporal calibration. The denoising enhancement uses an electromagnetic interference adaptive denoising algorithm to process the lidar point cloud for denoising, and also applies this to the infrared thermal imaging. A temperature-weighted enhancement algorithm is used to enhance pixel values; the spatiotemporal calibration employs a federated Kalman filter (FKF) to achieve multi-sensor time synchronization, and a QR code-point cloud joint calibration method is used to achieve spatial calibration, establishing a transformation matrix from the sensor coordinate system to the robot body coordinate system. ; S3. Multimodal Data Layered Fusion: Construct a three-level fusion architecture of 'physical layer - feature layer - decision layer', including: physical layer fusion, feature layer fusion and decision layer fusion; S4. Deep Analysis of Fuded Data: Based on the output fused data, three types of analysis models are constructed, including: an equipment defect detection model using a temporal attention CNN-LSTM network; a personnel behavior monitoring model using fused posture keypoint and distance data; and a construction risk prediction model using an energy consumption-risk coupling model; and S5. Decision Output: Generate three types of decision results from the analysis results, including: real-time alarms, 3D SLAM maps, and automatic inspection reports.

[0006] Preferably, the purity of the point cloud after denoising in step S2 is... satisfy: ,in, , To be affected by electromagnetic interference The number of varying noise points, The total number of point clouds; the infrared thermal imaging A temperature-weighted enhancement algorithm is used to enhance pixel values. for: , The average operating temperature of the substation equipment is taken as 35℃.

[0007] Preferably, the time synchronization error in step S2 The Federal Kalman Filter (FKF) time update equation is: ,in , For electromagnetic interference adaptive time synchronization gain, This is the sensor's original timestamp. Used as the BeiDou reference time stamp; Transformation matrix for: The coordinate transformation formula is: ,in To calibrate the residuals, , The rotation matrix is ​​calculated from the QR code's attitude angle; The translation vector is determined by minimizing the point cloud matching error.

[0008] Preferably, the physical layer fusion in step S3 is spatiotemporal unification, specifically including: generating a unified state vector of the robot's 'position-attitude-velocity' based on multi-source trust-weighted EKF fusion motion state data. The EKF state update formula is: ; in, , For location, For attitude angle, For speed, Here is the state transition matrix. To control the input matrix, For sensor measurement vectors, For sensor trust weights, LiDAR Beidou , For the noise matrix, It is Gaussian noise. , Let be the noise covariance matrix.

[0009] More preferably, the feature layer fusion is a geometric-semantic association, including constructing a dynamic attention feature fusion model to fuse geometric features and semantic features: Geometric feature extraction: The LiDAR point cloud is voxelized, and the surface plane of the device is fitted by an improved RANSAC algorithm to output geometric feature vectors. , S For the area of ​​a plane, L For the perimeter, For surface roughness, The height deviation from the standard plane, where the roughness , Point in the plane z coordinate, For average z coordinate; Semantic feature extraction: A dual-branch ViT-Base encoder is used to process dual-light images, and the visible light branch outputs semantic labels for devices / personnel. , For equipment type, For personnel status, infrared branch output temperature characteristics , The highest temperature, The lowest temperature, For temperature variance; Dynamic attention fusion: through attention weights Assign geometric and semantic feature contributions and fuse features for: ; in, The stronger the electromagnetic interference and the closer the temperature is to the abnormal threshold, the higher the weight of the geometric feature.

[0010] More preferably, the decision layer is fused into semantic-rule reasoning, specifically including: introducing a large-scale visual language model for substations (Substation-VLM), combined with an operation and maintenance rule base. To achieve semantic reasoning and reasoning confidence. The corrected formula is: ,in The original confidence level of the model. This is an indicator function; it returns 1 if the condition is met, and 0 otherwise; it outputs a structured scene description. , For device status, For personnel behavior, It is classified as a risk level.

[0011] Preferably, the equipment defect detection model in step S4 employs a temporal attention CNN-LSTM network, with input fused features. time series The defect determination formula is: ,in, , This is the time-series decay factor, with a value of 5. An adaptive defect threshold for electromagnetic interference is used, resulting in a defect identification accuracy of ≥96%. Among them, the fusion features of the last 3 frames Assigned attention weights .

[0012] Preferably, the formula for determining violations in the personnel behavior monitoring model described in step S4 is: ,in The distance between personnel and energized equipment is calculated using lidar. The tilt angle of the person's posture is extracted by OpenPose; For personnel height, the response time to violations should be ≤0.8s; Risk levels in the construction risk prediction model The calculation is as follows: ; in, , where is the path overlap rate; , representing the robot's energy consumption deviation rate. Actual energy consumption This represents the average energy consumption. The environmental risk coefficient has a risk prediction accuracy rate of ≥92%.

[0013] Further preferably, the visible light branch and infrared branch of the dual-branch ViT-Base encoder share the underlying weights, and the infrared branch introduces a temperature attention mechanism to address temperatures higher than [the specified temperature]. Higher attention weights are assigned to regions .

[0014] Secondly, embodiments of the present invention also provide a system for multimodal data fusion and analysis for intelligent substation inspection to implement the above-mentioned method, comprising: The data acquisition module is configured to collect lidar point cloud data in substation inspection scenarios using a quadruped inspection robot equipped with a multi-sensor system. Dual-light camera image data Beidou positioning data Temperature collected by environmental sensors ,humidity and electromagnetic interference intensity The dual-light camera image data Including visible light images With infrared thermal imaging ; The data preprocessing module includes a denoising and enhancement unit and a spatiotemporal calibration unit. The denoising and enhancement unit uses an electromagnetic interference adaptive denoising algorithm to process the lidar point cloud for denoising, and performs denoising on the infrared thermal image. A temperature-weighted enhancement algorithm is used to enhance pixel values; the spatiotemporal calibration unit uses a federated Kalman filter (FKF) to achieve multi-sensor time synchronization, and a QR code-point cloud joint calibration method to achieve spatial calibration, establishing a transformation matrix from the sensor coordinate system to the robot body coordinate system. ; The data layering module is configured to build a three-level fusion architecture of 'physical layer-feature layer-decision layer', including: physical layer fusion unit, feature layer fusion unit and decision layer fusion unit; The fusion data deep analysis module is configured to construct three types of analysis models based on the output fusion data, including: an equipment defect detection model using a temporal attention CNN-LSTM network; a personnel behavior monitoring model using fused posture keypoint and distance data; and a construction risk prediction model using an energy consumption-risk coupling model; and... The decision output module is configured to generate three types of decision results from the analysis results, including: real-time alarms, 3D SLAM maps, and automatic inspection reports.

[0015] Compared with the prior art, the beneficial results of the present invention are as follows: (1) Significantly improved environmental adaptability: Through customized algorithms such as electromagnetic interference adaptive denoising and temperature weighted enhancement, the point cloud purity is ≥92% and the semantic recognition misjudgment rate is ≤5% in electromagnetic interference environments of -20℃~50℃ and 40dB~120dB, adapting to the special working conditions of substations.

[0016] (2) Spatiotemporal synchronization accuracy doubled: The Federal Kalman filter and QR code-point cloud joint calibration are adopted, with time deviation ≤8ms and spatial error ≤3cm, which improves the accuracy by more than 50% compared with the existing technology, laying the foundation for deep integration.

[0017] (3) Breakthrough in integration depth and scenario coverage: The three-level integration architecture realizes the full-link association of "physical-feature-decision". The Substation-VLM large model, combined with operation and maintenance rules, covers the full scenario of "equipment-personnel-construction", which improves the scenario coverage rate by 67% compared with the existing technology.

[0018] (4) The analysis performance has been greatly improved: the accuracy of equipment defect identification is ≥96%, the response time of personnel violations is ≤0.8s, and the accuracy of construction risk prediction is ≥92%, which is 10%~20% higher than the existing technology.

[0019] (5) Strong decision-making practicality: The closed-loop output of three-dimensional visualization map + priority ranking report + real-time alarm improves the efficiency of operation and maintenance by more than 70% and reduces labor costs and safety risks. Attached Figure Description

[0020] The accompanying drawings are included to provide a further understanding of the embodiments and are incorporated in and constitute a part of this specification. The drawings illustrate embodiments and, together with the description, serve to explain the principles of the invention. Other embodiments and many anticipated advantages of the embodiments will be readily recognized as they become better understood through reference to the following detailed description. Elements in the drawings are not necessarily to scale. The same reference numerals refer to corresponding similar parts.

[0021] Figure 1 This is a flowchart illustrating the method for multimodal data fusion and analysis for intelligent substation inspection according to an embodiment of the present invention. Figure 2 This is a schematic diagram of the architecture of a system for multimodal data fusion and analysis for intelligent substation inspection, according to an embodiment of the present invention. Detailed Implementation

[0022] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings.

[0023] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0024] This invention aims to solve the core problems of existing intelligent substation inspection multimodal data fusion technologies, namely, poor environmental adaptability, low spatiotemporal synchronization accuracy, insufficient fusion depth, incomplete scene coverage, and weak analysis robustness. Specifically, it includes: It enables reliable acquisition and denoising of multimodal data under special operating conditions such as strong electromagnetic interference and high and low temperatures in substations; improves the spatiotemporal synchronization accuracy of multi-sensor data, ensuring time deviation ≤8ms and spatial error ≤3cm; constructs a full-link deep fusion architecture to achieve deep correlation between geometry, semantics, and rules; covers full-scenario analysis of equipment defects, personnel behavior, and construction risks, improving recognition accuracy and real-time performance; and outputs accurate and practical decision-making results to support efficient operation and maintenance.

[0025] In a first aspect, embodiments of the present invention disclose a method for multimodal data fusion and analysis for intelligent substation inspection, such as... Figure 1 As shown, the method includes the following steps: S1. Multimodal Data Acquisition: Based on a quadrupedal inspection robot equipped with a multi-sensor system, LiDAR point cloud data is collected in substation inspection scenarios. Dual-light camera image data Beidou positioning data Temperature collected by environmental sensors ,humidity and electromagnetic interference intensity The dual-light camera image data Including visible light images With infrared thermal imaging ; S2. Multimodal data preprocessing: This includes denoising enhancement and spatiotemporal calibration. The denoising enhancement uses an electromagnetic interference adaptive denoising algorithm to process the lidar point cloud for denoising, and also applies this to the infrared thermal imaging. A temperature-weighted enhancement algorithm is used to enhance pixel values; the spatiotemporal calibration employs a federated Kalman filter (FKF) to achieve multi-sensor time synchronization, and a QR code-point cloud joint calibration method is used to achieve spatial calibration, establishing a transformation matrix from the sensor coordinate system to the robot body coordinate system. ; S3. Multimodal Data Layered Fusion: Construct a three-level fusion architecture of 'physical layer - feature layer - decision layer', including: physical layer fusion, feature layer fusion and decision layer fusion; S4. Deep Analysis of Fuded Data: Based on the output fused data, three types of analysis models are constructed, including: an equipment defect detection model using a temporal attention CNN-LSTM network; a personnel behavior monitoring model using fused posture keypoint and distance data; and a construction risk prediction model using an energy consumption-risk coupling model; and S5. Decision Output: Generate three types of decision results from the analysis results, including: real-time alarms, 3D SLAM maps, and automatic inspection reports.

[0026] Specifically, the method includes the following steps: S10: Multimodal data acquisition Based on a quadruped inspection robot equipped with a multi-sensor system, various types of data are collected in substation scenarios, including: LiDAR point cloud data The system acquires equipment outline, obstacle location, and terrain information through front and rear dual solid-state LiDARs with a resolution of 0.04m and a scanning frequency of 25Hz. Dual-light camera image data Images containing visible light (Resolution 1920×1080@30fps) and infrared thermal imaging (640×512 resolution @ 20fps), used for semantic recognition and temperature anomaly detection respectively; IMU attitude data Sampling frequency 250Hz, robot roll angle Pitch angle Yaw angle and acceleration information; BeiDou positioning data Positioning accuracy is 0.8m, and the robot's global position and reference timestamp are collected. Equipment status data: including concrete strength data collected collaboratively by the robotic arm and rebound hammer. (Detection accuracy ±1), meter readings collected by a ULP low-power camera. (Error ≤ 0.5%) Environmental parameter data: temperature (-20℃~50℃, accuracy ±0.5℃), humidity (20% RH~90% RH), Electromagnetic Interference Intensity (40dB~120dB, shielding effectiveness ≥40dB).

[0027] S20: Multimodal data preprocessing It includes two sub-steps: noise reduction and enhancement, and spatiotemporal calibration. 1) Noise reduction and enhancement: 1. LiDAR point cloud denoising: Employing an adaptive electromagnetic interference denoising algorithm, the purity of the point cloud after denoising is improved. satisfy: ,in, , To be affected by electromagnetic interference The number of varying noise points, This represents the total number of point clouds.

[0028] when hour, ; 2. Infrared Image Enhancement: A temperature-weighted enhancement algorithm is used to enhance pixel values. for: , The average operating temperature of the substation equipment is taken as 35℃.

[0029] 35℃ is the average operating temperature of the equipment, which improves the identification of abnormal temperature areas.

[0030] 2) Spatiotemporal calibration: 1. Time synchronization: Federated Kalman filtering (FKF) is used, and the time update equation is:

[0031] in, , For electromagnetic interference adaptive time synchronization gain, This is the sensor's original timestamp. BeiDou reference timestamp; time synchronization error .

[0032] 2. Spatial Calibration: A QR code-point cloud joint calibration method was adopted. The calibration board consisted of a 3×3 array of QR codes (5cm spacing). A transformation matrix from the sensor coordinate system to the robot body coordinate system was established. The coordinate transformation formula is: ,in To calibrate the residuals, , The rotation matrix is ​​calculated from the QR code's attitude angle; The translation vector is determined by minimizing the point cloud matching error.

[0033] S30: Multimodal Data Layered Fusion Construct a three-tiered fusion architecture: "Physical Layer - Feature Layer - Decision Layer". S301: Physical Layer Fusion (Spatiotemporal Unification): Based on multi-source trust-weighted EKF fusion of motion state data, a unified state vector of the robot's 'position-attitude-velocity' is generated. The EKF state update formula is: ; in, , For location, For attitude angle, For speed, The state transition matrix ( dimension), To control the input matrix ( dimension), For sensor measurement vectors, For sensor trust weights, LiDAR Beidou , For the noise matrix, It is Gaussian noise. , Here is the noise covariance matrix. The fused positioning accuracy is ≤4cm.

[0034] S302: Feature Layer Fusion (Geometric-Semantic Association): Constructing a dynamic attention feature fusion model that integrates geometric and semantic features. 1. Geometric feature extraction: Voxelization of the LiDAR point cloud (resolution) By improving the RANSAC algorithm to fit the surface plane of the device, geometric feature vectors are output. , S For the area of ​​a plane, L For the perimeter, For surface roughness, The height deviation from the standard plane, where the roughness , Point in the plane z coordinate, For average z coordinate; 2. Semantic Feature Extraction: A dual-branch ViT-Base encoder is used to process dual-light images, with the visible light branch outputting semantic labels for devices / personnel. , For equipment type, For personnel status, infrared branch output temperature characteristics , The highest temperature, The lowest temperature, For temperature variance; 3. Dynamic attention fusion: through attention weights Assign geometric and semantic feature contributions and fuse features for: ; in, The stronger the electromagnetic interference and the closer the temperature is to the anomaly threshold, the higher the weight of the geometric feature. This adapts to the reliability of features under different environments.

[0035] In one specific embodiment, the visible light branch and infrared branch of the dual-branch ViT-Base encoder share the underlying weights, and the infrared branch introduces a temperature attention mechanism to address temperatures higher than [the specified temperature]. Higher attention weights are assigned to regions .

[0036] S303: Decision-making layer fusion (semantic-rule reasoning): Introducing the Substation Visual Language Model (Substation-VLM) and combining it with the operation and maintenance rule base. To achieve semantic reasoning and reasoning confidence. The corrected formula is: ,in, The original confidence level of the model. This is an indicator function; it returns 1 if the condition is met, and 0 otherwise; it outputs a structured scene description. , For device status, For personnel behavior, It is classified as a risk level.

[0037] In one specific embodiment, the pre-training data includes 100,000 images of substation equipment and 2,000 operation and maintenance rules, resulting in an operation and maintenance rule base. .

[0038] S40: Deep Data Fusion Analysis Construct three types of enhanced analysis models: Equipment defect detection model: Temporal attention CNN-LSTM network, input fused features time series The defect determination formula is: ,in, , This is the time-series decay factor, with a value of 5. An adaptive defect threshold for electromagnetic interference is used, resulting in a defect identification accuracy of ≥96%. In one specific embodiment, the fusion features of the last 3 frames are... Assigned attention weights .

[0039] Personnel behavior monitoring model: Integrating posture key points and distance data, the violation judgment formula is as follows: ,in The distance between personnel and energized equipment is calculated using lidar. The tilt angle of the person's posture is extracted by OpenPose; For personnel height, the response time to violations should be ≤0.8s; Construction risk prediction model: Energy consumption-risk coupling model, risk level The calculation is as follows: ; in, , where is the path overlap rate; , representing the robot's energy consumption deviation rate. Actual energy consumption This represents the average energy consumption. The environmental risk coefficient has a risk prediction accuracy rate of ≥92%.

[0040] S50: Decision Output Three types of decision outcomes are generated: Real-time alarms: audible and visual alarms (≥85dB), delay ≤2.5s, distinguishing between three types of alarms: equipment defects, personnel violations, and construction risks; 3D Dynamic SLAM Map: Integrates point clouds and semantic tags, supports web access (http: / / 192.168.0.2). Accessed via / ), map update frequency 1Hz, defect location error ≤4cm; Structured inspection report: generated within 4 minutes, supports PDF / Excel export, includes defect type, BeiDou coordinates, risk level, handling suggestions, and priority ranking. .

[0041] Further reference Figure 2 As an implementation of the methods shown in the above figures, this application provides an embodiment of a system for multimodal data fusion and analysis for intelligent substation inspection. This system embodiment is similar to... Figure 1 Corresponding to the method embodiments shown, the system can be specifically applied to various electronic devices.

[0042] In a second aspect, embodiments of the present invention also disclose a system for implementing the method described in any of the first aspects for intelligent substation inspection, comprising: a data acquisition module 21, a data preprocessing module 22, a data layering module 23, a fused data deep analysis module 24, and a decision output module 25.

[0043] In one specific embodiment, the data acquisition module 21 is configured to collect lidar point cloud data in a substation inspection scenario using a quadruped inspection robot equipped with a multi-sensor system. Dual-light camera image data Beidou positioning data Temperature collected by environmental sensors ,humidity and electromagnetic interference intensity The dual-light camera image data Including visible light images With infrared thermal imaging ; The data preprocessing module 22 includes a denoising and enhancement unit and a spatiotemporal calibration unit. The denoising and enhancement unit uses an electromagnetic interference adaptive denoising algorithm to process the lidar point cloud for denoising, and performs denoising on the infrared thermal image. A temperature-weighted enhancement algorithm is used to enhance pixel values; the spatiotemporal calibration unit uses a federated Kalman filter (FKF) to achieve multi-sensor time synchronization, and a QR code-point cloud joint calibration method to achieve spatial calibration, establishing a transformation matrix from the sensor coordinate system to the robot body coordinate system. ; Data layering module 23 is configured to construct a three-level fusion architecture of 'physical layer-feature layer-decision layer', including: physical layer fusion unit, feature layer fusion unit, and decision layer fusion unit; fused data deep analysis module 24 is configured to construct three types of analysis models based on the output fused data, including: a device defect detection model using a temporal attention CNN-LSTM network; a personnel behavior monitoring model using fused posture keypoint and distance data; and a construction risk prediction model using an energy consumption-risk coupling model; and The decision output module 25 is configured to generate three types of decision results from the analysis results, including: real-time alarms, 3D SLAM maps, and automatic inspection reports.

[0044] The functions and methods of the above modules correspond to each other, and will not be repeated here.

[0045] The following is an example of a 500kV substation renovation and expansion project. The substation includes 2 transformers, 8 circuit breakers, and 32 sets of insulators. The construction area is divided into an equipment installation area, a cable laying area, and a high-altitude operation area. The quadruped inspection robot performs inspection tasks along a preset route (covering all equipment and construction areas), with an inspection radius of 500m and a single inspection time of 2 hours.

[0046] (I) Equipment parameters and experimental data 1. Data acquisition module parameters: LiDAR: Resolution 0.04m, scanning frequency 25Hz, 360° scanning range; Dual-light camera: Visible light 1920×1080@30fps, infrared 640×512@20fps, 25x optical zoom; IMU: Sampling frequency 250Hz, attitude accuracy 0.1°; Environmental sensors: temperature measurement error ±0.3℃, electromagnetic interference detection range 40dB~120dB.

[0047] 2. Pretreatment effect: When the electromagnetic interference E=100dB, the noise removal rate of the lidar point cloud is 93%, which is 23% higher than that of traditional median filtering; After infrared image enhancement, the contrast of abnormal temperature areas (≥45℃) is improved by 40%, which facilitates defect identification; Spatiotemporal synchronization error: time deviation 7.2ms, spatial error 2.8cm, meeting the accuracy requirements.

[0048] 3. Fusion and Analysis Results: After physical layer fusion, the robot's positioning accuracy is 3.5cm, and it operates stably in scenarios involving climbing a 30° slope and crossing stairs (18cm in height). After feature layer fusion, the geometric-semantic association accuracy reached 94%, an improvement of 18% compared to single feature recognition. Equipment defect detection: Detected bulges in the transformer casing (6mm deformation) and two instances of insulator damage, with an accuracy rate of 97%. Personnel behavior monitoring: Identifies one violation of "approaching live electrical equipment (within 3.2m) without wearing a safety helmet," with a response time of 0.6s; Construction risk prediction: The risk level of path overlap between the equipment installation area and the cable laying area is predicted to be "high", which is consistent with the actual working conditions.

[0049] 4. Decision output effect: Real-time alarms are delayed by 2.3 seconds, allowing maintenance personnel to promptly address violations. The 3D SLAM map clearly marks the location of defects, and the web access latency is ≤1 second; The automatic inspection report is generated in 3.8 minutes, and the defect priority ranking matches the on-site handling needs.

[0050] (II) Verification of Implementation Steps 1. Data Acquisition: The robot starts from the autonomous charging station and collects multimodal data along the preset route. The ULP camera collects meter readings every 10m, and the robotic arm collects strength data every 5m in the concrete equipment area. 2. Preprocessing: The collected data is denoised and spatiotemporally calibrated. After adaptive denoising, the effective point ratio of the lidar point cloud is 93%. 3. Layered Fusion: The physical layer fuses to generate the robot's state vector, the feature layer fuses to output associated features, and the decision layer fuses to infer "transformer temperature is normal, personnel violated regulations once, and there is one construction risk"; 4. In-depth analysis: The personnel violation was confirmed as "approaching live equipment without wearing a safety helmet," and the construction risk was "path overlap and collision hazard." 5. Decision output: Trigger audible and visual alarms, generate 3D maps and inspection reports, and maintenance personnel handle violations and risks according to the report priority.

[0051] The embodiments of this invention are adapted to special operating conditions of substations, have high spatiotemporal synchronization accuracy, cover the entire scenario of "equipment-personnel-construction", and significantly improve the accuracy of analysis and the practicality of decision-making, which can effectively support intelligent inspection and operation and maintenance decision-making of substations.

[0052] The above description is merely a preferred embodiment of the present invention and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention is not limited to the specific combination of the above-described technical features, but also includes other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in this invention.

Claims

1. A method for multimodal data fusion and analysis for intelligent substation inspection, characterized in that, The method includes the following steps: S1. Multimodal Data Acquisition: Based on a quadrupedal inspection robot equipped with a multi-sensor system, LiDAR point cloud data is collected in substation inspection scenarios. Dual-light camera image data Beidou positioning data Temperature collected by environmental sensors ,humidity and electromagnetic interference intensity The dual-light camera image data Including visible light images With infrared thermal imaging ; S2. Multimodal data preprocessing: This includes denoising enhancement and spatiotemporal calibration. The denoising enhancement uses an electromagnetic interference adaptive denoising algorithm to process the lidar point cloud for denoising, and also applies this to the infrared thermal imaging. A temperature-weighted enhancement algorithm is used to enhance pixel values; the spatiotemporal calibration employs a federated Kalman filter (FKF) to achieve multi-sensor time synchronization, and a QR code-point cloud joint calibration method is used to achieve spatial calibration, establishing a transformation matrix from the sensor coordinate system to the robot body coordinate system. ; S3. Multimodal Data Layered Fusion: Construct a three-level fusion architecture of 'physical layer - feature layer - decision layer', including: physical layer fusion, feature layer fusion and decision layer fusion; S4. Deep Analysis of Fuded Data: Based on the output fused data, three types of analysis models are constructed, including: an equipment defect detection model using a temporal attention CNN-LSTM network; a personnel behavior monitoring model using fused posture keypoint and distance data; and a construction risk prediction model using an energy consumption-risk coupling model; and S5. Decision Output: Generate three types of decision results from the analysis results, including: real-time alarms, 3D SLAM maps, and automatic inspection reports.

2. The method for multimodal data fusion and analysis according to claim 1, characterized in that, Point cloud purity after denoising in step S2 satisfy: ,in, , To be affected by electromagnetic interference The number of varying noise points The total number of point clouds; the infrared thermal imaging A temperature-weighted enhancement algorithm is used to enhance pixel values. for: , The average operating temperature of the substation equipment is taken as 35℃.

3. The method for multimodal data fusion and analysis according to claim 1, characterized in that, Time synchronization error in step S2 The Federal Kalman Filter (FKF) time update equation is: ,in , For electromagnetic interference adaptive time synchronization gain, This is the sensor's original timestamp. Used as the BeiDou reference time stamp; Transformation matrix for: The coordinate transformation formula is: ,in To calibrate the residuals, , The rotation matrix is ​​calculated from the QR code's attitude angle; The translation vector is determined by minimizing the point cloud matching error.

4. The method for multimodal data fusion and analysis according to claim 1, characterized in that, The physical layer fusion described in step S3 is a spatiotemporal unification, specifically including: generating a unified state vector of the robot's 'position-attitude-velocity' based on multi-source trust-weighted EKF fusion motion state data. The EKF state update formula is: ; in, , For location, For attitude angle, For speed, Here is the state transition matrix. To control the input matrix, For sensor measurement vectors, For sensor trust weights, LiDAR Beidou , For the noise matrix, It is Gaussian noise. , Let be the noise covariance matrix.

5. The method for multimodal data fusion and analysis according to claim 4, characterized in that, The feature layer is fused into a geometric-semantic association, including constructing a dynamic attention feature fusion model that fuses geometric and semantic features: Geometric feature extraction: The LiDAR point cloud is voxelized, and the surface plane of the device is fitted by an improved RANSAC algorithm to output geometric feature vectors. S is the area of ​​the plane, and L is the perimeter. For surface roughness, The height deviation from the standard plane, where the roughness , Let z be the z-coordinate of a point in the plane. The average z-coordinate; Semantic feature extraction: A dual-branch ViT-Base encoder is used to process dual-light images, and the visible light branch outputs semantic labels for devices / personnel. , For equipment type, For personnel status, infrared branch output temperature characteristics , The highest temperature, The lowest temperature, For temperature variance; Dynamic attention fusion: through attention weights Assigning geometric and semantic feature contributions and fusing features for: ; in, The stronger the electromagnetic interference and the closer the temperature is to the abnormal threshold, the higher the weight of the geometric feature.

6. The method for multimodal data fusion and analysis according to claim 5, characterized in that, The decision-making layer is integrated into semantic-rule reasoning, specifically including: introducing a large-scale visual language model for substations (Substation-VLM), combined with an operation and maintenance rule base. To achieve semantic reasoning and reasoning confidence. The corrected formula is: ,in The original confidence level of the model. This is an indicator function; it returns 1 if the condition is met, and 0 otherwise; it outputs a structured scene description. , For device status, For personnel behavior, Risk level.

7. The method for multimodal data fusion and analysis according to claim 1, characterized in that, The equipment defect detection model described in step S4 employs a temporal attention CNN-LSTM network, with input fused features. time series The defect determination formula is: ,in, , This is the time-series decay factor, with a value of 5. An adaptive defect threshold for electromagnetic interference is used, resulting in a defect identification accuracy of ≥96%. Among them, the fusion features of the last 3 frames Assigned attention weights .

8. The method for multimodal data fusion and analysis according to claim 1, characterized in that, The formula for determining violations in the personnel behavior monitoring model described in step S4 is as follows: ,in The distance between personnel and live equipment is calculated by lidar. The tilt angle of the person's posture is extracted by OpenPose; For personnel height, the response time to violations should be ≤0.8s; Risk levels in the construction risk prediction model The calculation is as follows: ; in, , where is the path overlap rate; , representing the robot's energy consumption deviation rate. Actual energy consumption This represents the average energy consumption. The environmental risk coefficient has a risk prediction accuracy rate of ≥92%.

9. The method for multimodal data fusion and analysis according to claim 5, characterized in that, The dual-branch ViT-Base encoder shares underlying weights between its visible light branch and infrared branch, and the infrared branch introduces a temperature attention mechanism to address temperatures higher than [the specified temperature range]. Higher attention weights are assigned to regions .

10. A system for implementing the method described in any one of claims 1-9 for multimodal data fusion and analysis of intelligent substation inspection, characterized in that, include: The data acquisition module is configured to collect lidar point cloud data in substation inspection scenarios using a quadruped inspection robot equipped with a multi-sensor system. Dual-light camera image data Beidou positioning data Temperature collected by environmental sensors ,humidity and electromagnetic interference intensity The dual-light camera image data Including visible light images With infrared thermal imaging ; The data preprocessing module includes a denoising and enhancement unit and a spatiotemporal calibration unit. The denoising and enhancement unit uses an electromagnetic interference adaptive denoising algorithm to process the lidar point cloud for denoising, and performs denoising on the infrared thermal image. A temperature-weighted enhancement algorithm is used to enhance pixel values; the spatiotemporal calibration unit uses a federated Kalman filter (FKF) to achieve multi-sensor time synchronization, and a QR code-point cloud joint calibration method to achieve spatial calibration, establishing a transformation matrix from the sensor coordinate system to the robot body coordinate system. ; The data layering module is configured to build a three-level fusion architecture of 'physical layer-feature layer-decision layer', including: physical layer fusion unit, feature layer fusion unit and decision layer fusion unit; The fusion data deep analysis module is configured to construct three types of analysis models based on the output fusion data, including: an equipment defect detection model using a temporal attention CNN-LSTM network; a personnel behavior monitoring model using fused posture keypoint and distance data; and a construction risk prediction model using an energy consumption-risk coupling model; and... The decision output module is configured to generate three types of decision results from the analysis results, including: real-time alarms, 3D SLAM maps, and automatic inspection reports.