Visual substation intelligent inspection system based on multi-modal large language model
By fusing image, temperature, vibration, and text data through a multimodal large language model, a dynamic knowledge base is constructed to achieve intelligent substation inspection, solving the problems of low efficiency and insufficient intelligence in traditional inspections, and improving inspection efficiency and accuracy.
Patent Information
- Application Number
- CN202510272703.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-11-07
AI Technical Summary
Traditional substation inspections are inefficient, have a high rate of missed inspections, and lack sufficient intelligence. Existing deep learning-based solutions lack multimodal data fusion, have lagging knowledge updates, and insufficient interactive capabilities, making them difficult to adapt to inspection needs in complex scenarios.
A multimodal large language model (LLM) is used to fuse image, temperature, vibration and text data to build a dynamic knowledge base that supports natural language interaction, generates adaptive diagnostic rules, and provides interpretable fault reports and maintenance suggestions through augmented reality (AR).
It significantly improves inspection efficiency by more than 40%, reduces the missed detection rate to below 1%, supports rapid diagnosis of more than 95% of new types of faults, reduces manual intervention, and shortens the decision-making time of operation and maintenance personnel by 50%.
Smart Images

Figure CN120912166A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of power system intelligence, and specifically relates to a substation intelligent inspection system and method fusing a multimodal large language model (Multimodal LLM), three-dimensional visualization and deep learning technology, realizing multi-dimensional perception of equipment state, dynamic knowledge reasoning and adaptive decision support, and improving the intelligent level of substation inspection. BACKGROUND
[0002] Traditional substation inspection relies on manual or single video monitoring systems, and has problems such as low efficiency, high omission rate and insufficient intelligence. Although the existing deep learning-based scheme can realize image recognition, it has the following three defects: first, weak data fusion capability: image, temperature, vibration and other multi-modal data are analyzed independently, and lack cross-modal correlation; second, knowledge update lag: fault diagnosis relies on static models, which are difficult to adapt to new equipment or complex scenarios; third, insufficient interaction capability: lack of natural language interaction and dynamic explanation capability, making it difficult for operation and maintenance personnel to quickly understand system decisions.
[0003] In recent years, with the development of artificial intelligence technology, large language models (LLM) based on natural language processing, multi-modal reasoning and knowledge generalization have shown significant advantages and have been gradually applied to the field of power equipment detection. However, the existing technology still has the following shortcomings: lack of full use of three-dimensional spatial information of substations; unable to realize real-time analysis and feedback of inspection data; limited intelligence of the system, making it difficult to meet the inspection needs in complex scenarios.
[0004] Therefore, there is an urgent need for a visual substation intelligent inspection method that fuses a multimodal large language model (LLM) to improve inspection efficiency and accuracy. SUMMARY
[0005] The application proposes a substation intelligent inspection system that fuses a multimodal large language model (LLM), which solves the above problems through the following innovations:
[0006] Multimodal data fusion: combining the cross-modal understanding capability of LLM, realizing joint analysis of image, temperature, vibration and text data;
[0007] Dynamic knowledge base construction: using LLM to continuously learn power equipment knowledge and historical fault cases, generating adaptive diagnostic rules;
[0008] Natural language interaction and decision explanation: supporting voice / text interaction, and generating explainable fault reports and maintenance recommendations based on LLM.
[0009] TECHNICAL SCHEME
[0010] The system includes the following core modules:
[0011] 1. Multi-modal data acquisition module
[0012] Sensor cluster (high-definition camera, infrared thermal imager, vibration sensor, voiceprint collector) collects device image, temperature, vibration and noise data in real time;
[0013] The data preprocessing unit completes denoising, spatio-temporal alignment and multi-modal data fusion to generate structured input.
[0014] 2. Three-dimensional dynamic modeling module
[0015] Based on laser point cloud and RGB-D camera data, a high-precision three-dimensional semantic model of the substation is constructed;
[0016] Fusion of LLM semantic understanding ability, automatic labeling of device attributes (such as model, rated parameters, historical maintenance records).
[0017] 3. Multi-modal large language model (LLM) engine
[0018] Cross-modal feature extraction: Align image and text features using visual-linguistic pre-training models (such as CLIP); Dynamic knowledge base: integrate power equipment manuals, international standards (such as IEC 61850) and historical fault data, and generate device diagnosis knowledge graph through LLM;
[0019] Adaptive inference formula-based dynamic adjustment of fault probability:
[0020] P fault =α×f LLM (Q text ,D sensor )+β×P CNN (I)+γ×P LSTM (T,V)
[0021] Where f LLM is the semantic inference probability output by LLM, Q text is the natural language query, α, β, γ are adaptive weight coefficients, respectively controlling the contribution of LLM inference, CNN image classification, LSTM time series analysis, 0-1, and satisfy α+β+γ=1; P fault is the comprehensive fault probability of the device, reflecting the possibility of the current device failure, 0-1, the larger the value, the higher the risk of failure; Q text Natural language query (such as "check the status of insulator A of transformer A"); D sensor Textual description of sensor data; P CNN (I) is the fault probability output by the image classification model based on convolutional neural network (CNN), I is the device image data (such as infrared thermal image, visible light image); P LSTM(T, V) is the failure probability output by the time series analysis model based on long short-term memory network (LSTM), T is the device temperature time series data, and V is the device vibration time series data.
[0022] 4. Augmented Reality (AR) Interaction Module
[0023] Overlay real-time device status (such as temperature heat map, vibration spectrum) through AR glasses; integrate LLM voice assistant, support natural language instructions (such as "display the last 3 abnormal records of transformer A").
[0024] 5. Self-explaining Decision Support Module
[0025] Generate readability reports based on LLM, including fault causes (such as "insulator contamination level exceeds threshold, recommend cleaning within 48 hours"); provide maintenance scheme deduction function, simulate the impact of different decisions on device life.
[0026] 6. Communication and Feedback Module
[0027] Realize real-time uploading and cloud storage of inspection data, ensure the safety and traceability of data; support remote monitoring and alarm functions, timely notify relevant personnel to handle abnormal situations.
[0028] Technical Details
[0029] Deep Integration of LLM and Fault Diagnosis
[0030] Step 1: Convert sensor data into natural language description (such as "circuit breaker B phase temperature rises to 85℃, exceeds baseline value by 10%"), input LLM for context analysis;
[0031] Step 2: LLM calls dynamic knowledge base to match similar cases, outputs fault hypotheses and confidence;
[0032] Step 3: Combine CNN image classification results and LLM semantic reasoning, determine the final diagnosis conclusion through weighted decision.
[0033] Dynamic Knowledge Base Update Mechanism
[0034] Based on the incremental learning ability of LLM, automatically analyze the latest maintenance records and industry literature, update the knowledge graph nodes; through reinforcement learning to optimize weight coefficients α, β, γ, improve the generalization ability in complex scenarios.
[0035] Data Acquisition Module: The data acquisition module includes various sensors such as high-definition cameras, infrared imagers, vibration sensors, etc., for collecting multi-dimensional data such as images, temperatures, vibrations, etc. of substation equipment. The sensors are connected to the data acquisition terminal through wireless or wired means, and the data acquisition terminal performs preliminary processing on the collected data, such as data format conversion, data compression, etc., to ensure efficient data transmission.
[0036] Three-dimensional Modeling Module: The three-dimensional modeling module uses three-dimensional laser scanning technology to perform high-precision three-dimensional reconstruction of the substation environment and equipment. Laser scanning equipment (such as 3D scanners) obtains three-dimensional point cloud data of substation equipment by emitting laser beams and measuring the time difference of reflected laser beams. Then, deep learning algorithms are used to process the point cloud data to generate high-precision three-dimensional models. The model can be dynamically updated to reflect the real-time state of the substation equipment.
[0037] Artificial Intelligence Algorithm Module: The artificial intelligence algorithm module uses deep learning technology, convolutional neural networks (CNN), and object detection algorithms (such as YOLO) to identify power equipment failures. The specific steps are as follows:
[0038] Step 1: Data preprocessing: Perform denoising, normalization, and other preprocessing operations on the collected multi-dimensional data such as images, temperatures, vibrations, etc. to improve data quality.
[0039] Step 2: Feature extraction: Extract image features (such as edges, textures), temperature features (such as temperature distribution), and vibration features (such as vibration frequency) from the preprocessed data.
[0040] Step 3: Model training: Train the deep learning model using labeled training data. During the training process, adjust the model parameters to optimize the model performance.
[0041] Step 4: Model classification: Apply the trained model to real-time data to classify equipment states and identify abnormal states (such as insulator contamination, conductor heating, equipment loosening, etc.).
[0042] Step 1: Data Analysis - Comprehensive analysis of the collected multi-dimensional data, including image analysis, temperature analysis, vibration analysis, etc.
[0043] Step 2: Inspection Report Generation - Generate detailed inspection reports based on analysis results, including device status, abnormal points, fault prediction, etc.
[0044] Step 3: Fault Prediction - Use machine learning algorithms to predict future possible faults of the device based on historical data.
[0045] Step 4: Optimization Suggestions - Provide optimization suggestions based on fault prediction results to assist maintenance personnel in developing maintenance plans.
[0046] Communication and Feedback Module - The communication and feedback module enables real-time uploading and cloud storage of inspection data. This module supports remote monitoring and alarm functions, notifying relevant personnel to handle abnormal situations in a timely manner. The specific steps are as follows:
[0047] Step 1: Data Upload - Real-time upload of collected inspection data to cloud servers through wireless networks.
[0048] Step 2: Cloud Storage - Store inspection data in cloud servers to ensure data security and traceability.
[0049] Step 3: Remote Monitoring - Maintenance personnel can view the running status of substation equipment in real time through remote terminals.
[0050] Step 4: Alarm Function - When abnormal situations are detected, the system automatically sends alarm information to relevant personnel, notifying them to handle it in a timely manner.
[0051] Innovation Points
[0052] 1. Multi-modal LLM-driven: First to deeply integrate visual-linguistic large models with power data, realizing cross-modal semantic reasoning;
[0053] 2. Dynamic knowledge evolution: Construct a self-evolving device knowledge base through LLM, supporting zero-sample diagnosis of new faults;
[0054] 3. Human-machine collaborative decision-making: Combine AR visualization and natural language interaction to reduce the technical threshold of maintenance personnel;
[0055] 4. Self-explanatory output: Generate decision basis based on LLM to improve system transparency and user trust. Brief Description of the Drawings
[0056] Figure 1 : System architecture diagram (showing data flow of LLM engine and various modules);
[0057] Figure 2: Multimodal data fusion flowchart (sensor data→LLM semantic mapping→decision output);
[0058] Figure 3 : AR interaction interface schematic diagram (including voice instruction input and three-dimensional fault labeling);
[0059] Figure 4 : Dynamic knowledge base updating mechanism (incremental learning and knowledge graph expansion). DETAILED DESCRIPTION
[0060] The specific embodiment of the present application includes the following steps:
[0061] Step one: deploy a sensor cluster and configure a multimodal data acquisition terminal;
[0062] Step two: train a visual-language large model and build an initial power equipment knowledge graph;
[0063] Step three: develop an AR interaction interface and integrate an LLM voice assistant;
[0064] Step four: optimize LLM weight parameters through historical data to realize online updating of the fault diagnosis model.
[0065] Expected effects
[0066] The implementation of the present application can significantly improve the substation inspection efficiency by more than 40%, reduce the missed detection rate to less than 1%, support the rapid diagnosis of more than 95% of new faults, and reduce manual intervention; through self-explaining reports, the decision-making time of operation and maintenance personnel is shortened by 50%.
[0067] Other explanations
[0068] The embodiments of the present application are not limited to the above description, and any improvements and optimizations based on the present application are within the protection scope of the present application.
Claims
1. A multi-modal large language model-based visual substation intelligent inspection system, characterized in that, Comprising: Multimodal data acquisition module, through the sensor cluster (high-definition camera, infrared thermal imager, vibration sensor, voiceprint collector) real-time acquisition of substation equipment image, temperature, vibration and noise data, and denoising, space-time alignment and multimodal fusion; Three-dimensional dynamic modeling module: based on laser point cloud and RGB-D camera data to construct high-precision three-dimensional semantic model of substation, and fuse LLM semantic understanding ability to automatically label equipment attributes (model, historical maintenance record, etc.); Multimodal large language model (LLM) engine, including: cross-modal feature extraction unit, using visual-linguistic pre-training model (such as CLIP) to align image and text features; Dynamic knowledge base, integrating device manual, international standards and historical fault data, generating diagnosis knowledge graph through LLM, supporting incremental learning and reinforcement learning to optimize weight coefficients (α, β, γ); Self-adaptive reasoning unit based on formula to dynamically adjust fault probability: P fault = a x f LLM (Q text , D sensor ) + b x P CNN (I) + g x P LSTM (T, V) wherein f LLM is the semantic inference probability output by the LLM, Q text is the natural language query, a, b, g are adaptive weight coefficients, respectively controlling the contribution of LLM inference, CNN image classification, LSTM time series analysis, 0-1, and satisfy a+b+g=1; P fault is the device comprehensive failure probability, reflecting the possibility of the current device failure, 0-1, the larger the value, the higher the risk of failure; Q text Natural language query (such as "check the insulator state of transformer A"); D sensor Textual description of sensor data; P CNN (I) The failure probability output by the image classification model based on convolutional neural network (CNN), I is the device image data (such as infrared thermal image, visible light image); P LSTM (T, V) is the failure probability output by the time series analysis model based on long short-term memory network (LSTM), T is the device temperature time series data; V is the device vibration time series data. Augmented reality (AR) interaction module: through AR glasses to superimpose real-time state of equipment (temperature heat map, vibration frequency spectrum curve, fault labeling box), integrate LLM voice assistant to support natural language instructions; Self-explaining decision support module: based on LLM to generate readable fault report (including fault reason, confidence, maintenance time window) and maintenance scheme deduction; Communication and feedback module: for real-time uploading, cloud storage and remote alarm of inspection data.
2. A multi-modal large language model-based visual substation intelligent inspection method, characterized in that Comprising the following steps: Step one: collect multimodal sensor data; Step two: build three-dimensional semantic model and label equipment attributes; Step three: through LLM cross-modal reasoning, CNN image classification and LSTM time series analysis, combine dynamic weight calculation to integrate fault probability; Step four: through AR interface to real-time display equipment state and fault labeling, and support natural language interaction: Step five: generate interpretable fault report and maintenance suggestion, and real-time upload data to cloud.
Citation Information
Cited By
Unmanned aerial vehicle inspection control system and method based on acoustic imaging gas pipeline leakage detection
CN121165788A
LLM-based mobile energy storage system autonomous inspection method and intelligent robot
CN121643237A
Intelligent operation and maintenance auxiliary system for water turbine of hydropower station
CN121810269A
Transformer substation fault interactive reasoning and auxiliary decision-making method and system based on natural language processing, electronic equipment and storage medium
CN121938408A
A device maintenance management system and method based on a low-code platform and AI intelligent glasses
CN122492182A