Substation fault detection method and system based on images and digital twinning
By performing block pixel structure modeling of substation images and dynamic behavior simulation of digital twin models, the problems of environmental interference and insufficient accuracy in substation fault detection are solved, and high-precision fault identification and verification are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANDONG QICHENG ELECTRIC TECH CO LTD
- Filing Date
- 2026-04-21
- Publication Date
- 2026-07-17
Smart Images

Figure CN122415495A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of visual inspection technology for power equipment, and in particular to a method and system for substation fault detection based on images and digital twins. Background Technology
[0002] Currently, substation fault detection mostly relies on comparing visual features from single images and matching fixed historical templates. While some technologies incorporate digital twin models, these are only used for offline visualization of equipment status and are not integrated with real-time image detection processes. Conventional detection methods depend on grayscale differences and single-pixel value comparisons to determine equipment anomalies, without conducting dedicated statistical modeling of pixel structures for segmented images of the equipment, nor building a corresponding pixel structure benchmark model based on the real-time detection period.
[0003] Existing detection methods are susceptible to interference from environmental factors such as lighting and shooting angle, lack specificity in structural feature comparison, and suffer from significant misjudgments of abnormal images in the initial screening stage. Digital twin models cannot achieve spatial coordinate mapping with abnormal images, cannot trigger dynamic 3D model driving based on abnormal images to simulate device physical behavior, and cannot synchronously verify simulated behavior with actual physical behavior. This invention aims to construct a pixel structure benchmark model specific to the device's current detection period and complete the initial screening of anomalies through structural deviation, while simultaneously achieving spatial coordinate mapping between abnormal images and digital twins, dynamic behavior simulation, and cross-validation with actual physical behavior, thus solving the problem of insufficient accuracy in existing detection methods. Summary of the Invention
[0004] The purpose of this invention is to overcome the shortcomings of existing technologies and to propose a substation fault detection method and system based on images and digital twins.
[0005] To achieve the above objectives, the present invention adopts the following technical solution: a substation fault detection method based on image and digital twin, comprising: Collect a set of synchronous images of the substation, and segment out equipment block images containing image regions of specific power equipment. Perform statistical modeling on the pixel structure of the equipment block images to generate a pixel structure benchmark model of the equipment in the current detection period. Extract a new image frame sequence from the substation synchronization image set, and segment the images in the new image frame sequence into block images of the equipment to be inspected according to the equipment type; Calculate the real-time pixel structure features of the block image of the device under inspection, and measure the deviation between the real-time pixel structure features and the pixel structure benchmark model of the corresponding device type. When the measured deviation exceeds the preset structural deviation threshold, the block image of the device under inspection is determined to be an abnormal image in the initial screening. For the power equipment corresponding to the initial screening abnormal image, its digital twin model is activated, the initial screening abnormal image is mapped to the spatial coordinates of the digital twin model, and the three-dimensional model driving is triggered to simulate the dynamic behavior of the power equipment in physical space and generate a simulated behavior sequence. The device behavior state presented by the simulated behavior sequence is cross-validated with the physical behavior state extracted synchronously from the actual video stream; When the cross-validation results do not match, the initial screening abnormal image is confirmed as the final fault characterization image, and all information related to the final fault characterization image is recorded.
[0006] As a further aspect of the present invention, a set of synchronous images of a substation is acquired, and equipment block images containing specific power equipment image regions are segmented, including: Receive a set of synchronous images of the substation acquired simultaneously by multiple fixed-view camera devices during the detection period; The system invokes a pre-built substation equipment knowledge graph to locate equipment in each original image of the substation synchronous image set, segmenting the image into blocks containing specific power equipment regions. Specifically, this includes: The substation equipment knowledge graph includes the preset three-dimensional spatial coordinates of each type of power equipment, the topological connection relationship between equipment, and the standard appearance texture features. For each original image in the substation synchronous image set, the preset three-dimensional spatial coordinates and camera parameters in the knowledge graph are used to back-project the three-dimensional spatial coordinates onto the two-dimensional plane of the original image to obtain the estimated image area of the power equipment. Within the estimated image region, image template matching is performed based on the standard appearance texture features in the knowledge graph to locate the image boundary of the power equipment and segment out the equipment block image containing the power equipment.
[0007] As a further aspect of the present invention, statistical modeling of the pixel structure of the device's segmented image to generate a baseline model of the device's pixel structure during the current detection period includes: The device block images are categorized according to device type to generate device image clusters corresponding to each device type; The pixel structure baseline model includes a primary color ratio vector, an edge distribution matrix, and a texture co-occurrence trend. For each device type, the device image is clustered, and all device block images are decomposed into pixel channels. The pixel value distribution of each color channel is calculated to form the main color ratio vector. Edges are extracted from the block images of all devices, the position density of edge points is calculated, and an edge distribution matrix is generated. Texture analysis is performed on the block images of all devices to calculate the consistency of texture direction and repetition period, and to form a texture symbiosis trend; The statistical average of the main color ratio vector, edge distribution matrix, and texture co-occurrence trend is used as the pixel structure benchmark model for the corresponding device type in the current detection period.
[0008] As a further aspect of the present invention, the real-time pixel structure features of the block image of the device under inspection are calculated, and the deviation between the real-time pixel structure features and the pixel structure benchmark model of the corresponding device type is measured, including: The pixel channel decomposition is performed on the block image of the device under inspection, and its real-time main color ratio vector is calculated; Edge extraction is performed on the block image of the device under inspection, and its real-time edge distribution matrix is calculated; Perform texture analysis on the block image of the device under inspection and calculate its real-time texture co-occurrence trend; The color deviation is obtained by calculating the cosine similarity between the real-time main color ratio vector and the main color ratio vector in the pixel structure baseline model. The edge distribution deviation is obtained by calculating the matrix norm distance between the real-time edge distribution matrix and the edge distribution matrix in the pixel structure baseline model. The real-time texture co-occurrence trend is compared with the texture co-occurrence trend in the pixel structure baseline model to calculate the trend correlation and obtain the texture trend deviation. The color deviation, edge distribution deviation, and texture trend deviation are weighted and fused to obtain the comprehensive deviation. The overall deviation is compared with the structural deviation threshold.
[0009] As a further aspect of the present invention, for the power equipment corresponding to the initial screening abnormal image, its digital twin model is activated, the initial screening abnormal image is mapped to the spatial coordinates of the digital twin model, and a three-dimensional model driving is triggered, including: The digital twin model includes the physical property parameters, kinematic constraints, and stress-deformation rules of the internal components of the power equipment. Based on the timestamp of the initial screening abnormal image, obtain all the status parameters of the power equipment operation at the time of the timestamp recorded by the substation monitoring system; All the state parameters are used as driving inputs to the digital twin model, which is then run from the initial state to simulate the continuous behavior of the power equipment in the physical space starting from the timestamp. Record the dynamic deformation and displacement sequence of the three-dimensional component surface corresponding to the equipment area in the initial screening abnormal image during the simulation operation of the digital twin model. The dynamic deformation and displacement sequence is the simulation behavior sequence.
[0010] As a further aspect of the present invention, the physical behavior states synchronously extracted from the actual video stream include: Starting from the timestamp of the initial screening abnormal image, extract the subsequent image frame sequence of the power equipment in physical space from the actual video stream corresponding to the original acquired substation synchronous image set; The subsequent image frame sequence is subjected to device region tracking and segmentation to extract the actual contour shape and spatial position of the physical device in each frame image, forming a physical behavior state sequence.
[0011] As a further aspect of the present invention, cross-validation is performed between the device behavior state presented by the simulated behavior sequence and the physical behavior state synchronously extracted from the actual video stream, including: For each simulated frame in the simulated behavior sequence, calculate the simulated contour shape and simulated spatial position of the corresponding three-dimensional component surface in the digital twin model; For each physical frame in the physical behavior state sequence, the actual contour shape and actual spatial position of the physical device at the corresponding time are extracted; The overlap between the simulated contour shape and the actual contour shape at the same time is compared, and the Euclidean distance between the simulated spatial position and the actual spatial position at the same time is calculated. A mismatch event is recorded when the overlap is lower than the shape overlap threshold or the Euclidean distance exceeds the displacement deviation threshold.
[0012] As a further aspect of the present invention, when the cross-validation results do not match, the initial screening abnormal image is confirmed as the final fault characterization image, and all information related to the final fault characterization image is recorded, including: The total number of mismatch events is counted, and when the total number exceeds a preset event counting threshold, the cross-validation result is determined to be mismatched. The initial screening anomaly image that triggers cross-validation is confirmed as the final fault characterization image; All information related to the final fault characterization image is recorded in the fault file. The information includes: the initial screening abnormal image and its timestamp, the corresponding device type, the measured comprehensive deviation, the simulated behavior sequence, the physical behavior state sequence, the time and type of the mismatch event, the total number of mismatch events, the panoramic view image of the original image where the final fault characterization image is located, and the system environment parameters at the time of image acquisition.
[0013] As a further aspect of the present invention, it also includes: Historical fault cases are periodically extracted from the fault archives, and the final fault characterization images and information in the historical fault cases are back-injected into the substation equipment knowledge graph. In the substation equipment knowledge graph, a new fault texture feature library is added for the corresponding power equipment nodes; The new fault texture feature library contains image features of the fault region in the final fault characterization image, which is used to assist in identifying the same or similar faults in the subsequent image template matching step of device localization.
[0014] As a further aspect of the present invention, the present invention also includes a substation fault detection system based on images and digital twins, the system including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor, when executing the computer program, implements the steps of the substation fault detection method based on images and digital twins as described above.
[0015] Compared with the prior art, the advantages and positive effects of the present invention are as follows: Pixel structure statistical modeling is performed on the segmented device block images to generate a pixel structure benchmark model corresponding to the current detection period of the device. The real-time pixel structure features of the device block images to be inspected are calculated, and the deviation of these features from the pixel structure benchmark model of the corresponding device type is measured. The deviation exceeding the preset structural deviation threshold is used as the basis for initial screening of abnormal images. This can get rid of the constraints of fixed historical templates and form a unique comparison benchmark based on the visual structure features of the device in real time detection period. It avoids environmental interference caused by simple pixel difference comparison, making the judgment criteria for initial screening of abnormal images more in line with the actual visual state of the device and weakening the influence of feature deviation caused by non-fault factors.
[0016] The initial screening of abnormal images is mapped to the spatial coordinates of the corresponding digital twin model of the power equipment. This triggers the 3D model to simulate the dynamic behavior of the equipment in the physical space and generate a simulated behavior sequence. The equipment behavior state presented by the simulated behavior sequence is cross-validated with the physical behavior state extracted synchronously from the actual video stream. The mismatch of the verification results is used as the basis for confirming the final fault characterization image. This can break the application form of digital twins that are only statically displayed, realize the precise spatial binding of abnormal images and digital twin models, restore the real operating state of the equipment through dynamic behavior simulation, and eliminate false anomalies at the image level by relying on bidirectional behavior verification, so as to achieve accurate identification of fault characterization and distinguish between image noise and the actual fault state of the equipment. Attached Figure Description
[0017] Figure 1 This is a flowchart of the substation fault detection method based on image and digital twin as described in this invention; Figure 2 A flowchart for generating a pixel structure baseline model by statistically modeling the pixel structure of a device block image; Figure 3 A comparative analysis diagram of device edge and texture features; Figure 4 A diagram showing the positional discrepancy between the digital twin and the physical device; Figure 5 This is a cross-validation diagram of physical devices and digital twin models. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0019] In the description of this invention, it should be understood that the terms "length," "width," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientation or positional relationships, are based on the orientation or positional relationships shown in the accompanying drawings and are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, in the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0020] See Figure 1 This invention provides a substation fault detection method based on image and digital twin, the method comprising: A set of synchronous images from a substation is acquired, and equipment block images containing specific power equipment image regions are segmented. The pixel structure of these equipment block images is statistically modeled to generate a pixel structure baseline model for the equipment within the current detection period. A new image frame sequence is extracted from the substation synchronous image set, and the images in this new sequence are segmented into equipment block images according to equipment type. The real-time pixel structure features of the equipment block images are calculated, and the deviation between these features and the corresponding equipment type's pixel structure baseline model is measured. When the measured deviation exceeds a preset structural deviation threshold, the equipment block image is determined to be an initial screening abnormal image. For the power equipment corresponding to the initial screening abnormal image, its digital twin model is activated, mapping the initial screening abnormal image to the spatial coordinates of the digital twin model. This triggers a 3D model drive to simulate the dynamic behavior of the power equipment in physical space, generating a simulated behavior sequence. The equipment behavior state presented by the simulated behavior sequence is cross-validated with the physical behavior state synchronously extracted from the actual video stream. When the cross-validation results do not match, the initial screening abnormal image is confirmed as the final fault characterization image, and all information related to the final fault characterization image is recorded.
[0021] In one embodiment of the present invention, a set of synchronous images of a substation, acquired simultaneously by multiple fixed-view camera devices during a detection period, is received. A pre-constructed substation equipment knowledge graph is invoked to locate the equipment in each original image of the synchronous image set, segmenting the image into blocks containing specific power equipment image regions. The substation equipment knowledge graph contains preset three-dimensional spatial coordinates, topological connections between equipment, and standard appearance texture features for each type of power equipment. For each original image in the synchronous image set, the preset three-dimensional spatial coordinates from the knowledge graph and camera parameters are used to back-project the three-dimensional spatial coordinates onto the two-dimensional plane of the original image, obtaining a predicted image region of the power equipment. Within the predicted image region, image template matching is performed based on the standard appearance texture features from the knowledge graph to locate the image boundary of the power equipment, and segmenting the image into blocks containing the power equipment.
[0022] The device image blocks are categorized according to device type, generating device image clusters corresponding to each device type. The pixel structure baseline model includes a dominant color ratio vector, an edge distribution matrix, and a texture co-occurrence trend. For each device image cluster, pixel channel decomposition is performed on all device image blocks, calculating the pixel value distribution of each color channel to form a dominant color ratio vector. Edge extraction is performed on all device image blocks, and the position density of edge points is statistically analyzed to generate an edge distribution matrix. Texture analysis is performed on all device image blocks, calculating the consistency and repetition period of texture directions to form a texture co-occurrence trend. The statistical average of the dominant color ratio vector, edge distribution matrix, and texture co-occurrence trend is used as the pixel structure baseline model for the corresponding device type in the current detection period.
[0023] In practice, the system receives a set of synchronized images of the substation from multiple fixed-view cameras captured simultaneously during the detection period. These images are captured by multiple high-definition network cameras deployed at different locations within the substation, triggered simultaneously by signals, thus obtaining multi-view image data of the substation's panoramic view at the same time. A pre-built substation equipment knowledge graph is invoked. This graph contains preset 3D spatial coordinates for each type of power equipment, the topological connections between equipment, and standard appearance texture features. The preset 3D spatial coordinates are determined based on substation design drawings and on-site survey data. The standard appearance texture features are derived from standard images of the equipment at the time of manufacture or high-definition sample images captured under fault-free conditions. For each original image in the substation synchronized image set, the preset 3D spatial coordinates and camera parameters (including the camera intrinsic and extrinsic parameter matrices) from the substation equipment knowledge graph are used to back-project the 3D spatial coordinates onto the 2D plane of the original image, resulting in an estimated image region of the power equipment, which is a rectangular bounding box. Within the estimated image region, image template matching is performed based on the standard appearance texture features in the substation equipment knowledge graph to locate the image boundaries of the power equipment and segment the equipment block image containing the power equipment. The image template matching adopts the normalized cross-correlation method.
[0024] See Figure 2The device images are categorized according to device type, generating device image clusters for each type, including circuit breakers, disconnectors, and transformer bushings. The pixel structure baseline model includes a dominant color ratio vector, an edge distribution matrix, and a texture co-occurrence trend. For each device image cluster, pixel channel decomposition is performed on all device images, calculating the pixel value distribution for each color channel to form a dominant color ratio vector. In some embodiments, the dominant color ratio vector is a three-dimensional vector, with its components representing the proportion of pixel values in the device image falling into the red, green, and blue dominant color intervals in a specific color space. Edge extraction is performed on all device images, statistically analyzing the position density of edge points to generate an edge distribution matrix. The Canny operator is used for edge extraction. Texture analysis is performed on all device images, calculating the consistency and repetition period of texture directions to form a texture co-occurrence trend. This trend describes the arrangement pattern of texture primitives within a local area of the image. The statistical average of the dominant color ratio vector, edge distribution matrix, and texture co-occurrence trend is used as the pixel structure baseline model for the corresponding device type in the current detection period.
[0025] In some embodiments, the consistency of texture orientation is quantified by calculating the peak intensity of the local gradient orientation histogram of the image; optionally, a texture orientation consistency index is defined. as follows: in: It represents the total number of device block images belonging to the same device type within the current detection period. Indicates the first After the image is divided into blocks by the device, the calculated first block is... The statistical values of each gradient direction interval in the gradient direction histogram. This represents the maximum value in the histogram, reflecting the intensity of the most dominant edge directions in the image. Texture Direction Consistency Index This is the average intensity of the principal direction across all device-block images. The texture repetition period is obtained by performing a two-dimensional Fourier transform on the device-block images and analyzing the periodic peak intervals of their power spectra in specific directions. The texture co-occurrence trend is ultimately determined by the texture orientation consistency index. Together with the calculated repetition period value, it represents the characteristics. This representation process is part of statistical modeling and is used to establish a baseline of the texture features of the device under normal conditions.
[0026] In one embodiment of the present invention, pixel channel decomposition is performed on the block image of the device under inspection to calculate its real-time dominant color ratio vector. Edge extraction is performed on the block image of the device under inspection to calculate its real-time edge distribution matrix. Texture analysis is performed on the block image of the device under inspection to calculate its real-time texture co-occurrence trend. Cosine similarity is calculated between the real-time dominant color ratio vector and the dominant color ratio vector in the pixel structure benchmark model to obtain the color deviation. Matrix norm distance is calculated between the real-time edge distribution matrix and the edge distribution matrix in the pixel structure benchmark model to obtain the edge distribution deviation. Trend correlation is calculated between the real-time texture co-occurrence trend and the texture co-occurrence trend in the pixel structure benchmark model to obtain the texture trend deviation. The color deviation, edge distribution deviation, and texture trend deviation are weighted and fused to obtain a comprehensive deviation. The comprehensive deviation is compared with the structural deviation threshold.
[0027] In the specific implementation, pixel channel decomposition is performed on the block image of the disconnector switch type equipment under inspection. The real-time dominant color ratio vector of the block image of the equipment under inspection is calculated. In the specific implementation, pixel channel decomposition is performed in the RGB color space. The R, G, and B component values of all pixels in the block image of the equipment under inspection are counted, and the proportion of pixels whose pixel values fall into the preset red main interval, green main interval, and blue main interval is calculated to the total number of pixels in the image. These three proportions constitute a three-dimensional real-time dominant color ratio vector. Edge extraction is performed on the block image of the disconnector switch type equipment under inspection, and the real-time edge distribution matrix of the block image of the equipment under inspection is calculated. Edge extraction also uses the Canny operator. Then, the image is evenly divided into an M-row N-column grid in space, and the number of edge pixels in each grid cell is counted to form an M×N matrix as the real-time edge distribution matrix. Texture analysis is performed on the block images of the equipment under inspection, which is a type of disconnect switch. The real-time texture co-occurrence trend of the block images of the equipment under inspection is calculated. The texture analysis process includes calculating the local gradient direction histogram of the image to obtain the main direction intensity, and obtaining the texture repetition period through Fourier spectrum analysis. These two values together characterize the real-time texture co-occurrence trend.
[0028] The color deviation is obtained by calculating the cosine similarity between the real-time main color ratio vector of the block image of the disconnector type device under test and the main color ratio vector in the pixel structure benchmark model of the disconnector type. It can be understood that the closer the cosine similarity value is to 1, the more similar the color distribution. The color deviation is defined as the difference between 1 and the cosine similarity value. The edge distribution matrix of the block image of the disconnector type device under test is then calculated by performing matrix norm distance calculation between the real-time edge distribution matrix of the block image of the disconnector type device under test and the edge distribution matrix in the pixel structure benchmark model of the disconnector type. In some embodiments, the matrix norm distance is calculated using the Frobenius norm, which is the square root of the sum of the squares of the differences between corresponding elements of the two matrices. Finally, the texture trend deviation is obtained by calculating the trend correlation between the real-time texture symbiosis trend of the block image of the disconnector type device under test and the texture symbiosis trend in the pixel structure benchmark model of the disconnector type. Optionally, the trend correlation calculation is obtained by comparing the main direction intensity and the repetition period value, calculating their relative deviations respectively, and then summing them.
[0029] The color deviation, edge distribution deviation, and texture trend deviation are weighted and fused to obtain the comprehensive deviation. In some embodiments, the comprehensive deviation is defined. The calculation formula is as follows: in: Indicates the overall deviation. This represents the cosine similarity value between the real-time primary color ratio vector and the baseline primary color ratio vector. This represents the matrix norm distance between the real-time edge distribution matrix and the baseline edge distribution matrix. This represents the deviation value calculated from the trend correlation between the real-time texture co-occurrence trend and the baseline texture co-occurrence trend. , , These are the preset weighting coefficients corresponding to color deviation, edge distribution deviation, and texture trend deviation, respectively, and they satisfy... It can be understood that the weighted fusion process unifies the feature deviations of three different dimensions into a single scalar value. The calculated overall deviation is compared with a pre-set structural deviation threshold for the disconnector type. When the overall deviation exceeds the structural deviation threshold, the current image block of the device under inspection is determined to be an abnormal image in the initial screening for that disconnector. In specific implementation, the structural deviation threshold is set after statistical analysis of the overall deviation of historical normal images.
[0030] In one embodiment of the present invention, the digital twin model includes physical property parameters, kinematic constraints, and stress-deformation rules of the internal components of the power equipment. Based on the timestamp of the initial screening anomaly image, all state parameters of the power equipment's operation at the time of the substation monitoring system are obtained. These state parameters are used as the driving input to the digital twin model, causing it to run from its initial state and simulate the continuous behavior of the power equipment in physical space from the timestamp. During the simulation, the dynamic deformation and displacement sequence of the three-dimensional component surfaces corresponding to the equipment area in the initial screening anomaly image are recorded; this dynamic deformation and displacement sequence is the simulated behavior sequence.
[0031] In practical implementation, the digital twin model includes the physical property parameters, kinematic constraints, and stress-deformation rules of the internal components of the power equipment. The physical property parameters include the elastic modulus, Poisson's ratio, density, and coefficient of thermal expansion of the component materials. Kinematic constraints define the permissible relative motion relationships between components, and stress-deformation rules, based on the principles of material mechanics, define the relationship between deformation and stress under load. Based on the timestamp of the initial screening anomaly image of the disconnector switch, all state parameters of the disconnector switch operation at the timestamp recorded by the substation monitoring system are obtained. These state parameters include the real-time opening and closing position signals of the disconnector switch, the current value through the main conductive circuit, the oil or air pressure of the operating mechanism, and the ambient temperature. All state parameters are used as the driving input to the digital twin model, enabling it to operate from the physical state corresponding to the timestamp, simulating the continuous behavior of the disconnector switch in physical space from the timestamp. In practical implementation, the driving input is converted into boundary conditions and loads acting on the corresponding components within the digital twin model.
[0032] The dynamic deformation and displacement sequence of the three-dimensional component surface corresponding to the equipment area in the initial screening anomaly image is recorded during the simulation operation of the digital twin model. This dynamic deformation and displacement sequence is the simulated behavior sequence. In some embodiments, the dynamic deformation of the three-dimensional component surface... Based on its material properties and the stress it is subjected to, the optional dynamic deformation can be calculated. With stress The relationship satisfies the physical rules built into the digital twin model, for example, it can be represented as a linear elastic relationship: ,in Represents dynamic deformation variables. This represents the surface stress of the component calculated by the digital twin model. This represents the elastic modulus of the component material. The digital twin model advances the simulation according to a preset time step. At the end of each simulation step, the changes in the three-dimensional coordinates of each feature point on the surface of the target three-dimensional component are recorded, thus forming a displacement sequence ordered by time. At the same time, the strain caused by stress on the surface of the component is recorded, forming a deformation sequence.
[0033] In some embodiments, the hydraulic pressure of the operating mechanism in all state parameters of the disconnecting switch is converted into a load that drives the piston movement of the hydraulic cylinder in the digital twin model. The current value is used to calculate the Joule thermal load of the conductive rod and induce thermal deformation simulation. It is understood that by using real-time system state parameters as driving inputs, the simulated behavior sequence of the digital twin model should theoretically be consistent with the real behavior of the physical world device under the same initial conditions and external stimuli. The simulation operation of the digital twin model covers a preset future time period, thereby generating a continuous behavior sequence for subsequent comparison. In specific implementations, the duration of the simulation operation is consistent with the duration of extracting the physical behavior state sequence from the actual video stream. The simulated behavior sequence records not only macroscopic displacement but also minute deformations. Optionally, for the contact surface of the disconnecting switch, the digital twin model simulates the minute expansion deformation caused by overheating due to increased current and records this deformation process in the dynamic deformation sequence.
[0034] See Figure 3 This is a comparative analysis chart of device edge and texture features, used to visually present the differences in edge and texture features among different samples. The edge and texture feature values of normal samples are highly close to the baseline model, with minimal fluctuations, indicating stable and normal device appearance. The feature values of the samples under inspection show a significant increase, especially sample 2, where the edge feature value exceeds 25 and the texture feature value approaches 17, far exceeding the normal range, suggesting possible abnormal changes in the device appearance. The larger the feature difference between the sample under inspection and the baseline model, the more significant the deviation of the image pixel structure from the normal state, serving as direct input for subsequent calculations of edge distribution deviation and texture trend deviation. When both types of features increase significantly simultaneously, an initial anomaly detection can be triggered, providing a preliminary basis for subsequent cross-validation using the digital twin model. The feature distribution of normal samples can be used to update the pixel structure baseline model, improving long-term detection accuracy.
[0035] In one embodiment of the present invention, starting from the timestamp of the initially screened abnormal image, a sequence of subsequent image frames of the power equipment in physical space is extracted from the actual video stream corresponding to the originally acquired set of synchronous images of the substation. Equipment region tracking and segmentation are performed on the subsequent image frame sequence to extract the actual contour shape and spatial position of the physical equipment in each frame, forming a physical behavior state sequence. For each simulated frame in the simulated behavior sequence, the simulated contour shape and simulated spatial position of the corresponding three-dimensional component surface in the digital twin model are calculated. For each physical frame in the physical behavior state sequence, the actual contour shape and actual spatial position of the physical equipment at the corresponding time are extracted. The overlap between the simulated contour shape and the actual contour shape at the same time is compared, and the Euclidean distance between the simulated spatial position and the actual spatial position at the same time is calculated. When the overlap is lower than the shape overlap threshold, or the Euclidean distance exceeds the displacement deviation threshold, a mismatch event is recorded.
[0036] In the specific implementation, starting from the timestamp of the initial screening abnormal image of the disconnector switch, subsequent image frame sequences of the power equipment in physical space are extracted from the actual video stream corresponding to the original acquired substation synchronous image set. In this implementation, the extracted image frame sequence covers all consecutive video frames within a preset time period after the timestamp, with the frame rate consistent with the original video stream. Equipment region tracking and segmentation are performed on the extracted subsequent image frame sequence to extract the actual contour shape and spatial position of the physical equipment in each frame, forming a physical behavior state sequence. Equipment region tracking uses a kernel correlation filter-based tracking algorithm to lock the disconnector switch region in the image. Segmentation uses a semantic segmentation network to perform pixel-level segmentation of the tracking region to obtain an accurate equipment contour. The actual contour shape is defined by the binary mask obtained from the segmentation, and the actual spatial position is characterized by calculating the coordinates of the center point of the minimum bounding rectangle of the actual contour shape in the image pixel coordinate system. Refer to Table 1, which shows a simplified example of physical behavior state sequence data containing key information extracted from the actual video stream.
[0037] Table 1: Sequence List of Physical Behavior States For each simulated frame in the simulated behavior sequence, the simulated contour shape and simulated spatial position of the corresponding 3D component surface in the digital twin model are calculated. In some embodiments, the simulated contour shape is obtained by projecting the 3D model of the digital twin model at the corresponding simulation moment onto a 2D image plane according to the same camera parameters, and extracting the projected 2D contour. The simulated spatial position is obtained by calculating the coordinates of the center point of the minimum bounding rectangle of the simulated contour shape on the image plane. For each physical frame in the physical behavior state sequence, the actual contour shape and actual spatial position of the physical device at the corresponding moment are extracted. The actual contour shape and actual spatial position are derived from the data recorded in the physical behavior state sequence at the corresponding moment. The overlap between the simulated contour shape and the actual contour shape at the same moment is compared. It can be understood that the overlap comparison is used to quantify the consistency of the simulated contour and the real contour in spatial coverage. Optionally, the overlap... The calculation formula is defined as follows: in: Indicates the degree of overlap. This represents the pixel area of the simulated contour shape on the image plane. This represents the pixel area of the actual contour shape on the image plane. This represents the pixel area of the region where the simulated and actual contour shapes intersect. The Euclidean distance is calculated between the simulated and actual spatial positions at the same time point. Euclidean distance directly calculates the straight-line distance between the center points of the simulated and actual contours in the image pixel coordinate system. The calculated overlap... A mismatch event is recorded when the shape overlap is below a preset threshold (which is a decimal between 0 and 1) or when the calculated Euclidean distance exceeds a preset displacement deviation threshold. In some embodiments, the mismatch event is labeled with the specific time of its occurrence and the type of violation causing the recording, such as "low overlap" or "large displacement deviation." The cross-validation process is performed frame-by-frame and time-by-time, generating a series of mismatch event records.
[0038] See Figure 4This is a chart analyzing the positional discrepancy between a digital twin and a physical device, used to quantify the spatial consistency between the digital twin model and the physical device. From 0.0 to 0.5 seconds, the Euclidean distance remains stable within 12 pixels, showing a significant difference from the threshold and indicating a high degree of positional matching. From 0.5 to 0.6 seconds, the distance first exceeds the 15-pixel threshold, triggering a positional mismatch event. After 0.6 seconds, the distance rapidly increases, reaching nearly 47 pixels at 0.9 seconds, indicating a sharp increase in the degree of deviation. A smaller Euclidean distance indicates a higher degree of fidelity in the digital twin model's reproduction of the physical device's position; a distance exceeding the threshold represents a discrepancy between the model's behavior and that of the real device. According to the implementation rules, when the Euclidean distance exceeds 15.0 pixels, it is marked as a positional mismatch event, which is one of the core bases for subsequent determination of the final fault. The accelerated increase in deviation after 0.5 seconds indicates that the physical device may have undergone displacement beyond the model's expectations, requiring further cross-validation based on contour overlap.
[0039] In one embodiment of the present invention, the total number of mismatch events is counted. When the total number exceeds a preset event counting threshold, the cross-validation result is determined to be mismatched. The initial screening abnormal image that triggers cross-validation is confirmed as the final fault characterization image. All information related to the final fault characterization image is recorded in the fault archive. The information includes: the initial screening abnormal image and its timestamp, the corresponding equipment type, the measured comprehensive deviation, the simulated behavior sequence, the physical behavior state sequence, the time and type of the mismatch event, the total number of mismatch events, the panoramic view image of the original image where the final fault characterization image is located, and the system environment parameters at the time of image acquisition. Historical fault cases are periodically extracted from the fault archive, and the final fault characterization images and their information in the historical fault cases are back-injected into the substation equipment knowledge graph. In the substation equipment knowledge graph, a new fault texture feature library is added for the corresponding power equipment node. The new fault texture feature library contains the image features of the fault area in the final fault characterization image, which is used to assist in identifying the same or similar faults in the subsequent image template matching step of equipment positioning.
[0040] In specific implementation, the total number of all mismatch events recorded during cross-validation is counted. When the total number of mismatch events exceeds a preset event counting threshold, the cross-validation result is determined to be mismatched. In some embodiments, the event counting threshold is set based on the total duration of the simulated behavior sequence and the sampling frequency. The initial screening abnormal image of the disconnector that triggered cross-validation is confirmed as the final fault characterization image. It can be understood that the confirmation process marks the transition from preliminary abnormal indication to final fault determination. All information related to the final fault characterization image is recorded in the fault archive, which is a structured database. In specific implementation, the recorded information includes: the initial screening abnormal image and its timestamp, the corresponding equipment type being a disconnector, the measured comprehensive deviation, the simulated behavior sequence, the physical behavior state sequence, the specific time and type of the mismatch event (low overlap or large displacement deviation), the total number of mismatch events, the panoramic view image of the original image containing the final fault characterization image, and the system environmental parameters at the time of image acquisition, including ambient temperature and light intensity.
[0041] Historical fault cases are periodically extracted from fault archives. These historical fault cases contain all previously confirmed final fault characterization images and their associated information. The extraction operation is performed automatically according to a preset time cycle. The final fault characterization images and their information from the historical fault cases are then back-injected into the substation equipment knowledge graph. Back-injection refers to adding fault-related feature data as new knowledge nodes or attributes under the corresponding device nodes in the knowledge graph. In the substation equipment knowledge graph, a new fault texture feature library is added to the corresponding power equipment nodes. It can be understood that each device node can be associated with multiple fault texture feature libraries, and each library corresponds to a fault mode that has occurred. The new fault texture feature library contains the image features of the fault area in the final fault characterization image. In some embodiments, the fault area is obtained by differential segmentation by comparing the fault image with a baseline image of the equipment in normal state. Its image features include the local color histogram, texture feature vector, and shape descriptor of the fault area. In the image template matching step for subsequent equipment location, this aids in identifying identical or similar faults. Optionally, during the image template matching stage, the identification is performed not only by matching with standard appearance texture features in the substation equipment knowledge graph, but also by calculating similarity with various recorded fault texture features. If the similarity with a certain fault feature exceeds a predetermined threshold, the potential fault type is directly indicated. This can be understood as defining the total number of mismatch events. The calculation formula is: in: This indicates the total number of non-matching events. This represents the total number of frames compared during the cross-validation process. It is an indicator variable, when the first When the frame comparison result is determined to be a mismatch event ,otherwise .when If the number of events exceeds the preset event count threshold, the cross-validation results are considered mismatched.
[0042] See Figure 5 This is a cross-validation chart of a physical device and a digital twin model, used to determine whether the behavior of the digital twin model matches that of the physical device. Shape overlap is the pixel overlap ratio between the simulated and actual contours; a higher value indicates better shape consistency (red dashed line: 50%). Values below this are considered shape mismatches. Position distance is the Euclidean pixel distance between the center points of the simulated and actual contours; a lower value indicates better position consistency (green dashed line: 5 pixels). Values exceeding this are considered position mismatches. In frames 1-15, the shape overlap slowly decreases from approximately 95% to approximately 50%, while the position distance gradually increases from 0 pixels to approximately 5 pixels, both approaching the threshold, indicating that the model and physical device behavior are basically consistent. In frames 15-16, the shape overlap first falls below the 50% threshold, and the position distance first exceeds the 5-pixel threshold, triggering mismatch conditions simultaneously, indicating that a fault is beginning to appear. In frames 16-30, the shape overlap continues to decrease rapidly to approximately 2%, while the position distance continues to climb to approximately 10 pixels, with the deviation increasing dramatically, indicating a complete mismatch between the model and physical device behavior.
[0043] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
Claims
1. A substation fault detection method based on image and digital twin, characterized in that, include: Collect a set of synchronous images of the substation, and segment out equipment block images containing image regions of specific power equipment. Perform statistical modeling on the pixel structure of the equipment block images to generate a pixel structure benchmark model of the equipment in the current detection period. Extract a new image frame sequence from the substation synchronization image set, and segment the images in the new image frame sequence into block images of the equipment to be inspected according to the equipment type; Calculate the real-time pixel structure features of the block image of the device under inspection, and measure the deviation between the real-time pixel structure features and the pixel structure benchmark model of the corresponding device type. When the measured deviation exceeds the preset structural deviation threshold, the block image of the device under inspection is determined to be an abnormal image in the initial screening. For the power equipment corresponding to the initial screening abnormal image, its digital twin model is activated, the initial screening abnormal image is mapped to the spatial coordinates of the digital twin model, and the three-dimensional model driving is triggered to simulate the dynamic behavior of the power equipment in physical space and generate a simulated behavior sequence. The device behavior state presented by the simulated behavior sequence is cross-validated with the physical behavior state extracted synchronously from the actual video stream; When the cross-validation results do not match, the initial screening abnormal image is confirmed as the final fault characterization image, and all information related to the final fault characterization image is recorded.
2. The substation fault detection method based on image and digital twin as described in claim 1, characterized in that, Acquire a set of synchronous images of the substation, segment the images into blocks containing specific power equipment image regions, including: Receive a set of synchronous images of the substation acquired simultaneously by multiple fixed-view camera devices during the detection period; The system invokes a pre-built substation equipment knowledge graph to locate equipment in each original image of the substation synchronous image set, segmenting the image into blocks containing specific power equipment regions. Specifically, this includes: The substation equipment knowledge graph includes the preset three-dimensional spatial coordinates of each type of power equipment, the topological connection relationship between equipment, and the standard appearance texture features. For each original image in the substation synchronous image set, the preset three-dimensional spatial coordinates and camera parameters in the knowledge graph are used to back-project the three-dimensional spatial coordinates onto the two-dimensional plane of the original image to obtain the estimated image area of the power equipment. Within the estimated image region, image template matching is performed based on the standard appearance texture features in the knowledge graph to locate the image boundary of the power equipment and segment out the equipment block image containing the power equipment.
3. The substation fault detection method based on image and digital twin as described in claim 1, characterized in that, Statistical modeling of the pixel structure of the device's segmented images is performed to generate a baseline model of the device's pixel structure within the current detection period, including: The device block images are categorized according to device type to generate device image clusters corresponding to each device type; The pixel structure baseline model includes a primary color ratio vector, an edge distribution matrix, and a texture co-occurrence trend. For each device type, the device image is clustered, and all device block images are decomposed into pixel channels. The pixel value distribution of each color channel is calculated to form the main color ratio vector. Edges are extracted from the block images of all devices, the position density of edge points is calculated, and an edge distribution matrix is generated. Texture analysis is performed on the block images of all devices to calculate the consistency of texture direction and repetition period, and to form a texture symbiosis trend; The statistical average of the main color ratio vector, edge distribution matrix, and texture co-occurrence trend is used as the pixel structure benchmark model for the corresponding device type in the current detection period.
4. The substation fault detection method based on image and digital twin as described in claim 1, characterized in that, Calculate the real-time pixel structure features of the block image of the device under inspection, and measure the deviation between the real-time pixel structure features and the pixel structure benchmark model of the corresponding device type, including: The pixel channel decomposition is performed on the block image of the device under inspection, and its real-time main color ratio vector is calculated; Edge extraction is performed on the block image of the device under inspection, and its real-time edge distribution matrix is calculated; Perform texture analysis on the block image of the device under inspection and calculate its real-time texture co-occurrence trend; The color deviation is obtained by calculating the cosine similarity between the real-time main color ratio vector and the main color ratio vector in the pixel structure baseline model. The edge distribution deviation is obtained by calculating the matrix norm distance between the real-time edge distribution matrix and the edge distribution matrix in the pixel structure baseline model. The real-time texture co-occurrence trend is compared with the texture co-occurrence trend in the pixel structure baseline model to calculate the trend correlation and obtain the texture trend deviation. The color deviation, edge distribution deviation, and texture trend deviation are weighted and fused to obtain the comprehensive deviation. The overall deviation is compared with the structural deviation threshold.
5. The substation fault detection method based on image and digital twin according to claim 1, characterized in that, For the power equipment corresponding to the initial screening abnormal image, activate its digital twin model, map the initial screening abnormal image to the spatial coordinates of the digital twin model, and trigger 3D model driving, including: The digital twin model includes the physical property parameters, kinematic constraints, and stress-deformation rules of the internal components of the power equipment. Based on the timestamp of the initial screening abnormal image, obtain all the status parameters of the power equipment operation at the time of the timestamp recorded by the substation monitoring system; All the state parameters are used as driving inputs to the digital twin model, which is then run from the initial state to simulate the continuous behavior of the power equipment in the physical space starting from the timestamp. Record the dynamic deformation and displacement sequence of the three-dimensional component surface corresponding to the equipment area in the initial screening abnormal image during the simulation operation of the digital twin model. The dynamic deformation and displacement sequence is the simulation behavior sequence.
6. The substation fault detection method based on image and digital twin according to claim 1, characterized in that, Physical behavior states extracted synchronously from the actual video stream include: Starting from the timestamp of the initial screening abnormal image, extract the subsequent image frame sequence of the power equipment in physical space from the actual video stream corresponding to the original acquired substation synchronous image set; The subsequent image frame sequence is subjected to device region tracking and segmentation to extract the actual contour shape and spatial position of the physical device in each frame image, forming a physical behavior state sequence.
7. The substation fault detection method based on image and digital twin according to claim 1, characterized in that, Cross-validation is performed between the device behavior state presented by the simulated behavior sequence and the physical behavior state synchronously extracted from the actual video stream, including: For each simulated frame in the simulated behavior sequence, calculate the simulated contour shape and simulated spatial position of the corresponding three-dimensional component surface in the digital twin model; For each physical frame in the physical behavior state sequence, the actual contour shape and actual spatial position of the physical device at the corresponding time are extracted; The overlap between the simulated contour shape and the actual contour shape at the same time is compared, and the Euclidean distance between the simulated spatial position and the actual spatial position at the same time is calculated. A mismatch event is recorded when the overlap is lower than the shape overlap threshold or the Euclidean distance exceeds the displacement deviation threshold.
8. The substation fault detection method based on image and digital twin according to claim 1, characterized in that, When the cross-validation results do not match, the initial screening abnormal image is confirmed as the final fault characterization image, and all information related to the final fault characterization image is recorded, including: The total number of mismatch events is counted, and when the total number exceeds a preset event counting threshold, the cross-validation result is determined to be mismatched. The initial screening anomaly image that triggers cross-validation is confirmed as the final fault characterization image; All information related to the final fault characterization image is recorded in the fault file. The information includes: the initial screening abnormal image and its timestamp, the corresponding device type, the measured comprehensive deviation, the simulated behavior sequence, the physical behavior state sequence, the time and type of the mismatch event, the total number of mismatch events, the panoramic view image of the original image where the final fault characterization image is located, and the system environment parameters at the time of image acquisition.
9. The substation fault detection method based on image and digital twin as described in claim 8, characterized in that, Also includes: Historical fault cases are periodically extracted from the fault archives, and the final fault characterization images and information in the historical fault cases are back-injected into the substation equipment knowledge graph. In the substation equipment knowledge graph, a new fault texture feature library is added for the corresponding power equipment nodes; The new fault texture feature library contains image features of the fault region in the final fault characterization image, which is used to assist in identifying the same or similar faults in the subsequent image template matching step of device localization.
10. A substation fault detection system based on image and digital twin, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the substation fault detection method based on image and digital twin as described in any one of claims 1 to 9.