An artificial intelligence-based power distribution cabinet state real-time monitoring method

CN122598103APending Publication Date: 2026-08-18HANGZHOU LINAN JINGCHENG ELECTRIC APPLIANCES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610750271.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-28
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

[0003]然而,在实际工业现场环境中,设备基础与建筑结构存在的持续微幅振动,以及工况变化导致的柜体本身热形变等因素,会使得视觉传感器与柜体间的实际空间位姿发生缓慢时变,导致系统赖以进行状态判读的视觉参考系发生未知偏移,破坏了分析算法的前置几何条件,依赖于静态标定的视觉监测系统,其感知结果的准确性建立在可能已失效的空间基准之上,不仅会造成对部件位移、松脱等状态的误判,更严重的是系统对于因基准漂移而导致的图像模糊、特征提取失败等自身感知能力退化问题无法自知,从而漏检真正的异常状态,使得监测系统的可靠性与可信度在长期运行中面临本质挑战

Benefits of technology

[0041]1. By establishing a complete set of visual self-perception and self-calibration logic, the monitoring system can actively identify and distinguish whether image changes are caused by sensor pose drift or by the normal movement of movable parts inside the cabinet. By dynamically analyzing the motion patterns of preset reference points and introducing a cross-verification mechanism of semantic roles and background light flow field, an introspective judgment of the state of the visual reference system is realized. This enables the system to have the cognitive ability to distinguish between self-motion and object motion, thereby avoiding the misjudgment of the overall drift of the reference coordinate system caused by slow sensor displacement or vibration as a fault of equipment inside the cabinet. It also prevents the incorrect classification of the actual movement of parts inside the cabinet as a sensor problem and the unnecessary image correction, significantly improving the logical rationality and source credibility of the state judgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122598103A_ABST
    Figure CN122598103A_ABST
Patent Text Reader

Abstract

This invention discloses an artificial intelligence-based real-time monitoring method for the status of power distribution cabinets, specifically relating to the field of industrial equipment image processing and monitoring technology. It addresses the problem of misjudgments and missed detections caused by slow changes in the relative pose between the sensor and the cabinet, leading to a shift in the visual reference frame. The method involves continuously acquiring monitoring image sequences and extracting the coordinates of preset reference points; analyzing the motion patterns of the reference points to preliminarily determine the cause of the changes; when sensor pose changes are suspected, resolving the semantic role of the reference points and calculating the background optical flow field; constructing two hypotheses—overall sensor pose change and independent movement of movable parts—and cross-validating them using semantic roles and the background optical flow field; finally, based on the validation results, determining whether to perform coordinate compensation on the images, and completing the power distribution cabinet status identification on this basis, thereby achieving reliable monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing and monitoring technology for industrial equipment, and more specifically, to a method for real-time monitoring of the status of power distribution cabinets based on artificial intelligence. Background Technology

[0002] In the field of intelligent operation and maintenance of power distribution equipment, fixedly installed visual sensors can be used to collect image or video data, and the status inside the cabinet can be automatically analyzed based on computer vision algorithms. Such methods usually rely on a basic assumption: the relative spatial position relationship between the visual sensor and the monitored power distribution cabinet is fixed and constant after installation. Under this assumption, the system establishes a mapping relationship between the image coordinate system and the physical coordinate system of the cabinet through pre-calibration, thereby realizing the identification and monitoring of component positions, instrument readings, and appearance status.

[0003] However, in actual industrial environments, factors such as continuous micro-vibrations of equipment foundations and building structures, as well as thermal deformation of the cabinet itself caused by changes in operating conditions, can cause the actual spatial pose between the vision sensor and the cabinet to change slowly over time. This leads to an unknown shift in the visual reference frame upon which the system relies for state interpretation, disrupting the pre-existing geometric conditions of the analysis algorithm. Vision monitoring systems that rely on static calibration rely on a potentially invalid spatial reference for accurate perception results. This not only causes misjudgments of component displacement and loosening, but more seriously, the system is unable to recognize its own perception degradation caused by reference drift, such as image blurring and feature extraction failure. Consequently, it misses out on truly abnormal states, posing a fundamental challenge to the reliability and trustworthiness of the monitoring system during long-term operation. Summary of the Invention

[0004] In order to overcome the above-mentioned defects of the prior art, the present invention provides a method for real-time monitoring of the status of power distribution cabinets based on artificial intelligence to solve the problems mentioned in the background art.

[0005] To achieve the above objectives, the present invention provides the following technical solution:

[0006] A method for real-time monitoring of the status of a power distribution cabinet based on artificial intelligence includes the following steps:

[0007] S1. Continuously acquire monitoring image sequences of the power distribution cabinet through a visual sensor fixed at the monitoring point;

[0008] S2. Extract the image coordinates of multiple preset reference points on the power distribution cabinet from the monitoring image sequence;

[0009] S3. Based on the change pattern of image coordinates at multiple reference points at adjacent time points, determine whether the change pattern is more likely to be caused by the overall pose change of the sensor.

[0010] S4. When this is the case, analyze the semantic role of each reference point among multiple reference points, and calculate the background optical flow field of non-reference areas in the monitoring image sequence;

[0011] S5. Construct the overall pose change hypothesis of the sensor and the independent motion hypothesis of the movable parts respectively, and verify the overall pose change hypothesis of the sensor and the independent motion hypothesis of the movable parts based on semantic roles and background optical flow field.

[0012] S6. When the verification result confirms the assumption of overall sensor pose change, coordinate compensation is performed based on the image coordinate changes of multiple reference points, and then the power distribution cabinet status is identified based on the compensated image; when the verification result confirms the assumption of independent movement of movable parts, the power distribution cabinet status is identified directly based on the monitoring image sequence.

[0013] Furthermore, a sequence of monitoring images of the distribution cabinet is continuously acquired using a visual sensor fixed at the monitoring point, including:

[0014] Images of the inside of the power distribution cabinet are acquired at a preset cycle under visible light conditions;

[0015] And when the ambient light is insufficient, the supplementary lighting device is activated and images of the inside of the power distribution cabinet under supplementary lighting conditions are collected, thereby obtaining a monitoring image sequence.

[0016] Furthermore, the image coordinates of multiple preset reference points on the distribution cabinet are extracted from the monitoring image sequence, including:

[0017] In the first frame of the monitored image sequence, the initial image coordinates of each preset reference point are identified and recorded using a feature point detection algorithm;

[0018] For each subsequent frame in the monitored image sequence, the feature points detected in the current frame are matched with the feature points corresponding to the recorded initial image coordinates using the feature descriptor matching method, so as to track and obtain the image coordinates of each reference point in the current frame.

[0019] The correctness of the current frame image coordinates obtained through matching and tracking is verified using the principle of multi-view geometric constraints. Incorrect matching and tracking results are eliminated, thereby extracting the effective image coordinates of each reference point in each frame.

[0020] Furthermore, the setting and maintenance of multiple preset reference points include: selecting physical feature points with high texture contrast on the surface of the power distribution cabinet and key internal components as candidate reference points; uniquely identifying the candidate reference points and initially calibrating them with the cabinet's physical coordinate system; and in subsequent monitoring, if the image coordinates of a certain reference point cannot be stably tracked or extracted in multiple consecutive frames of images, then using a backup candidate reference point to replace it and updating the preset reference point information.

[0021] Furthermore, based on the change patterns of image coordinates at multiple reference points at adjacent time points, it is determined whether the change patterns are more likely to be caused by changes in the overall pose of the sensor, including:

[0022] Based on the changes in the image coordinates of multiple reference points at adjacent time points, calculate a pose transformation model that describes the overall rigid body motion of all reference points.

[0023] Calculate the residual between the coordinates of each reference point after transformation according to the pose transformation model and its actual image coordinates at adjacent time points;

[0024] Analyzing the distribution characteristics of the residuals, if the residuals are generally small and uniformly distributed, the change pattern is more likely to be caused by changes in the overall pose of the sensor.

[0025] If the residuals exhibit a localized increase in clustering associated with a specific physical component, then the change pattern is more likely caused by the independent movement of the movable component.

[0026] Furthermore, when this is the case, the semantic role of each of the multiple reference points is resolved, and the background optical flow field of non-reference regions in the monitored image sequence is calculated, including:

[0027] Based on predefined reference point identification information, each reference point is assigned a semantic role that represents it as a structural fixed point or a point associated with a movable part.

[0028] Between adjacent frames of the monitored image sequence, pixels with significant texture features are selected from the non-reference area after excluding all reference points from the preset neighborhood.

[0029] The background optical flow field is calculated by comparing the spatial position movement of pixels with significant texture features in adjacent frames.

[0030] Furthermore, based on the predefined reference point identification information, a semantic role is assigned to each reference point, including: the reference point fixed to the cabinet frame of the distribution cabinet, the busbar connection, or the static insulating component is assigned the semantic role of a structural fixing point; the reference point fixed to the circuit breaker operating handle, the disconnecting switch contact, or the withdrawable component is assigned the semantic role of a movable component association point.

[0031] Furthermore, hypotheses regarding the overall pose change of the sensor and the independent motion of the movable parts are constructed separately, and these hypotheses are validated based on semantic roles and background optical flow fields, including:

[0032] A hypothesis about the overall pose change of the sensor is constructed. A global rigid body transformation is calculated based on the image coordinate changes of all reference points. This transformation is then used to predict the optical flow changes that should occur in non-reference areas, thus obtaining the hypothetical optical flow field.

[0033] The assumption of independent motion of movable parts is constructed. Local motion is calculated based only on the image coordinate changes of reference points whose semantic role is the associated point of the movable part. Combined with the image coordinate invariance of reference points whose semantic role is the fixed point of the structure, the optical flow changes that should be generated in non-reference areas are predicted to obtain the competing optical flow field.

[0034] Calculate the degree of agreement between the assumed optical flow field and the competing optical flow field and the background optical flow field, respectively;

[0035] If the fit of the optical flow field is higher, the verification result of the overall sensor pose change hypothesis is valid; if the fit of the competing optical flow field is higher, the verification result of the independent motion hypothesis of the movable part is valid.

[0036] Furthermore, calculating a global rigid body transformation based on the image coordinate changes of all reference points includes: using the set of image coordinates of all reference points at adjacent time points, and fitting an optimal affine transformation matrix using the least squares method. The affine transformation matrix describes the global coordinate mapping relationship under the assumption of the overall pose change of the sensor.

[0037] Furthermore, when the verification result confirms the assumption of overall sensor pose change, coordinate compensation is performed based on the image coordinate changes of multiple reference points, and then the power distribution cabinet status is identified based on the compensated image; when the verification result confirms the assumption of independent movement of movable parts, the power distribution cabinet status is directly identified based on the monitored image sequence, including:

[0038] If the verification result confirms the assumption of overall sensor pose change, then a coordinate transformation relationship from the current image coordinate system to the reference image coordinate system is calculated based on the image coordinate changes of multiple reference points. The coordinate transformation relationship is then used to perform geometric correction on the monitoring image sequence to obtain the compensated image. After that, the power distribution cabinet status is identified by analyzing the component positions and instrument readings on the compensated image.

[0039] If the verification result confirms the assumption that the movable parts move independently, the status of the distribution cabinet can be identified directly by analyzing the position of the parts and the instrument readings on the current frame of the monitoring image sequence.

[0040] Compared with the prior art, the present invention has the following beneficial effects:

[0041] 1. By establishing a complete set of visual self-perception and self-calibration logic, the monitoring system can actively identify and distinguish whether image changes are caused by sensor pose drift or by the normal movement of movable parts inside the cabinet. By dynamically analyzing the motion patterns of preset reference points and introducing a cross-verification mechanism of semantic roles and background light flow field, an introspective judgment of the state of the visual reference system is realized. This enables the system to have the cognitive ability to distinguish between self-motion and object motion, thereby avoiding the misjudgment of the overall drift of the reference coordinate system caused by slow sensor displacement or vibration as a fault of equipment inside the cabinet. It also prevents the incorrect classification of the actual movement of parts inside the cabinet as a sensor problem and the unnecessary image correction, significantly improving the logical rationality and source credibility of the state judgment.

[0042] 2. It ensures the consistency and reliability of the condition monitoring results during long-term operation. Because the system can perform precise coordinate compensation after confirming changes in sensor pose, it effectively restores the stable mapping relationship between the image coordinate system and the physical world. This ensures that subsequent visual analysis algorithms based on fixed position templates, such as component identification and instrument readings, always work under a geometrically consistent standard perspective, significantly reducing false alarms and missed alarms caused by reference frame drift. At the same time, when the system determines that the component is moving on its own, it directly analyzes the original image, avoiding errors and delays introduced by over-processing. This not only improves the automation and intelligence level of the monitoring system, but more importantly, it endows it with the core capability of long-term autonomous, stable, and reliable operation in complex industrial environments, solving the inherent vulnerability of traditional static visual monitoring methods. Attached Figure Description

[0043] Figure 1 This is a flowchart of a method for real-time monitoring of the status of a power distribution cabinet based on artificial intelligence, according to the present invention. Detailed Implementation

[0044] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0045] Example: Figure 1 This invention presents a method for real-time monitoring of the status of a power distribution cabinet based on artificial intelligence, which includes the following steps:

[0046] S1. Continuously acquire monitoring image sequences of the power distribution cabinet through a visual sensor fixed at the monitoring point;

[0047] S2. Extract the image coordinates of multiple preset reference points on the power distribution cabinet from the monitoring image sequence;

[0048] S3. Based on the change pattern of image coordinates at multiple reference points at adjacent time points, determine whether the change pattern is more likely to be caused by the overall pose change of the sensor.

[0049] S4. When this is the case, analyze the semantic role of each reference point among multiple reference points, and calculate the background optical flow field of non-reference areas in the monitoring image sequence;

[0050] S5. Construct the overall pose change hypothesis of the sensor and the independent motion hypothesis of the movable parts respectively, and verify the overall pose change hypothesis of the sensor and the independent motion hypothesis of the movable parts based on semantic roles and background optical flow field.

[0051] S6. When the verification result confirms the assumption of overall sensor pose change, coordinate compensation is performed based on the image coordinate changes of multiple reference points, and then the power distribution cabinet status is identified based on the compensated image; when the verification result confirms the assumption of independent movement of movable parts, the power distribution cabinet status is identified directly based on the monitoring image sequence.

[0052] S1. Continuously acquire monitoring image sequences of the power distribution cabinet using a visual sensor fixed at the monitoring point. Specifically, this is implemented as follows:

[0053] Images of the power distribution cabinet's interior are acquired at preset intervals under visible light conditions. The preset interval is determined based on the shortest possible time interval between changes in critical states within the cabinet and the image acquisition and processing capabilities of the vision sensor. For example, when monitoring instrument readings or indicator light states, considering that changes in these states are usually directly related to circuit operation and may occur on a timescale of seconds, the preset interval can be set to 1 second to ensure rapid state changes are captured. Conversely, when monitoring temperature distribution or slow deformation of mechanical structures within the cabinet, the preset interval can be set to 10 seconds or even longer to balance real-time monitoring with system processing load. The preset interval is set during system initialization based on the primary monitoring tasks and can be adjusted subsequently through the configuration interface. The vision sensor triggers an image acquisition at the end of each preset interval, acquiring a digital image reflecting the current instantaneous state inside the power distribution cabinet. This image is stored in a standard format, such as RGB, with a resolution set to 1920 pixels × 1080 pixels to ensure clear visual details of instrument panel scales, switch positions, and critical connection components within the cabinet.

[0054] When ambient light is insufficient, a supplementary lighting device is activated, and images of the inside of the distribution cabinet under supplementary lighting conditions are acquired, thus obtaining a monitoring image sequence. The determination of insufficient ambient light is achieved as follows: an ambient light sensor is integrated near the vision sensor to continuously measure the ambient light intensity of the environment where the distribution cabinet is located, measured in lux. A light intensity threshold is set, based on the requirement that the signal-to-noise ratio of key areas in the images of the distribution cabinet's interior acquired without supplementary lighting meets the minimum requirements for subsequent image feature extraction and state recognition algorithms. The specific value of this light intensity threshold is determined experimentally. For example, in the target distribution cabinet installation environment, the ambient light intensity is gradually reduced while images are acquired. By analyzing the decrease in contrast and sharpness in the preset reference point areas of the images, the ambient light intensity value corresponding to the point where image quality begins to significantly deteriorate is set as the light intensity threshold. When the ambient light intensity measured by the ambient light sensor is lower than this light intensity threshold, the system determines that the ambient light is insufficient. The activation control and image acquisition process of the supplementary lighting device are as follows: The system sends an activation command to the supplementary lighting device, which is an LED light source group installed around the vision sensor with its light direction pointing towards the monitoring area inside the power distribution cabinet; upon receiving the command, the supplementary lighting device lights up, and its luminous intensity is pre-configured to be sufficient to raise the illuminance of the key areas inside the power distribution cabinet to a level higher than the illuminance threshold; after the supplementary lighting device is lit and stabilized, the vision sensor performs an image acquisition, obtaining an image of the inside of the power distribution cabinet illuminated under supplementary lighting conditions; after the image acquisition is completed, the system sends a shutdown command to the supplementary lighting device to save energy and avoid the potential impact of continuous heat generation from the supplementary lighting device. Through the above process, whether it is periodic acquisition under sufficient natural light conditions or supplementary lighting acquisition triggered under insufficient light conditions, all the obtained images of the inside of the power distribution cabinet are arranged sequentially in time, forming a monitoring image sequence for subsequent analysis. Each frame in this monitoring image sequence is accompanied by timestamp information, and for acquisition frames triggered by supplementary lighting, a supplementary lighting identifier is marked in the image metadata so that subsequent processing flows know the lighting conditions of that frame.

[0055] S2. Extract the image coordinates of multiple preset reference points on the distribution cabinet from the monitoring image sequence. Specifically, this is implemented as follows:

[0056] In the first frame of the monitored image sequence, the initial image coordinates of each preset reference point are identified and recorded using a feature point detection algorithm. Specifically, the feature point detection algorithm employs either a scale-invariant feature transform algorithm or a directional FAST and rotation BRIEF algorithm. Before applying the feature point detection algorithm, the first frame image is preprocessed to enhance the detectability of feature points. Preprocessing includes grayscale conversion and Gaussian filtering to smooth noise. Algorithm parameters are configured based on typical texture features of the images inside the distribution cabinet. For example, the contrast threshold for detecting feature points in the scale-invariant feature transform algorithm is set to, for example, 0.03, and the edge threshold is set to, for example, 10. The contrast threshold is set based on testing on a large number of distribution cabinet sample images. Adjusting the contrast threshold ensures that the algorithm can stably detect key features such as bolt head corners and label edges, while avoiding excessive false feature points generated by noise or uniform surfaces. The edge threshold is set to limit the algorithm's detection of feature points on strong edges, preventing feature points from being overly concentrated on high-contrast lines such as cabinet door edges. The specific value of the edge threshold is determined by observing the uniformity of feature point distribution on key components inside the cabinet. The feature point detection algorithm searches for pixel regions in the entire image that match corner or blob features and outputs a series of feature point coordinates. To filter out feature points corresponding to preset reference points, the detected feature point positions need to be compared with pre-stored information. This pre-stored information is established during system initialization and contains the approximate area expected to be located for each preset reference point in the first frame of the image. This area is typically a square region centered at the expected coordinates with a side length of, for example, 100 pixels. For each preset reference point, the algorithm searches for feature points output by the feature point detection algorithm within the area defined by the corresponding pre-stored information. If a feature point is found within the area, its image coordinates are recorded as the initial image coordinates of the preset reference point. If multiple feature points are found within the area, the feature point with the largest response value is selected. The response value is calculated by the feature point detection algorithm and represents the degree of difference between the feature point and its surrounding area. If no feature points are found within the area defined by the pre-stored information, manual assignment or template matching is used to assist in localization within the area, and the final determined pixel coordinates are recorded as the initial image coordinates of the reference point. The initial image coordinates of all preset reference points are recorded in a reference point coordinate list, which is also associated with a unique identifier for each reference point.

[0057] For each subsequent frame in the monitored image sequence, a feature descriptor matching method is used to match the detected feature points in the current frame with the feature points corresponding to the recorded initial image coordinates, thereby tracking the image coordinates of each reference point in the current frame. Specifically, the process involves using the same feature point detection algorithm and parameters as the first frame to detect feature points in the current frame, obtaining the feature point set and the corresponding feature descriptor for that set. A feature descriptor is a mathematical vector used to characterize the texture information of the image patch surrounding the feature point; for example, the descriptor for the scale-invariant feature transform algorithm is a 128-dimensional vector. Simultaneously, the image coordinates successfully tracked for each reference point in the previous frame are obtained from the recorded reference point coordinate list, and an image patch centered on these coordinates is extracted. The feature descriptor for this patch is calculated and used as the reference descriptor for the reference point. The matching process is performed independently for each reference point: the Euclidean distance between the descriptor of each feature point in the current frame's feature point set and the reference descriptor of the reference point is calculated. Among all calculated Euclidean distances, the minimum and second-minimum Euclidean distance values ​​are found. A distance ratio threshold is set, with a common value being, for example, 0.8. This threshold is based on experience and is used to determine the degree of difference between the best and second-best matches. A smaller threshold indicates stricter matching requirements and can effectively eliminate fuzzy matches. If the smallest Euclidean distance value divided by the second-smallest Euclidean distance value is less than the distance ratio threshold, the feature point in the current frame corresponding to the smallest Euclidean distance is considered to have successfully matched the reference point, and the image coordinates of the feature point in the current frame are used as the image coordinates of the reference point in the current frame. If the match fails, the image coordinates of the reference point in the current frame are marked as temporarily lost. This process iterates through all reference points, completing coordinate tracking from the previous frame to the current frame.

[0058] For the current frame image coordinates obtained through matching tracking, the correctness is verified using the principle of multi-view geometric constraints, eliminating erroneous matching tracking results, thereby extracting the valid image coordinates of each reference point in each frame. The principle of multi-view geometric constraints is based on the fact that in the same rigid scene and when the camera only changes pose, the matched feature point pairs in images from different viewpoints must satisfy specific geometric relationships. The verification implementation steps are as follows: First, from all successfully matched reference points in the current frame, randomly select, for example, 8 pairs of matching points. A matching point is a coordinate pair formed by the image coordinates of a reference point in the previous frame and the image coordinates obtained through matching tracking in the current frame. Using the 8 pairs of matching points, a fundamental matrix is ​​calculated using the direct linear transformation algorithm. The fundamental matrix is ​​a 3x3 matrix that describes the epipolar geometric relationship between two views. Then, using the calculated fundamental matrix, it is checked whether all successfully matched reference point pairs meet the geometric constraints. For each pair of matching points, the image coordinates of the previous frame are expressed in homogeneous coordinate form and multiplied on the left by the fundamental matrix to obtain an epipolar line equation. The distance from the matching point in the current frame to the epipolar line is calculated; this distance is called the reprojection error. A reprojection error threshold is set, typically based on the positioning accuracy of the image coordinates, for example, 2 pixels. This threshold is set because feature point positioning usually has an error of 1 to 2 pixels; setting it to 2 pixels allows for normal positioning fluctuations while filtering out incorrect matches that significantly deviate from geometric constraints. If the reprojection error of a pair of matching points exceeds the threshold, the match is considered incorrect. If the proportion of incorrect matches exceeds a certain threshold (e.g., 50%), the currently calculated fundamental matrix is ​​discarded, and 8 pairs of matching points are randomly selected for recalculation and verification. This process is repeated up to, for example, 100 times. The threshold is set to ensure that correct matches constitute the majority of the matching points used to calculate the fundamental matrix; typically, the proportion of correct matches should exceed 50%. If a set of matching points is found where the proportion of incorrect matches is below the threshold, this set is retained, and all reference point pairs judged as incorrect matches are marked as invalid in the current frame. After this verification step, the image coordinates not marked as invalid are the valid image coordinates of each reference point in the current frame. Valid image coordinates will update the list of reference point coordinates for tracking matching in the next frame.

[0059] The setting and maintenance of multiple preset reference points includes the following specific operations: Selecting physical feature points with high texture contrast on the surface and key internal components of the distribution cabinet as candidate reference points. The selection process considers the stable and identifiable characteristics of physical feature points in images under various lighting conditions and viewing angles, such as the junction of bolt heads and metal plates, corner points of printed characters on nameplates, and specific marking points on insulating materials. Each candidate reference point is uniquely identified and initially calibrated with the cabinet's physical coordinate system. Unique identification involves assigning a unique number to each candidate reference point. Initial calibration determines the three-dimensional coordinates of each candidate reference point in the actual physical space of the distribution cabinet. Initial calibration is completed during the installation and commissioning phase of the distribution cabinet using measuring equipment such as a laser tracker or total station. The measured physical coordinates are then correlated with the image coordinates of the candidate reference points in images acquired by a vision sensor in a standard pose, and the camera's external parameters are calculated. The initial calibration data, along with the candidate reference point number, physical coordinates, and standard image coordinates, are stored in the preset reference point information database. In subsequent monitoring, if the image coordinates of a reference point cannot be stably tracked or extracted in multiple consecutive frames, a backup candidate reference point is activated for replacement, and the preset reference point information is updated. The criterion for unstable tracking or extraction is: in, for example, five consecutive frames, the image coordinates of a reference point are marked as temporarily lost or marked as invalid after correctness verification. The number of consecutive frames in the criterion is set to balance system response speed and resistance to accidental interference; five consecutive frames mean that the system still has a chance to recover tracking after a brief period of occlusion or sudden changes in illumination, avoiding premature replacement. When the criterion is met, the system queries the preset reference point information database for predefined backup candidate reference point information near the physical area where the faulty reference point is located. Activating a backup candidate reference point means adding the backup candidate reference point's number, physical coordinates, and initial image coordinates obtained in the first frame image using the same feature point detection and region matching method to the currently active reference point coordinate list and tracking process to replace the failed reference point. Simultaneously, the association relationships in the preset reference point information database are updated to reflect this replacement, ensuring the continuity of subsequent maintenance operations.

[0060] S3. Based on the change pattern of image coordinates at multiple reference points at adjacent time points, determine whether the change pattern is more likely to be caused by the overall pose change of the sensor. Specifically, this is implemented as follows:

[0061] Based on the changes in the image coordinates of multiple reference points at adjacent time points, a pose transformation model describing the overall rigid body motion of all reference points is calculated. The pose transformation model is represented by a two-dimensional affine transformation model, which can describe translation, rotation, scaling, and shearing motions within the image plane. The mathematical form of the two-dimensional affine transformation model is a formula that transforms a point from coordinates (xi, yi) to coordinates (xi', yi'), containing six unknown parameters. The specific calculation process is as follows: the image coordinates of all valid reference points at the previous time point are used as the source point set, and the image coordinates of the corresponding reference point at the current time point are used as the target point set. The six parameters of the two-dimensional affine transformation model are solved using the least squares method to minimize the overall error between the coordinates obtained after transforming the source point set according to the two-dimensional affine transformation model and the target point set. The overall error is defined as the sum of the squares of the Euclidean distances between all corresponding point pairs. The solution process constitutes a system of linear equations, and the optimal solution for the six parameters is obtained through standard matrix operations. The six parameters define a specific two-dimensional affine transformation matrix. The two-dimensional affine transformation matrix is ​​the pose transformation model that describes the overall rigid body motion of all reference points.

[0062] Calculate the residual between the coordinates of each reference point after transformation according to the pose transformation model and its actual image coordinates at adjacent time points. For each valid reference point, substitute its image coordinates at the previous time step into a two-dimensional affine transformation matrix, and calculate the predicted coordinates of the reference point after transformation according to the pose transformation model through matrix multiplication. The predicted coordinates are two-dimensional coordinate values. Then, calculate the Euclidean distance between the predicted coordinates and the actual image coordinates of the same reference point at the current time step; this Euclidean distance is the residual of the reference point. The residual is a scalar value, representing the degree of deviation between the actual motion of the reference point and the prediction of the overall rigid body motion model. Repeat this calculation process for all valid reference points to obtain a set of residual values, with one residual value corresponding to each valid reference point.

[0063] Analyzing the distribution characteristics of the residuals, if the residuals are generally small and uniformly distributed, the change pattern is more likely caused by changes in the overall sensor pose. The analysis process first sets a residual threshold to define a generally small residual distribution. The residual threshold is set based on statistical analysis of residual values ​​from historical monitoring data where the sensor was confirmed to be stationary or experiencing only slight overall drift; for example, the 95th percentile of the historical residual value distribution is used as the residual threshold. Historical monitoring data refers to the set of residual values ​​collected and calculated during the calibration phase when the sensor and cabinet are known to be relatively stationary or experiencing only slow overall translation. The condition for generally small residuals is that more than a certain proportion, such as 90% of the effective reference points, have calculated residual values ​​less than the residual threshold. Secondly, the spatial uniformity of the residual distribution is analyzed. The image plane is divided into a regular grid, for example, 100 grids of 10 rows by 10 columns. The average residual of the reference points falling within each grid is calculated. The standard deviation of the average residuals across all grids is then calculated. A distribution uniformity threshold is set, which can be determined based on the tolerance for fluctuations in the average residual between grids, for example, set to 20% of the residual threshold. If the standard deviation of the average residual between grids is less than the distribution uniformity threshold, the residual distribution is considered uniform. When both conditions of generally small residuals and uniform residual distribution are met simultaneously, the current change pattern is determined to be more likely caused by changes in the overall pose of the sensor.

[0064] If the residuals exhibit a localized clustering increase characteristic associated with specific physical components, the change pattern is more likely caused by the independent movement of movable parts. Identification of this localized clustering increase characteristic is achieved through the following steps: First, identify a subset of reference points whose residual values ​​are significantly greater than the overall level. Set a high residual threshold, which can be a multiple of the overall residual threshold, such as 3 times. Filter out all reference points whose residual values ​​are greater than the high residual threshold; these reference points are called high residual reference points. Second, analyze the spatial distribution of high residual reference points in the image. Perform spatial clustering analysis on the image coordinates of the high residual reference points, for example, using a density-based noise-applied spatial clustering algorithm. Set a clustering radius threshold, which is determined based on the pixel size range occupied by typical movable parts within the distribution cabinet in the image. This range is obtained by measuring the projected size of the actual movable parts in the camera's field of view; for example, the clustering radius threshold can be set to 50 pixels. If one or more spatial clusters exist, and one cluster contains more than a minimum cluster point count threshold (e.g., 3 points), and these high-residual-reference points are identified as belonging to the same physical component based on pre-stored reference point semantic role information (e.g., all belonging to the same circuit breaker operating mechanism), then the residuals are determined to exhibit a localized clustering increase characteristic associated with a specific physical component. Finally, this characteristic is verified: the consistency of the residual direction or motion pattern of the reference points within the cluster is checked, for example, by calculating the average direction of the displacement vector of the reference points within the cluster from the previous frame to the current frame. If the verification passes, the current change pattern is determined to be more likely caused by the independent movement of movable components.

[0065] S4. When this is the case, analyze the semantic role of each reference point among multiple reference points, and calculate the background optical flow field of non-reference areas in the monitored image sequence. Specifically, the implementation is as follows:

[0066] When step S3 determines that the change pattern is more likely caused by a change in the overall pose of the sensor, step S4 is executed. Based on predefined reference point identification information, a semantic role is assigned to each reference point, characterizing it as a structural fixed point or a movable component associated point. The predefined reference point identification information is stored in a reference point information database, which is established during system initialization through manual configuration or a semi-automatic calibration tool. The reference point information database records an identification entry for each preset reference point, which includes at least the reference point's unique number, a description of its actual installation location on the distribution cabinet, and its component type classification. The process of assigning semantic roles is executed automatically by the program: the system reads the unique numbers of all currently valid reference points, queries the reference point information database based on the unique number, and obtains the corresponding component type classification field. If the value of the component type classification field represents a category of fixed structure, such as cabinet frame, busbar connection point, or static insulation support, then the semantic role of that reference point is assigned as a structural fixed point. If the value of the component type classification field represents the category of an active component, such as a circuit breaker operating handle, a disconnector switch moving contact, or a point on the moving frame of a withdrawable circuit breaker, then the semantic role of that reference point is assigned as an associated point of an active component. In this way, a clear semantic role is assigned to each reference point that is successfully tracked and verified as valid in the current monitoring image sequence.

[0067] Between adjacent frames of a monitored image sequence, pixels with significant texture features are selected from non-reference areas after excluding all reference points' surrounding preset neighborhoods. The specific implementation includes the following steps: 1. Determine the range of the preset neighborhood. The preset neighborhood is defined as a circular area with a specific radius centered on the current image coordinates of each reference point. The radius of the preset neighborhood is set to ensure coverage of the image patch range used for reference point feature extraction, while leaving margins to prevent background texture selection from being interfered with by the reference points' own features. The preset neighborhood radius can be set, for example, to 15 pixels. 2. The program marks these preset neighborhood areas on the image based on the current image coordinates of all valid reference points. 3. The non-reference area refers to the image area remaining after removing all these preset neighborhood areas from the entire image. 4. Select pixels with significant texture features within the non-reference area. 5. Use a corner detection algorithm to probe within the non-reference area, such as the Harris corner detection algorithm. The corner response function threshold in the Harris corner detection algorithm needs to be set. The corner response function threshold is set based on testing on a large number of sample images in non-reference areas. Adjusting the threshold allows the algorithm to stably detect corners in texture-rich regions while avoiding excessive noise in smooth areas. A feasible method is to calculate the statistical characteristics of the gradient in the non-reference image and set the corner response function threshold as a percentage of the squared gradient magnitude, such as 2%. The corner detection algorithm outputs the coordinates of a set of pixels; these are the selected pixels with significant texture features. To improve computational efficiency and optical flow reliability, the number of selected points is controlled, retaining a few of the strongest responses using non-maximum suppression, for example, retaining 500 points.

[0068] The background optical flow field is calculated by comparing the spatial movement of pixels with significant texture features in adjacent frames. The calculation of the background optical flow field uses the Lucas-Kanade optical flow method. The specific calculation process is as follows: The input consists of two adjacent image frames, i.e., the current frame and the previous frame, and a set of coordinates of pixels with significant texture features selected within the non-reference region of the previous frame. For each pixel in this set, the algorithm defines a small image window around it, such as a square window of size 7 pixels × 7 pixels. Based on the assumptions of constant brightness and spatial consistency, the algorithm estimates the displacement vector of the pixel from the previous frame to the current frame through iterative solutions. The displacement vector is a two-dimensional vector containing a horizontal displacement component u and a vertical displacement component v. The core of the solution is to construct and solve a system of linear equations about u and v, with the coefficients of the system calculated from the spatial gradient within the image window. The algorithm iterates multiple times until the change in the displacement vector is less than a preset iteration termination threshold, such as 0.01 pixels. The iteration termination threshold is set to balance computational accuracy and speed. The convergence of the displacement estimate during iteration is observed experimentally, and convergence is considered achieved when the iteration update is less than 0.01 pixels. To handle potentially large displacements, an image pyramid method is used, which performs optical flow calculations from coarse to fine across multiple image scales. The number of pyramid layers and scaling factor are determined based on the image resolution and the maximum possible displacement estimate; for example, a 3-layer pyramid is constructed, with a scaling factor of 0.5 for each layer. Finally, for each input feature pixel, the algorithm outputs a two-dimensional displacement vector. All successfully calculated displacement vectors and their corresponding previous frame pixel coordinates together constitute a sparse background optical flow field.

[0069] Based on predefined reference point identification information, a semantic role is assigned to each reference point, including more specific physical location mapping rules. Reference points fixed to the cabinet frame are assigned the semantic role of structural fixing points; the cabinet frame refers to the metal frame that constitutes the main supporting structure of the distribution cabinet, such as the connection point between vertical columns and horizontal reinforcing beams. Reference points fixed to busbar connections are assigned the semantic role of structural fixing points; busbar connections refer to the bolt fixing points of the main copper busbars used for transmitting electrical energy within the cabinet. Reference points fixed to stationary insulating components are assigned the semantic role of structural fixing points; stationary insulating components refer to ceramic insulators or epoxy resin partitions that support or enclose live components. Reference points fixed to the circuit breaker operating handle are assigned the semantic role of movable part association points; the circuit breaker operating handle is a mechanical component used for manual opening and closing. Reference points fixed to the disconnector contacts are assigned the semantic role of movable part association points; the disconnector contacts are the movable conductive parts that isolate the circuit. Reference points fixed to removable components are semantically assigned as associated points for movable components. Removable components refer to functional modules, such as drawer-type circuit breakers, that can be pushed in or pulled out along guide rails. These mapping rules are pre-stored as knowledge in the component type classification field of the reference point information database. In practical implementation, the system maintains a list of component types, with each item clearly labeled as belonging to either the fixed structure category or the movable component association category. When assigning semantic roles to reference points, the system queries the component type associated with that reference point and compares it with the component type list to complete the assignment.

[0070] S5. Construct hypotheses on the overall pose change of the sensor and the independent motion of the movable parts, respectively, and verify these hypotheses based on semantic roles and background optical flow fields. The specific implementation is as follows:

[0071] A hypothesis about the overall pose change of the sensor is constructed. Based on the image coordinate changes of all reference points, a global rigid body transformation is calculated, and this transformation is used to predict the optical flow changes that should occur in non-reference areas, thus obtaining the hypothetical optical flow field. The calculation of the global rigid body transformation specifically involves using the image coordinate sets of all reference points at adjacent time points and calculating an optimal affine transformation matrix through least squares fitting. The input is the target point set composed of the image coordinates of all reference points at the current time, and the source point set composed of the image coordinates of these reference points at the previous time. Each reference point corresponds to one source point coordinate and one target point coordinate. The affine transformation matrix is ​​a 3x3 matrix, with the last row fixed at values ​​0, 0, and 1; therefore, six parameters need to be solved. These six parameters define the two-dimensional affine transformation from the source point set to the target point set, including translation, rotation, scaling, and shearing components. The least squares fitting method is used to solve for the six parameters that minimize the sum of squared errors between the coordinates of all reference points after the affine transformation and the coordinates of the target point. The sum of squared errors is calculated as the sum of the squares of the Euclidean distances between the transformed coordinates of each reference point and the target coordinates. The solution process is accomplished by constructing and solving a system of normal equations, which can be stably solved by matrix inversion or singular value decomposition. The resulting affine transformation matrix describes the global coordinate mapping relationship under the assumption of overall sensor pose change. When using this transformation to predict the optical flow changes that should occur in non-reference regions, it is necessary to perform this calculation on each pixel with significant texture features selected during the background optical flow field calculation. For each pixel with significant texture features selected in the non-reference region of the previous frame, its coordinates are used as the source point. The calculated affine transformation matrix is ​​applied to calculate the predicted coordinates of that point in the current frame. The displacement vector between the predicted coordinates and the source point coordinates is the predicted optical flow vector for that point. This process is repeated for all selected pixels to obtain a set of predicted optical flow vectors. These vectors have the same structure as the background optical flow field, i.e., each point corresponds to a two-dimensional displacement vector. The field formed by this set of predicted optical flow vectors is the assumed optical flow field.

[0072] The assumption of independent motion of movable parts is constructed. Local motion is calculated only based on the image coordinate changes of reference points whose semantic role is that of movable part associated points. Combined with the image coordinate invariance of reference points whose semantic role is that of structural fixed points, the optical flow changes that should occur in non-reference areas are predicted, resulting in a competing optical flow field. The specific implementation includes the following steps: Identify all reference points whose semantic role is that of movable part associated points, and all reference points whose semantic role is that of structural fixed points. Based on the image coordinate changes of these reference points at adjacent time points, establish a local motion model for the movable part associated points. The local motion model can adopt an affine transformation model similar to the global rigid body transformation, but only uses the coordinate set of movable part associated points for fitting. Using the image coordinates of all movable part associated points at the current time as the target point set, and the image coordinates of these points at the previous time as the source point set, a local affine transformation matrix is ​​calculated using the least squares method. This local affine transformation matrix describes the overall motion followed by these movable part associated points when it is assumed that only the movable parts move independently. For structural fixed points, under the assumption of independent motion of movable parts, their image coordinates should remain unchanged, i.e., the displacement vector is zero. Next, the optical flow changes in non-reference regions are predicted. Pixels in non-reference regions may or may not be affected by the movement of the movable part. To predict, the spatial relationship between each pixel and the movable part needs to be determined. One approach is to define an influence region centered on each movable part's associated point, such as a circular region with a radius of 100 pixels. The radius of the influence region is set based on estimating the range of background areas that the movable part's movement might affect, which is determined by analyzing the physical dimensions of the movable part and its projection ratio in the image. For pixels with significant texture features selected in the non-reference region of the previous frame, it is checked whether they fall within the influence region of any movable part's associated point. If a pixel falls within the influence region of a movable part's associated point, the local affine transformation matrix corresponding to that movable part is applied to transform the pixel from its coordinates in the previous frame to its predicted coordinates in the current frame, thus obtaining the predicted optical flow vector for that point. If a pixel does not fall within the influence region of any movable part's associated point, under the assumption of independent movement of the movable part, that point should be considered part of the static background, and its predicted optical flow vector should be zero. Perform this judgment and calculation on all selected pixels to obtain a set of predicted optical flow vectors. The field formed by this set of vectors is the competing optical flow field.

[0073] The fit between the hypothetical optical flow field and the competing optical flow field and the background optical flow field is calculated separately. The fit calculation needs to be performed on each pixel in the background optical flow field where an actual optical flow vector was successfully calculated. The background optical flow field provides a set of actually observed displacement vectors, each corresponding to a pixel location. For the hypothetical and competing optical flow fields, they also provide predicted displacement vectors at the same pixel locations. The fit can be quantified using the average endpoint error. The average endpoint error is calculated as follows: for each pixel, the Euclidean distance between the predicted and actual displacement vectors is calculated, and then the average of the Euclidean distances across all pixels is taken. Specifically, the calculation process is as follows: First, ensure the set of pixels being compared is consistent, i.e., only consider pixels that have valid displacement vectors in the background, hypothetical, and competing optical flow fields. For the hypothetical optical flow field, the Euclidean distance between its predicted displacement vector and the actual displacement vector in the background optical flow field is calculated to obtain the error value for each point. The error values ​​of all points are summed and then divided by the total number of points to obtain the average endpoint error of the hypothetical optical flow field. Similarly, for the competing optical flow field, the Euclidean distance between its predicted displacement vector and the actual displacement vector of the background optical flow field is calculated, and the average value is obtained to obtain the average endpoint error of the competing optical flow field. The smaller the average endpoint error, the better the prediction matches the reality. Therefore, the degree of agreement can be defined as the reciprocal of the average endpoint error, or the average endpoint error can be used directly as a measure, in which case the smaller the value, the higher the degree of agreement. In actual comparison, the magnitude of the two average endpoint errors can be directly compared. In addition, to enhance robustness, other degree of agreement measures can be used, such as the inlier ratio. The calculation of the inlier ratio requires setting an error threshold, such as 2 pixels. Points with an error less than this error threshold are considered inliers. The error threshold is set based on the pixel-level error that usually exists in optical flow calculation itself, and is determined by analyzing the error distribution of the optical flow algorithm in a static scene. The inlier ratio is the proportion of the number of inliers to the total number of points. The inlier ratios corresponding to the hypothetical optical flow field and the competing optical flow field can be calculated separately. The higher the inlier ratio, the better the degree of agreement. The specific measure method used needs to be set in advance during implementation and kept consistent throughout the verification process.

[0074] If the fit of the assumed optical flow field is higher, the verification result of the overall sensor pose change hypothesis is valid. If the fit of the competing optical flow field is higher, the verification result of the independent motion hypothesis of the movable part is valid. The specific judgment rules are as follows: If the average endpoint error is used as the fit index, the average endpoint error of the assumed optical flow field and the average endpoint error of the competing optical flow field are compared. If the average endpoint error of the assumed optical flow field is less than the average endpoint error of the competing optical flow field, the assumed optical flow field is considered to have a higher fit, and the verification result of the overall sensor pose change hypothesis is valid. Conversely, if the average endpoint error of the competing optical flow field is less than the average endpoint error of the assumed optical flow field, the competing optical flow field is considered to have a higher fit, and the verification result of the independent motion hypothesis of the movable part is valid. If the interior point ratio is used as the fit index, the interior point ratio of the assumed optical flow field and the competing optical flow field are compared. If the interior point ratio of the assumed optical flow field is greater than the interior point ratio of the competing optical flow field, the assumed optical flow field is considered to have a higher fit, and the verification result of the overall sensor pose change hypothesis is valid. Conversely, if the proportion of interior points in the competing optical flow field is greater than that in the assumed optical flow field, the matching degree of the competing optical flow field is considered higher, and the verification result of the independent motion assumption of the movable part is valid. During comparison, if the two matching degree indices are very close, for example, the difference in average endpoint error is less than a preset tolerance threshold (which can be set to, for example, 0.1 pixels), or the difference in interior point proportion is less than, for example, 5%, then it can be considered that a clear judgment cannot be made. In this case, the decision can be delayed and the accumulated information of subsequent frames can be waited for, or an uncertain state processing procedure can be triggered, such as issuing a prompt requiring manual verification. Otherwise, the verification result is clearly given according to the above rules. The verification result will be used to guide the decision in step S6, determining whether to perform coordinate compensation.

[0075] Calculating a global rigid body transformation based on the image coordinate changes of all reference points involves: using the set of image coordinates of all reference points at adjacent time points, and fitting an optimal affine transformation matrix using the least squares method. The affine transformation matrix describes the global coordinate mapping relationship under the assumption of overall sensor pose change. The specific fitting process is as follows: The image coordinates of all valid reference points at the previous time step are taken as the source point set, and the image coordinates of the reference points at the current time step are taken as the target point set. The two-dimensional affine transformation model describes a linear mapping relationship from the source point coordinates to the target point coordinates. This mapping relationship is determined by six transformation parameters to be determined, which together constitute the core elements of the affine transformation matrix. The purpose of fitting using the least squares method is to find a set of optimal transformation parameters such that after performing an affine transformation on the source point set according to this set of parameters, the overall error between the predicted point set and the true target point set is minimized. The overall error is usually measured by the sum of the squares of the positional deviations between all corresponding point pairs. This fitting problem is reduced to a standard least squares optimization problem. Mathematically, solving this problem is equivalent to solving two systems of linear equations based on point coordinate data. One system solves for three parameters related to the horizontal coordinate transformation, and the other solves for three additional parameters related to the vertical coordinate transformation. These two systems of linear equations are directly derived from the source and target coordinate data of all valid reference points. Solving such systems of linear equations is a well-known technique in the field and can be achieved using standard linear algebraic methods, such as normal equation solving in matrix operations, or robust matrix decomposition methods in numerical computation, such as QR decomposition. The six transformation parameters obtained from the final solution determine a specific affine transformation matrix, which is the optimal global rigid body transformation obtained through fitting.

[0076] S6. When the verification result confirms the assumption of overall sensor pose change, coordinate compensation is performed based on the image coordinate changes of multiple reference points, and then the power distribution cabinet status is identified based on the compensated image; when the verification result confirms the assumption of independent movement of movable parts, the power distribution cabinet status is directly identified based on the monitoring image sequence, specifically as follows:

[0077] If the verification result confirms the assumption of overall sensor pose change, a coordinate transformation relationship from the current image coordinate system to the reference image coordinate system is calculated based on the image coordinate changes of multiple reference points. This coordinate transformation relationship is then used to perform geometric correction on the monitoring image sequence to obtain a compensated image. Subsequently, the status of the power distribution cabinet is identified by analyzing component positions and instrument readings on the compensated image. The input for calculating the coordinate transformation relationship is the image coordinates of all valid reference points used for verification in step S5 at the current moment, as well as the pre-stored reference coordinates of these reference points in the reference image coordinate system. The reference image coordinate system is a stable image coordinate system established during the initial calibration phase. The reference image coordinate system is typically selected from the coordinate system corresponding to the first image acquired after system installation and debugging, when the relative pose relationship between the vision sensor and the power distribution cabinet is confirmed to be correct. The reference coordinates of each reference point in the reference image coordinate system are obtained by manually marking on the images acquired during the initial calibration phase or by automatically extracting and recording feature points. The obtained reference coordinates are stored along with the unique number of the reference point. Calculating the coordinate transformation relationship involves solving a two-dimensional spatial transformation model that maps the image coordinates of the reference point at the current moment to its corresponding reference coordinates as accurately as possible. The two-dimensional spatial transformation model used is a two-dimensional affine transformation. By fitting using the least squares method, the six parameters of this two-dimensional affine transformation can be solved. The specific fitting process is consistent with the method for calculating the global rigid body transformation in step S5, but here the target point set is the reference coordinates of the reference points. The obtained affine transformation matrix defines the coordinate transformation relationship from the current image coordinate system to the reference image coordinate system.

[0078] Geometric correction is performed on the monitoring image sequence using coordinate transformation relationships to obtain a compensated image. Geometric correction involves applying a calculated affine transformation matrix to the current frame of the monitoring image sequence to eliminate geometric distortions caused by sensor pose changes. The specific implementation process involves image resampling. Based on the affine transformation matrix, the theoretical position of each pixel in the reference image coordinate system is calculated for each pixel in the current frame. Since the transformed pixel position may be a non-integer coordinate, an interpolation algorithm is needed to determine the pixel's grayscale or color value at that position. The interpolation algorithm used is bilinear interpolation. Bilinear interpolation uses the grayscale values ​​of the four nearest neighbor pixels around the target point and calculates the grayscale value of the target point through two linear interpolations. After this resampling calculation is completed for all pixels in the current frame, a new image is generated. The coordinate system of this new image is aligned with the reference image coordinate system; this new image is the compensated image. In the compensated image, the visual appearance of the distribution cabinet should be consistent with the initial calibration state.

[0079] The status of the distribution cabinet is then identified by analyzing component positions and instrument readings on the compensated image. Component position analysis focuses on key movable or indicating components within the distribution cabinet, such as the circuit breaker's open / close indicator, the location of the grounding switch, and the orientation of the operating handle. For each component to be monitored, a region of interest is predefined in the reference image coordinate system. The region of interest is defined based on the possible location and range of the component in the reference image, and is typically a rectangular bounding box. Sub-images within the predefined region of interest are extracted from the compensated image. An image template matching algorithm is used to compare this sub-image with a pre-stored template image representing the component's normal state. The template image is an image patch of the component in its standard state, acquired in the reference image coordinate system during system initialization. The image template matching algorithm calculates the similarity between the sub-image and the template image, using a similarity metric including the normalized cross-correlation coefficient. A similarity threshold is set, taking into account minor changes in illumination and image noise. A large number of image patches of the component in its normal state are collected and matched with a template image. The distribution of similarity values ​​is statistically calculated, and a critical value that can reliably distinguish between correct and incorrect matches is chosen as the similarity threshold, for example, 0.85. If the calculated similarity is higher than the threshold, the component is determined to be in a normal position; if the calculated similarity is lower than the threshold, the component may be in an abnormal position, triggering an alert. Instrument reading analysis mainly targets pointer or digital instruments. For pointer instruments, the instrument dial area and pointer rotation center are predefined in the reference image coordinate system. On the compensated image, the instrument dial area is extracted, and the pointer line is detected using an image processing algorithm. The image processing algorithm used for detecting the pointer line is the Hough transform. The angle of the pointer line relative to the predefined zero-point baseline is calculated. Based on the calibration relationship between the angle and the instrument range, the current reading is obtained. The calibration relationship between the angle and the instrument range is established during system initialization by recording the pointer angle at multiple known readings. For digital instruments, a predefined digital display area is established within a reference image coordinate system. This area is then binarized and segmented into characters, and optical character recognition (OCR) technology is used to identify the numerical strings, thus obtaining the readings. After analyzing all predefined component locations and instrument readings, the results are synthesized to form the final distribution cabinet status identification conclusion.

[0080] If the verification result confirms the assumption of independent movement of movable parts, the distribution cabinet status can be identified directly by analyzing the component positions and instrument readings on the current frame of the monitoring image sequence. Since the verification result indicates that the image changes originate from the movement of the movable parts within the cabinet, rather than sensor drift, geometric correction is unnecessary, and the original acquired current frame image can be used directly for analysis. The method for analyzing component positions and instrument readings is the same as the principle used for analysis on the compensated image, but all predefined areas of interest, template images, and dial areas must be based on a unified dynamic coordinate system corresponding to the current monitoring image sequence. The dynamic coordinate system is typically established using the coordinate system established by the first frame of monitoring image acquired after system startup. In actual implementation, component position analysis also uses image template matching, but the template image used is a standard state image block of the component acquired in the dynamic coordinate system. During instrument reading analysis, the coordinates used to locate the dial and character areas are also predefined in the dynamic coordinate system. When performing matching or recognition directly on the current frame image, the algorithm needs to have a certain scale or slight deformation robustness. This can be achieved by using image features that are insensitive to scale changes for template matching, such as using SIFT features for matching, or by normalizing the size of the character image before optical character recognition. Finally, based on the results of direct analysis, the status of the power distribution cabinet is comprehensively judged.

[0081] All calculations involved in the embodiments are dimensionless numerical calculations, and the preset parameters and thresholds in the calculations are set by those skilled in the art according to the actual situation.

[0082] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, in the form of a computer program product.

[0083] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and inventive constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0084] In addition, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.

[0085] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.

[0086] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0087] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for real-time monitoring of the status of a power distribution cabinet based on artificial intelligence, characterized in that, Includes the following steps: S1. Continuously acquire monitoring image sequences of the power distribution cabinet through a visual sensor fixed at the monitoring point; S2. Extract the image coordinates of multiple preset reference points on the power distribution cabinet from the monitoring image sequence; S3. Based on the change pattern of image coordinates at multiple reference points at adjacent time points, determine whether the change pattern is more likely to be caused by the overall pose change of the sensor. S4. When this is the case, analyze the semantic role of each reference point among multiple reference points, and calculate the background optical flow field of non-reference areas in the monitoring image sequence; S5. Construct the overall pose change hypothesis of the sensor and the independent motion hypothesis of the movable parts respectively, and verify the overall pose change hypothesis of the sensor and the independent motion hypothesis of the movable parts based on semantic roles and background optical flow field. S6. When the verification result confirms the assumption of overall sensor pose change, coordinate compensation is performed based on the image coordinate changes of multiple reference points, and then the power distribution cabinet status is identified based on the compensated image; when the verification result confirms the assumption of independent movement of movable parts, the power distribution cabinet status is identified directly based on the monitoring image sequence.

2. The method for real-time monitoring of the status of a power distribution cabinet based on artificial intelligence according to claim 1, characterized in that, The monitoring image sequence of the power distribution cabinet is continuously acquired by a visual sensor fixed at the monitoring point, including: Images of the inside of the power distribution cabinet are acquired at a preset cycle under visible light conditions; And when the ambient light is insufficient, the supplementary lighting device is activated and images of the inside of the power distribution cabinet under supplementary lighting conditions are collected, thereby obtaining a monitoring image sequence.

3. The method for real-time monitoring of the status of a power distribution cabinet based on artificial intelligence according to claim 1, characterized in that, Extract the image coordinates of multiple preset reference points on the distribution cabinet from the monitoring image sequence, including: In the first frame of the monitored image sequence, the initial image coordinates of each preset reference point are identified and recorded using a feature point detection algorithm; For each subsequent frame in the monitored image sequence, the feature points detected in the current frame are matched with the feature points corresponding to the recorded initial image coordinates using the feature descriptor matching method, so as to track and obtain the image coordinates of each reference point in the current frame. The correctness of the current frame image coordinates obtained through matching and tracking is verified using the principle of multi-view geometric constraints. Incorrect matching and tracking results are eliminated, thereby extracting the effective image coordinates of each reference point in each frame.

4. The method for real-time monitoring of the status of a power distribution cabinet based on artificial intelligence according to claim 3, characterized in that, The setting and maintenance of multiple preset reference points include: selecting physical feature points with high texture contrast on the surface of the power distribution cabinet and key internal components as candidate reference points; uniquely identifying the candidate reference points and initially calibrating them with the cabinet's physical coordinate system; and in subsequent monitoring, if the image coordinates of a certain reference point cannot be stably tracked or extracted in multiple consecutive frames of images, then using a backup candidate reference point to replace it and updating the preset reference point information.

5. The method for real-time monitoring of the status of a power distribution cabinet based on artificial intelligence according to claim 1, characterized in that, Based on the change patterns of image coordinates at multiple reference points at adjacent time points, determine whether the change patterns are more likely to be caused by changes in the overall sensor pose, including: Based on the changes in the image coordinates of multiple reference points at adjacent time points, calculate a pose transformation model that describes the overall rigid body motion of all reference points. Calculate the residual between the coordinates of each reference point after transformation according to the pose transformation model and its actual image coordinates at adjacent time points; Analyzing the distribution characteristics of the residuals, if the residuals are generally small and uniformly distributed, the change pattern is more likely to be caused by changes in the overall pose of the sensor. If the residuals exhibit a localized increase in clustering associated with a specific physical component, then the change pattern is more likely caused by the independent movement of the movable component.

6. The method for real-time monitoring of the status of a power distribution cabinet based on artificial intelligence according to claim 1, characterized in that, At that time, the semantic role of each of the multiple reference points is analyzed, and the background optical flow field of non-reference areas in the monitored image sequence is calculated, including: Based on predefined reference point identification information, each reference point is assigned a semantic role that represents it as a structural fixed point or a point associated with a movable part. Between adjacent frames of the monitored image sequence, pixels with significant texture features are selected from the non-reference area after excluding all reference points from the preset neighborhood. The background optical flow field is calculated by comparing the spatial position movement of pixels with significant texture features in adjacent frames.

7. The method for real-time monitoring of the status of a power distribution cabinet based on artificial intelligence according to claim 6, characterized in that, Based on the predefined reference point identification information, a semantic role is assigned to each reference point, including: reference points fixed to the cabinet frame of the distribution cabinet, busbar connection, or static insulating parts are assigned the semantic role of structural fixing points; reference points fixed to the circuit breaker operating handle, disconnector contact, or withdrawable parts are assigned the semantic role of movable part association points.

8. The method for real-time monitoring of the status of a power distribution cabinet based on artificial intelligence according to claim 1, characterized in that, Hypotheses regarding the overall pose change of the sensor and the independent motion of its moving parts are constructed separately. These hypotheses are then validated based on semantic roles and the background optical flow field, including: A hypothesis about the overall pose change of the sensor is constructed. A global rigid body transformation is calculated based on the image coordinate changes of all reference points. This transformation is then used to predict the optical flow changes that should occur in non-reference areas, thus obtaining the hypothetical optical flow field. The assumption of independent motion of movable parts is constructed. Local motion is calculated based only on the image coordinate changes of reference points whose semantic role is the associated point of the movable part. Combined with the image coordinate invariance of reference points whose semantic role is the fixed point of the structure, the optical flow changes that should be generated in non-reference areas are predicted to obtain the competing optical flow field. Calculate the degree of agreement between the assumed optical flow field and the competing optical flow field and the background optical flow field, respectively; If the fit of the optical flow field is higher, the verification result of the overall sensor pose change hypothesis is valid; if the fit of the competing optical flow field is higher, the verification result of the independent motion hypothesis of the movable part is valid.

9. The method for real-time monitoring of the status of a power distribution cabinet based on artificial intelligence according to claim 8, characterized in that, Calculating a global rigid body transformation based on the image coordinate changes of all reference points involves: using the set of image coordinates of all reference points at adjacent time points, and fitting an optimal affine transformation matrix using the least squares method. The affine transformation matrix describes the global coordinate mapping relationship under the assumption of the overall pose change of the sensor.

10. The method for real-time monitoring of the status of a power distribution cabinet based on artificial intelligence according to claim 1, characterized in that, When the verification result confirms the assumption of overall sensor pose change, coordinate compensation is performed based on the image coordinate changes of multiple reference points, and then the power distribution cabinet status is identified based on the compensated image; when the verification result confirms the assumption of independent movement of movable parts, the power distribution cabinet status is directly identified based on the monitored image sequence, including: If the verification result confirms the assumption of overall sensor pose change, then a coordinate transformation relationship from the current image coordinate system to the reference image coordinate system is calculated based on the image coordinate changes of multiple reference points. The coordinate transformation relationship is then used to perform geometric correction on the monitoring image sequence to obtain the compensated image. After that, the power distribution cabinet status is identified by analyzing the component positions and instrument readings on the compensated image. If the verification result confirms the assumption that the movable parts move independently, the status of the distribution cabinet can be identified directly by analyzing the position of the parts and the instrument readings on the current frame of the monitoring image sequence.