Power distribution cabinet intelligent operation and maintenance full life cycle data processing method and system

By constructing a multi-dimensional data sampling window and topological relationship, and combining a hybrid storage architecture of time-series database and unstructured data management engine, the problem of correlation of multi-source heterogeneous data in power distribution cabinets is solved, and the accurate correction and standardized encapsulation of data are realized, thereby improving the operational reliability and stability of power distribution cabinets.

CN121901707APending Publication Date: 2026-04-21SHANDONG RONGQING INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANDONG RONGQING INFORMATION TECH CO LTD
Filing Date
2026-01-09
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively characterize the inherent relationships between parameters under equipment operating conditions when processing multi-source heterogeneous data from power distribution cabinets, resulting in insufficient accuracy and timeliness in condition assessment and fault prediction.

Method used

A multi-dimensional data sampling window is constructed, the topological relationship between core feature parameters is established, the data is standardized and encapsulated through data quality calibration parameters, and a hybrid storage architecture combining time-series database and unstructured data management engine is adopted to realize hierarchical, compressed and associated storage of data, and to build a digital twin of the power distribution cabinet for health status assessment and fault prediction.

Benefits of technology

It improves the accuracy, consistency and availability of data, optimizes the efficiency of storage resource utilization, reduces storage costs, enables early prediction and proactive intervention of the operating status of power distribution cabinets, and enhances the reliability and stability of equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121901707A_ABST
    Figure CN121901707A_ABST
Patent Text Reader

Abstract

The invention provides a power distribution cabinet intelligent operation and maintenance full-life-cycle data processing method and system, and relates to the technical field of data processing, and the method comprises the steps: 2, building a multi-dimensional data sampling window based on key state parameters in an original full-life-cycle data set, and building a topological relation between core characteristic parameters in the sampling window; step 3, based on the topological relation, performing structured topological partitioning on the data in the sampling window, and calculating the spatial shape feature similarity of data distribution of each partition; and generating a data quality calibration parameter according to the data distribution characteristics, the association strength and the shape consistency index of each partition. Real-time sensing and intelligent fault prediction of the operation state of the power distribution cabinet are realized, and the operation and maintenance efficiency and the data full-life-cycle management level are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method and system for intelligent operation and maintenance of power distribution cabinets throughout their entire lifecycle. Background Technology

[0002] With the application of intelligent technology in the field of power equipment operation and maintenance, the operation and maintenance mode of distribution cabinets is gradually shifting from periodic inspection and post-maintenance to condition monitoring and predictive maintenance. Most existing intelligent operation and maintenance solutions rely on the deployment of sensors and monitoring devices to collect various status parameters of distribution cabinets, such as current, voltage, temperature, and vibration. However, in actual application, due to the diverse sources and structures of monitoring data, as well as the uneven data quality, existing technologies face some limitations in processing these multi-source heterogeneous full life cycle data.

[0003] For example, when attempting to conduct long-term health status assessments and early fault warnings for a distribution cabinet, existing methods mostly process and analyze current signals collected by different sensors, cabinet temperature distribution data, partial discharge data, and switch action records as independent data streams. This approach may struggle to effectively characterize the inherent and dynamic correlations between these parameters under actual equipment operating conditions. Due to the lack of in-depth exploration and structured representation of the collaborative changes among multiple parameters under specific operating conditions or time windows, relying solely on the judgment of single parameter thresholds or simple data fusion can sometimes lead to misjudgments of the overall equipment status. Existing methods may be insufficient in capturing such cross-parameter, time-series correlation characteristics, which could affect the accuracy of status assessments and the timeliness of fault predictions. Summary of the Invention

[0004] The technical problem to be solved by this invention is to provide a method and system for intelligent operation and maintenance of power distribution cabinets throughout their entire lifecycle, so as to realize real-time perception of the operating status of power distribution cabinets and intelligent fault prediction, thereby improving operation and maintenance efficiency and the level of data lifecycle management.

[0005] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows: Firstly, a method for intelligent operation and maintenance of power distribution cabinets throughout their entire lifecycle data processing, the method comprising: Step 1: Collect multi-source heterogeneous data from the power distribution cabinet to form the original full lifecycle dataset; Step 2: Based on the key state parameters in the original full lifecycle dataset, construct a multi-dimensional data sampling window and establish the topological relationship between the core feature parameters within the sampling window; Step 3: Based on topological relationships, perform structured topological partitioning of the data within the sampling window, calculate the spatial shape feature similarity of the data distribution in each partition, and generate data quality calibration parameters based on the data distribution characteristics, correlation strength, and shape consistency index of each partition. Step 4: Using data quality calibration parameters, perform data transformation and standardization encapsulation on the original full lifecycle dataset to obtain a standardized transmission dataset; Step 5: Upload the standardized transmission dataset through a high-speed communication network. Based on the time-series characteristics and value density of the standardized transmission dataset, adopt a hybrid storage architecture that combines a time-series database and an unstructured data management engine on the central data processing side to achieve data hierarchical, compressed, and associated storage, forming a full lifecycle data asset. Step 6: Based on the full lifecycle data assets, construct a digital twin of the distribution cabinet, and use data analysis models to conduct health status assessment and fault prediction analysis on the digital twin to generate intelligent operation and maintenance decisions. Secondly, the intelligent operation and maintenance full lifecycle data processing system for distribution cabinets includes: The data acquisition module is used to collect multi-source heterogeneous data from the power distribution cabinet to form a raw full lifecycle dataset. The building module is used to construct a multi-dimensional data sampling window based on the key state parameters in the original full life cycle dataset, and to establish the topological relationship between the core feature parameters within the sampling window; The generation module is used to perform structured topological partitioning of the data within the sampling window based on topological relationships, calculate the spatial shape feature similarity of the data distribution in each partition, and generate data quality calibration parameters based on the data distribution characteristics, correlation strength, and shape consistency index of each partition. The encapsulation module is used to transform and standardize the original full lifecycle dataset using data quality calibration parameters to obtain a standardized transmission dataset. The management module is used to upload standardized transmission datasets through a high-speed communication network. Based on the time-series characteristics and value density of the standardized transmission datasets, a hybrid storage architecture combining a time-series database and an unstructured data management engine is adopted on the central data processing side to achieve data hierarchical, compression, and associated storage, forming full lifecycle data assets. The decision-making module is used to construct a digital twin of the power distribution cabinet based on the full life cycle data assets, and to use data analysis models to conduct health status assessment and fault prediction analysis on the digital twin to generate intelligent operation and maintenance decisions.

[0006] The above-described solution of the present invention has at least the following beneficial effects: By constructing multi-dimensional data sampling windows and topological relationships between core feature parameters, the inherent correlation patterns between different types and stages of data are effectively mined. Based on the structured topological partitioning and data distribution feature analysis of topological relationships, the distribution differences, correlation strength, and shape consistency of data in each partition are identified. The generated data quality calibration parameters can accurately correct and standardize the original data, improving the accuracy, consistency, and usability of the data. A hybrid storage architecture combining time-series databases and unstructured data management engines is adopted. Combining the time-series characteristics and value density of data, the hierarchical, compressed, and associated storage of data is achieved. This not only adapts to the diverse characteristics of data throughout its entire lifecycle, optimizes the utilization efficiency of storage resources, and reduces the overall cost of data storage, but also enables the orderly management and rapid retrieval of data assets, ensuring the long-term traceability and reusability of data assets. By constructing a digital twin of the power distribution cabinet and conducting health status assessment and fault prediction analysis, the entire lifecycle data assets can be transformed into the core basis for intelligent operation and maintenance decisions, enabling early prediction and proactive intervention of the power distribution cabinet's operating status, and improving the reliability and stability of the power distribution cabinet's operation. Attached Figure Description

[0007] Figure 1 This is a flowchart illustrating the intelligent operation and maintenance full lifecycle data processing method for power distribution cabinets provided in an embodiment of the present invention.

[0008] Figure 2 This is a schematic diagram of the intelligent operation and maintenance full life cycle data processing system for power distribution cabinets provided in an embodiment of the present invention. Detailed Implementation

[0009] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0010] like Figure 1 As shown in the figure, an embodiment of the present invention proposes a method for intelligent operation and maintenance of power distribution cabinets throughout their entire lifecycle data processing. The method includes the following steps: Step 1: Collect multi-source heterogeneous data from the power distribution cabinet to form the original full lifecycle dataset; Step 2: Based on the key state parameters in the original full lifecycle dataset, construct a multi-dimensional data sampling window and establish the topological relationship between the core feature parameters within the sampling window; Step 3: Based on topological relationships, perform structured topological partitioning of the data within the sampling window, calculate the spatial shape feature similarity of the data distribution in each partition, and generate data quality calibration parameters based on the data distribution characteristics, correlation strength, and shape consistency index of each partition. Step 4: Using data quality calibration parameters, perform data transformation and standardization encapsulation on the original full lifecycle dataset to obtain a standardized transmission dataset; Step 5: Upload the standardized transmission dataset through a high-speed communication network. Based on the time-series characteristics and value density of the standardized transmission dataset, adopt a hybrid storage architecture that combines a time-series database and an unstructured data management engine on the central data processing side to achieve data hierarchical, compressed, and associated storage, forming a full lifecycle data asset. Step 6: Based on the full lifecycle data assets, construct a digital twin of the power distribution cabinet, and use data analysis models to conduct health status assessment and fault prediction analysis on the digital twin to generate intelligent operation and maintenance decisions.

[0011] In this embodiment of the invention, by constructing a multi-dimensional data sampling window and the topological relationship between core feature parameters, the inherent correlation patterns between different types and stages of data are effectively mined. Based on the structured topological partitioning and data distribution feature analysis of the topological relationship, the distribution differences, correlation strength, and shape consistency of data in each partition are identified. The generated data quality calibration parameters can achieve accurate correction and standardized encapsulation of the original data, improving the accuracy, consistency, and usability of the data. A hybrid storage architecture combining a time-series database and an unstructured data management engine is adopted. Combining the time-series characteristics and value density of the data, the hierarchical, compressed, and associated storage of the data is realized. This not only adapts to the diverse characteristics of data throughout the entire life cycle, optimizes the utilization efficiency of storage resources, and reduces the overall cost of data storage, but also enables the orderly management and rapid retrieval of data assets, ensuring the long-term traceability and reusability of data assets. By constructing a digital twin of the power distribution cabinet and conducting health status assessment and fault prediction analysis, the data assets throughout the entire life cycle can be transformed into the core basis for intelligent operation and maintenance decisions, enabling early prediction and proactive intervention of the power distribution cabinet's operating status, and improving the reliability and stability of the power distribution cabinet's operation.

[0012] In a preferred embodiment of the present invention, step 1 above, which involves collecting multi-source heterogeneous data from the power distribution cabinet to form an original full lifecycle dataset, may include: In this embodiment of the invention, the coverage scope of the power distribution cabinet's full lifecycle data is first determined, encompassing the complete stages from factory commissioning, normal operation, maintenance and repair to decommissioning. Based on the power distribution cabinet's electrical structure, such as main circuits, branch circuits, key components, and operational requirements, the types of multi-source data to be collected are determined, specifically including electrical performance data, mechanical condition data, environmental condition data, and operation and maintenance-related text data. Subsequently, suitable sensing and monitoring equipment is selected according to the data collection requirements for each type of data. For electrical performance data, current sensors, voltage sensors, power transmitters, and harmonic analyzers are used; for mechanical condition data, vibration sensors, displacement sensors, and contact temperature sensors are used, installed at busbar joints, switch contacts, and other connection points; for environmental condition data, temperature and humidity sensors, dust... Dust concentration sensors are deployed inside the distribution cabinet and in the surrounding environment. A data acquisition terminal is also provided to ensure compatibility with various sensors and support multi-protocol data access. During deployment, the equipment installation positions are rationally planned according to the electrical circuit distribution, component locations, and environmental monitoring points to ensure precise alignment between sensors and monitored objects. For example, vibration sensors are attached to the circuit breaker housing, and temperature sensors are aligned with key connection points. The acquisition terminal is fixed in a safe area inside the cabinet to ensure stable data transmission. Electrical performance data acquisition involves real-time collection of parameters such as current, voltage, active / reactive power, and harmonic content during the operation of the distribution cabinet using sensors. The acquisition frequency is set according to the required maintenance accuracy, typically 1-5 seconds per acquisition under normal operating conditions, increasing to 0 seconds per acquisition when there are large load fluctuations or equipment malfunctions.1-1 second / time to ensure the dynamic change trajectory of captured parameters; Mechanical status data acquisition: Vibration sensors continuously collect vibration intensity, vibration frequency, and displacement of the cabinet and key components; temperature sensors monitor temperature changes at connection points in real time; the acquisition frequency is synchronized with electrical performance data, focusing on capturing mechanical status fluctuations under scenarios such as switch actions and sudden load changes; Environmental condition data acquisition: Temperature, humidity, and dust concentration sensors continuously monitor temperature, relative humidity, and dust concentration data inside and around the distribution cabinet 24 hours a day, with a acquisition frequency of 10-30 seconds / time, recording long-term trends of environmental parameters; Maintenance text data acquisition: Collect text data such as equipment factory information, maintenance records, alarm event logs, and fault diagnosis reports through manual input or system integration. Data is entered in real-time upon occurrence to ensure its timeliness and completeness. All data acquired by the collection devices is aggregated through the collection terminal. First, single-source data undergoes preliminary verification to remove obviously invalid data, such as negative values ​​caused by sensor malfunctions or outliers exceeding reasonable ranges. Then, it is labeled according to unified rules, adding a unique identifier to each data entry, including the distribution cabinet number, collection device ID, and collection timestamp, to determine the corresponding collection location and data type. Subsequently, data from different sources and with different structures are categorized and integrated according to timestamps and device identifiers to construct the basic framework of the original dataset. Ultimately, this forms a raw, full-lifecycle dataset covering the entire lifecycle of the distribution cabinet and containing multiple types of heterogeneous data. The dataset must retain the original data format and collection information completely.

[0013] In a preferred embodiment of the present invention, step 2 above, which involves constructing a multi-dimensional data sampling window based on the key state parameters in the original full lifecycle dataset and establishing the topological relationship between core feature parameters within the sampling window, may include: In this embodiment of the invention, step 220 involves extracting key state parameters reflecting the electrical performance, mechanical condition, and environmental conditions of the distribution cabinet from the original full lifecycle dataset, thus forming a set of key state parameters. Specifically, this includes determining the criteria for selecting key state parameters. First, the core operational requirements and common causes of failures of the distribution cabinet are analyzed. Then, relevant power industry operation and maintenance standards, such as the safe operation procedures for distribution cabinets and the rated technical parameters of the equipment, such as rated current, rated voltage, and allowable temperature rise range, are considered to determine the selection principles. Only parameters that directly characterize the operational status assessment and fault early warning of the distribution cabinet and reflect the core performance of the equipment or environmental influencing factors are extracted, while parameters that exclude others are excluded. Redundant and irrelevant auxiliary data, such as the sensor's own power supply voltage and other parameters not directly related to the device status, are excluded. Three categories of key status parameters are extracted: electrical performance key parameters are extracted by filtering four core parameters—current, voltage, power, and harmonics—from the electrical data collected in the original full lifecycle dataset. Specifically, current parameters are extracted from real-time operating current data of each circuit; voltage parameters are extracted from bus and branch circuit voltage data; power parameters are extracted from active and reactive power data; and harmonic parameters are extracted from total harmonic distortion (THD) of voltage, total harmonic distortion (THD) of current, and the content of each harmonic. During extraction, the corresponding acquisition location must be considered, such as main circuit or branch circuit. To ensure parameters correspond to the equipment's electrical structure, key mechanical parameters are extracted from the original dataset's mechanical data, including vibration intensity, displacement, and connection point temperature. Vibration intensity parameters are extracted from the cabinet and key components, such as circuit breakers and disconnectors, based on vibration acceleration data. Displacement parameters are extracted from the displacement data of the switch mechanism. Connection point temperature parameters are extracted from the real-time temperature data of key connection points such as busbar joints, switch contacts, and terminal blocks. During extraction, the specific mechanical components corresponding to the parameters must be labeled to ensure compatibility with the equipment's physical structure. Key environmental parameters are also extracted from the original dataset's environmental data, including ambient temperature... For the two types of parameters, humidity and dust concentration, the environmental temperature and relative humidity data of the internal cavity and surrounding environment of the distribution cabinet are extracted, and the dust concentration data of the internal dust content of the distribution cabinet are extracted. During extraction, it is necessary to associate the environmental location information of the collection point, such as the top of the cabinet and the surrounding area outside the cabinet. A set of key state parameters is constructed by summarizing all the key state parameters extracted above and organizing them according to the classification dimensions of electrical performance, mechanical state and environmental conditions. A unique identifier is added to each parameter, and the identifier includes the parameter type, collection location and corresponding equipment component, forming a structured set of key state parameters. At the same time, the association index between each parameter and the original dataset is retained.

[0014] Step 221: Based on the set of key state parameters, synchronous sampling and organization of key state parameters are performed using a preset fixed time interval as the time base and an alarm event exceeding a threshold as the event trigger base, respectively, to construct a synchronous multi-dimensional data sampling window. Specifically, this includes: setting dual-baseline sampling parameters and setting a preset fixed time interval. Combining the distribution cabinet's operation and maintenance accuracy requirements, the changing characteristics of key parameters, and data redundancy control objectives, the fixed time interval is determined through trial operation verification. Under normal operating conditions, when the load rate is below 60% of the rated load, the time interval is set to 5 seconds, i.e., full parameter synchronous sampling is triggered every 5 seconds. When the distribution cabinet is under high load conditions (60%-80% load rate), the interval is adjusted to 2 seconds; when under heavy load conditions (load rate exceeding 80%), the interval is adjusted to 1 second. After setting the interval, the timed triggering module of the sampling control system is entered. The system also sets linkage conditions for interval adjustments, such as automatically triggering interval switching by judging the load rate in real time through the main circuit current parameter; alarm event threshold settings are implemented, with two levels of thresholds set for each parameter in the key status parameter set, combined with the equipment's rated parameters, industry safety standards, and historical fault data; for example, if the main circuit rated current is 630A, the current warning threshold is set to 567A (90% of the rated current) and the alarm threshold is set to 693A (1.1 times the rated current); if the allowable temperature rise at the circuit breaker connection point is 40K (the maximum allowable temperature is 65℃ when the ambient temperature is 25℃), the temperature warning threshold is set to 55℃ and the alarm threshold is set to 65℃; the ambient humidity warning threshold is 75%RH and the alarm threshold is 85%RH. All thresholds are entered into the threshold judgment mode of the sampling control system, and the thresholds are reviewed to ensure their rationality and safety.

[0015] Dual-reference synchronous sampling is implemented. Time-reference synchronous sampling involves the sampling control system sending synchronous acquisition commands to all sensors collecting key parameters at fixed time intervals. These commands include a sampling timestamp, parameter identifier, and acquisition accuracy requirements. Upon receiving the command, each sensor initiates data acquisition at a unified timestamp. After acquisition, each sensor transmits data back to the sampling control system in real time. The system organizes the data in a unified timestamp + parameter identifier + acquired value + sensor status format, forming multi-dimensional parameter data segments for a single time period. Each segment contains complete data for all key parameters at that time point. Event-triggered reference synchronous sampling involves the sampling control system receiving the acquired data for each key parameter in real time. The system compares the data one by one using a threshold judgment mode. When the value of any parameter exceeds its set alarm threshold, an alarm event is immediately triggered. The system records the precise timestamp of the alarm event, the parameter that triggered the alarm, and its value. The sampling period is determined by extending forward 10 seconds from the reference time point to cover the parameter changes before the alarm and backward 30 seconds to cover the parameter fluctuation trend after the alarm. The sampling frequency of this period is also recorded. The sampling rate is increased to 0.5 seconds / time. The system sends high-frequency synchronous acquisition commands to all sensors, marking alarm-related sampling identifiers in the commands to ensure that the sensors collect data synchronously at high-frequency intervals. After acquisition, the system organizes all parameter data within the time period into alarm event ID + unified timestamp + parameter identifier + acquired value format to form alarm-related multi-dimensional parameter data segments, and marks the alarm type corresponding to the segment, such as current overload alarm or temperature overheat alarm. A synchronous multi-dimensional data sampling window is constructed to integrate and deduplicate the data segments obtained from the above two benchmarks. If a time period is covered by both time benchmark and event trigger benchmark, the high-frequency data segment of the event trigger benchmark is retained. Each data segment is an independent synchronous multi-dimensional data sampling window, and the window naming format is sampling benchmark-start timestamp-end timestamp. Each window contains complete data of all key status parameters within the same time period, and all data has a unified millisecond-level timestamp synchronization identifier to ensure the time consistency and integrity of multi-parameter data within the window, ultimately forming a set of synchronous multi-dimensional data sampling windows covering all scenarios such as normal operation, load fluctuation, and alarm anomalies.

[0016] Step 222: Within the synchronous multi-dimensional data sampling window, based on the electrical principles and physical layout information of the distribution cabinet, determine the electrical connection logic and physical location association between parameters; simultaneously calculate the time-domain and frequency-domain correlation of the time series of each parameter within the window; comprehensively consider the electrical connection logic, physical location association, and time-series correlation to generate a directed or undirected topology with key state parameters as nodes and the mutual influence and constraint relationships between parameters as edges; specifically including: sorting out the basic information of the distribution cabinet and determining the parameter association relationship; collecting and verifying basic information; collecting detailed electrical schematic diagrams and physical layout drawings of the target distribution cabinet; simultaneously verifying the sampling location corresponding to each parameter in the key state parameter set to ensure that the parameters accurately correspond to the circuits and components in the drawings; determining the electrical connection logic; based on the improved electrical schematic diagram, sorting out the circuits and components corresponding to the key state parameters one by one to determine the electrical connection relationship between parameters. For example, the main circuit current parameter and the main circuit voltage parameter correspond to the same main circuit and are connected in series through the busbar and circuit breaker, and there is a series correlation between the current and voltage; the current parameter of the same branch circuit and the active power parameter of the same circuit are related by power The parameters are derived from the current and voltage parameters of the same circuit, showing a correlation between current and power. Current parameters of different branch circuits are all connected to the same busbar, indicating a parallel relationship. The cabinet grounding resistance parameter is related to the grounding protection of all electrical circuit parameters. During the review process, each parameter relationship must be recorded individually to determine the relationship type: series / parallel / derived / protection. Physical location relationships must be determined. Based on the physical layout drawings and on-site survey results, the spatial relationship between the sensor installation locations corresponding to key status parameters and equipment components must be determined, thus establishing the physical location relationships between parameters. For example, the sensor corresponding to the circuit breaker vibration intensity parameter is installed on the circuit breaker housing, and the sensor corresponding to the circuit breaker connection point temperature parameter is installed at the circuit breaker's terminals, with a spacing of less than 5 cm, indicating a proximity relationship within the same component. The sensor corresponding to the cabinet's internal temperature and humidity parameter is installed in the middle of the cabinet, with a spacing of 10-30 cm from all components within the cabinet, indicating a peripheral environment relationship. Sensors corresponding to parameters of different distribution cabinets are installed in independent cabinets, indicating no direct physical location relationship. When recording relationships, the relationship type and corresponding spatial spacing must be noted.

[0017] The time-domain and frequency-domain correlation of each parameter's time series within the calculation window is determined by extracting complete time-series data for any two parameters from a single synchronous multi-dimensional data sampling window. This ensures a one-to-one correspondence between the timestamps of the two time series. If any timestamps are missing, they are padded using the mean of adjacent valid data. The padded time-series data undergoes preliminary processing to remove obvious anomalous jumps. For example, data with values ​​suddenly exceeding 10 times the normal range are considered to be caused by sensor malfunctions, removed, and padded again. The time-domain correlation is then calculated for the preprocessed time series of the two parameters using the following steps: First, calculate the correlation for each... The change in the first parameter at each timestamp is calculated as the value of the first parameter at the current timestamp minus the value of the first parameter at the previous adjacent timestamp. The change in the second parameter at each timestamp is calculated using the same method. The second step is to calculate the difference between the changes in the two parameters at each timestamp, which is the difference between the changes in the first and second parameters, and take the absolute value of this difference. The third step is to sum the absolute values ​​of the differences across all timestamps to obtain the total difference. The fourth step is to divide the total difference by the total number of timestamps to obtain the average difference value. A smaller average difference value indicates a more consistent trend in the changes of the two parameters, and a higher temporal correlation. Conversely, the lower the time-domain correlation, the lower the correlation. Simultaneously, correlation levels are categorized based on the average difference value: less than 0.5 indicates high correlation, 0.5-1.5 indicates moderate correlation, and greater than 1.5 indicates low correlation. For frequency domain correlation calculation, for time-series data of the same two parameters, the correlation is calculated using the following steps: First, identify the peak values ​​of the two parameters. The peak value is determined by the current timestamp being greater than the values ​​of the three adjacent timestamps before and after it. Second, count the number of peak values ​​for each parameter and the number of times the peak timestamps completely overlap, with an overlap timetamp deviation not exceeding 100 milliseconds. Third, calculate... Peak overlap rate is calculated by dividing the number of overlapping peak points by the average of the total number of peak points for both parameters. The fourth step involves analyzing the fluctuation periods of the two parameters. By calculating the time interval between two consecutive peak points, the fluctuation period of each parameter is obtained. The degree of overlap between the fluctuation periods of the two parameters is then statistically analyzed. If the difference in the periods of the two parameters is less than 10% of the average period, it is considered that the periods overlap. The frequency domain correlation is determined by combining the peak overlap rate and the degree of period overlap. A peak overlap rate greater than 60% and period overlap indicate high correlation; a peak overlap rate of 30%-60% or near-perfect period overlap indicates moderate correlation; and a peak overlap rate less than 30% and no period overlap indicates low correlation.

[0018] Generate topology relationships and determine topology nodes. Each parameter in the key state parameter set is used as an independent node in the topology relationship. Each node is labeled with core information, including parameter name, parameter type, corresponding acquisition location, and corresponding device component. All nodes are initially laid out according to the classification dimensions of electrical performance, mechanical state, and environmental conditions. The topology edges and directions are determined. First, considering electrical connection logic, physical location association, and temporal correlation, it is determined whether there are mutual influences or constraints between parameters. If two parameters have electrical connection logic or physical location association, and the temporal correlation is high or medium, then a valid correlation exists, and a connection edge needs to be constructed. If only a single correlation exists, it is determined as a weak correlation, and a connection edge is not constructed at this time. Second, the direction of the edge is determined. If the correlation is a unidirectional influence, such as a change in current parameter leading to a change in power parameter, and a change in power parameter... If the change does not have a reverse effect on the current parameter, then a directed edge is constructed, with the starting node of the directed edge being the source parameter and the ending node being the affected parameter. If the relationship is a two-way constraint, such as a change in the circuit breaker vibration intensity parameter affecting the connection point temperature parameter, and conversely, an excessively high connection point temperature also aggravates the vibration, and the temporal correlation between the two is highly correlated, then an undirected edge is constructed. The third step is to label the attributes of the edges, labeling the association type and temporal correlation level on each edge to determine the core basis for the association between parameters. A complete topological relationship is formed by integrating all nodes and their corresponding directed / undirected edges according to the above rules to construct a complete topological relationship network. The network is verified to ensure that there are no isolated nodes. If there are isolated nodes, the association determination process is reviewed again to confirm whether any associations are omitted and whether there are any incorrect edge directions. Finally, a structured topological relationship is generated with key parameters as nodes and mutual influence / constraint relationships as edges.

[0019] The dual-reference synchronous sampling mechanism ensures continuous data coverage under normal operating conditions through fixed-time sampling, while event-triggered high-frequency sampling captures dynamic parameter changes under abnormal operating conditions. The combination of the two achieves seamless coverage of different operating scenarios throughout the entire lifecycle.

[0020] In a preferred embodiment of the present invention, step 3 above, based on topological relationships, performs structured topological partitioning of the data within the sampling window, calculates the spatial shape feature similarity of the data distribution in each partition, and generates data quality calibration parameters based on the data distribution characteristics, correlation strength, and shape consistency index of each partition, which may include: In this embodiment of the invention, step 330 involves, based on the electrical connections and logical associations determined by the topological relationships, structurally dividing the multi-dimensional operating parameter data collected within the sampling window into several data partitions with strong internal correlations, according to the associated equipment groups or electrical circuit units. Specifically, this includes: First, refining the division criteria by retrieving the topological relationship document between the core feature parameters established in step 2, and extracting the electrical connection relationships and logical associations corresponding to each group of core feature parameters from the document. The electrical connection relationships need to determine the specific component connection methods, such as the series relationship between circuit breakers and contactors, or the parallel relationship between multiple branch circuits and the main circuit. The logical association relationships need to clarify the influence logic between parameters, such as the causal relationship that changes in current parameters directly lead to changes in temperature parameters, and the accompanying change relationship between vibration parameters and displacement parameters. At the same time, these extracted relationships are organized into an association relationship comparison table, which clearly lists the associated parameters, association types, and association strength descriptions corresponding to each parameter. Second, matching the sampling window data with the equipment components. The process involves meticulously reviewing all multi-dimensional operational parameter data collected within the current sampling window. Each parameter is identified by a source identifier, determining which specific sensor collected it, which specific equipment component within the distribution cabinet it corresponds to, and its associated electrical circuit number. For parameters whose attribution is not directly clear, the equipment ledger and sensor deployment list of the distribution cabinet are consulted to ensure each parameter accurately corresponds to a specific equipment component and electrical circuit. The third step involves structured partitioning. Based on actual maintenance needs, partitioning can be performed by associated equipment groups or electrical circuit units. If partitioning by associated equipment groups, all parameter data corresponding to equipment components belonging to the same functional module or having a direct collaborative working relationship are grouped into one category. For example, the current, voltage, vibration intensity, and connection point temperature parameters corresponding to circuit breaker No. 1, as well as the current parameters of the fuse directly connected in series with this circuit breaker, are integrated into a single category for circuit breaker No. 1. The data is partitioned by equipment group. If the partitioning is chosen to be based on electrical circuit units, the parameter data corresponding to all equipment components belonging to the same electrical circuit will be grouped into one category. For example, the current, voltage, and power parameters corresponding to all circuit breakers, contactors, and loads in power circuit L1, as well as the environmental temperature and humidity parameters corresponding to that circuit, will be integrated into a single power circuit L1 data partition. The fourth step is to verify the partition correlation and add identifiers. After the partitioning is completed, check whether the parameters in each partition have direct electrical connections or logical relationships, ensuring strong correlation within the partition, by referring to the previously compiled correlation table. At the same time, check the parameter correlation between different partitions to ensure that the correlation between partitions is weak. After verification, add a unique partition identifier to each partition. The identifier must fully include the partitioning dimension, the corresponding equipment group name or circuit number, the data acquisition time window number, the number of parameters contained in the partition, and the parameter names.

[0021] Step 331: Analyze the distribution characteristics of all data samples in the multidimensional feature space within each data partition, determine the main direction, range, and density of the point set spatial distribution, and construct a geometric structure to represent the overall spatial distribution morphology of the point set. Specifically, this includes: First, constructing a multidimensional feature space matching the partition parameters. For a single data partition, first determine all parameter dimensions contained within that partition. For example, if a partition contains three parameters (current, voltage, and temperature), then construct a three-dimensional feature space; if it contains four parameters (current, voltage, temperature, and vibration), then construct a four-dimensional feature space. Then, map each parameter dimension to a coordinate axis in the multidimensional feature space, labeling each coordinate axis with the corresponding parameter name. For example, the X-axis in the three-dimensional feature space corresponds to current, the Y-axis to voltage, and the Z-axis to temperature, ensuring that each parameter has a unique corresponding coordinate axis. Second, map the data samples to a point set in the multidimensional feature space. According to the data sample acquisition time sequence, extract the specific values ​​of each parameter in each data sample one by one. Based on the correspondence between parameters and coordinate axes, map each... Each data sample is mapped into a multidimensional feature space to form a data point. For example, if a data sample has a current value of 5A, a voltage value of 220V, and a temperature value of 35℃, the corresponding positions X=5, Y=220, and Z=35 in the three-dimensional feature space are found and marked as a data point. After all data samples in this partition are mapped, all data points together constitute a complete point set. The third step is to analyze the main directions of the spatial distribution of the point set. By visually observing the overall extension trend of the point set in the multidimensional feature space, the extension length of the point set in each coordinate axis direction is compared one by one. Specifically, the farthest data point of the point set in each coordinate axis direction is found, the distance of the data point to the origin of the coordinate is measured, and the coordinate axis direction with the largest distance is determined as the main direction of the point set distribution. If there are two or more coordinate axis directions with similar extension distances and both being the maximum value, these directions are all determined as the main directions. For example, if the extension distance of a point set in the X-axis and Y-axis directions is significantly greater than that in the Z-axis direction, and the difference between the two is very small, the X-axis and Y-axis directions are both taken as the main directions.

[0022] The fourth step is to analyze the spatial distribution range of the point set. For each coordinate axis of the multidimensional feature space, find the maximum and minimum values ​​of all data points on that axis. Subtract the minimum value from the maximum value to obtain the distribution range of the point set along that coordinate axis. For example, if the maximum value of all data points on the X-axis (current) is 10A and the minimum value is 2A, then the distribution range along the X-axis is 8A; if the maximum value on the Y-axis (voltage) is 230V and the minimum value is 210V, the distribution range is 20V. By summing up the distribution ranges along all coordinate axes, we obtain the overall distribution range of the point set in the entire multidimensional feature space. The fifth step is to analyze the density of the spatial distribution of the point set. First, determine a minimum closed region that can completely enclose all data points. Divide this closed region evenly into several small cubes of equal volume. The size of the small cubes is determined based on the total number of data points, ensuring that each small cube contains at least one data point. If the number of data points is small, some small cubes may be empty. Then, count the data points contained in each small cube. The first step is to calculate the local density of a region within a small cube by dividing the number of data points within that cube by the volume of the cube. The local densities of all small cubes are then compared to identify the regions with the highest and lowest densities. The density range of the region containing the majority of data points is also recorded, completing the analysis of the point set distribution density. The sixth step involves constructing a geometric structure representing the distribution pattern of the point set. Based on the main directions, ranges, and density characteristics of the point set distribution obtained from the previous analysis, a suitable geometric shape is selected to construct the structure. If the point set exhibits an approximately ellipsoidal distribution, an elliptical geometric structure is constructed with the center of the point set as the center and the distribution range along each coordinate axis as the semi-axis length. This ensures that the elliptical can completely enclose all data points, and that the major and minor axes of the elliptical are aligned with the main distribution directions of the point set. If the point set exhibits an approximately polyhedral distribution, the vertices of the polyhedron are determined based on the positions of the edge data points. These vertices are then connected to form a polyhedral geometric structure, ensuring that the polyhedron fits the distribution contour of the point set and accurately reflects its spatial distribution pattern.

[0023] Step 332: For the geometric structure, extract a set of geometric feature parameters to quantitatively describe the extension range, directionality, and discreteness of the geometric structure in space, and construct the geometric feature parameter set into a corresponding feature vector; specifically, this includes: First, determining the extraction dimensions and specific types of geometric feature parameters, determining to extract parameters from three core dimensions: extension range, directionality, and discreteness, and specifying the specific parameter type for each dimension. Parameters extracted for the extension range dimension include the maximum extension span of the geometric structure in each coordinate axis direction and the overall volume of the geometric structure; parameters extracted for the directionality dimension include the angle between the main extension direction of the geometric structure and each coordinate axis, and the number of main extension directions; parameters extracted for the discreteness dimension include the average distance from all data points in the point set to the center of the geometric structure, the maximum distance from the data point to the center, and the degree of difference in distance from the data point to the center; Second, extracting the geometric feature parameters for the extension range dimension, for each coordinate axis direction... The maximum extension span in the direction is directly taken as the distribution range of the point set in the direction of the coordinate axis analyzed in step 331 as the maximum extension span of the geometric structure in that direction. For the overall volume of the geometric structure, if it is an ellipsoidal structure, the overall volume is obtained by measuring the lengths of the major and minor semi-axes of the ellipsoid and combining the relationship between the volume of the ellipsoid and the length of the semi-axes, according to the calculation logic of the volume of the ellipsoid. If it is a polyhedron structure, the overall volume is obtained by measuring the area and corresponding height of each face of the polyhedron and combining the calculation logic of the volume of the polyhedron. In the third step, the geometric feature parameters of the directional dimension are extracted. For the angle between the main extension direction of the geometric structure and each coordinate axis, the size of the angle is determined by observing the relative position of the main extension direction of the geometric structure and each coordinate axis and combining the basic logic of angle judgment. For the number of main extension directions, the number of the main distribution directions of the point set determined in step 331 is directly counted as the number of main extension directions of the geometric structure.

[0024] The fourth step is to extract geometric feature parameters of the discreteness dimension. First, determine the center of the geometric structure. The center of the elliptic structure is the center of the sphere of the elliptic, and the center of the polyhedron structure is the intersection of the lines connecting the vertices of the polyhedron. Then, measure the straight-line distance from each data point to the center of the geometric structure one by one. Add up the distances of all data points and divide by the total number of data points to obtain the average distance. Find the largest distance among all data points and take it as the maximum distance. Compare the distances of all data points with the average distance and count the difference between the distances. The larger the difference, the higher the degree of difference. Use three levels (high, medium, and low) or specific textual descriptions to characterize the degree of difference. The fifth step is to standardize the extracted geometric feature parameters. Collect multiple sets of geometric feature parameter data of the data partition under the historical normal operation state, calculate the average value and difference range of each indicator in each set of parameters. For each geometric feature parameter extracted now, subtract the average value of the corresponding parameter in the historical normal data from the actual value of the parameter, and then divide the resulting difference by the difference range of the corresponding parameter in the historical normal data to obtain the standardized parameter value. This process eliminates the influence of different parameters due to their different dimensions and numerical ranges, ensuring that all parameters are within the same comparable range. The sixth step involves constructing feature vectors by pre-setting a fixed parameter arrangement order. For example, parameters for the extension range dimension are arranged first, with the maximum extension span arranged in coordinate axis order, followed by the overall volume. Then, parameters for the directionality dimension are arranged, with the included angle arranged in coordinate axis order, followed by the number of main extension directions. Finally, parameters for the dispersion dimension (average distance, maximum distance, degree of difference) are arranged. Following this fixed order, all standardized geometric feature parameters are arranged sequentially to form an ordered parameter combination. This ordered combination is the feature vector corresponding to the data partition, and each feature vector uniquely represents the spatial distribution characteristics of the data point set in that partition.

[0025] Step 333: By calculating the similarity measure between feature vectors, the feature vectors of the same data partition across different sampling windows are compared to quantify the similarity of spatial shape features. Specifically, this includes: First, filtering and matching feature vectors from different sampling windows of the same data partition. First, determine the two different sampling windows to be compared, ensuring that these two sampling windows belong to the same data partition and cover the same parameter dimensions. Then, retrieve the feature vectors of the data partition under these two sampling windows from the data storage area, denoted as vector C and vector D respectively. After retrieval, check the order and type of parameters in the two vectors one by one to ensure that vector C and vector D have the same number of dimensions and that the parameters at corresponding positions are consistent. If the data types are completely identical, and there are cases where parameters are missing or their order is inconsistent, first, the missing parameters are supplemented or their order is adjusted to make the two vectors meet the comparison conditions. Second, a similarity measurement method is selected and the calculation logic is clarified. Cosine similarity is chosen as the similarity measurement method, and the calculation logic is as follows: the degree of similarity is determined by judging the angle between two feature vectors in multidimensional space. The smaller the angle, the closer the directions of the two vectors are, and the more similar their corresponding spatial shape features. The larger the angle, the greater the difference in direction, and the lower the similarity of spatial shape features. If the cosine similarity value is close to 1, it means the angle is close to 0 degrees, indicating extremely high similarity; if it is close to 0, it means the angle is close to 90 degrees, indicating extremely low similarity.

[0026] The third step involves performing similarity calculations. Following the logic of cosine similarity, first, the parameters at corresponding positions in vectors C and D are identified. Each pair of corresponding parameters is multiplied to obtain several product results, which are then summed to obtain a total value. Next, the overall lengths of vectors C and D are calculated. When calculating the magnitude of vector C, all parameters in vector C are squared, and the square root of these squares is the magnitude of vector C. The magnitude of vector D is calculated using the same method. Finally, the sum of the previously obtained product values ​​is divided by the product of the magnitudes of vectors C and D to obtain the result of the similarity calculation for vectors C and D. The fourth step involves setting a similarity threshold and quantifying the judgment result. This involves collecting cosine similarity data between feature vectors of different sampling windows under historical normal operating conditions for the data partition, statistically analyzing the distribution range of these data, setting the lower limit of the distribution range as the similarity threshold, and comparing the currently calculated similarity value with this threshold. If the similarity value is greater than or equal to the threshold, it indicates that the spatial shape feature similarity of the same data partition between these two sampling windows is high, and the data distribution pattern is stable. If the similarity value is less than the threshold, it indicates that the spatial shape feature similarity is low, and the data distribution pattern has changed significantly. Simultaneously, the similarity value, threshold, and judgment result are recorded in detail to provide a basis for data quality assessment.

[0027] Step 334: Calculate the distribution dispersion characteristics of time-series monitoring data within each partition to obtain a first evaluation index characterizing the degree of data fluctuation and dispersion; simultaneously, calculate the matching degree between the data sequence of the partition and the state space shape characteristics of the typical equipment corresponding to the partition to obtain a second evaluation index characterizing the data self-consistency within the partition; integrate the first and second evaluation indices to generate preliminary data quality assessment results; specifically including: First, calculate the first evaluation index. For a single data partition, first extract all time-series monitoring data within the partition, and organize the data into an ordered time-series data sequence according to the chronological order of data collection. Then, calculate... The dispersion of this time-series data sequence is determined by first identifying the maximum and minimum values, calculating the difference between them to obtain the range; then calculating the mean, and for each data point, calculating the difference between the mean and the mean, summing the absolute values ​​of each difference, and dividing by the total number of data points to obtain the average deviation; using the range and average deviation as core indicators to comprehensively characterize the fluctuation and dispersion of the data, the larger the range and average deviation, the more drastic the data fluctuation and the higher the dispersion, and the combined result of these two indicators is used as the first evaluation indicator; the second step is to construct the state space shape characteristics of typical equipment. The library collects historical data of the equipment corresponding to the data partition under different typical states, including normal operation, minor abnormality, moderate abnormality, and severe abnormality. For the historical data of each typical state, the corresponding geometric structure is constructed and feature vectors are extracted according to the methods in steps 331 and 332. These feature vectors are then classified and organized according to typical states to form a typical equipment state space shape feature library. The library clearly marks the equipment state, data acquisition conditions, and other information corresponding to each feature vector. The third step is to calculate the second evaluation index, retrieve the feature vector of the current data partition, and compare it with each feature vector in the typical equipment state space shape feature library. For a typical state's standard feature vector, the matching degree is calculated according to the similarity calculation method in step 333. After the calculation, the typical state with the highest matching degree is found and the highest matching degree value is recorded. If the typical state corresponding to the highest matching degree value is a normal operating state and the matching degree value is high, it indicates that the data sequence of the current partition is highly consistent with the spatial shape characteristics of the normal state, and the data self-consistency is good. If the typical state corresponding to the highest matching degree value is an abnormal state or the matching degree value is low, it indicates that the current data sequence does not match the characteristics of any typical state, and the data self-consistency is poor. The highest matching degree value is used as the second evaluation index to characterize the data self-consistency.

[0028] The fourth step involves integrating the two evaluation indicators to generate a preliminary data quality assessment result. Based on the actual needs of distribution cabinet operation and maintenance, and considering the impact of historical data quality issues on operation and maintenance decisions, weights are assigned to the first and second evaluation indicators respectively. For example, if the degree of data fluctuation and dispersion has a greater impact on decisions, a higher weight (e.g., 0.6) is assigned to the first evaluation indicator, and a lower weight (e.g., 0.4) to the second evaluation indicator. Subsequently, the value of the first evaluation indicator is multiplied by its corresponding weight to obtain the first weighted value; the value of the second evaluation indicator is multiplied by its corresponding weight to obtain the second weighted value; and the first and second weighted values ​​are added together to obtain the comprehensive result. The fifth step is to determine the preliminary data quality level and preset the grading range of the comprehensive evaluation value. For example, a comprehensive evaluation value between 0.8 and 1.0 is excellent, between 0.6 and 0.8 is good, between 0.4 and 0.6 is average, and between 0 and 0.4 is poor. Based on the currently calculated comprehensive evaluation value, determine the corresponding preliminary data quality level. At the same time, organize and summarize the specific values ​​of the first evaluation indicator (range, average deviation), the specific values ​​of the second evaluation indicator (highest matching degree value), the weights of the two indicators, the comprehensive evaluation value, the preliminary quality level, and other information to form a complete preliminary data quality evaluation result.

[0029] Step 335: Based on the preliminary data quality assessment results, and combined with the electrical connection topology and physical layout relationships of the equipment units within the distribution cabinet, construct a partitioned correlation analysis framework. Perform data feature consistency analysis on partitions with correlation relationships, identify and quantify contradictions and conflicts between related partitions in terms of data trends, amplitudes, or states, and generate a partitioned correlation contradiction diagnosis report. Specifically, this includes: First, sorting out partitioned correlation relationships and selecting key analysis objects. First, retrieve the electrical connection topology diagram and physical layout diagram of the distribution cabinet. Combined with the partitioning criteria in Step 330, sort out the correlation relationships between all data partitions one by one. If the equipment groups corresponding to two partitions have... Direct electrical connections are identified as electrically associated partition pairs; if the equipment components corresponding to two partitions are physically adjacent, they are identified as physically associated partition pairs. All associated partition pairs are compiled into an associated partition pair list, determining the association type and specific basis for each pair. Simultaneously, preliminary data quality assessment results for each partition are retrieved, and partitions with a quality level of "average" or "poor" are selected as key analysis targets. The second step involves constructing a partition association analysis framework. Using the key analysis targets as the core and combining the associated partition pair list, a partition association analysis framework is built. Within this framework, all associated partitions corresponding to each key analysis target are identified, along with the partitions requiring analysis. The specific dimensions include three core dimensions: data trend consistency, data amplitude consistency, and data state consistency. Simultaneously, the framework defines analytical standards for each dimension. For example, the standard for data trend consistency is that the time-series data of two partitions change in the same direction; the standard for data amplitude consistency is that the parameter amplitudes of the two partitions conform to electrical or physical correlation laws; and the standard for data state consistency is that the initial data quality levels of the two partitions are similar. The third step involves conducting data trend consistency analysis. For each pair of related partitions, the time-series monitoring data sequences of the two partitions are extracted, and the changing trends of the two sequences are plotted in chronological order. For example, describing the first... The time-series data for the first partition shows a slow upward trend from minute 1 to minute 10, remains stable from minute 10 to minute 20, and shows a slow downward trend from minute 20 to minute 30. The time-series data for the second partition is also described. By comparing the two trend descriptions, it is determined whether the directions of change are consistent. If the data for both partitions show the same rhythm of rising, stabilizing, and falling, the trends are consistent. If one rises and the other falls, or the rhythms are completely different, a trend contradiction exists. Furthermore, the degree of trend contradiction is quantified: if the two trends are completely opposite in direction, it is considered a serious contradiction; if the rhythms are different but the directions are consistent, it is considered a minor contradiction.

[0030] The fourth step is to conduct data amplitude consistency analysis. Based on the electrical or physical correlation between the parameters of the two related partitions, the normal amplitude ratio range between them is determined. For example, for current and voltage partitions in the same electrical circuit, the normal current and voltage amplitude ratio range is determined according to the correlation law corresponding to Ohm's law; for two adjacent temperature sensor partitions, the normal temperature amplitude difference range is determined according to the conduction law of ambient temperature. Then, the parameter amplitudes of the two partitions at corresponding times are extracted, the actual ratio or difference is calculated, and it is determined whether it is within the normal range. If it is within the normal range, it means that the amplitudes are consistent; if it exceeds the normal range, it means that there is an amplitude contradiction. At the same time, the degree of amplitude contradiction is quantified by calculating the difference between the actual value and the upper or lower limit of the normal range. The larger the difference, the more serious the contradiction. The fifth step is to conduct data state consistency analysis, comparing the two related partitions. The initial data quality level is determined as follows: if both partitions are rated as excellent or good, the status is consistent; if one partition is rated as excellent or good and the other as average or poor, a status contradiction exists. Simultaneously, the degree of contradiction is quantified based on the size of the level difference: a difference of two levels indicates a serious contradiction; a difference of one level indicates a minor contradiction. The sixth step involves generating a segmented correlation contradiction diagnosis report. This summarizes the three-dimensional analysis results for each pair of correlated partitions, determining whether a contradiction exists, what type of contradiction exists, the specific manifestation of the contradiction, and the severity of the contradiction. The analysis results for all correlated partition pairs are sorted from highest to lowest severity and organized into a structured segmented correlation contradiction diagnosis report. The report must also include information such as the correlation type, analysis basis, and data source for each partition pair to ensure the report's completeness and traceability.

[0031] Step 336 involves integrating the preliminary data quality assessment results with the inter-regional correlation contradiction diagnosis reports to form a diagnostic evidence set. This set undergoes multi-factor comprehensive analysis to achieve comprehensive tracing, location, and type diagnosis of data anomaly patterns. Based on the final diagnostic conclusion, data quality calibration parameters are quantified and generated. Specifically, this includes: First, integrating and forming the diagnostic evidence set. This involves collecting all preliminary data quality assessment results and inter-regional correlation contradiction diagnosis reports for each region, sorting out all information, and removing duplicates. Then, the data is categorized and integrated by region, creating a dedicated diagnostic evidence subset for each region. Each subset contains three parts: first, the preliminary data quality assessment information for that region itself; second, the data quality assessment information for that region; third, the data quality assessment information for that region itself; fourth, the data quality assessment information for that region itself; fifth, the data quality assessment information for that region itself; and sixth, the data quality assessment information for that region itself. The first step involves analyzing the contradictions with all other related partitions; the second step involves identifying the factors and logic for multi-factor comprehensive judgment, including the preliminary data quality level of the partition itself, the number and severity of contradictions between the partition and related partitions, historical fault records of sensors, sensor calibration cycles, historical operational fault records of equipment components, and current operating conditions. The judgment logic involves combining and analyzing the specific circumstances of each factor, and using elimination and causal deduction methods to determine the cause, location, and type of data anomalies. The third step involves conducting comprehensive judgment to achieve anomaly tracing, location, and type diagnosis. For each partition's diagnostic evidence subset, operators analyze each judgment factor one by one. If the partition's initial quality level is poor and there are serious contradictions with multiple related partitions, and the corresponding sensor has recent fault records that are within the calibration validity period, then the source of the anomaly is determined to be a sensor fault. If the partition's initial quality level is average and there are only minor contradictions with a few related partitions, the sensor has no fault records and is within the calibration validity period, but the current equipment is under high load and special operating conditions, then the source of the anomaly is determined to be normal data fluctuations caused by special operating conditions. During anomaly localization, based on the contradiction analysis information and sensor fundamentals... Information is used to determine the specific partition, parameter, and corresponding sensor where the abnormal data is located. For example, if the current parameter of the associated equipment group of circuit breaker No. 1 is abnormal, the corresponding sensor is current sensor number C1. When diagnosing the type of abnormality, the type is determined based on the source cause. Abnormalities caused by sensor failure are classified as equipment failure anomalies; abnormalities caused by interference during data transmission are classified as transmission interference anomalies; normal fluctuations caused by special operating conditions are classified as operating condition-affected fluctuations; abnormalities caused by missing data are classified as missing data anomalies; and data values ​​that deviate from the normal range without a clear cause of failure are classified as random error anomalies.

[0032] The fourth step involves quantifying and generating three types of data quality calibration parameters. The first type is a confidence coefficient vector. For each parameter collected by a sensor, a confidence coefficient is assigned based on the anomaly diagnosis results. If the data source has no anomalies, the confidence coefficient is set to 0.9-1.0; if there are minor anomalies, the confidence coefficient is set to 0.6-0.8; if there are severe anomalies, the confidence coefficient is set to 0.1-0.3; if data is missing, the confidence coefficient is set to 0. The confidence coefficients of all data sources are arranged in a preset data source order to form a confidence coefficient vector. The second type is an anomaly pattern judgment threshold set. Historical normal data corresponding to this partition is collected, and the normal value range of each parameter is statistically analyzed. The maximum value of the normal range is set as the upper threshold, and the minimum value is set as the lower threshold. Simultaneously, based on the fluctuation characteristics of the anomaly data, a fluctuation frequency threshold is set. The upper and lower thresholds and fluctuation frequency thresholds of all parameters are arranged to form an anomaly pattern judgment threshold set. The third type is a suggested calibration benchmark value and direction vector. For random... For error-type anomaly parameters, the average value of historical normal data for that parameter is used as the correction benchmark. If the current anomaly data is greater than the average, the direction vector is set to decrease, meaning the anomaly data needs to be adjusted downwards towards the average value; if the current anomaly data is less than the average, the direction vector is set to increase. For sensor fault-type anomaly parameters, the correction benchmark value is calculated by referring to the values ​​of the corresponding normal parameters within the same associated partition and combining the normal correlation ratio between the two. The direction vector is determined based on the direction of the difference between the faulty sensor data and the normal parameter values. For data missing anomalies, the correction benchmark value is calculated based on the data trend of the same parameter in adjacent time windows and the values ​​of the corresponding parameters in associated partitions. The direction vector is determined based on the parameter trend in adjacent time windows. The fifth step is to integrate and form complete data quality calibration parameters. The generated confidence coefficient vector, anomaly mode judgment threshold set, suggested correction benchmark value, and direction vector are summarized, and clear explanatory information is added to each type of parameter to form complete data quality calibration parameters.

[0033] This method enables the orderly classification of multi-source heterogeneous data, strengthens the correlation between data within partitions, avoids the problem of low analysis efficiency caused by disordered data, and transforms the spatial shape characteristics of data distribution into quantifiable parameters by constructing geometric structures and extracting feature vectors, thereby achieving the representation of data distribution morphology.

[0034] In a preferred embodiment of the present invention, step 4 above, which involves using data quality calibration parameters to perform data transformation and standardized encapsulation on the original full lifecycle dataset to obtain a standardized transmission dataset, may include: In this embodiment of the invention, step 440 involves applying data quality calibration parameters to perform weighted correction processing on the multi-source heterogeneous data in the original full lifecycle dataset to obtain calibrated data. Specifically, this includes: First, clarifying the correspondence between calibration parameters and original data. This involves retrieving the data quality calibration parameters generated in step 3 and determining the data sources corresponding to the confidence coefficient vector, correction reference value, and direction vector. That is, each parameter corresponds to a specific data item collected by a particular sensor in the original full lifecycle dataset. Simultaneously, the multi-source heterogeneous data in the original full lifecycle dataset is categorized and organized according to data source, forming original data subsets based on data source. Each subset is labeled with the corresponding sensor model, acquisition parameter name, data acquisition time range, etc., to ensure accurate matching between calibration parameters and original data. Second, performing a weighted calculation operation. For each original data subset, each original data point is extracted, and the data point is multiplied by its corresponding confidence coefficient to obtain weighted data. The specific calculation logic is as follows: [The text abruptly ends here, so the translation stops as well.] The first step involves multiplying the initial data point's value by the corresponding data source's reliability coefficient. A reliability coefficient closer to 1 indicates higher data source reliability, resulting in weighted data that is closer to the original data. Conversely, a lower reliability coefficient indicates lower data source reliability, leading to a more significant weakening of the weighted data. The second step involves adjusting the weighted data using the calibration benchmark and direction vector. If the direction vector is adjusted upwards, the weighted data is added to the benchmark value, bringing it closer to the benchmark. If the direction vector is adjusted downwards, the weighted data is subtracted from the benchmark value, achieving the same effect. The third step involves summarizing the calibrated data. All original data points, after weighting and adjustment, are organized and categorized according to their acquisition time and data source, forming a calibrated dataset. Each data point is labeled with its corresponding calibration basis.

[0035] Step 441: Perform dimension unification processing on the calibrated data to obtain dimension-unified data. Imput the missing values ​​in the dimension-unified data to obtain a complete and standardized dataset. Specifically, this includes: First, perform dimension unification processing. First, identify the dimension types of all parameters in the calibrated data to determine the dimension differences between different parameters. Then, select a normalization method to achieve dimension unification. The specific operation logic is as follows: For each type of parameter, first collect the corresponding historical normal operation data and find the maximum and minimum values ​​of the historical normal data. Then, subtract the minimum value of the historical normal data from each calibrated data point of the parameter to obtain the first difference. Then, subtract the minimum value of the historical normal data from the maximum value of the historical normal data to obtain the second difference. Finally, divide the first difference by the second difference to obtain the normalized result of the data. After the above processing, all parameters are converted into dimensionless data between 0 and 1, forming dimension-unified data. Second, identify missing values ​​in the dimension-unified data, according to the data source and collection time. The process involves three steps: First, each data point in the unified-dimensional data set is checked sequentially. If the parameter data corresponding to a certain collection time point is missing (i.e., there is no corresponding value), it is determined to be a missing value, and the specific location information of the missing value is recorded to form a missing value list. Second, imputation processing is performed on the missing values. Based on the location of the missing value and the characteristics of the corresponding parameter, different imputation methods are selected. For continuously collected time-series parameters, if the missing value is missing at a single time point, adjacent time point data is used for imputation, that is, the average of the two adjacent valid data before and after the missing value is taken as the imputation value. If the missing value is missing at multiple consecutive time points, trend extension imputation is used, that is, the corresponding data for the missing time period is calculated based on the trend of the change of the valid data before the missing value as the imputation value. For parameters with strong correlation, if a parameter has a missing value, correlation parameter imputation is used, that is, the valid data of the correlation parameter is multiplied by the normal correlation ratio of the two parameters to obtain the imputation value. After imputing all missing values, the dataset is integrated to obtain a complete and standardized dataset.

[0036] Step 442: The complete standardized dataset is encapsulated into a standardized data packet according to the preset communication protocol and data structure specifications. This includes a data body composed of the complete standardized dataset, a quality label generated based on data quality calibration parameters, and a topology identifier used to identify the location relationships of the data source distribution cabinet and components. Specifically, this includes: First, determining the preset communication protocol and data structure specifications, retrieving the pre-set communication protocol, and determining the encoding method, transmission rate, verification rules, and other requirements for data transmission; simultaneously, determining the preset data structure specifications, specifying the components of the standardized data packet, the order of arrangement of each component, data format, field length, etc.; Second, constructing the data body of the data packet, organizing the complete standardized dataset according to the data body format requirements in the preset data structure specifications, arranging the data according to data source classification and collection time order, adding a data type identifier to each data point to ensure that the data in the data body is arranged in an orderly manner, with a unified format, and conforms to the encoding requirements of the communication protocol; Third, generating the quality label of the data packet based on the data generated in step 3. The process involves five steps: First, generating a data quality calibration parameter. This parameter extracts key quality information to generate quality labels, including the credibility level of each data source (based on credibility coefficients, such as 0.9-1.0 for Level 1, 0.6-0.8 for Level 2, and 0.1-0.3 for Level 3), whether the data has undergone calibration, the calibration type, and data integrity indicators. This information is then organized according to a preset format to form quality labels. Second, generating a topology identifier for the data packet. This involves extracting key location information from the data source to generate a topology identifier, including the unique number of the distribution cabinet, the name and number of the internal components of the distribution cabinet corresponding to the data, the installation location of the components, and the installation location and number of the acquisition sensors. This location information is organized according to preset encoding rules to form a topology identifier, ensuring that the source device and location of the data can be accurately located through the topology identifier. Third, completing standardized encapsulation. According to preset data structure specifications, the constructed data body, quality labels, and topology identifiers are combined sequentially, encoded using a preset communication protocol encoding method, and a data check code is added to ultimately form a standardized data packet.

[0037] Step 443: Integrate all the encapsulated standardized data packets to form a standardized transmission dataset. Specifically, this includes: First, identifying the relationships between standardized data packets. Each encapsulated standardized data packet is checked individually. Based on the distribution cabinet number, component number in the topology identifier of each data packet, and the data source information in the quality label, the relationships between data packets are identified. Data packets belonging to the same distribution cabinet are grouped together. Data packets belonging to the same component within the same distribution cabinet are further categorized, forming data packet groups based on distribution cabinets and components. Second, sorting the data packets chronologically. Within each distribution cabinet and component data packet group, the data acquisition time information corresponding to each standardized data packet is extracted. The data packets are sorted in chronological order to ensure that the time-series data of the same component is arranged in an orderly manner. Third... The first step involves adding an overall identifier to the integrated dataset. This identifier includes the dataset's generation time, the time range it covers, the number and number of distribution cabinets, and the number and type of components. This creates a dataset identifier list, which is then combined with the sorted data packets. The fourth step is to verify the integrity and consistency of the integration. This involves checking the data packets in each group to ensure they are complete and free of omissions or duplicates. Simultaneously, the format of all data packets is checked for consistency, ensuring they conform to the preset communication protocol and data structure specifications. If any omissions are found, the corresponding data packets are promptly added. If duplicates are found, redundant data packets are deleted. If inconsistent formats are found, the packets are re-encapsulated. After verification, all data packets from each group are integrated according to the hierarchical structure of distribution cabinet number - component number - acquisition time to form a standardized transmission dataset.

[0038] By using data quality calibration parameters for weighted correction, various anomalies in the original data are specifically corrected to improve data reliability. The data is standardized and packaged according to preset specifications to form data packets with a unified structure and complete information, thereby improving data manageability and traceability. The standardized data packets are integrated through association grouping and time sorting to form a standardized transmission dataset with a clear structure and distinct hierarchy.

[0039] In a preferred embodiment of the present invention, step 5 above involves uploading the standardized transmission dataset via a high-speed communication network. Based on the time-series characteristics and value density of the standardized transmission dataset, a hybrid storage architecture combining a time-series database and an unstructured data management engine is adopted on the central data processing side to achieve data tiering, compression, and associated storage, forming a full lifecycle data asset. This may include: In this embodiment of the invention, step 550 involves uploading the standardized transmission dataset to the central data processing side via a high-speed communication network to obtain the uploaded standardized transmission dataset. Specifically, this includes: First, conducting pre-upload preparation and verification. This involves confirming the completeness of the standardized transmission dataset by checking the dataset identifier list, verifying the number of distribution cabinets, components, data packets, and data coverage time range to ensure no data packets are missing, duplicated, or formatted incorrectly. Simultaneously, checking the connection status of the preset high-speed communication network ensures that the network transmission rate, stability, and bandwidth meet the data upload requirements. Other network processes that may interfere with transmission are closed to ensure a smooth transmission channel. Second, configuring transmission parameters and establishing a connection. According to the preset communication protocol requirements, configuring parameters such as the data transmission encoding method, verification rules, and transmission port ensures that the transmission parameters are consistent with the receiving parameters of the central data processing side. Then, through the configured high-speed communication network, establishing a transmission connection between the local data storage terminal and the central data processing side, sending a connection request, and waiting. The central side responds, and upon receiving a connection confirmation signal from the central side, establishes the transmission link. The third step involves data uploading and real-time monitoring. Following the hierarchical order of distribution cabinet-component-collection time, various data packets from the standardized transmission dataset are uploaded sequentially to the central data processing side via the established transmission link. During the upload process, the transmission status is monitored in real-time, recording the upload progress, transmission time, and any transmission errors for each data packet. If a transmission interruption or error occurs, a retransmission mechanism is immediately triggered to re-upload the corresponding data packet until it is successfully uploaded. The fourth step verifies the integrity of the uploaded data and confirms receipt. After all data packets are uploaded, the central data processing side verifies the integrity of the received data packets, comparing the number of received data packets and the data coverage with the dataset identifier list at the local sending end. If both are completely consistent, the central side sends a confirmation signal indicating that the upload is complete and the data is intact. Upon receiving the confirmation signal, the local end completes the upload operation and identifies the dataset received by the central side as the uploaded standardized transmission dataset.

[0040] Step 551 involves analyzing the data in the uploaded standardized transmission dataset and generating a data grading strategy based on the data's temporal characteristics, access frequency, and predictive analysis value density. Specifically, this includes: First, extracting and analyzing the temporal characteristics of the data. This involves traversing the uploaded standardized transmission dataset, extracting the collection time information corresponding to each data packet, determining the collection cycle of various data types, whether the data is arranged continuously in time sequence, and whether there are time sequence breakpoints. The data is then categorized into short-cycle time-series data, medium-cycle time-series data, and long-cycle time-series data according to the collection cycle, and recording the proportion of each category and its corresponding component type. Second, statistically analyzing the data access frequency. This involves retrieving historical data access records from the central data processing side, statistically analyzing the number of times, access duration, and access frequency distribution of various data types in the uploaded dataset within the past preset time period. For new data without historical access records, the possible future access frequency is predicted based on the operational and maintenance requirements. Third, evaluating and analyzing the predictive analysis value density of the data. This involves considering the core requirements of intelligent operation and maintenance of distribution cabinets, namely health status assessment and fault prediction. The first step involves assessing the value of various data types for predictive analytics. Data that directly reflects the core operating status of equipment and plays a crucial supporting role in fault prediction is classified as high-value-density data. Data that has some auxiliary role in predictive analytics but is not core is classified as medium-value-density data. Data that is only used for historical tracing and has minimal impact on current predictive analytics is classified as low-value-density data. The fourth step involves determining the weights of the grading dimensions and generating a grading strategy. Preset weights are assigned to the three grading dimensions of time-series characteristics, access frequency, and predictive analytics value density. For example, predictive analytics value density has the highest weight, followed by access frequency, and then time-series characteristics. For each type of data, a performance score is assigned on each of the three dimensions. For example, high-value density scores 10 points, medium-value density scores 6 points, and low-value density scores 2 points. The score of each dimension is multiplied by the corresponding weight, and the weighted scores of the three dimensions are summed to obtain a comprehensive score for each type of data. Based on the comprehensive score, a grading threshold is set, and the data is divided into three levels: high-level, medium-level, and low-level. The judgment criteria, corresponding data types, and storage priorities for each level of data are determined, forming a complete data grading strategy.

[0041] Step 552: Based on the data grading strategy, perform data grading processing on the uploaded standardized transmission dataset to obtain data subsets including data of different levels. Specifically, this includes: First, interpreting the grading strategy and clarifying the grading standards. The generated data grading strategy is thoroughly analyzed, and the corresponding judgment conditions for high-level, medium-level, and low-level data are determined. For example, high-level data must simultaneously meet the criteria of short-cycle time-series data, high-frequency access, and high value density; medium-level data must meet the criteria of medium-cycle time-series data, medium-frequency access, and medium value density; and low-level data must meet the criteria of long-cycle time-series data, low-frequency access, and low value density. Simultaneously, the attribution rules for various types of edge data are determined. Second, traversing the data and determining the level of each data type. Following the order of distribution cabinet-component-data type, all data in the uploaded standardized transmission dataset is traversed. The level of each data type and each data packet is determined according to the judgment standards in the grading strategy. During the determination process, the grading of each data type is recorded. The criteria for classifying a data packet as high-level data include, for example, a collection cycle of 10 seconds (short cycle), a historical access frequency of 5 times per week (high frequency), and support for fault prediction (high value density). The third step involves categorizing the data into different level subsets. All data classified as high-level is integrated to form a high-level data subset. Similarly, medium-level and low-level data are integrated to form medium-level and low-level data subsets. Each data subset is labeled with its corresponding level identifier, included data types, covered distribution cabinets and components, and data collection time range. A classification judgment list is also attached, recording the classification criteria for each data in the subset. The fourth step verifies the accuracy of the classification. A portion of data from each data subset is randomly selected to review whether the classification judgment results conform to the classification strategy. If a classification error is found, the level of the corresponding data is adjusted promptly, and the classification judgment list is corrected. After verification, the validity of the three levels of data subsets is confirmed.

[0042] Step 553: For time-series monitoring data continuously collected at preset short periods within the data subset, compress and store it using a time-series database; for unstructured data in the data subset classified as event logs, topology data, and analysis reports, store it using an unstructured data management engine. Specifically, this includes: First, filtering two types of target data: From each level of data subset, filter out time-series monitoring data continuously collected at preset short periods. This type of data is mainly concentrated in high-level and mid-level data subsets. Simultaneously, filter out unstructured data such as event logs, topology data, and analysis reports. This type of data is mostly distributed in mid-level and low-level data subsets. Second, configure the time-series database and compression parameters: For short-period time-series monitoring data, select a suitable time-series database, configure the database storage path and partitioning rules, and configure data compression parameters. Use a dedicated compression method for time-series data. For continuously collected time-series data, if the difference between two adjacent data points is less than a preset threshold, merge and store these two data points, replacing the original two values ​​with the average of the two data points; if the difference is greater than the threshold... The first step involves preserving the original data points to reduce data storage space. The second step is to compress and store the time-series data. The selected short-cycle time-series monitoring data is written to the configured time-series database according to time partitioning rules. During the writing process, a preset compression algorithm is executed simultaneously to complete the data compression and storage. After storage, the compression ratio, storage location, and corresponding time partition information for each segment of time-series data are recorded for easy retrieval. The third step is to configure the unstructured data management engine and storage parameters. For unstructured data, the corresponding unstructured data management engine is selected, and parameters such as storage path, data indexing rules, and access permissions are configured. Simultaneously, the unstructured data is standardized, such as converting analysis reports of different formats to a preset document format to ensure the data format meets storage requirements. The fifth step is to store the unstructured data. The standardized unstructured data is written to the unstructured data management engine according to the configured indexing rules. After storage, a data storage list is generated, recording the storage location, index information, and data size of each unstructured data point.

[0043] Step 554 establishes a unified data identifier and association index for the data stored in the time-series database and the data stored in the unstructured data management engine. Based on the unified data identifier and association index, it constructs the association relationship between different types of data to form a full lifecycle data asset. Specifically, this includes: First, designing a unified data identifier rule. Combining key information such as distribution cabinet, component, data type, and acquisition time, a unified data identifier rule is designed. The data identifier consists of four parts: distribution cabinet number, component number, data type code, and acquisition timestamp. The distribution cabinet number and component number use a preset unique code, and the data type code is based on time sequence. Data, event logs, topology relationships, and analysis reports are each assigned a unique code, with collection timestamps accurate to the second. For example, PD-001-QL-01-SX-20260105103000 represents data collected at 10:30:00 on January 5, 2026, for distribution cabinet No. 1 and circuit breaker No. 1. The second step involves assigning a unified identifier to both types of databases. All time-series data stored in the time-series database is traversed, and a unique unified data identifier is assigned to each data point and each segment of time-series data according to the unified identifier rules. Similarly, each piece of unstructured data stored in the unstructured data management engine is assigned a unique identifier. The first step involves assigning a unique, unified data identifier. After allocation, the data identifier is associated with and stored in relation to the corresponding data, ensuring that the specific data can be directly located through the data identifier. The second step is to establish a correlation index to analyze the inherent relationships between time-series data and unstructured data. For example, time-series monitoring data for a specific period may correspond to device alarm logs for the same period; time-series data for a specific component may correspond to the topological relationship data of that component; and quality assessment results of time-series data may correspond to relevant analysis reports. Based on these relationships, a correlation index table is established with the unified data identifier as the core. This table clearly records the correspondence between the unified identifiers of different types of data, such as time-series data identifiers. The fourth step involves mapping the data to the corresponding event log identifiers and topology relationship identifiers. This involves building relationships and forming data assets. Based on the association index table, relationships are established between time-series database data and unstructured data management engine data to enable cross-database association queries for different types of data. Subsequently, all stored data, unified data identifiers, and association indexes are integrated to form a complete data set with clear data types and relationships covering the entire lifecycle of the power distribution cabinet, from installation and commissioning to operation monitoring and maintenance. This is the full lifecycle data asset, and a data asset list is generated, recording information such as the data types, coverage, and association rules included in the assets.

[0044] Based on the time-series characteristics, access frequency, and predictive analysis value density of data, a hierarchical strategy is generated to achieve data classification, enabling data storage resources to be tilted towards high-value, high-frequency access data, thereby improving the utilization efficiency of storage resources. Through the establishment of unified data identification and associated indexes, a close relationship is built between different types of data, making the scattered data into an organic whole. The resulting full lifecycle data assets facilitate the rapid retrieval and management of data.

[0045] In a preferred embodiment of the present invention, step 6 above, which involves constructing a digital twin of the power distribution cabinet based on full lifecycle data assets and using a data analysis model to perform health status assessment and fault prediction analysis on the digital twin to generate intelligent operation and maintenance decisions, may include: In this embodiment of the invention, step 660 involves driving and constructing a digital twin that operates synchronously with the physical distribution cabinet in real time or is reconstructed at selected time points, based on the full lifecycle data assets. Specifically, this includes: screening and preprocessing the full lifecycle data assets; firstly, selecting data directly related to the physical structure, electrical performance, mechanical state, and operating history of the distribution cabinet from the constructed full lifecycle data assets, including the size parameters, material properties, electrical connection topology data, historical operating sequence monitoring data, fault record data, maintenance history data, and environmental parameter data of each component of the distribution cabinet; preprocessing the screened data to remove duplicate data and correct identified abnormal data to ensure data integrity and consistency; based on the physical parameters and electrical schematic diagram of the distribution cabinet, constructing a virtual 3D model that matches the physical distribution cabinet 1:1 using 3D modeling technology; accurately reproducing the cabinet structure, installation positions of internal components, electrical connection relationships, and mechanical transmission structure of the distribution cabinet in the virtual 3D model, providing a physical carrier for subsequent data mapping and state restoration; establishing a one-to-one correspondence between the full lifecycle data assets and each component of the virtual 3D model, and constructing the data... The system employs a mapping rule base to directly associate pre-processed static data with the corresponding component attributes of the virtual model, serving as the basic static parameters of the virtual model. Dynamic time-series data is synchronized and linked to the dynamic attribute fields of the corresponding components via a data interface, achieving precise binding between data and virtual components. Through a high-speed communication link, the system collects and uploads the latest operational data from the physical distribution cabinet to the full lifecycle data asset, pushing it to the digital twin system in real-time at a preset synchronization frequency. Based on the data mapping rules, the system automatically fills the dynamic attributes of the corresponding components in the virtual 3D model with real-time data, driving the virtual model's operational status to maintain real-time synchronization with the physical distribution cabinet, achieving a virtual real-time restoration of the physical state. The system receives a target time point input by the user, extracts the complete operational data, component status data, and environmental data corresponding to that time point from the full lifecycle data asset, and imports the data from that time point in batches into the virtual 3D model according to the data mapping rules. It resets the static and dynamic attributes of each component in the virtual 3D model to the state of the target time point, completing the precise reconstruction of the distribution cabinet's operational status at that time point, and allowing users to review historical states.

[0046] Step 661: Using the data analysis model deployed on the digital twin, analyze the real-time operating status of the digital twin, calculate the comprehensive health assessment index, complete the quantitative assessment of the current health status of the distribution cabinet, and generate a health status assessment result. Specifically, this includes: extracting historical operating data, fault data, and maintenance record data from multiple similar distribution cabinets from the full lifecycle data assets as sample datasets for training the data analysis model; labeling the sample data according to normal operating status, minor abnormal status, severe abnormal status, and fault status; dividing the labeled sample data into training and validation sets in a 7:3 ratio; and, based on the distribution cabinet's operating mechanism, selecting characteristic variables that reflect the operating status of each component and the overall system from the real-time operating parameters associated with the digital twin, including current stability characteristics, voltage fluctuation characteristics, component temperature characteristics, vibration intensity characteristics, and switch action response characteristics; and further analyzing the selected data. The feature variables are normalized to eliminate the impact of dimensional differences on the training of the data analysis model. A deep learning algorithm is used to construct the data analysis model. The input layer of the data analysis model is the selected normalized feature variables, and the hidden layer is set with 3-5 layers of neural network to explore the deep correlation between feature variables. The output layer is the health status assessment result. The training set data is input into the data analysis model for iterative training. After each round of training, the model output accuracy is verified using validation set data. The performance of the data analysis model is optimized by adjusting parameters such as the number of network layers, the number of neurons, and the learning rate, until the evaluation accuracy of the data analysis model on the validation set reaches the preset requirements, thus completing the construction of the data analysis model. The trained data analysis model is deployed to the digital twin system, and a real-time calling interface for the data analysis model and the status data of each component of the digital twin is configured to ensure that the data analysis model can obtain the running status data of the virtual model in real time.

[0047] Through the interface of the digital twin system, real-time operational status data of each component of the virtual model is collected, including real-time current and voltage data of each electrical circuit, real-time temperature data of each connection point, real-time vibration data of key mechanical components, and real-time action status data of switches. The collected real-time status data is input into the deployed data analysis model, which performs real-time data parsing, extracting the current values ​​and trends of each characteristic variable. First, the health weight corresponding to each characteristic variable is determined, based on historical fault data statistics. Higher weights are assigned to characteristic variables with a greater impact from faults, such as connection point temperature (0.25), current stability (0.2), vibration intensity (0.18), voltage fluctuation (0.15), switch action response (0.12), and other auxiliary characteristics (0.1). Then, the weights of each characteristic variable are calculated. The individual health score for each variable is based on its normal value range. If the current value is within the normal range, the individual score is 100 points. If it exceeds the normal range, points are deducted proportionally. For example, 10 points are deducted for exceeding the upper limit of the normal range by 10%, 20 points for exceeding it by 20%, and so on. Finally, a weighted summation is used to calculate the comprehensive health assessment index. The calculation method is to multiply the individual health score of each variable by its corresponding weight, and then add all the products together. The sum is the comprehensive health assessment index. A grading standard for the comprehensive health assessment index is preset. The calculated comprehensive health assessment index is compared with the grading standard to determine the current health level of the distribution cabinet. At the same time, combined with the abnormal characteristic variables parsed by the data analysis model, a health status assessment result containing the health level, the name of the abnormal characteristic variable, and the degree of abnormality is generated.

[0048] Step 662: Based on the health status assessment results and combined with the evolution patterns of accumulated time-series operational data and performance degradation analysis architecture in the full lifecycle data assets, predict and analyze the probability of failure and development trend of key components of the distribution cabinet, and generate failure trend prediction analysis results. Specifically, this includes: extracting abnormal characteristic variables and corresponding health level data from the health status assessment results; extracting accumulated time-series operational data, historical performance degradation data of key components, and historical failure data of the distribution cabinet and similar distribution cabinets from the full lifecycle data assets; organizing the extracted data, sorting them chronologically to form a coherent data sequence with time as the axis; calling the preset distribution cabinet performance degradation analysis architecture, which is built based on long-term operational data and failure statistics of similar distribution cabinets, including the performance degradation path, degradation rate, and typical failure mode mapping relationship of each key component under different health levels; adapting the current distribution cabinet's health status assessment results, accumulated time-series operational data, and performance degradation analysis architecture to determine the current performance degradation stage of the distribution cabinet; and selecting historical data samples corresponding to the adapted performance degradation stage. This sample consists of operational data from similar distribution cabinets at the same degradation stage. Combined with the current distribution cabinet's abnormal characteristic variable data, the data analysis model constructed in step 661 is used for extended prediction. Specifically, the data analysis model compares the similarity between the current data and historical pre-fault data, and, combined with the baseline value of the fault occurrence probability corresponding to the degradation stage, calculates the probability of occurrence for each fault type in the current state. For example, if the similarity between the current data and pre-fault data for a certain type is 80%, and the baseline probability of occurrence for this type of fault in this degradation stage is 30%, then the current fault occurrence probability is the product of the similarity and the baseline probability, then corrected. Based on the accumulated time-series operational data, time change curves for each abnormal characteristic variable are plotted. The changing trend of the abnormal characteristic variable is analyzed in conjunction with the corresponding degradation rate in the performance degradation analysis architecture. Based on the changing trend and the fault occurrence probability, the possible time range and severity level of the fault are predicted. The above analysis results are integrated to generate a fault trend prediction analysis result containing the name of key components, possible fault types, fault occurrence probability, predicted time range, fault development trend level, and impact range.

[0049] Step 663 integrates the quantitative health assessment results with the fault trend prediction analysis results, and generates intelligent operation and maintenance decisions through preset decision rules, including suggestions on the timing and content of preventive maintenance, fault risk warnings at different levels, and operational parameter optimization and adjustment strategies. Specifically, this includes: associating and integrating the health status assessment results generated in step 661 with the fault trend prediction analysis results generated in step 662, removing duplicate information, and forming a unified status and prediction fusion dataset; for example, associating a poor health level with a 60% probability of terminal overheating faults and the possibility of occurrence within one week to determine the correspondence between the current status and potential faults; and calling preset intelligent operation and maintenance rules. A decision rule base is established, built upon distribution cabinet operation and maintenance specifications and fault handling priorities. It includes operation and maintenance handling logic corresponding to different health levels, fault probabilities, and fault types. The fused dataset is matched with the decision rule base to determine the corresponding operation and maintenance handling priorities and directions. For example, the rule base pre-determines that a dangerous health level or a fault occurrence probability ≥80% corresponds to the highest priority, requiring immediate handling; a poor health level with a fault occurrence probability of 50%-80% corresponds to high priority, requiring handling within 3 days; and a moderate health level with a fault occurrence probability of 30%-50% corresponds to medium priority, requiring a maintenance plan to be developed within 1 week. Based on the matched handling priorities and fault probabilities... The time range is measured to determine the timing of preventative maintenance; combined with fault type and abnormal characteristic variables, the maintenance content is determined. For example, a 60% probability of terminal overheating fault corresponds to maintenance content such as checking the tightness of the terminal, cleaning the oxide layer on the terminal surface, and testing the insulation performance of the terminal; abnormal circuit breaker vibration corresponds to maintenance content such as checking the circuit breaker fixing bolts, replacing aging buffer components, and calibrating the operating mechanism; based on the probability of fault occurrence and the development trend level, different levels of fault risk warnings are generated; each warning level includes warning information such as the warning cause, potential fault impact, and emergency handling tips; combined with historical best operating parameter data and current operating parameters in the full life cycle data asset, For abnormal characteristic variables, optimization and adjustment strategies for operating parameters are generated. For example, if the current fluctuation is too large and the health level drops, the optimization strategy is to adjust the load distribution, transferring 20% ​​of the current circuit load to the backup circuit to stabilize the current within the normal range. If the ambient temperature is too high and the component temperature rises, the optimization strategy is to turn on the cabinet heat dissipation device and adjust the operating power of the heat dissipation device to control the cabinet temperature within the range of 25-35℃. Preventive maintenance suggestions, fault risk warnings, and operating parameter optimization and adjustment strategies are integrated into a complete intelligent operation and maintenance decision document, which determines the implementing body, execution time limit, and execution requirements of each decision content, and generates standardized fault trend prediction and analysis results.

[0050] By building a digital twin based on full lifecycle data assets, the virtual and accurate restoration of the physical distribution cabinet's operating status and the retrospective of its historical status can be achieved. This clarifies the direction and requirements of operation and maintenance work, improves the pertinence and efficiency of operation and maintenance work, extends the service life of the distribution cabinet, and reduces operation and maintenance costs and downtime losses.

[0051] like Figure 2 As shown, embodiments of the present invention also provide a data processing system for the entire lifecycle of intelligent operation and maintenance of power distribution cabinets, including: The data acquisition module is used to collect multi-source heterogeneous data from the power distribution cabinet to form a raw full lifecycle dataset. The building module is used to construct a multi-dimensional data sampling window based on the key state parameters in the original full life cycle dataset, and to establish the topological relationship between the core feature parameters within the sampling window; The generation module is used to perform structured topological partitioning of the data within the sampling window based on topological relationships, calculate the spatial shape feature similarity of the data distribution in each partition, and generate data quality calibration parameters based on the data distribution characteristics, correlation strength, and shape consistency index of each partition. The encapsulation module is used to transform and standardize the original full lifecycle dataset using data quality calibration parameters to obtain a standardized transmission dataset. The management module is used to upload standardized transmission datasets through a high-speed communication network. Based on the time-series characteristics and value density of the standardized transmission datasets, a hybrid storage architecture combining a time-series database and an unstructured data management engine is adopted on the central data processing side to achieve data hierarchical, compression, and associated storage, forming full lifecycle data assets. The decision-making module is used to construct a digital twin of the power distribution cabinet based on the full life cycle data assets, and to use data analysis models to conduct health status assessment and fault prediction analysis on the digital twin to generate intelligent operation and maintenance decisions.

[0052] It should be noted that this system is a system corresponding to the above method. All implementation methods in the above method embodiments are applicable to this embodiment and can achieve the same technical effect.

[0053] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for intelligent operation and maintenance of power distribution cabinets throughout their entire lifecycle, characterized in that: The method includes: Step 1: Collect multi-source heterogeneous data from the power distribution cabinet to form the original full lifecycle dataset; Step 2: Based on the key state parameters in the original full lifecycle dataset, construct a multi-dimensional data sampling window and establish the topological relationship between the core feature parameters within the sampling window; Step 3: Based on topological relationships, perform structured topological partitioning of the data within the sampling window, calculate the spatial shape feature similarity of the data distribution in each partition, and generate data quality calibration parameters based on the data distribution characteristics, correlation strength, and shape consistency index of each partition. Step 4: Using data quality calibration parameters, perform data transformation and standardization encapsulation on the original full lifecycle dataset to obtain a standardized transmission dataset; Step 5: Upload the standardized transmission dataset through a high-speed communication network. Based on the time-series characteristics and value density of the standardized transmission dataset, adopt a hybrid storage architecture that combines a time-series database and an unstructured data management engine on the central data processing side to achieve data hierarchical, compressed, and associated storage, forming a full lifecycle data asset. Step 6: Based on the full lifecycle data assets, construct a digital twin of the power distribution cabinet, and use data analysis models to conduct health status assessment and fault prediction analysis on the digital twin to generate intelligent operation and maintenance decisions.

2. The method for intelligent operation and maintenance full lifecycle data processing of power distribution cabinets according to claim 1, characterized in that, Based on the key state parameters in the original full lifecycle dataset, a multi-dimensional data sampling window is constructed, and the topological relationships between core feature parameters are established within the sampling window, including: From the original full life cycle dataset, key state parameters reflecting the electrical performance, mechanical condition and environmental conditions of the distribution cabinet are extracted to form a set of key state parameters. Based on the set of key state parameters, the key state parameters are synchronously sampled and organized using a preset fixed time interval as the time base and an alarm event exceeding a threshold as the event trigger base, to form a synchronous multi-dimensional data sampling window. Within the synchronous multi-dimensional data sampling window, the electrical connection logic and physical location association between parameters are determined based on the electrical principle and physical layout information of the power distribution cabinet. At the same time, the time domain and frequency domain correlation of the time series of each parameter within the window are calculated. By combining the electrical connection logic, physical location association, and time series correlation, a directed or undirected topological relationship is generated with key state parameters as nodes and the mutual influence and constraint relationship between parameters as edges.

3. The method for intelligent operation and maintenance full lifecycle data processing of power distribution cabinets according to claim 2, characterized in that, The key electrical performance parameters include at least current, voltage, power, and harmonics; the key mechanical parameters include at least vibration intensity, displacement, and connection point temperature; and the key environmental parameters include at least ambient temperature, humidity, and dust concentration.

4. The method for intelligent operation and maintenance full lifecycle data processing of power distribution cabinets according to claim 3, characterized in that, Based on topological relationships, the data within the sampling window is structured into topological partitions, and the spatial shape feature similarity of the data distribution in each partition is calculated, including: Based on the electrical connections and logical relationships determined by the topological relationships, the multi-dimensional operating parameter data collected within the sampling window are structured and divided into several data partitions with strong internal correlations according to the associated equipment groups or electrical circuit units. The distribution characteristics of all data samples in each data partition in the multidimensional feature space are analyzed to determine the main direction, range and density of the spatial distribution of the point set, and to construct a geometric structure for overall characterization of the spatial distribution of the point set. For the geometric structure, a set of geometric feature parameters are extracted to quantitatively describe the extent, directionality and discreteness of the geometric structure in space, and the set of geometric feature parameters is constructed into the corresponding feature vector; By calculating the similarity measure between feature vectors, the feature vectors of the same data partition are compared across different sampling windows, thus quantifying the similarity of spatial shape features.

5. The method for intelligent operation and maintenance full lifecycle data processing of power distribution cabinets according to claim 4, characterized in that, Based on the data distribution characteristics, correlation strength, and shape consistency indices of each partition, data quality calibration parameters are generated, including: The distribution dispersion characteristics of time-series monitoring data within each partition are calculated to obtain the first evaluation index characterizing the degree of data fluctuation and dispersion. At the same time, the matching degree between the data sequence of the partition and the state space shape characteristics of the typical equipment corresponding to the partition is calculated to obtain the second evaluation index characterizing the data self-consistency within the partition. The first evaluation index and the second evaluation index are integrated to generate preliminary data quality evaluation results. Based on the preliminary data quality assessment results, and combined with the electrical connection topology and physical layout relationship of the equipment units in the distribution cabinet, a partition correlation analysis framework is constructed. Data feature consistency analysis is performed on the partitions with correlation relationships to identify and quantify the contradictions and conflicts between the correlated partitions in terms of data trends, amplitudes or states, and generate a partition correlation contradiction diagnosis report. The preliminary data quality assessment results are integrated with the inter-regional correlation contradiction diagnosis report to form a diagnostic evidence set. The diagnostic evidence set is then subjected to multi-factor comprehensive analysis to achieve comprehensive source tracing, location and type diagnosis of data anomaly patterns. Based on the final diagnostic conclusion, data quality calibration parameters are quantitatively generated.

6. The method for intelligent operation and maintenance full lifecycle data processing of power distribution cabinets according to claim 5, characterized in that, The data quality calibration parameters include a credibility coefficient vector characterizing the relative credibility of each data source, a set of anomaly pattern determination thresholds, and suggested calibration benchmark values ​​and direction vectors.

7. The method for intelligent operation and maintenance full lifecycle data processing of power distribution cabinets according to claim 6, characterized in that, Using data quality calibration parameters, the original full lifecycle dataset is transformed and standardized to obtain a standardized transport dataset, including: By applying data quality calibration parameters, weighted correction is performed on multi-source heterogeneous data in the original full lifecycle dataset to obtain calibrated data. The calibrated data is subjected to dimension unification processing to obtain dimension unified data. Missing values ​​in the dimension unified data are imputed to obtain a complete and standardized dataset. The complete and standardized dataset is packaged into a standardized data packet according to the preset communication protocol and data structure specifications. The packet includes a data body consisting of the complete and standardized dataset, a quality label generated based on the data quality calibration parameters, and a topology identifier used to identify the location relationship of the data source distribution cabinet and components. All the encapsulated standardized data packets are integrated to form a standardized transmission dataset.

8. The intelligent operation and maintenance full lifecycle data processing method and system for power distribution cabinets according to claim 7, characterized in that, Standardized transmission datasets are uploaded via high-speed communication networks. Based on the time-series characteristics and value density of these datasets, a hybrid storage architecture combining a time-series database and an unstructured data management engine is adopted on the central data processing side. This enables data tiering, compression, and associative storage, forming a full lifecycle data asset, including: The standardized transmission dataset is uploaded to the central data processing side via a high-speed communication network to obtain the uploaded standardized transmission dataset; The data included in the uploaded standardized transmission dataset is analyzed, and a data grading strategy is generated based on the data's temporal characteristics, access frequency, and predictive analysis value density. Based on the data classification strategy, the uploaded standardized transmission dataset is processed to classify the data, resulting in data subsets that include data at different levels. Time-series monitoring data, which is classified into data collected continuously at preset short cycles, is compressed and stored using a time-series database. Unstructured data, which is classified into event logs, topology relationships, and analysis reports, is stored using an unstructured data management engine. Establish unified data identifiers and association indexes for data stored in time-series databases and data stored in unstructured data management engines. Based on the unified data identifiers and association indexes, build association relationships between different types of data to form full lifecycle data assets.

9. The method for intelligent operation and maintenance of power distribution cabinets throughout their entire lifecycle data processing according to claim 8, characterized in that, Based on full lifecycle data assets, a digital twin of the power distribution cabinet is constructed, and a data analysis model is used to conduct health status assessment and fault prediction analysis on the digital twin to generate intelligent operation and maintenance decisions, including: Based on full lifecycle data assets, drive and build a digital twin that operates in real time synchronously with the physical distribution cabinet or is reconstructed at selected time points; Using a data analysis model deployed on the digital twin, the real-time operating status of the digital twin is analyzed, a comprehensive health assessment index is calculated, a quantitative assessment of the current health status of the power distribution cabinet is completed, and a health status assessment result is generated. Based on the health status assessment results, and combined with the time-series operation data accumulated in the full life cycle data assets and the evolution law of the performance degradation analysis architecture, the probability of failure and development trend of key components of the power distribution cabinet are predicted and analyzed, and the failure trend prediction analysis results are generated. By integrating the quantitative assessment results of health metrics with the results of fault trend prediction analysis, intelligent operation and maintenance decisions are generated through preset decision rules, including suggestions on the timing and content of preventive maintenance, early warnings of different levels of fault risks, and strategies for optimizing and adjusting operating parameters.

10. A data processing system for the entire lifecycle of intelligent operation and maintenance of power distribution cabinets, wherein the system implements the method as described in any one of claims 1 to 9, characterized in that, include: The data acquisition module is used to collect multi-source heterogeneous data from the power distribution cabinet to form a raw full lifecycle dataset. The building module is used to construct a multi-dimensional data sampling window based on the key state parameters in the original full life cycle dataset, and to establish the topological relationship between the core feature parameters within the sampling window; The generation module is used to perform structured topological partitioning of the data within the sampling window based on topological relationships, calculate the spatial shape feature similarity of the data distribution in each partition, and generate data quality calibration parameters based on the data distribution characteristics, correlation strength, and shape consistency index of each partition. The encapsulation module is used to transform and standardize the original full lifecycle dataset using data quality calibration parameters to obtain a standardized transmission dataset. The management module is used to upload standardized transmission datasets through a high-speed communication network. Based on the time-series characteristics and value density of the standardized transmission datasets, a hybrid storage architecture combining a time-series database and an unstructured data management engine is adopted on the central data processing side to achieve data hierarchical, compression, and associated storage, forming full lifecycle data assets. The decision-making module is used to construct a digital twin of the power distribution cabinet based on the full life cycle data assets, and to use data analysis models to conduct health status assessment and fault prediction analysis on the digital twin to generate intelligent operation and maintenance decisions.