Intelligent electric meter electricity utilization behavior interpretability analysis method based on causal atlas
Through multi-source data fusion and multi-time-scale causal relationship discovery, multi-level causal map is constructed, and the causal map is dynamically updated and the real-time effective path index is solved, which solves the problems of causal map calculation complexity and index maintenance overhead, and realizes efficient interpretability analysis of the power consumption behavior of smart meters.
Patent Information
- Application Number
- CN202510756548.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-09
AI Technical Summary
The existing technology has high computational complexity in the dynamic maintenance of causal maps and high maintenance overhead in the analysis of power consumption behavior of smart meters, resulting in insufficient real-time and scalability, affecting the engineering application of causal maps in real-time power consumption behavior analysis.
Through multi-source data fusion and multi-time-scale causal relationship discovery, multi-level causal maps are built, and the causal map and real-time effective path index are dynamically updated, and the multi-level causal map and real-time effective path index are combined, dynamic weight updates and hierarchical resource allocation are adopted to optimize the calculation and index maintenance of the causal map.
It reduces the complexity of causal map update and index maintenance overhead, improves query efficiency, and realizes efficient interpretability analysis of the power consumption behavior of smart meters.
Smart Images

Figure CN120278399A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to power electronics technology, especially an interpretable analysis method for electricity consumption behavior of smart meters based on a causal graph. Background Art
[0002] With the rapid development of smart grid technology and the increasing complexity of user electricity consumption behavior, smart meters have become an important infrastructure for the digital transformation of the power system. The interpretable analysis method for electricity consumption behavior of smart meters based on a causal graph can not only accurately identify abnormal electricity consumption, but more importantly, can provide a transparent and understandable explanation mechanism for users and power companies. Interpretable analysis is of great significance for improving the level of user electricity consumption management, optimizing the operation efficiency of the power grid, and supporting smart grid decision-making. Especially in the context of carbon neutrality, accurately understanding the internal causal mechanism of user electricity consumption behavior helps to formulate more accurate energy conservation and emission reduction strategies and promote the formation of a green electricity consumption model. At the same time, interpretable electricity consumption behavior analysis is also a key technical support for building user trust and promoting the popularization and application of smart grid technology.
[0003] Currently, in the field of smart meter electricity consumption behavior analysis, there are mainly two categories of methods: those based on statistical learning and those based on deep learning. Statistical learning methods mainly use traditional statistical techniques such as time series analysis and regression analysis, and model and analyze electricity consumption data by constructing ARIMA models, support vector machines, etc. Deep learning methods mainly use neural network architectures such as LSTM and CNN, and identify electricity consumption patterns and abnormal behaviors through the training of a large amount of historical electricity consumption data. In terms of causal relationship identification, existing research mainly uses Granger causality test as a single causal identification method to judge the causal relationship between variables through a lag regression model. In terms of graph construction, most methods adopt a static graph structure, pre-defining the relationship between nodes and edges, lacking dynamic adaptability. In terms of anomaly detection, existing methods mainly rely on threshold judgment or clustering analysis to identify abnormal electricity consumption behavior by setting fixed thresholds or distance metrics.
[0004] However, the existing technical solutions face two key technical bottlenecks in practical applications: one is the computational complexity problem of dynamic maintenance of the causal graph. When the electricity consumption behavior changes, traditional methods need to recalculate the relationships between all nodes and edges in the entire causal graph, and the computational complexity is O(n 2)Level, which leads to serious lack of real-time performance in the scenario of large-scale electricity consumption data; second, the problem of excessive maintenance overhead of pre-computed path indexes. To improve the efficiency of causal path queries, existing methods usually adopt a full-scale pre-computed index strategy. However, when the causal graph changes, all pre-computed paths need to be updated synchronously, and the index maintenance overhead accounts for more than 30% of the total computing resources, seriously affecting the overall performance and scalability of the system. These two problems directly restrict the engineering application of causal graph methods in real-time electricity consumption behavior analysis. Summary of the Invention
[0005] The object of the invention is to provide an interpretable analysis method for the electricity consumption behavior of smart meters based on a causal graph, in order to solve at least one technical problem existing in the prior art.
[0006] Technical solution: An interpretable analysis method for the electricity consumption behavior of smart meters based on a causal graph includes the following steps: Read the original electricity consumption data of smart meters, and construct a primary multi-level causal graph through multi-source data fusion and multi-time scale causal relationship discovery; When an electricity consumption anomaly is detected, identify the influence range of the abnormal node in the primary multi-level causal graph, maintain and update it, and generate a dynamic multi-level causal graph and a real-time effective path index; Receive a user query request, search sequentially from the key and important path indexes in the real-time effective path index, and if not found, conduct a real-time search to obtain a set of target causal paths; Read the set of target causal paths, verify and interpret the output through counterfactual reasoning, and generate a verified causal explanation report.
[0007] Furthermore, through multi-source data fusion and multi-time scale causal relationship discovery, construct a multi-level causal graph, including: Read the original electricity consumption data of smart meters, perform standardization and stratification processing, and generate a stratified time window data set; Based on the stratified time window data set, calculate MI, GC, and TC and fuse them to generate a multi-dimensional causal intensity vector; Read the multi-dimensional causal intensity vector, calculate each causal relationship and verify its stability at different time scales, and generate a multi-scale causal intensity matrix; Based on the multi-scale causal intensity matrix, construct and store the causal relationships at different time scales in layers, and generate a multi-level causal graph and the corresponding basic causal relationship index.
[0008] Furthermore, generate a multi-dimensional causal intensity vector, including: Based on the stratified time window data set, calculate the mutual information entropy MI, Granger causality GC, and time consistency coefficient TC respectively, and perform normalization processing respectively; Calculate the dynamic weights α, β, and γ of MI, GC, and TC based on the window stability coefficient and the temporal correlation strength; Read the normalized MI, GC, TC, and dynamic weights α, β, γ, and calculate the fusion strength K = (α × MI 2 + β × GC 2 + γ × TC 2 ) 0.5 , and generate a multi-dimensional causal strength vector.
[0009] Furthermore, generate a real-time effective path index, including: Combine the multi-level causal graph, identify abnormal nodes from the current real-time power consumption data stream, calculate the local influence propagation range, and generate an affected area identifier; Based on the affected area identifier, through a three-layer computing resource intelligent allocation and time slice round-robin scheduling strategy, perform distributed asynchronous update scheduling on the real-time layer, quasi-real-time layer, and background layer to generate an optimized hierarchical update task queue; For the hierarchical update task queue and the real-time power consumption data stream, perform weight update and accuracy correction to generate a dynamically maintained multi-level causal graph and a real-time effective path index.
[0010] Furthermore, generate a dynamically maintained multi-level causal graph and a real-time effective path index, including: Based on the optimized hierarchical update task queue and the real-time power consumption data stream, calculate the new edge weight Q = old weight - (expired data contribution / window size) + (new data contribution / window size), and perform incremental calculation based on this to obtain an incremental update weight table; For the incremental update weight table, record the continuous update times and error change trends of each edge, and when it exceeds the preset value, mark it as an edge that needs to be corrected; For the edges that need to be marked, recalculate the edge weights and correct them based on the complete historical data to generate a dynamically maintained multi-level causal graph and a real-time effective path index.
[0011] Furthermore, calculate the local influence propagation range and generate an affected area identifier, including: Read the real-time power consumption data stream and the multi-level causal graph, identify the nodes with power changes exceeding the threshold, obtain the abnormal node list and the abnormal intensity level; based on this, calculate the influence area and generate an influence intensity distribution map through the influence propagation formula; The influence propagation formula is: Influence range I = initial intensity I0 × e -λd × cumulative edge weight norm × decay adjustment factor; Where: cumulative edge weight norm = Σ(edge weight i) / sqrt(path length); d is the propagation hop count, limited to 3 hops forward and 3 hops backward, and the adaptive attenuation coefficient λ = 0.3 + 0.2×(abnormal intensity level / 5).
[0012] Further, search sequentially from the critical and important path indexes in the real-time valid path index. If not found, perform real-time search, including: In response to the user query request, identify the source node set and the target node set through query type classification and key entity extraction and output them; Read the dynamic multi-level causal graph, and divide the pre-computed paths in the causal relationship basic index into critical path indexes, important path indexes, and background path indexes to generate a real-time valid path index; For the source node set and the target node set, first search in the critical and important path indexes. If not found, perform real-time search in the background path index.
[0013] Further, generating a real-time valid path index includes: Read the causal relationship basic index and the dynamic multi-level causal graph, calculate the node importance and path importance, and accordingly store the pre-computed paths into the critical, important, and background path indexes respectively to generate a hierarchical path index architecture; Read the hierarchical path index architecture, add a timestamp and a valid flag to each pre-computed path, update the critical paths immediately, update the important paths with a delay, and verify when querying the background paths to generate a real-time valid path index.
[0014] Further, after first searching in the critical and important path indexes and performing real-time search in the background path index if not found, it also includes: Based on the search results, form a candidate path set and a search time consumption record; For the candidate path set, in combination with the dynamic multi-level causal graph, parallelly check the current causal strength and connectivity of each edge in multiple candidate paths, and eliminate the paths containing invalid edges to form a set of valid paths; For each valid path, perform path scoring and sorting to obtain a set of target causal paths; Among them, path scoring = (causal strength weight × a + (path length weight × b) + (historical success rate × c); a + b + c = 1.
[0015] Further, adopt counterfactual reasoning verification and explanation output to generate a verified causal explanation report including: Based on the target causal path set and the dynamic multi-level causal graph, remove or modify the key causal edges, combine with parameter perturbation simulation, construct the counterfactual scenario set and the expected result range, and combine with the original power consumption data of the smart meter. The engine calculates the changes in the power consumption pattern under each scenario, compares with the historical actual data for consistency, and forms an interpretability credibility score and a verification result report; For the verification result report, adopt the logical organization of the causal chain and the quantification of technical indicators to generate a causal explanation report including the causal relationship strength, the impact time delay, and the confidence interval.
[0016] Further, removing or modifying the key causal edges includes: When the edge in the causal path fails, select an alternative path from the alternative paths, or recalculate a new valid path.
[0017] Beneficial effects: reduce the update complexity and the index maintenance overhead, improve the query efficiency, and provide an efficient interpretable technical solution for the power consumption behavior analysis of smart meters. Description of the Drawings
[0018] Figure 1 is the flowchart of the present invention.
[0019] Figure 2 is the flowchart of constructing the multi-level causal graph of the present invention.
[0020] Figure 3 is the flowchart of generating the multi-dimensional causal strength vector of the present invention.
[0021] Figure 4 is the flowchart of generating the real-time valid path index of the present invention. Detailed Embodiment
[0022] As Figures 1 to 4 shown, a method for interpretable analysis of the power consumption behavior of a smart meter based on a causal graph is provided, including the following steps: S1. Read the original power consumption data of the smart meter, the basic user information, and the external environment data, and construct a multi-level causal graph and a causal relationship basic index through data preprocessing and an improved multi-window causal strength calculation method.
[0023] S11. Read the original power consumption data of the smart meter (such as the power time series at intervals of 1, 3, and 15 minutes, and the sampling periods of different meters are different), the basic user information (equipment list, house characteristics), and the external environment data (temperature, humidity, weather), and obtain a synchronized and standardized data matrix through timestamp alignment and data type unification processing.
[0024] S12. Parallel calculation of the causal strength in multiple time windows S121. Read the synchronized standardized data matrix, and simultaneously construct an A-second high-frequency response window, a B-second medium-frequency stable window, and a C-second low-frequency trend window through a triple nested window mechanism. Each window independently maintains a data buffer to obtain a hierarchical time window dataset and a window synchronization flag. The traditional method uses a single fixed window. The three-layer nested mechanism of this method can capture causal features at different time scales simultaneously.
[0025] S122. Calculate the mutual information entropy-Granger fusion causal strength, including: reading the hierarchical time window dataset, and through the comprehensive causal strength = α·MI(X;Y|Z) + β·GC(X→Y) + γ·TC(X,Y,t); where: MI is the mutual information entropy, GC is the Granger causality, TC is the time consistency coefficient, α + β + γ = 1, and the weights are dynamically adjusted according to the window type to obtain a multi-dimensional causal strength vector. Existing methods use Granger causality or mutual information alone. This method fuses three causal metrics to improve the robustness of causal recognition.
[0026] It solves the problem of insufficient recognition accuracy of traditional single causal metric methods in the complex electricity consumption environment of smart meters. Specifically, MI can capture non-linear causal relationships (such as the non-linear response of temperature to air conditioner electricity consumption), GC can identify time-series causal relationships (such as the delay effect of device startup), and TC can verify the time stability of causal relationships (such as the consistency of periodic electricity consumption patterns). Through the adaptive weight mechanism of α = 0.4 + 0.2 × window stability coefficient, when the electricity consumption data fluctuates greatly, the MI weight is automatically increased to improve the non-linear recognition ability, and when the time-series correlation is strong, the GC weight is increased to highlight the time causal features. The Euclidean norm fusion ensures the geometric meaning of the comprehensive strength, improving the causal recognition accuracy from 65% of the traditional single method to 82%.
[0027] S123. Calculate the causal relationship stability score, including: reading the multi-dimensional causal strength vector and the window synchronization flag, and calculating the stability degree of each causal relationship at different time scales through a cross-window consistency verification algorithm to obtain the causal stability score and the multi-scale causal strength matrix.
[0028] S13. Generate a hierarchical causal graph construction and index, including: reading the multi-scale causal strength matrix and the causal stability score, and storing the causal relationships at different time scales in layers through a hierarchical graph construction algorithm to generate a multi-level causal graph (including a real-time layer, a quasi-real-time layer, and a background layer) and a causal relationship basic index (pre-compute the top 3 optimal paths between key nodes).
[0029] S2: Read the real-time electricity consumption data stream and the multi-level causal graph, and through a load-balanced distributed update mechanism, realize a dynamically updated causal graph and an affected area identifier.
[0030] S21: Intelligent Identification Mechanism for Affected Areas S211: Fast Abnormal Node Location Algorithm, including: reading real-time power consumption data stream and multi-level causal graph, identifying nodes with power change exceeding 200% of the baseline through a threshold dynamic adjustment mechanism (calculated based on 3σ standard deviation of user historical baseline), obtaining a list of abnormal nodes and abnormal intensity levels.
[0031] S212: Calculate the local influence propagation range, including: reading the list of abnormal nodes, abnormal intensity levels and multi-level causal graph, through the influence propagation attenuation formula: influence range I(d) = initial intensity I0 × e^(-λd) × cumulative edge weight; where: d is the propagation distance, λ is the attenuation coefficient, restricting the propagation depth to 3 hops forward and 3 hops backward, obtaining the affected area identifier and influence intensity distribution map. Traditional methods require full graph traversal calculation, with a complexity of O(n 2 ). This method reduces the complexity to O(k·logn) through influence propagation limitation, where k is the number of affected nodes, greatly improving the calculation efficiency.
[0032] Restrict the causal graph update range within 3 hops forward and 3 hops backward of the abnormal nodes, solving the problem of computational complexity explosion caused by traditional full graph traversal updates. In the application scenario of smart meters, when an abnormal situation occurs in a certain electrical device, its influence usually has local characteristics. For example, an air conditioner failure mainly affects the temperature control loop and power supply loop directly related to it, rather than independent device loops such as water heaters or washing machines. Through the adaptive attenuation coefficient λ = 0.3 + 0.2×(abnormal intensity level / 5), the higher the abnormal degree, the wider the propagation range. The cumulative normalization of edge weights avoids the problem of numerical instability in long-distance propagation, and the attenuation adjustment factor min(1.0, 2.0 / d) ensures reasonable attenuation of propagation intensity. This method reduces the update calculation complexity from O(n 2 ) to O(k•logn). In a smart meter network with a scale of 1000 nodes, the calculation amount is reduced from 1,000,000 times to 24 times, and the calculation efficiency is increased by 99.998%, achieving real-time causal graph update ability at the millisecond level.
[0033] S213: Dynamic Influence Weight Allocation Mechanism, including: reading the influence intensity distribution map and causal stability score, through the inverse proportional adjustment algorithm of influence weight and stability, assigning higher update priorities to unstable causal relationships, obtaining a weighted influence area map.
[0034] S22: Distributed Asynchronous Update Scheduling Mechanism S221: Three-layer computing resource intelligent allocation algorithm, including: reading the weighted influence area map and the abnormal intensity level, and through the resource allocation strategy: allocate 15% CPU for the real-time layer (processing high-priority causal edges), 20% CPU for the quasi-real-time layer (processing medium-priority causal edges), and 10% CPU for the background layer (processing low-priority causal edges), dynamically adjust the allocation ratio according to the severity of the abnormality to obtain a hierarchical resource allocation table. The existing method uses synchronous batch updates, resulting in calculation load spikes. The three-layer asynchronous resource allocation of this method avoids calculation competition and ensures real-time response.
[0035] S222: Implement the time slice round-robin scheduling strategy, including: reading the hierarchical resource allocation table and the affected area identifier, and through the time slice round-robin mechanism: the real-time layer continuously executes a 50ms time slice, the quasi-real-time layer executes once every 30 seconds, and the background layer executes once every 5 minutes, and adopt preemptive scheduling (abnormal events can interrupt low-priority tasks) to obtain a hierarchical update task queue and a scheduling execution time table. The time slice round-robin mechanism based on priority solves the resource competition problem of traditional synchronous updates and realizes load balancing.
[0036] S223: Dynamic load monitoring and adaptive adjustment, including: reading the scheduling execution time table and the system's real-time CPU usage rate, and through the load monitoring algorithm, real-time detect the computing load of each layer. When the total CPU usage rate exceeds 50%, automatically degrade the execution strategy (pause the background layer update and extend the quasi-real-time layer interval) to obtain an optimized hierarchical update task queue.
[0037] It solves the problems of calculation load spikes and response delays caused by traditional synchronous batch updates. In the real-time monitoring scenario of smart meters, different types of causal relationships have different timeliness requirements: critical equipment failures require real-time response (such as abnormal meter tripping), general equipment status changes can be processed quasi-real-time (such as normal load switching), and historical trend analysis can be processed in the background (such as monthly electricity consumption pattern analysis). Dynamically adjust the CPU allocation ratio according to the abnormal intensity level through the resource intelligent allocation algorithm to ensure the priority processing of key tasks. The dynamic load monitoring mechanism automatically triggers the degradation strategy when the total CPU usage rate exceeds 50%, pauses the background layer update and extends the quasi-real-time layer interval, avoiding system overload.
[0038] S23: Incremental causal edge weight correction algorithm S231: Sliding window incremental weight update, including: reading the optimized hierarchical update task queue and the real-time electricity data stream, and through the incremental update formula: new edge weight = old weight - (expired data contribution / window size) + (new data contribution / window size); perform incremental calculation on each causal edge that needs to be updated to obtain an incremental update weight table.
[0039] S232: Cumulative error real-time detection mechanism, including: reading the incrementally updated weight table, recording the continuous update times and error change trends of each edge through the error accumulation tracking algorithm, when the error rate exceeds the 5% threshold or the continuous update exceeds 100 times, marking the edge that needs to be corrected, and obtaining the error monitoring report and correction requirement mark.
[0040] S233: Perform adaptive precision correction, including: reading the correction requirement mark and the multi-level causal graph, performing local recalculation on the marked edges (recalculating the edge weights using the complete historical data), resetting the error counter after the correction is completed, and obtaining the dynamically updated causal graph. Traditional incremental updates have the problem of error accumulation that cannot be solved. This method controls the error within 5% while maintaining the update speed through the adaptive correction mechanism.
[0041] It solves the problem that the error accumulation of the traditional incremental update method cannot converge. In the long-term operation scenario of smart meters, the causal relationship needs to be continuously updated according to new power consumption data. Traditional methods either use full recalculation (huge computational overhead) or simple incremental updates (continuous error accumulation). This embodiment records the continuous update times and error change trends of each edge, and automatically triggers local recorrection when the continuous update exceeds 100 times or the error rate exceeds 5%, recalculating the edge weight using the complete historical data and resetting the error counter. It not only avoids frequent full recalculation but also prevents the infinite accumulation of errors, ensuring the long-term stability of causal relationship recognition. In the 30-day continuous operation test, the update accuracy of the edge weights remained above 95%, and the update speed was 15 times higher than that of full recalculation, providing technical support for the reliability maintenance of the causal graph of smart meters.
[0042] S3: Read the user query request, the dynamically updated causal graph, and the causal relationship basic index, and generate a set of target causal paths through the fast path search algorithm maintained by hierarchical indexing.
[0043] S31: Query intent parsing and path requirement recognition, including: reading the user query request, and identifying the source node set, target node set, and query priority label through the query type classifier (real-time anomaly explanation, historical pattern analysis, predictive analysis) and key entity extraction.
[0044] S32: Hierarchical index dynamic maintenance mechanism S321: Build a three - level path index architecture, including: reading the causal relationship basic index and the dynamically updated causal graph. Through a path importance evaluation algorithm (based on node access frequency and causal strength), pre - calculated paths are classified into: critical path index (paths related to core power - consuming devices), important path index (paths related to secondary devices), and background path index (other paths), obtaining a hierarchical path index architecture. Existing pre - calculated indexes use a flat structure with high maintenance overhead. The three - level hierarchical index of this method reduces the maintenance overhead from 30% to less than 10%.
[0045] S322: Implement a lazy update strategy, including: reading the hierarchical path index architecture and the affected area identifier. Through a lazy update mechanism: critical paths are updated immediately, important paths are updated with a delay, and background paths are only verified for validity when queried. And a path timeliness marker (recording the last verification time of the path) is adopted, obtaining a path index with timeliness markers. Traditional indexes need to synchronously maintain all paths. This method greatly reduces the maintenance overhead through the lazy strategy while ensuring the path availability during query.
[0046] S323: Intelligent verification of index consistency, including: reading the path index with timeliness markers and the dynamically updated causal graph. Checking the current state of each edge in the path through a path validity verification algorithm, and replacing or deleting invalid paths, obtaining a real - time valid path index and an index update log.
[0047] S33: Fast path reconstruction and verification algorithm S331: Multi - level degraded path search strategy, specifically: reading the source node set, target node set, real - time valid path index, and query priority label. Through a three - layer degraded search mechanism: the first layer searches from the critical path index (target < 50ms), the second layer searches from the important path index and verifies (target < 100ms), and the third layer conducts a real - time graph search (target < 200ms), obtaining a candidate path set and a search time consumption record.
[0048] In the scenario of real-time abnormal analysis of smart meters, users expect to obtain the explanation results immediately after detecting abnormal electricity consumption. However, the response time of traditional path search methods is uncontrollable, ranging from dozens of milliseconds to several seconds, which cannot meet the real-time requirements. This method adopts a multi-level degradation strategy. First, it attempts to quickly match the target path from the pre-computed critical path index. If not found, it enters the important path index for validity verification, and finally performs real-time graph search as a fallback guarantee. The path quality score = (causal strength weight × 0.4) + (path length weight × 0.3) + (historical success rate × 0.3) ensures the reliability of the returned path and avoids the problem of sacrificing quality for the sake of speed. Traditional path search adopts a single strategy with uncontrollable response time. This method ensures that 95% of the queries are completed within 100 ms through a multi-level degradation mechanism, while ensuring the path quality.
[0049] S332: Path validity parallel verification mechanism Read the candidate path set and the dynamically updated causal graph, and simultaneously check the current causal strength and connectivity of each edge in multiple candidate paths through a parallel verification algorithm, eliminate the paths containing invalid edges, and obtain the valid path set.
[0050] S333: Comprehensive path quality score calculation Read the valid path set and the search time consumption record. Through the comprehensive score formula: path score = (causal strength weight × 0.4) + (path length weight × 0.3) + (historical success rate × 0.3); perform quality scoring and sorting on each valid path to obtain the target causal path set and the path credibility distribution.
[0051] S4: Read the target causal path set, the dynamically updated causal graph, and the original electricity consumption data of the smart meter. Through counterfactual scenario simulation and credibility evaluation, generate a verified causal explanation report.
[0052] S41: Automatically generate counterfactual scenarios, including: read the target causal path set and the dynamically updated causal graph, and generate a set of counterfactual scenarios and the expected result range through causal chain break simulation (removing or modifying key causal edges) and parameter perturbation simulation.
[0053] S42: Simulation verification and credibility quantification evaluation, including: read the set of counterfactual scenarios, the expected result range, and the original electricity consumption data of the smart meter, calculate the change in the electricity consumption pattern under each scenario through a counterfactual simulation engine, and compare it with the historical actual data for consistency to generate an explanation credibility score and a verification result report.
[0054] S43: Generate a comprehensive interpretation report, including: reading the set of target causal paths, the path credibility distribution, the interpretation credibility score, and the verification result report, and generating a verified causal interpretation report containing the causal relationship strength, the impact time delay, and the confidence interval through the logical organization of the causal chain and the quantification of technical indicators.
[0055] According to one aspect of the present application, S321, constructing a three-level path index architecture, specifically: S3211, Quantitatively evaluating the node importance: Read the basic causal relationship index and the dynamically updated causal graph, and use the node importance calculation formula: Node importance = (Degree centrality × 0.3) + (Access frequency × 0.4) + (Sum of causal strengths × 0.3); Where: The degree centrality is the number of edges connected to the node, the access frequency is the number of times the node is involved in historical queries, and the sum of causal strengths is the accumulation of the weights of all connected edges, to obtain the node importance score table.
[0056] S3212, Comprehensively calculating the path importance: Read the node importance score table and the basic causal relationship index, and use the path importance evaluation algorithm: Perform a weighted average of the importance scores of all nodes in the path (the weights of the start and end points are 0.4, and the weights of the intermediate nodes are 0.2), and correct it in combination with the historical query frequency of the path, to obtain the path importance ranking table. Traditional indexes do not distinguish the path importance. This method realizes the intelligent grading of paths through the comprehensive evaluation of node centrality and query frequency.
[0057] S3213, Hierarchical storage of three-level indexes: Read the path importance ranking table, and divide it through the importance threshold: Paths with an importance score > 0.8 are stored in the critical path index, those with 0.5 - 0.8 are stored in the important path index, and those < 0.5 are stored in the background path index. Each level of index uses different data structures for optimization (hash table for critical paths, B+ tree for important paths, and compressed storage for background paths), to obtain the hierarchical path index architecture.
[0058] According to one aspect of the present application, S322, Refining the lazy update strategy, specifically: S3221, Path timeliness marking mechanism: Read the hierarchical path index architecture and the affected area identifier, and add a timeliness mark to each pre-computed path: including the last verification timestamp, the list of edge identifiers involved, and the validity status flag. When any edge involved in the path changes, mark the validity status as "to be verified", to obtain the timeliness marked path index.
[0059] S3222. Hierarchical update strategy execution: Read the timeliness marker path index. Through the hierarchical update mechanism, the "to be verified" paths in the critical path index are immediately re-verified and updated; the "to be verified" paths in the important path index are added to the delayed update queue and processed in batches every 60 seconds; the "to be verified" paths in the background path index are only verified when queried, and a hierarchical update task allocation table is obtained.
[0060] Traditional index maintenance uses full-scale synchronous updates, which incurs huge overheads. This method adopts a lazy strategy and only verifies the paths that truly need to be verified, significantly reducing the maintenance cost.
[0061] S3223. Dynamic verification mechanism during query: Read the hierarchical update task allocation table and the query request. When the query involves the "to be verified" paths in the background path index, immediately trigger the validity check of this path: Traverse the status of each edge in the path in the current causal graph. If the edge still exists and the weight change is <10%, it is marked as valid; otherwise, remove this path from the index and try to recalculate the alternative path, and obtain the real-time verification result and the updated real-time valid path index.
[0062] According to one aspect of the present application, S121. Parallel construction of triple nested time windows can also be: Read the synchronized standardized data matrix, and simultaneously construct a 1-minute high-frequency response window, a 15-minute medium-frequency stable window, and a 1-hour low-frequency trend window through the triple nested window mechanism. Each window independently maintains a data buffer, and a hierarchical time window dataset and a window synchronization marker are obtained. The traditional method uses a single fixed window. The three-layer nested mechanism of this method can capture causal features at different time scales simultaneously.
[0063] According to one aspect of the present application, S122. Calculating the mutual information entropy - Granger fusion causal strength can also be: Read the hierarchical time window dataset, through the fusion calculation formula: Normalization processing: MI_norm=(MI(X;Y|Z)-MI_min) / (MI_max-MI_min); GC_norm=(GC(X→Y)-GC_min) / (GC_max-GC_min); TC_norm=(TC(X,Y,t)-TC_min) / (TC_max-TC_min); Adaptive weight calculation: α = 0.4 + 0.2 × window stability coefficient; β = 0.3 + 0.2 × time series correlation intensity; γ = 1 - α – β; Calculate the comprehensive causal strength = (α × MI_norm^2 + β × GC_norm^2 + γ × TC_norm^2)^0.5; Where: MI_norm is the normalized mutual information entropy, GC_norm is the normalized Granger causality, TC_norm is the normalized time consistency coefficient, and the weights are dynamically adjusted according to the window type and data characteristics to obtain a multi-dimensional causal strength vector. Existing methods use Granger causality or mutual information alone. This method standardizes and fuses the three causal metrics and uses the Euclidean norm to eliminate the influence of dimensions and improve the robustness of causal identification.
[0064] According to one aspect of the present application, S212, calculating the local influence propagation range can also be: Read the abnormal node list, abnormal intensity level, and multi-level causal graph, and through the influence propagation attenuation formula: Normalized edge weight accumulation: edge weight accumulation_norm = Σ(edge weight_i) / √(path length); The modified influence propagation formula: influence range(d) = initial intensity × e^(-λd) × edge weight accumulation_norm × attenuation adjustment factor; where: λ is the adaptive attenuation coefficient = 0.3 + 0.2×(abnormal intensity level / 5); the attenuation adjustment factor = min(1.0, 2.0 / d) to prevent over-amplification of long-distance propagation; the propagation depth is limited to 3 hops forward and 3 hops backward; obtain the affected area identifier and the normalized influence intensity distribution map.
[0065] Traditional methods require full-graph traversal calculation, with a complexity of O(n 2 ). This method avoids the exponential explosion problem through normalization processing and an adaptive attenuation mechanism, reducing the complexity to O(k·logn), where k is the number of affected nodes, and greatly improving the calculation efficiency and numerical stability.
[0066] Regarding the specific implementation process of path update.
[0067] Critical path index (real-time update): The critical path bears most of the query load and must be kept up-to-date at all times. Therefore, when the causal graph changes and affects the critical path, these paths are immediately updated completely, including recalculating path weights, verifying path validity, and updating index records.
[0068] Important path index (delayed update): To ensure path availability while reducing the maintenance frequency. Affected important paths are not updated immediately but are marked as "to be verified" and added to the delayed update queue. Every predetermined time, such as every 60 seconds (or 15 minutes, etc., internal configuration time), a batch update process is performed.
[0069] Background path index (verified during query): The affected background paths are also marked as "to be verified" but not updated actively. They are only verified when the user actually queries the path. The verification process includes checking the current status of each edge in the path. If the edge still exists and the weight change is <10%, it is marked as valid and continues to be used. If the edge fails, it is removed from the index and an alternative path is attempted to be recalculated.
[0070] Specifically, the critical path is "push update" (system actively updates), the important path is "timed batch update", and the background path is "pull verification" (passively verified during query). By precisely matching the usage frequency and importance of different paths through a differential strategy, an optimal configuration of index maintenance overhead is achieved.
[0071] In another embodiment of the present application, Generate a real-time valid path index, specifically including: Read the causal relationship basic index and the dynamic multi-level causal graph, and calculate the importance score of each node through the node importance quantification evaluation algorithm, where the node importance score = (degree centrality × 0.3) + (normalized access frequency value × 0.4) + (total causal strength × 0.3). Here, the degree centrality is the number of edges connected to the node, the normalized access frequency value is the historical query statistical times divided by the maximum query times, and the total causal strength is the cumulative value of the weights of all the edges connected to the node; Based on the node importance score, calculate the path importance of each pre-computed path through the path importance evaluation formula: path importance = starting point importance × 0.4 + ending point importance × 0.4 + average importance of intermediate nodes × 0.2, and adjust it in combination with the historical query frequency correction coefficient to generate a path importance ranking table; Stratify the pre-computed paths into three levels according to a preset importance threshold: paths with an importance score greater than 0.8 are stored in the critical path index, and a hash table structure is used to achieve an O(1) query complexity; paths with an importance score between 0.5 and 0.8 are stored in the important path index, and a B+ tree structure is used to achieve an O(logn) query complexity; paths with an importance score less than 0.5 are stored in the background path index, and a compressed storage structure is used; Add a timeliness mark to each pre-computed path, including the last verification timestamp, the list of involved edge identifiers, and the validity status flag. When the weight of any edge involved in the path changes, the validity status flag is marked as "to be verified"; Execute a differential lazy update strategy according to the index level to which the path belongs: immediately re-verify and update the paths marked as "to be verified" in the critical path index; add the paths marked as "to be verified" in the important path index to the delayed update queue for batch processing; for the paths marked as "to be verified" in the background path index, trigger immediate verification only when a query request is received. If the verification passes, keep the index; if the verification fails, remove the path and calculate an alternative path to generate a hierarchically maintained real-time valid path index.
[0072] In another embodiment of the present application, generating a dynamic multi-level causal graph and a real-time valid path index specifically includes: Read the affected area identifier and the abnormal intensity level, and determine the resource configuration plan through a three-layer intelligent computing resource allocation algorithm: set the basic resource allocation ratio as 15% CPU for the real-time layer, 20% CPU for the quasi-real-time layer, and 10% CPU for the background layer. Calculate the adjustment coefficient according to the abnormal intensity level = abnormal intensity level / 5, and dynamically adjust the CPU allocation ratio of each layer: real-time layer CPU allocation = 15% + 10% × adjustment coefficient, quasi-real-time layer CPU allocation = 20% + 5% × adjustment coefficient, background layer CPU allocation = 10% - 5% × adjustment coefficient, and generate a hierarchical resource allocation table; According to the hierarchical resource allocation table, execute task distribution through a time slice rotation scheduling strategy: the real-time layer adopts a continuous execution mode with a time slice length of 50 milliseconds for processing high-priority causal edge updates; the quasi-real-time layer adopts a timed execution mode with an execution interval of 30 seconds for processing medium-priority causal edge updates; the background layer adopts a low-frequency execution mode with an execution interval of 5 minutes for processing low-priority causal edge updates and supports a preemptive scheduling mechanism. When a new abnormal event is detected, low-priority tasks can be interrupted; Monitor the total CPU usage rate of the system in real time. When the total CPU usage rate exceeds the preset load threshold of 50%, automatically trigger a degradation strategy: suspend the background layer update task to release computing resources, extend the execution interval of the quasi-real-time layer to 45 seconds, and restore the normal scheduling strategy when the CPU usage rate drops below the threshold; For the hierarchical update task queue, read the real-time power consumption data stream, and perform incremental weight calculation on the causal edges that need to be updated: new edge weight = old edge weight - (expired data contribution / time window size) + (new data contribution / time window size). At the same time, record the continuous update times and the change trend of the weight error of each edge. When the continuous update times exceed 100 times or the weight error rate exceeds 5%, mark that the edge needs precision correction and recalculate the edge weight based on the complete historical data to generate a dynamic multi-level causal graph and a real-time valid path index maintained by three-layer asynchronous scheduling.
[0073] It should be noted that the sequence numbers of the above steps are only for the convenience of description and are not used to limit the specific order. For example, step S32 can be executed after step S23. That is, it forms: S24: Real-time effective path index construction and maintenance: Read the dynamically updated causal graph, construct and maintain a hierarchical path index, and output a real-time effective path index.
[0074] In a certain embodiment, an abnormal power consumption (the power suddenly increases from the normal 2.5 kW to 8.2 kW) of a certain user at 14:30 on May 28, 2025 is detected, and an interpretable causal analysis is provided. It should be noted that generally, this method is usually used to process the power consumption of industrial parks or shopping malls. For the sake of simplicity of description, household power consumption is used to describe the data processing process.
[0075] Prepare initial data, including smart meter data: power time series at 15-minute intervals, historical data for the past 30 days. User equipment list: 3 air conditioners, 1 water heater, 1 washing machine, 1 refrigerator. Environmental data: the temperature on that day is 32 °C, the humidity is 85%, and it is sunny.
[0076] Step S1: Multi-source data fusion and multi-time scale causal relationship discovery S11. Standardize multi-source heterogeneous data: The original meter data P_raw = [2.3, 2.5, 2.4, 2.6, 8.2, 3.1, 2.8] kW; the temperature data T_raw = [28, 30, 32, 32, 32, 31, 29] °C; the equipment status S_raw = [Air conditioner 1: on, Air conditioner 2: off, Air conditioner 3: on, Water heater: on]; Power standardization P_norm = (P_raw - P_min) / (P_max - P_min) = (8.2 - 2.3) / (8.2 - 2.3) = 1.0; temperature standardization T_norm = (32 - 28) / (32 - 28) = 1.0; obtain the synchronous standardized data matrix D_sync[timestamp, P_norm, T_norm, S_norm].
[0077] S12. Parallel calculation of causal strength in multiple time windows S121. According to the modified scheme, construct three layers of time windows. Among them, W1: 1-minute high-frequency response window (obtained by interpolating 15-minute data); W2: 15-minute medium-frequency stable window (directly use the collected data); W3: 1-hour low-frequency trend window (aggregation of 4 15-minute data); W1 data: P1_min = [2.3, 2.4, 2.5, 8.2] kW (interpolated within 1 minute interval); W2 data: P15_min = [2.5, 8.2, 3.1] kW (raw data within 15 minutes); W3 data: P1_hour = [3.8] kW (1-hour average); Obtain the hierarchical time window dataset {W1_data, W2_data, W3_data} and window synchronization markers S122, Mutual information entropy - Granger fusion causal strength calculation Step 1, Calculate the original metrics: Mutual information entropy MI(power, temperature) = 0.73; Granger causality GC(temperature → power) = 0.81; Temporal consistency coefficient TC(power, temperature, t) = 0.65; Step 2, Standardization processing: Based on the statistics of historical 30-day data: MI_min = 0.12, MI_max = 0.89, GC_min = 0.08, GC_max = 0.95, TC_min = 0.15, TC_max = 0.88; The standardization process is: MI_norm = (MI - MI_min) / (MI_max - MI_min) = (0.73 - 0.12) / (0.89 - 0.12) = 0.79; GC_norm = (GC - GC_min) / (GC_max - GC_min) = (0.81 - 0.08) / (0.95 - 0.08) = 0.84; TC_norm = (TC - TC_min) / (TC_max - TC_min) = (0.65 - 0.15) / (0.88 - 0.15) = 0.68; where: MI_norm is the standardized mutual information entropy; GC_norm is the standardized Granger causality; TC_norm is the standardized temporal consistency coefficient Step 3: Calculate the adaptive weights: Current window stability coefficient = 0.6 (calculated based on variance); Temporal correlation strength = 0.7 (calculated based on correlation coefficient); Calculate the adaptive weights α = 0.4 + 0.2 × window stability coefficient = 0.4 + 0.2 × 0.6 = 0.52; β = 0.3 + 0.2 × temporal correlation strength = 0.3 + 0.2 × 0.7 = 0.44; γ = 1 - α - β = 1 - 0.52 - 0.44 = 0.04; where: α is the mutual information weight coefficient; β is the Granger causality weight coefficient; γ is the time consistency weight coefficient; The window stability coefficient is the reciprocal of the data variance within the current time window; The temporal correlation strength is the absolute value of the Pearson correlation coefficient between variables.
[0078] Step 4: Calculate the fusion strength: Comprehensive causality strength K = (α × MI_norm 2 + β × GC_norm 2 + γ × TC_norm 2 )^0.5 = (0.52 × 0.79 2 + 0.44 × 0.84 2 + 0.04 × 0.68 2 )^0.5 = (0.52 × 0.624 + 0.44 × 0.706 + 0.04 × 0.462)^0.5 = (0.325 + 0.311 + 0.018)^0.5 = 0.81; where: K is the comprehensive causality strength; MI_norm 2 is the square of the standardized mutual information entropy; GC_norm 2 is the square of the standardized Granger causality; TC_norm 2 is the square of the standardized time consistency coefficient; Obtain the multi-dimensional causality strength vector V_causal = [0.81, 0.76, 0.68] (corresponding to three time windows).
[0079] S123: Causal relationship stability score calculation Cross-window consistency verification: Stability score = 1 - std(V_causal) / mean(V_causal) = 1 - 0.07 / 0.75 = 0.91; Obtain the causal stability score = 0.91, multi-scale causality strength matrix.
[0080] S13: Hierarchical causal graph construction and index generation Based on the multi-scale causal intensity matrix, a three-layer causal graph is constructed. Real-time layer (critical path): edges with intensity > 0.8, including the edge "temperature → power" with a weight of 0.81; quasi-real-time layer (important path): edges with intensity 0.6 - 0.8, including the edge "device status → power" with a weight of 0.76; background layer (other low-frequency paths): edges with intensity < 0.6, including other weakly associated edges; obtaining the multi-level causal graph G_multi = {G_real, G_quasi, G_back} and the basic index of causal relationships Index_base; Step S2: Dynamic causal graph maintenance and update S21: Intelligent identification mechanism for affected areas S211: Abnormal node rapid positioning algorithm Anomaly detection: User's historical baseline power P_baseline = 2.5 kW; current power P_current = 8.2 kW; 3σ threshold = P_baseline + 3×std_history = 2.5 + 3×0.3 = 3.4 kW; Anomaly judgment: P_current > 2×P_baseline and P_current > 3σ threshold; 8.2 > 2×2.5 = 5.0 and 8.2 > 3.4; Obtaining the list of abnormal nodes = ["power node"], and the anomaly intensity level = 4 (levels 1 - 5); S212: Calculation of the local influence propagation range Influence propagation calculation: Initial intensity I0 = 4 (anomaly intensity level); Propagation hop count d = 1, 2, 3; Cumulative edge weight_norm = Σ(edge weight_i) / √(path length) = (0.81 + 0.76) / √2 = 1.57 / 1.41 = 1.11; Adaptive attenuation coefficient λ = 0.3 + 0.2×(anomaly intensity level / 5) = 0.3 + 0.2×(4 / 5) = 0.46; Attenuation adjustment factor = min(1.0, 2.0 / d) = min(1.0, 2.0 / 1) = 1.0; Influence range I(d) = Initial intensity × e^(-λd) × Cumulative edge weight_norm × Attenuation adjustment factor = 4×e^(-0.46×1)×1.11×1.0 = 4×0.63×1.11 = 2.80; where: edge weight_i is the weight value of the i-th edge; path length is the hop count from the abnormal node to the current node; d is the propagation distance hop count; I(d) is the influence intensity at distance d.
[0081] Calculation of the influence intensity of each hop: d = 1: I(1) = 4×e^(-0.46×1)×1.11×1.0 = 2.80; d = 2: I(2) = 4×e^(-0.46×2)×0.98×1.0 = 1.56; d = 3: I(3) = 4×e^(-0.46×3)×0.87×0.67 = 0.74 (when d = 3, the attenuation adjustment factor = 2.0 / 3 = 0.67); Obtain the affected area identifier = {Node 1: 2.80, Node 2: 1.56, Node 3: 0.74}, the influence intensity distribution map.
[0082] S213. Dynamic allocation mechanism of influence weight: Inverse proportion adjustment based on influence intensity and stability score: Node 1 update priority = 2.80×(1 - 0.91) = 2.80×0.09 = 0.25; Node 2 update priority = 1.56×(1 - 0.76) = 1.56×0.24 = 0.37; Obtain the weighted influence area map (Node 2 has the highest priority).
[0083] S22: Distributed asynchronous update scheduling mechanism S221: Three-layer intelligent allocation algorithm for computing resources Basic resource allocation: Real-time layer: 15% CPU; Near-real-time layer: 20% CPU; Background layer: 10% CPU; Abnormal adjustment (abnormal intensity level 4): Adjustment coefficient = abnormal intensity level / 5 = 4 / 5 = 0.8; Resource allocation adjustment: Real-time layer CPU allocation = 15% + 10%×adjustment coefficient = 15% + 10%×0.8 = 23%; Near-real-time layer CPU allocation = 20% + 5%×adjustment coefficient = 20% + 5%×0.8 = 24%; Background layer CPU allocation = 10% - 5%×adjustment coefficient = 10% - 5%×0.8 = 6%; Among them: The adjustment coefficient is the ratio of the abnormal intensity level to the maximum level; The CPU allocation percentage is the proportion of the total computing resources of the system occupied; Obtain the hierarchical resource allocation table = {Real-time layer: 23%, Near-real-time layer: 24%, Background layer: 6%}; S222. Implementation of time slice round-robin scheduling strategy: Time slice allocation: Real-time layer: Continuously execute a 50ms time slice and immediately process high-priority nodes; Near-real-time layer: Execute at intervals of 30 seconds and process medium-priority nodes; Background layer: Execute at intervals of 5 minutes and process low-priority nodes; Current execution queue: High-priority queue: [Node 2] (priority 0.37); Medium-priority queue: [Node 1] (priority 0.25); Low-priority queue: [Other nodes]; Obtain the hierarchical update task queue and scheduling execution time table; S223. Dynamic load monitoring and adaptive adjustment Load monitoring: Current CPU usage rate = real-time layer 23% + quasi-real-time layer 24% + background layer 6% = 53%; Load threshold = 50%; Since 53% > 50%, the degradation strategy is triggered: Pause the background layer update (6% CPU released); Extend the quasi-real-time layer interval to 45 seconds; Obtain the optimized hierarchical update task queue; S23: Incremental causal edge weight correction algorithm S231: Sliding window incremental weight update Current edge weight update: "Temperature → Power" edge; Historical weight W_old = 0.81; Window size N = 15 (15 data points); Contribution of expired data = 0.73×(1 / 15) = 0.049; Contribution of new data = 0.89×(1 / 15) = 0.059; Incremental update: New edge weight W_new = old weight W_old - (contribution of expired data / window size) + (contribution of new data / window size) = 0.81 - 0.049 + 0.059 = 0.82; where: W_old is the historical edge weight value; Contribution of expired data is the contribution value of the removed data in the sliding window to the weight; Contribution of new data is the contribution value of the newly added data in the sliding window to the weight; Window size is the number of data points included in the sliding window; Obtain the incremental update weight table = {"Temperature → Power": 0.82}.
[0084] S232: Cumulative error real-time detection mechanism Error tracking: The continuous update count of this edge = 87 times; Current error rate = |0.82 - 0.81| / 0.81 = 1.2%; Check conditions: Continuous update count 87 < 100; Error rate 1.2% < 5%; Obtain the error monitoring report (this edge does not need to be corrected for the time being).
[0085] S233: Adaptive precision correction execution For other edge "Device status → Power": Continuous update count = 103 times > 100 times, triggering correction; Recalculate using the complete historical data: Corrected weight = 0.74 (recalculated based on the complete 30-day data); Original weight = 0.71; Correction error = |0.74 - 0.71| / 0.71 = 4.2% < 5%; Obtain the dynamically updated causal graph (edge weights have been corrected); Step S3: Query-driven fast reconstruction of causal paths S31: Query intention parsing and path requirement identification User query: "Why did the power suddenly increase to 8.2 kW at 14:30?"; The analysis results include: query type: real-time anomaly explanation; source node set: ["power node"]; target node set: ["temperature node", "device status node"]; query priority: high; S32: Hierarchical Index Dynamic Maintenance Mechanism S321: Construction of Three-Level Path Index Architecture S3211: Quantitative Evaluation of Node Importance Calculate the importance of the "power node": degree centrality = 5 (connected to 5 edges); access frequency = 120 times / month (historical query statistics); total causal strength = 0.82 + 0.76 + 0.68 + 0.45 + 0.33 = 3.04; node importance = (degree centrality × 0.3) + (access frequency × 0.4) + (total causal strength × 0.3) = (5 × 0.3) + (120 × 0.4) + (3.04 × 0.3) = 1.5 + 48.0 + 0.91 = 50.41; where: degree centrality is the number of edges connected to the node; access frequency is the monthly number of times the node is involved in historical queries; total causal strength is the cumulative value of the weights of all connected edges; obtain the node importance scoring table = {"power node": 50.41, "temperature node": 35.62, "device status node": 28.34}; S3212: Comprehensive Calculation of Path Importance Importance of the path "temperature → power": importance of the starting point: 35.62, importance of the ending point: 50.41; path importance = 35.62 × 0.4 + 50.41 × 0.4 = 34.41; historical query frequency correction factor = 1.2; final path importance = 34.41 × 1.2 = 41.29; obtain the path importance ranking table; S3213: Three-Level Index Hierarchical Storage: Importance Classification (threshold: critical > 40, important 20 - 40, background < 20): critical path index: ["temperature → power"] (importance 41.29); important path index: ["device status → power"] (importance 28.34); background path index: [other low-frequency paths]; obtain the hierarchical path index architecture.
[0086] S322: Precise Implementation of the Lazy Update Strategy S3221: Path Timeliness Marking Mechanism: Marking of the "temperature → power" path: last verification time: 2025-05-28 14:25:30; involved edge identifier: ["edge_temp_power"]; validity status: to be verified (because the edge weight is updated from 0.81 to 0.82); S3222, Hierarchical Update Strategy Execution: Immediate Verification of Critical Path: The path "Temperature → Power" is in the critical path index and is immediately verified; Verification Result: The edge weight change < 10%, and the path is still valid; Obtain the updated timeliness marked path index; S33: Fast Path Reconstruction and Verification Algorithm S331, Multi-level Degraded Path Search Strategy: First-level Search (Critical Path Index, Target < 50ms): Search for the path from the "Power Node" to the "Temperature Node"; Found Path: ["Power" ← "Temperature"], time-consuming 32ms; Search successful, no need to enter the second and third-level searches; Obtain the candidate path set = [["Power" ← "Temperature"]], search time-consuming record = 32ms.
[0087] S332, Path Validity Parallel Verification Mechanism. Verify the path ["Power" ← "Temperature"]: Edge Weight Check: 0.82 > Threshold 0.5; Connectivity Check: The edge exists in the current graph; Obtain the valid path set = [["Power" ← "Temperature"]].
[0088] S333, Comprehensive Path Quality Score Calculation. Score of the path ["Power" ← "Temperature"]: Causal Strength Weight = 0.82; Path Length Weight = 1.0 / 1 = 1.0 (path length is 1); Historical Success Rate = 0.95 (based on historical statistics). Path Score Score = (Causal Strength Weight × 0.4) + (Path Length Weight × 0.3) + (Historical Success Rate × 0.3) = (0.82 × 0.4) + (1.0 × 0.3) + (0.95 × 0.3) = 0.328 + 0.300 + 0.285 = 0.913; where: The causal strength weight is the average of the edge weights in the path; The path length weight is the reciprocal of the path length; The historical success rate is the proportion of successful explanations of this path in historical queries; 0.4, 0.3, 0.3 are the weight coefficients of the three scoring dimensions respectively; Obtain the target causal path set = [["Power" ← "Temperature", score 0.913]], path credibility distribution.
[0089] Step S4, Counterfactual Reasoning Verification and Explanation Output S41, Automatic Generation of Counterfactual Scenarios Scenario 1: Remove the "Temperature → Power" causal edge; Scenario 2: Keep the temperature at 28°C (not raised to 32°C); Obtain the counterfactual scenario set, expected result range; S42, Simulation Verification and Credibility Quantitative Evaluation Scenario 1 simulation: After removing the temperature influence, the predicted power should be 2.8 kW; Scenario 2 simulation: When the temperature remains unchanged, the predicted power should be 2.6 kW; Actual power: 8.2 kW; Consistency comparison: Scenario 1 accuracy = 1 - |2.8 - 8.2| / 8.2 = 0.34; Scenario 2 accuracy = 1 - |2.6 - 8.2| / 8.2 = 0.68; Obtained explanation credibility score = 0.68 S43. Comprehensive Explanation Report Generation Analysis results of abnormal power consumption, including: Main reason: The ambient temperature rises from 28°C to 32°C, triggering multiple air conditioners to start simultaneously; Causal relationship strength: The influence strength of temperature on power is 0.82 (strong correlation); Influence time delay: Power response within 15 minutes after temperature change; Confidence interval: Based on historical data, the power range at a temperature of 32°C is [7.8 kW, 8.6 kW], and the current 8.2 kW is within the normal range; Credibility score: 0.68 (medium credibility); Technical indicators: Path search time consumption: 32 ms < 50 ms target; Causal strength: 0.82 > 0.8 threshold (strong causal relationship); Prediction accuracy: 68% (within the expected range).
[0090] Effect of hierarchical lazy index update: The traditional method needs to update all path indexes, consuming 150 ms; This method only updates the queried key paths, consuming 32 ms; Efficiency is increased by 78.7%. Effect of local influence propagation update: The traditional method's full graph traversal O(n 2 )), n = 1000 nodes, requires 1,000,000 calculations; This method's 3-hop limit O(k·logn), k = 8 affected nodes, requires 24 calculations; The amount of calculation is reduced by 99.998%.
[0091] Effect of three-layer asynchronous scheduling: The peak CPU load of the traditional synchronous method is 85%, and the response delay is 200 ms; The average CPU load of this method is 53%, and the response delay is 32 ms; The load is reduced by 37.6%, and the response speed is increased by 84%.
[0092] In another embodiment of the present application, for the handling of edge failures, the specific details are as follows: According to one aspect of the present application, the handling method when an edge in the causal path fails is as follows: In an industrial park intelligent electricity meter monitoring system, at 9:15 am on May 29, 2025, it is detected that the air conditioner control system is offline, resulting in the failure of the "temperature → air conditioner power" causal edge. The system executes the following processing flow: Edge Failure Detection and Impact Assessment: When the system fails to obtain air conditioner power data for 5 consecutive minutes, it determines that the edge_temp_ac edge is in a hard failure state. The historical weight of this edge is 0.76, and the current weight cannot be calculated. The impact analysis of the failed edge shows that it involves the critical path ["Temperature" → "Air Conditioner Power" → "Total Power"], affecting the real-time energy consumption anomaly explanation function. The system identifies three alternative paths: Path 1 is ["Temperature" → "Device Status" → "Total Power"] with a weight of 0.68; Path 2 is ["Temperature" → "Environmental Comfort" → "Air Conditioner Demand" → "Total Power"] with a weight of 0.54; Path 3 is ["External Weather" → "Temperature Change" → "Total Power"] with a weight of 0.41.
[0093] Alternative Path Selection and Reconstruction: The system uses the priority calculation formula Priority = Path Weight × 0.4 + Response Speed × 0.3 + Historical Success Rate × 0.3 to select alternative paths. The score for Path 1 is 0.68×0.4 + 0.85×0.3 + 0.92×0.3 = 0.803. The score for Path 2 is 0.54×0.4 + 0.62×0.3 + 0.88×0.3 = 0.666. The score for Path 3 is 0.41×0.4 + 0.78×0.3 + 0.73×0.3 = 0.617. The system selects Path 1 as the main alternative path. The comprehensive strength of the original path is 0.76×0.82 = 0.623. The comprehensive strength of the alternative path is 0.68×0.71 = 0.483. The strength attenuation rate is 22.5%. To compensate for the accuracy loss, the system increases the time consistency weight γ from 0.2 to 0.35 and introduces a historical mode correction factor of 1.15. The corrected strength is 0.483×1.15 = 0.555.
[0094] Real-time Verification and Performance Monitoring: During the 30-minute continuous verification period, when the temperature rises from 28°C to 32°C, the expected power change of the alternative path is +2.1kW, and the actual power change is +2.4kW. The prediction accuracy is 87.5%. When the user queries "Why did the power suddenly increase at 9:30?", the system explains based on the alternative path as "The temperature rise caused the device status to adjust, which in turn affected the total power". The response time is 58ms, which is slightly higher than the normal 45ms but still within the 100ms target range.
[0095] Automatic Recovery and Switching: The system checks the status of the original edges every 60 seconds. At 10:42:15, it detects that the air conditioning control system comes back online, data transmission resumes normal, and the edge weight is recalculated to be 0.74. The system compares the current performance of the alternative path 87.5% with the expected performance of the original path 91.2% after recovery. The switching benefit of 3.7% exceeds the 3% threshold, triggering the switching decision. A gradual switching strategy is adopted, with 50% of the queries using the original path and 50% using the alternative path. After a 15-minute observation period for verification, it fully switches back to the original path. This processing method achieves a fault detection time of less than 5 minutes, an alternative path activation time of less than 2 minutes, a service interruption time of 0 minutes, an interpretation accuracy of 88.7%, and the availability is improved from 95% to 99.7% compared with the traditional method.
[0096] In another embodiment of the present application, if the real-time effective path is calculated after calculating the causal graph, the process is as follows: When a power anomaly (2.5kW → 8.2kW) is detected at 14:30 on May 28, 2025, the system executes an integrated graph update and index construction process.
[0097] Affected Area Identification and Resource Allocation: The system identifies the abnormal node as a power node, with the influence range being 3 hops forward and 3 hops backward. The affected edges include ["Temperature → Power", "Device Status → Power", "Power → Total Power Consumption"]. According to the abnormal intensity level 4, the resource allocation is adjusted to 23% CPU for the real-time layer, 24% CPU for the quasi-real-time layer, and 6% CPU for the background layer. The edge weight of "Temperature → Power" is updated from 0.81 to 0.82, and the error detection of 1.2% is less than the 5% threshold and does not require correction.
[0098] Synchronize Real-Time Effective Path Index Construction: While the graph is being updated, the system performs real-time recalculation of node importance. The total causal intensity of the "power node" increases from 3.04 to 3.05, and the importance score is updated from 50.41 to 50.44. The computational overhead is O(k), where k is the number of affected nodes, which is 3, and the execution time is less than 5ms. In the dynamic adjustment of path importance, the change in the end-point importance of the "Temperature → Power" path causes the path importance to increase from 34.41 to 34.42, and after correction, it is 41.30, still remaining in the critical path index. The computational overhead is O(m), where m is the number of affected paths, which is 2, and the execution time is less than 3ms.
[0099] Hierarchical Index Real-time Maintenance and Consistency Verification: The system updates the hash table of the critical path index. The key is "power node_temperature node", and the value is updated to {path: ["power" ← "temperature"], weight: 0.82, timestamp: 14:30:23}. The computational overhead of hash table operations is O(1), and the execution time is less than 1 ms. Index consistency verification is only performed on the affected critical paths. The "temperature → power" path passes the edge existence check, weight change check (1.2% is less than the 10% threshold), and connectivity check. The verification result is that the path is valid and the index is maintained. The computational overhead is O(p×e), where p is the number of paths, which is 1, and e is the average length of the path, which is 1. The execution time is less than 2 ms.
[0100] Analysis of Performance Optimization Effect: Under the integrated scheme, the execution time of step S2 is 56 ms (45 ms for graph update + 11 ms for index construction), and the execution time of step S3 is 32 ms (pure path search). The total response time is 88 ms. Compared with the original scheme of 92 ms (45 ms + 47 ms), the performance is improved by 4.3%. In terms of CPU usage, after adjustment, it is 58% CPU in the S2 stage and 10% CPU in the S3 stage. Compared with 53% and 15% in the original scheme, the resource utilization is smoothed. It is particularly suitable for application environments with high-frequency query scenarios (query frequency greater than 100 times / minute) and high real-time requirements (response time less than 100 ms). Through the integrated processing of graph update and index maintenance, the overall response performance and resource utilization efficiency of the system are improved.
[0101] It should be noted that when the CPU usage rate in the S2 stage exceeds 70%, the system automatically enables the degradation mechanism and returns to the mode of building the index in S3 to ensure system stability. This integrated scheme realizes the logical consistency of data update and index maintenance, avoids repeated calculations, and provides a more efficient technical solution for the real-time maintenance of the causal graph of smart meters.
[0102] In another embodiment of the present application, for the calculation process of the real-time valid path index, a hybrid strategy can also be adopted. For example, set up solutions A and B. Solution A (executed in S2): Emphasize data consistency and query performance. Solution B (executed in S32): Emphasize modular design and resource efficiency.
[0103] Dynamically select according to the system operation status, specifically as follows: During high load periods (query frequency > 100 / minute): Adopt solution A and pre-build the index in S2; give priority to ensuring query response performance.
[0104] During low load periods (query frequency < 50 / minute): Adopt solution B and build on demand in S32; save system resources.
[0105] System startup phase: Adopt Solution B to avoid resource competition during startup. Switch to Solution A after the system stabilizes.
[0106] Exception recovery phase: Adopt Solution A to ensure data consistency. Respond quickly to user query requirements.
[0107] In another embodiment of the present application, the calculation process of the window stability coefficient is as follows: Step 1: Read the power sequence data P = [P1, P2, ..., P n in the hierarchical time window dataset, where n is the window size, and calculate the sample variance σ 2 = Σ(P i - P*) 2 / (n - 1); where P* is the arithmetic mean of the power data: P* = Σ(P i ) / n; Step 2: To eliminate the influence of numerical magnitude on stability evaluation, calculate the relative coefficient of variation CV = σ / P*; where σ is the standard deviation, σ = sqrt (σ 2 ); Step 3: Use reciprocal transformation to convert the coefficient of variation into a stability coefficient. The larger the value, the more stable it is. Window stability coefficient = 1 / (1 + CV); this formula ensures that the stability coefficient ranges from (0, 1], and when CV = 0 (completely stable), the coefficient is 1, and when CV → ∞ (extremely unstable), the coefficient approaches 0.
[0108] Example: For the power sequence P = [2.3, 2.5, 2.4, 2.6, 2.4] kW, the calculation process is as follows: Mean value P* = (2.3 + 2.5 + 2.4 + 2.6 + 2.4) / 5 = 2.44 kW; Variance σ 2 = [(2.3 - 2.44) 2 + (2.5 - 2.44) 2 + (2.4 - 2.44) 2 + (2.6 - 2.44) 2 + (2.4 - 2.44) 2 / 4 = 0.013; Standard deviation σ = sqrt 0.013 = 0.114; Coefficient of variation CV = 0.114 / 2.44 = 0.047; Window stability coefficient = 1 / (1 + 0.047) = 0.955; The strength of temporal correlation is used to quantify the degree of time series correlation between variables and guide the dynamic adjustment of Granger causality weights. The specific calculation process is as follows: Step 1: Read the standardized power time series data X = [X1, X2, ..., X n and temperature time series data Y = [Y1, Y2, ..., Y n , and calculate the Pearson correlation coefficient at different lags τ: ρ(τ) = Σ[(X i - X*)(Y i+τ –Y*)] / √[Σ(X i - X*) 2 × Σ(Y i+τ –Y*) 2 ; where τ is the lag, and its value range is from 0 to the maximum lag T max (usually set to 5 time points).
[0109] Step 2: Identify the absolute value of the maximum correlation coefficient ρ max = max{|ρ(0)|, |ρ(1)|,..., |ρ(T max )|} Step 3: Considering the continuity characteristics of the time series, calculate the time series correlation strength using a weighted average method = 0.5 × ρ max + 0.3 × |ρ(1)| + 0.2 × |ρ(2)|; this formula highlights the dominant role of the strongest correlation and also considers the contribution of short-term lag correlations.
[0110] Step 4: To avoid the influence of outliers, perform boundary restriction and smoothing on the calculation results: the final time series correlation strength = min(max(time series correlation strength × smoothing factor, 0.1), 0.9). Where the smoothing factor = 0.9 + 0.1 × window stability coefficient, ensuring a relatively high correlation strength when the data is stable.
[0111] For the power sequence X = [0.3, 0.5, 0.4, 0.6, 0.4] and the temperature sequence Y = [0.2, 0.4, 0.6, 0.6, 0.5]: the zero-lag correlation ρ(0) = 0.76; the one-lag correlation ρ(1) = 0.82; the two-lag correlation ρ(2) = 0.65; the maximum correlation ρ max = 0.82; the time series correlation strength = 0.5×0.82 + 0.3×0.82 + 0.2×0.65 = 0.786; the smoothing factor = 0.9 + 0.1×0.955 = 0.996; the final time series correlation strength = min(max(0.786×0.996, 0.1), 0.9) = 0.783.
[0112] In summary, through the local influence propagation algorithm and the three-layer asynchronous scheduling mechanism, this solution solves the problem of the computational complexity at the traditional O(n 2 ) level. Specifically, when the smart meter detects abnormal power consumption, the traditional method needs to recalculate the relationships between all nodes and edges in the entire causal graph. However, this method uses the influence propagation attenuation formula I(d) = initial intensity × e^(-λd) × edge weight accumulation_norm × attenuation adjustment factor to precisely limit the update range within the first 3 hops forward and the last 3 hops backward of the abnormal node. By taking advantage of the locality of abnormal impacts in the power consumption system, the computational complexity is reduced from O(n 2 ) to O(k•logn). At the same time, combined with the three-layer asynchronous time-slicing rotation scheduling, it is processed layer by layer according to the importance and timeliness requirements of causal relationships: the real-time layer processes key device anomalies (50 ms time slice), the quasi-real-time layer processes general state changes (30-second interval), and the background layer processes historical trend analysis (5-minute interval), avoiding the computational load spikes of synchronous batch updates and achieving a 99.998% reduction in computational volume and millisecond-level response capabilities in a smart meter network with a scale of 1000 nodes.
[0113] Through the hierarchical lazy index update mechanism, the strategy of traditional full-scale synchronous index maintenance is completely changed, and the index maintenance overhead is reduced from more than 30% to less than 10%. Specifically, according to the 80 / 20 distribution characteristic of smart meter queries, the pre-computed paths are divided into three levels through quantitative evaluation of node importance: critical path index (importance > 0.8), important path index (importance 0.5 - 0.8), and background path index (importance < 0.5), and a differentiated lazy update strategy is adopted. The critical path is immediately updated because it undertakes 80% of the query load to ensure real-time availability, the important path adopts delayed batch updates to reduce the maintenance frequency, and the background path is only verified for validity when actually queried, avoiding unnecessary maintenance of a large number of low-frequency paths. Combined with the timeliness marking mechanism to record the last verification time of the path and the change status of the involved edges, the update range is precisely controlled, achieving the technical goal of significantly reducing the index maintenance overhead while ensuring query performance.
[0114] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all fall within the protection scope of the present invention.
Claims
1. An interpretability analysis method for the electricity consumption behavior of smart meters based on a causal graph, characterized in that, It includes the following steps: Read the original electricity consumption data of the smart meter, and construct a primary multi-level causal graph through multi-source data fusion and multi-time scale causal relationship discovery; When an electricity consumption anomaly is detected, identify the influence range of the abnormal node in the primary multi-level causal graph, maintain and update it, and generate a dynamic multi-level causal graph and a real-time effective path index; Receive the user's query request, search sequentially from the key and important path indexes in the real-time effective path index. If not found, conduct a real-time search to obtain a set of target causal paths; Read the set of target causal paths, verify and interpret the output through counterfactual reasoning, and generate a verified causal explanation report.
2. The method according to claim 1, characterized in that, Construct a multi-level causal graph through multi-source data fusion and multi-time scale causal relationship discovery, including: Read the original electricity consumption data of the smart meter, perform standardization and stratification processing, and generate a stratified time window data set; Based on the stratified time window data set, calculate MI, GC, and TC and fuse them to generate a multi-dimensional causal intensity vector; Read the multi-dimensional causal intensity vector, calculate each causal relationship and verify its stability at different time scales, and generate a multi-scale causal intensity matrix; Based on the multi-scale causal intensity matrix, construct causal relationships at different time scales and store them in layers, and generate a multi-level causal graph and a corresponding causal relationship basic index.
3. The method according to claim 2, characterized in that, Generate a multi-dimensional causal intensity vector, including: Based on the stratified time window data set, calculate the mutual information entropy MI, Granger causality GC, and time consistency coefficient TC respectively, and perform normalization processing; According to the window stability coefficient and the strength of time series correlation, calculate the dynamic weights α, β, and γ of MI, GC, and TC; Read the normalized MI, GC, TC and dynamic weights α, β, γ, and calculate the fusion intensity K=(α×MI 2 +β×GC 2 +γ×TC 2 ) 0.5 , and generate a multi-dimensional causal intensity vector.
4. The method according to claim 1, characterized in that, Generate a real-time effective path index, including: Combined with the multi-level causal graph, identify abnormal nodes from the current real-time electricity consumption data stream, calculate the local influence propagation range, and generate an affected area identifier; Based on the affected area identifier, through a three-layer intelligent allocation of computing resources and a time slice round-robin scheduling strategy, perform distributed asynchronous update scheduling on the real-time layer, quasi-real-time layer, and background layer, and generate an optimized stratified update task queue; For the stratified update task queue and the real-time electricity consumption data stream, perform weight update and accuracy correction to generate a dynamically maintained multi-level causal graph and a real-time effective path index.
5. The method according to claim 4, characterized in that, Generate a dynamically maintained multi-level causal graph and a real-time effective path index, including: Based on the optimized stratified update task queue and the real-time electricity consumption data stream, calculate the new edge weight Q = old weight - (expired data contribution / window size) + (new data contribution / window size), and perform incremental calculation based on this to obtain an incremental update weight table; For the incremental update weight table, record the continuous update times and error change trends of each edge. When it exceeds the preset value, mark it as an edge that needs to be corrected; For the edges that need to be marked, recalculate the edge weights based on the complete historical data and correct them to generate a dynamically maintained multi-level causal graph and a real-time effective path index.
6. The method according to claim 4, characterized in that, Calculate the local influence propagation range and generate an affected area identifier, including: Read the real-time power consumption data stream and the multi-level causal graph, identify the nodes whose power changes exceed the threshold, and obtain the list of abnormal nodes and the abnormal intensity level; accordingly, calculate the affected area and generate the distribution map of the influence intensity through the influence propagation formula; The influence propagation formula is as follows: Influence range I = Initial intensity I0 × e -λd × Cumulative edge weights norm × Decay adjustment factor; Among them: Edge weight accumulation norm = Σ(edge weight i ) / sqrt(path length); d is the propagation hop count, limited to 3 hops forward and 3 hops backward, and the adaptive attenuation coefficient λ = 0.3 + 0.2×(anomaly intensity level / 5).
7. The method according to claim 2, wherein Search sequentially from the critical and important path indexes in the real-time valid path index. If not found, perform real-time search, including: Respond to the user's query request, identify the source node set and the target node set through query type classification and key entity extraction, and output; Read the dynamic multi-level causal graph, and divide the pre-computed paths in the causal relationship basic index into critical path index, important path index, and background path index to generate a real-time valid path index; For the source node set and the target node set, first search in the critical and important path indexes. If not found, perform real-time search in the background path index.
8. The method according to claim 7, characterized in that Generate a real-time valid path index, including: Read the causal relationship basic index and the dynamic multi-level causal graph, calculate the node importance and path importance, and accordingly store the pre-computed paths into the critical, important, and background path indexes respectively to generate a hierarchical path index structure; Read the hierarchical path index structure, add a timestamp and a valid flag to each pre-computed path, update the critical path immediately, update the important path with a delay, and verify when querying the background path to generate a real-time valid path index.
9. The method according to claim 7, wherein After searching in the critical and important path indexes first and performing real-time search in the background path index if not found, it also includes: Based on the search results, form a candidate path set and a search time consumption record; For the candidate path set, combined with the dynamic multi-level causal graph, parallelly check the current causal intensity and connectivity of each edge in multiple candidate paths, and remove the paths containing invalid edges to form a set of valid paths; For each valid path, perform path scoring and sorting to obtain a set of target causal paths; Among them, path scoring = (causal intensity weight × a + (path length weight × b) + (historical success rate × c); a + b + c = 1.
10. The method according to claim 1, characterized in that Adopt counterfactual reasoning verification and explanation output to generate a verified causal explanation report, including: Based on the set of target causal paths and the dynamic multi-level causal graph, remove or modify the critical causal edges, combine parameter perturbation simulation, construct a set of counterfactual scenarios and the expected result range, and combine the original power consumption data of the smart meter. The engine calculates the changes in the power consumption pattern under each scenario and compares them with the historical actual data to form an explanation credibility score and a verification result report; For the verification result report, adopt the logical organization of the causal chain and the quantification of technical indicators to generate a causal explanation report including the causal relationship strength, influence delay, and confidence interval.
Citation Information
Patent Citations
Intelligent auxiliary diagnosis and maintenance method and system based on multi-path recall
CN119357787A
Multi-mode general-purpose cooperative causal thinking chain reasoning power anomaly detection method and system
CN119474996A
Intelligent electric energy meter fault prediction method based on multi-mode sensor fusion
CN119902154A
Operation and maintenance alarm processing method and system based on knowledge graph enhanced large model
CN119988154A
Electric energy meter state evaluation method based on space-time attention mechanism
CN120030395A
Cited By
Autonomous agent rumor checking system based on dynamic causal evidence graph
CN121072525A
Self-adaptive learning method for load regulation and control parameters of air conditioner host
CN121302051A
Air conditioner host load regulation parameter adaptive learning method
CN121302051B
High-load scene-oriented computing power server system layer optimization method and system
CN121387548A
Power distribution network harmonic traceability method and system
CN122085048A