Effective causal entropy-based upstream and downstream industry power consumption characteristic causal analysis method and system

By analyzing the causal relationship of electricity consumption characteristics of upstream and downstream industries in the power grid based on effective causal entropy, identifying key drivers and influencing nodes, and building a causal relationship network for electricity consumption characteristics, solving the problem of insufficient prediction and response capabilities of the power grid when electricity consumption demand fluctuates, realizing refined management of electricity consumption demand and optimized allocation of power resources.

CN120124871APending Publication Date: 2025-06-10STATE GRID JIANGSU ELECTRIC POWER CO LTD MARKETING SERVICE CENT
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510308990.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively analyze and optimize the causal relationship between the upstream and downstream industries of the industrial chain in the power grid, resulting in insufficient prediction and response capabilities in the power grid when facing fluctuations in electricity demand.

Method used

The causal analysis method for electricity consumption characteristics of upstream and downstream industries based on effective causal entropy is adopted. Through data collection, preprocessing, causal analysis and network analysis, key drivers and influential nodes in the industrial chain are identified to build a causal relationship network for electricity consumption characteristics.

Benefits of technology

Through the calculation of effective causal entropy and the application of eigenvector centrality and PageRank algorithm, the causal relationship of electricity consumption characteristics can be accurately identified, the prediction accuracy and response speed of the power grid for electricity consumption fluctuations, the allocation of power resources can be optimized, and the stability and economics of the power grid can be improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120124871A_ABST
    Figure CN120124871A_ABST
Patent Text Reader

Abstract

The invention discloses an effective causal entropy-based upstream and downstream industry power consumption characteristic causal analysis method and system. The method comprises the steps of obtaining industry chain upstream and downstream industry power consumption data including enterprise electrical quantity data and enterprise category data; performing data cleaning, data integration, abnormal value processing and data discretization preprocessing on the acquired data; effective causal entropy is calculated according to the preprocessed data, and a power utilization characteristic causal relationship network of the industry chain is formed; and respectively adopting a feature vector centrality algorithm and a PageRank algorithm, carrying out centrality analysis on the power consumption feature causal relationship network of the industrial chain based on the effective causal entropy, identifying key nodes in the network, and obtaining power consumption feature causal analysis results of the upstream and downstream industries. According to the method, key driving factors and influence nodes in an industrial chain can be identified, and a scientific basis is provided for stable operation of a power grid and formulation of a power supply strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of electricity consumption feature analysis, and relates to a method and system for causal analysis of electricity consumption features of upstream and downstream industries based on effective causal entropy. Background Art

[0002] With the continuous advancement of the global industrialization process, electricity, as a key element in the national economy and social development, has become increasingly prominent in its importance. The power grid, as a key link in power transmission and distribution, its stability and efficiency directly affect the economic operation of the entire society and the lives of residents. In this context, studying the causal relationship of electricity consumption features of upstream and downstream industries in the industrial chain, and conducting a detailed analysis of the electricity consumption growth of each industry have important theoretical and practical significance for the operation management and optimization of the power grid. By deeply studying the electricity consumption patterns of users, the economic situation of enterprises can be more accurately evaluated, and the development trend of industries can be predicted, which is an important basis for scientifically and reasonably evaluating market trends and guiding enterprises and the government in making economic decisions. By establishing a complete causal relationship framework for the industrial chain, the power grid can more accurately predict and respond to the electricity consumption demands of different industries and enterprises, thereby comprehensively optimizing the power grid structure. In-depth analysis of this causal relationship not only helps to improve the perception ability of the distribution network, enabling it to more sensitively identify and respond to changes in electricity demand, but also enhances the self-repair ability of the power grid. When a fault or demand fluctuation occurs in the power grid, this self-repair ability can quickly restore power supply, reduce the impact on enterprise operations, thereby constructing a high-quality and stable power grid system, ensuring that enterprises can efficiently obtain stable power supply, and providing a solid energy guarantee for the country's economic development.

[0003] Causal analysis of the electricity consumption features of upstream and downstream industries in the industrial chain plays an important role in power grid management, but there are still many challenges at present. First, the dynamic nature of electricity demand requires the power grid to have a high degree of flexibility and responsiveness. There are significant differences in the electricity demand of different industries during the production process, and it fluctuates with time and economic activities. For example, the electricity demand of the manufacturing industry will increase significantly during the production peak period, while the agricultural industry has a higher demand for electricity during the irrigation season. This demand fluctuation poses challenges to the load management and power dispatching of the power grid. Second, existing power decisions rely only on the electricity consumption of a single industry, ignoring the mutual influence between upstream and downstream industries in the industrial chain, resulting in the difficulty of government policies to achieve the expected results. Power grid enterprises need to conduct in-depth analysis of the electricity consumption features of different industries in the industrial chain, identify key factors, and discover potential causal relationships. By studying the electricity consumption features of upstream and downstream industries, the distribution and change laws of electricity demand can be better understood, thereby optimizing the allocation of power resources. Summary of the Invention

[0004] To address the deficiencies in the existing technologies, the present invention provides a method and system for causal analysis of electricity consumption characteristics of upstream and downstream industries based on effective causal entropy. Through data collection, preprocessing, causal analysis, and network analysis, it deepens the understanding of the dynamics of electricity demand and optimizes the allocation of power resources. In particular, through the application of improved eigenvector centrality and PageRank algorithms, it can identify key driving factors and influential nodes in the industrial chain, providing a scientific basis for the stable operation of the power grid and the formulation of power supply strategies. In addition, the calculation method of effective causal entropy provides a powerful tool for processing large-scale data and revealing causal mechanisms in complex systems, enabling the present invention to have broad application prospects in multiple fields such as power system management, energy optimization, and policy formulation.

[0005] The present invention adopts the following technical solutions.

[0006] In the first aspect of the present invention, a method for causal analysis of electricity consumption characteristics of upstream and downstream industries based on effective causal entropy is proposed, including:

[0007] Obtain electricity consumption data of upstream and downstream industries in the industrial chain, including enterprise electrical quantity data and enterprise category data;

[0008] Perform data cleaning, data integration, outlier processing, and data discretization preprocessing on the obtained data;

[0009] Calculate the effective causal entropy based on the preprocessed data to form a causal relationship network of electricity consumption characteristics of the industrial chain;

[0010] Respectively adopt the eigenvector centrality algorithm and the PageRank algorithm to perform centrality analysis on the causal relationship network of electricity consumption characteristics of the industrial chain based on the effective causal entropy, identify key nodes in the network, and obtain the causal analysis results of electricity consumption characteristics of upstream and downstream industries.

[0011] Preferably, the enterprise electrical quantity data includes the electricity consumption of the enterprise during peak and off-peak periods, the overall electricity load, and power quality indicators;

[0012] The enterprise category data includes the industry attribution of the enterprise, the positioning in the upstream and downstream industrial chains, and the enterprise scale.

[0013] Preferably, the performing data cleaning, data integration, outlier processing, and data discretization preprocessing on the obtained data includes:

[0014] Remove incorrect, duplicate, or incomplete data records, merge datasets from different sources to prevent data conflicts, delete, fill, or interpolate missing data, and discretize the data in the continuous dataset through a binning function.

[0015] Preferably, calculate the effective causal entropy based on the preprocessed data to form a causal relationship network of the electricity consumption characteristics of the industrial chain, specifically including:

[0016] Take the preprocessed data as the original data and calculate the causal entropy C Z→X|Y (k, l, m);

[0017] Randomly shuffle the sequence of the original data to obtain randomized data;

[0018] Calculate the causal entropy of the randomized data and combine it with the causal entropy C Z→X|Y (k, l, m) of the original data to obtain the effective causal entropy EC Z→X|Y (k, l, m).

[0019] Preferably, the calculation formula of the causal entropy C Z→X|Y (k, l, m) is as follows:

[0020]

[0021] where C Z→X|Y (k, l, m) is the causal entropy of the original data S t , that is, the causal entropy from process Z to process X based on process Y; k, l, m are the k, l, m time nodes before time t;

[0022] is the conditional entropy of the k - order time - lag subsequence of the known process X and the l - order time - lag subsequence of process Y ;

[0023] is the conditional entropy of the k - order time - lag subsequence of the known process X the l - order time - lag subsequence of process Y and the m - order time - lag subsequence of process Z .

[0024] Preferably,

[0025]

[0026] where p(,,) and p(|) are joint and conditional probability densities.

[0027] Preferably, the calculation formula of the effective causal entropy EC Z→X|Y (k, l, m) is as follows:

[0028]

[0029] In the formula, EC Z→X|Y(k, l, m) represents the effective causal entropy from process Z to process X based on process Y;

[0030] represents the causal entropy of randomized data.

[0031] The second aspect of the present invention proposes a causal analysis system for the electricity consumption characteristics of upstream and downstream industries based on effective causal entropy, including:

[0032] A data acquisition module for acquiring electricity consumption data of upstream and downstream industries in the industrial chain, including enterprise electrical quantity data and enterprise category data;

[0033] A data processing module for performing data cleaning, data integration, outlier processing, and data discretization preprocessing on the acquired data;

[0034] A causal entropy calculation module for calculating the effective causal entropy based on the preprocessed data to form a causal relationship network of electricity consumption characteristics of the industrial chain;

[0035] A centrality analysis module for respectively using the eigenvector centrality algorithm and the PageRank algorithm to perform centrality analysis on the causal relationship network of electricity consumption characteristics of the industrial chain based on the effective causal entropy, identifying key nodes in the network, and obtaining the causal analysis results of electricity consumption characteristics of upstream and downstream industries.

[0036] The third aspect of the present invention proposes a terminal, including a processor and a storage medium; the storage medium is used to store instructions; the processor is used to operate according to the instructions to execute the steps of the method.

[0037] The fourth aspect of the present invention proposes a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method are implemented.

[0038] Compared with the prior art, the beneficial effects of the present invention at least include:

[0039] 1. The present invention systematically processes the electricity consumption data of enterprises in upstream and downstream industries of the power grid industrial chain, obtains accurate information on key electricity consumption characteristics in the industrial chain, and these information provide a decision-making basis for power grid enterprises to optimize power resource allocation, enabling power supply to more accurately meet the key needs of the industrial chain and improving the overall operation efficiency and economy of the power grid.

[0040] 2. The effective causal entropy proposed by the present invention uses the concept of transfer entropy to quantify the non-linear causal influence of electricity consumption characteristics between upstream and downstream industries. By randomizing the original data, the non-random dependence between data is eliminated, thus accurately calculating the difference in causal entropy. The proposed effective causal entropy not only reveals the direct and indirect electricity consumption dependence relationships between industries, but also can quantify the intensity of these relationships. With the increase in the sample size, the recognition accuracy of causal relationships can be gradually improved, and the breadth of analysis can be broadened, providing a powerful analysis method for revealing potential causal connections in the industrial chain, thereby realizing the refined management of power grid loads.

[0041] 3. By constructing an industrial chain causal relationship network based on effective causal entropy, the obtained network is a weighted directed network, where the direction of the link reflects the direction of the causal relationship, and the weight of the link indicates the intensity of the causal relationship measured by effective causal entropy, significantly improving the prediction accuracy and response speed of the power grid to fluctuations in electricity demand.

[0042] 4. The present invention introduces effective causal entropy and combines eigenvector centrality and PageRank algorithm to deeply analyze the causal network. Through network analysis, the direct and indirect associations of electricity consumption characteristics between different industries can be revealed, and the importance of nodes in the causal relationship network can be quantified. The present invention can identify key industry nodes with high centrality and predict their potential impact on the stability of the power grid. These nodes may be the main driving factors of electricity demand and play a decisive role in the stability and efficiency of the industrial chain. It can help power grid enterprises prioritize the power supply to high-centrality industries that have a decisive impact on the stability and efficiency of the industrial chain and formulate corresponding power dispatching strategies to ensure the resilience and reliability of the power grid in the face of sudden demand changes, promote the synergy effect of the entire industrial chain, and ensure its stable and efficient operation. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 is the schematic diagram of the principle of the causal analysis method for electricity consumption characteristics of upstream and downstream industries based on effective causal entropy of the present invention;

[0044] Figure 2 is the flow chart of the causal analysis method for electricity consumption characteristics of upstream and downstream industries based on effective causal entropy of the present invention;

[0045] Figure 3 is the schematic diagram of the data preprocessing process of the present invention;

[0046] Figure 4 is the schematic diagram of the relationship between information entropy and conditional entropy of the present invention;

[0047] Figure 5 is the schematic diagram of the causal relationship network of the upstream and downstream industrial chains of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0048] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. The embodiments described in this application are only a part of the embodiments of the present invention, rather than all embodiments. Based on the spirit of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.

[0049] As Figures 1-4 shown, Embodiment 1 of the present invention provides a method for causal analysis of electricity consumption characteristics of upstream and downstream industries based on effective causal entropy, providing a brand-new perspective for the research on the causal relationship of electricity consumption characteristics of upstream and downstream industries in power grid enterprises. This method studies the potential causal relationship between upstream and downstream industries, and provides support for the coordinated development of upstream and downstream industries, industrial adjustment, and optimization of electricity consumption strategies according to the influence law of upstream and downstream industries. It mainly includes the collection of electricity consumption data of upstream and downstream industries in the industrial chain, data preprocessing, causal analysis, and network analysis. By collecting the electricity consumption data and enterprise data of each industry in the upstream and downstream of the industrial chain, this method analyzes the correlation relationship of electricity consumption characteristics of enterprises. This method uses data preprocessing technology to ensure data quality, and then applies the transfer entropy principle and data randomization strategy to calculate the effective causal entropy, revealing the non-linear causal relationship of electricity consumption characteristics between industries. Combining the eigenvector centrality and PageRank algorithm, the present invention can identify the key driving factors and influential nodes in the industrial chain, provide decision-making support for power grid enterprises, optimize the allocation of power resources, and improve the stability and economy of the power grid. Specifically as follows:

[0050] S1. Obtain the electricity consumption data of upstream and downstream industries in the industrial chain, including enterprise electrical quantity data and enterprise category data;

[0051] The collection of the electricity consumption data of upstream and downstream industries in the industrial chain mainly includes collecting and obtaining the legal electrical quantity data that can be obtained in power grid enterprises and the recorded enterprise category data. Among them, the electrical quantity data includes data such as peak / low valley electricity consumption, electricity load, and power, and the enterprise category data includes data such as industry category, upstream and downstream category, and enterprise scale. The purpose is to ensure that accurate and reliable data support can be obtained when performing subsequent data processing and causal analysis, which helps to ensure the integrity, credibility, and accuracy of the data used. The obtained key parameter data such as electricity consumption at peak and low valley periods, overall electricity load, and power quality indicators provide a direct view of the operation status of the power grid and are the basis for power grid analysis and optimization. The obtained classification information such as industry attribution, upstream and downstream industrial chain positioning, and enterprise scale helps to understand the electricity consumption behaviors and patterns of different industries and enterprises in the power grid.

[0052] S2. Perform data cleaning, data integration, outlier processing, and data discretization preprocessing on the obtained data;

[0053] After collecting the data, data preprocessing is required. The data preprocessing mainly includes processing and transforming the collected data, which contains four steps: data cleaning, data integration, outlier handling, and data discretization, to provide effective data for causal analysis:

[0054] 1) Data cleaning: Perform data cleaning on inaccurate data, identify and remove incorrect, duplicate, or incomplete data records to ensure that each data in the dataset is unique, thus avoiding bias and redundancy in data analysis.

[0055] 2) Data integration: Merge datasets from different sources and transform the format to ensure data consistency to prevent data conflicts.

[0056] 3) Outlier handling: Delete, fill, or interpolate missing data. Specifically, for missing values and outliers, these data points may be caused by measurement errors, data entry errors, or external interferences. Smooth or remove these abnormal data to improve data accuracy.

[0057] 4) Data discretization: Under the premise of ensuring data quality, perform preprocessing of data discretization to convert continuous variables into discrete categories for calculating the effective causal entropy. Since these data are to be used for calculating the effective causal entropy and discrete data must be used in the process of calculating the effective causal entropy, the focus of this step is to discretize the continuous time series, convert the continuous time series into a discrete series for the next step of effective causal entropy calculation.

[0058] Specifically, the method of binning (i.e., symbolic recoding) is adopted to convert the continuous dataset into a discrete dataset. In this process, the continuous data is divided into bins, and each value in the continuous data is assigned to a bin. To execute this process, the boundaries of the bins need to be determined. If these boundaries are determined as q 1 ,q 2 ,...,q n (q 1 <q 2 <...<q n ), then the continuous time series y t (only including the initial continuous dataset involved in causal analysis, such as enterprise electrical quantity data, including electricity consumption, overall electricity load, and power quality indicators during peak and off-peak periods of the enterprise) can be discretized by the symbolic recoding method, and the binning function S t is expressed as:

[0059]

[0060] At the end of this process, each value in the continuous data is assigned a number between 1 and n, where n is determined according to the length of the continuous time series.

[0061] S3. Calculate the effective causal entropy based on the preprocessed data to form a causal relationship network of the electricity consumption characteristics of the industrial chain.

[0062] Further preferably, after the data preprocessing is completed, the present invention introduces causal inference technology, which is an analytical method in statistics for identifying and quantifying the causal effects between variables. Its core goal is to extract the causal connections between events or variables from the observed data, that is, to analyze how one variable affects another variable, rather than just their correlation. By constructing causal models or causal diagrams, the potential relationships between variables can be formally described. These models can be directed acyclic graphs (DAGs) or other types of graphical representations. Generally, inference methods in probability theory are used to estimate the parameters in the causal model, which usually involves modeling and analyzing the probability distribution of potential outcomes. Different from traditional correlation analysis, causal inference technology attempts to go beyond the surface statistical associations and explore the direct causal paths between variables. For example, in power grid analysis, one may want to understand whether a change in the electricity consumption pattern of one industry will cause a change in the electricity consumption pattern of another industry, rather than just identifying their synchronous changes.

[0063] The traditional method for inferring the causal relationship between two random variables is Granger causality test. A major limitation of this test is that it can only provide information about the linear dependence between two variables, and thus cannot capture the inherent non-linear relationships commonly found in real-world systems. To overcome this difficulty, Schreiber proposed the concept of transfer entropy between two variables. Transfer entropy measures how the current and past states of one variable affect the future state of another variable, and it measures the flow of information. As an asymmetric measure, transfer entropy can be used to infer the directionality of information flow and further infer the causal relationship between two variables. In the industrial chain, there may be complex non-linear relationships in the electricity consumption characteristics of upstream and downstream industries. Transfer entropy can detect and quantify these non-linear causal relationships, while traditional linear causal analysis methods (such as Granger causality test) cannot. The causal analysis method based on transfer entropy can not only detect the non-linear causal relationship between variables, but also quantify the magnitude of information flow, providing a measure of the strength of the causal effect between the electricity consumption characteristics of upstream and downstream industries. This helps grid enterprises understand the specific impact degree between the electricity consumption characteristics of different industries. In addition, transfer entropy is conceptually equivalent to mutual information (MI), providing a method for measuring statistical independence. Its calculation does not depend on specific statistical models or distribution assumptions, which makes it applicable to a wide range of application scenarios and has broader universality, including data that does not conform to traditional statistical models.

[0064] Based on the concept of transfer entropy, the present invention proposes to explore the potential causal connections between upstream and downstream industries by conducting causal analysis on the electricity consumption characteristics of various industries in the industrial chain, and construct a causal graph of the upstream and downstream industrial chain, that is, a causal relationship network of the electricity consumption characteristics of the industrial chain, as Figure 5 shown, which can more accurately identify the causal relationship of electricity consumption characteristics between upstream and downstream industries, and be used to intuitively present the causal relationship path and correlation degree between various indicators, so as to provide decision-making support for grid enterprises to optimize the allocation of power resources. For example, identify which changes in the electricity consumption characteristics of industries will have a significant impact on other industries, and then adjust the power supply strategy. This method comprehensively analyzes the causal connections between various indicators in the upstream and downstream industrial chain, deepens the understanding of the operation rules of the industrial chain, and provides a scientific basis for the optimization and decision-making of the industrial chain.

[0065] Specifically, the key reason why transfer entropy cannot identify indirect coupling from direct coupling is that it is a pairwise measure between two processes. Therefore, the present invention proposes the concept of effective causal entropy based on transfer entropy and adds a data randomization link to overcome the causal entropy calculation error caused by data bias. By shuffling the original data, the dependence between variables is eliminated as much as possible. In this step, the order of the original data sequence is randomly shuffled, thus artificially destroying the inherent temporal dependence in the data. Subsequently, the causal entropy is calculated again for the shuffled data, and a measure reflecting the information flow under random conditions, namely effective causal entropy, is obtained based on the difference in causal entropy before and after. Expressing the effective causal entropy as the difference in causal entropy before and after can more accurately capture the true causal relationship between variables, rather than simple correlation or false signals caused by randomness in the data. As the sample size increases, the estimation of effective causal entropy will be more stable and reliable because it can better distinguish causal information flow from random fluctuations. In addition, the concept of effective causal entropy not only improves the accuracy of causal analysis but also expands its application scope. It enables researchers to perform causal inference on a wider range of systems and phenomena under different data scales and complexity conditions. In this process, the accuracy and breadth of causal relationships are gradually improved from a small number of samples to a large amount of data.

[0066] Specifically, effective causal entropy is derived from transfer entropy and is used to measure the reduction of uncertainty in a time series through the concept of entropy, reflecting the dynamics (causal relationship) of information flow, thereby achieving the purpose of quantifying the information flow and causal relationship between two systems.

[0067] Consider a discrete random variable X with its probability mass function represented by p(x) = Prob(X = x). To quantify the unpredictability of X, its (information) entropy can be calculated. The information entropy uses Shannon entropy, which is the average number of bits required to encode a discrete random variable, and is defined as

[0068] H(X) = -∑ x p(x) log 2 p(x) (2)

[0069] Now consider two random variables X and Y with a joint distribution, and the joint distribution is p(x,y) = Prob(X = x, Y = y).

[0070] The conditional distribution is p(x|y) = Prob(X = x|Y = y). Then the joint entropy H(X,Y) and conditional entropy H(X|Y) of X and Y are defined as

[0071] H(X,Y) = -∑ x,y p(x,y) log 2 p(x,y) (3)

[0072] H(X|Y) = -∑ y p(y)H(Y|X = x) = -∑ x,y p(x,y)log 2 p(x|y) (4)

[0073] Given the complete information about Y(X), the reduction in the uncertainty of X(Y) can be measured by the mutual information between X and Y, visualized as Figure 4 shown, i.e.,

[0074] I(X; Y) = H(X) - H(X|Y) = H(Y) - H(Y|X) (5)

[0075] Now turning to stochastic processes, for two stochastic processes {X t} and {Y t}, the transfer entropy between the two can be expressed as

[0076]

[0077] where the information entropy H X (k) represents the conditional entropy of the discrete random variable X under the probability distribution p(x) for the k - order time - lag subsequence of the known process X and is specifically expressed as

[0078]

[0079] H XY (k, l) is the conditional entropy of the k - order time - lag subsequence of the known process X and the l - order time - lag subsequence of the process Y and is

[0080]

[0081] Here, the present invention proposes a new metric, called the effective causal entropy. First, the causal entropy from Z to X (conditioned on X and Y) is defined as

[0082]

[0083] where, C Z→X|Y (k, l, m) is the original data S tThe causal entropy from process Z to process X based on process Y; where the causal entropy is a new concept established based on transfer entropy, describing the uncertainty or complexity of information in causal relationships; k, l, m are k, l, m time nodes before time t, indicating the analysis of causal relationships through past states; the enterprise electrical quantity data includes the electricity consumption, overall electrical load, and power quality indicators of the enterprise during peak and trough periods; the enterprise category data includes 6 items of data such as the industry attribution, upstream and downstream industrial chain positioning, and enterprise scale of the enterprise to calculate the causal entropy; processes Z, X, Y in this patent can refer to the electricity consumption data of a certain enterprise.

[0084] is the k - order time - lag subsequence of the known process X and the l - order time - lag subsequence of process Y of the conditional entropy; refers to the k - order time - lag subsequence of the variable after symbol - impulse coding pre - processing of the known process X, refers to the l - order time - lag subsequence of the variable after symbol - impulse coding pre - processing of the known process Y; the specific physical meaning can refer to the electricity consumption data of a certain enterprise. and The subscript of refers to a total of k time nodes from time t all the way back to t - k + 1;

[0085] is the k - order time - lag subsequence of the known process X the l - order time - lag subsequence of process Y and the m - order time - lag subsequence of process Z of the conditional entropy; refers to the m - order time - lag subsequence of the variable after symbol - impulse coding pre - processing; The subscript of represents a total of m time nodes from time t all the way back to t - m + 1;

[0086] The difference between the two is the causal entropy from Z to X.

[0087] Among them,

[0088]

[0089] is the joint and conditional probability density, and the probability function can be obtained by taking the derivative.

[0090] Another drawback of the causal entropy described above is that there is an error in the calculated causal entropy for sample bias. To solve this problem, the present invention proposes the concept of effective causal entropy. To calculate the effective causal entropy time, the present invention randomizes the data of the upstream and downstream industries in the industrial chain, eliminates non-random patterns or sequential effects that may exist in the data set, reduces mutual dependence, and ensures that the causal structure discovery process and analysis results are more fair and objective. From these shuffled data, a causal entropy is calculated. Then, the causal entropy obtained from the shuffled data is subtracted from the causal entropy obtained from the original data. This method can be expressed as:

[0091]

[0092] In the above expression, EC Z→X|Y (k, l, m) represents the effective causal entropy (i.e., causal strength) from process Z to process X based on process Y, represents the causal entropy calculated from the randomized data. Where Z shuffle is calculated in the same way from the randomized data obtained by shuffling the original data in order.

[0093] Through this process of data randomization, the dependence in Z and the dependence between X, Y, and Z are eliminated. The causal entropy is calculated again for these shuffled data to obtain a measure reflecting the information flow under random conditions. By implementing this method, the effective causal entropy can more accurately reveal the substantial causal chain between variables, distinguishing it from misleading signals generated solely based on data correlation or contingency. As the data sample size increases, the calculation result of the effective causal entropy becomes more stable and reliable, and it can effectively distinguish the real causal information flow from the fluctuations caused only by data randomness. This method provides a powerful analysis tool for deeply understanding the causal mechanism in complex systems, contributing to achieving a higher level of accuracy and depth in data analysis.

[0094] S4. Respectively adopt the eigenvector centrality algorithm and the PageRank algorithm to conduct centrality analysis on the causal relationship network of the electricity consumption characteristics of the industrial chain based on the effective causal entropy, identify the key nodes in the network, and obtain the causal analysis results of the electricity consumption characteristics of the upstream and downstream industries.

[0095] Further preferably, after obtaining the causal relationship network, in order to deeply understand the mutual relationship and the degree of mutual influence among various industries or nodes in the industrial chain, network analysis is carried out based on two efficient centrality measurement methods, Eigenvector Centrality and PageRank. These two methods play a crucial role in identifying and evaluating the importance of key nodes in the network. By combining the content features (such as text, etc.) of variable nodes and introducing effective causal entropy, the centrality and PageRank algorithms are improved, the weights are dynamically adjusted, and the importance of each variable in the causal network is further quantified, so as to identify the key nodes with high influence in the industrial chain. These features may be the key driving factors of electricity demand and have an important impact on the stability and efficiency of the entire industrial chain. By identifying industries with high centrality, the electricity supply to these industries can be preferentially guaranteed, the synergy effect of the entire industrial chain can be inferred, and the stable operation of the industrial chain can be ensured. Specifically, when implementing, the Eigenvector Centrality algorithm and the PageRank algorithm can be considered separately, or the weight value can be set to consider the node centrality by weighting;

[0096] Eigenvector Centrality is an index to measure the importance of a node. It not only considers the direct connection number of the node, that is, the number of neighbors of a node, but also considers the importance of the node's neighbors, that is, the centrality of these neighbor nodes themselves. Eigenvector Centrality takes these differences into account. If the centrality (importance) of a node's adjacent nodes is very high, then the centrality (importance) of this node should also be very high. Node x i 's Eigenvector Centrality is proportional to the sum of the centralities of its adjacent nodes. Eigenvector Centrality can be expressed by the following formula:

[0097]

[0098] In the above expression, κ -1 is the proportionality constant, x j is the neighbor node of x i ;

[0099] The above expression can be re-expressed using the adjacency matrix A ij of the network, and the following formula is obtained:

[0100]

[0101] The above formula can be expressed in matrix notation as Ax = κx (13)

[0102] where A is the adjacency matrix of the network, and its values are composed of effective causal entropy;

[0103] x is the eigenvector of A, and its elements are centrality values x i, represents the eigenvector centrality score (centrality value of the \(i\)-th node in the network, which can be iteratively calculated by a computer and converge to \(X\), and its value is the centrality value). It is based not only on the number of connections of the node but also on the importance of its neighbor nodes. If a node is connected to many nodes with high eigenvector centrality itself, then the eigenvector centrality of this node will also be high; \(x\) j is the eigenvector centrality score of the \(j\)-th node; \(x\) represents the eigenvector of the network, which is used to measure the eigenvector centrality of the network;

[0104] A ij is the element in the \(i\)-th row and \(j\)-th column of the adjacency matrix, representing the connection status between node \(i\) and node \(j\). If there is a connection relationship, the value is 1, otherwise it is 0;

[0105] \(\kappa\) is the largest eigenvalue of the adjacency matrix; it can be understood that this eigenvalue is different from the electricity consumption characteristics mentioned above. It is a basic concept in linear algebra, and eigenvector centrality is calculated by solving the eigenvector equation system;

[0106] \(n\) is the number of network nodes. Among them, a network node refers to a variable node in causal analysis, an industrial chain node is an enterprise in the industry, and electricity consumption characteristics refer to the quantifiable values of user electricity consumption behaviors, such as electricity consumption, etc.

[0107] Here, the effective causal entropy considered in S3 causal analysis is introduced into network analysis, and causal weights are assigned to the adjacency matrix, making the analysis of the industrial chain upstream and downstream network more accurate.

[0108] The PageRank algorithm is a centrality measurement method based on the network link structure, associated with eigenvector centrality, and is designed specifically for directed networks. To determine the centrality of a node, the PageRank algorithm considers three different factors, namely the number of nodes linking to the target, the PageRank centrality of the linking nodes, and the linking tendency of the linking nodes, that is, the number and quality of incoming links of the node, and the importance of the linking source nodes. The calculation formula is as follows:

[0109] \(x=(I - \alpha AD\) -1 ) -1 1 (14)

[0110] In the above expression, \(\alpha\) is a positive constant, 1 is a uniform vector \((1, 1, 1, \cdots)\), \(D\) is a diagonal matrix, and its elements where is the out-degree of the node, that is, in a directed graph, the number of edges starting from a node), \(A\) is the adjacency matrix, \(I\) is the identity matrix, and \(x\) is the network PageRank centrality vector.

[0111] Here, the effective causal entropy considered in the S3 causal analysis is introduced into the network analysis, and causal weights are assigned to the adjacency matrix, making the analysis results more accurate.

[0112] The above two methods have unique advantages in identifying key nodes in the network. By combining these two methods, the key driving factors and influential nodes in the industrial chain can be revealed. These nodes with high centrality may be the main sources of electricity demand, having a decisive impact on the stability and efficiency of the entire industrial chain. After identifying these key nodes, by monitoring and analyzing the electricity consumption behaviors of these key nodes, grid enterprises can better understand the changing trends of electricity demand. Grid enterprises and policymakers can allocate resources more effectively and adjust electricity supply strategies in a timely manner, such as giving priority to ensuring the electricity supply for these industries, so as to reduce potential supply chain risks and ensure the stable operation of the power grid and the efficient utilization of electricity resources. The analysis method of the present invention also helps grid enterprises to make more accurate predictions and response strategies in the face of electricity demand fluctuations.

[0113] Based on the above execution process, the causal relationships between industries can be dynamically learned using the historical electricity consumption data of a wide range of upstream and downstream industries. The causal network model can simulate the changes in electricity demand under different scenarios, and by understanding the interdependencies between industries, predict the impact of changes in the electricity consumption characteristics of industries on the entire industrial chain. This is crucial for formulating response strategies and conducting risk management.

[0114] Embodiment 2 of the present invention provides a causal analysis system for the electricity consumption characteristics of upstream and downstream industries based on effective causal entropy, including:

[0115] A data acquisition module, configured to acquire the electricity consumption data of upstream and downstream industries of the industrial chain including enterprise electrical quantity data and enterprise category data; the core task of the data acquisition module is to collect the electrical quantity data and enterprise category data of grid enterprises. This module will capture key electrical parameters including peak / low valley electricity consumption, electricity load, and power, etc., while recording classification information such as the industry to which the enterprise belongs and its position in the industrial chain. Through this process, the accuracy and credibility of the data are ensured, laying a solid foundation for subsequent data processing and causal analysis.

[0116] Specifically, by deploying an automated data acquisition system, the real-time nature and accuracy of the data are ensured. The enterprise category data will be integrated through a database management system, including information such as industry attribution, position in the industrial chain, and enterprise scale, etc., to ensure the consistency and accessibility of the data. After completing the data collection, preliminary analysis can also be carried out to evaluate the data quality and lay a solid foundation for subsequent data processing and causal analysis.

[0117] The data processing module is used to perform data cleaning, data integration, outlier handling, and data discretization preprocessing on the acquired data; the data preprocessing module focuses on the cleaning, integration, outlier handling, and discretization of the original data. It will remove errors, duplicates, and incomplete records in the data, and solve the format conflict problems brought about by the diversity of data sources. In addition, this module will handle missing values, ensuring the integrity of the data through methods such as deletion, filling, or interpolation. Finally, this module will convert continuous variables into discrete categories, providing structured data for the calculation of effective causal entropy.

[0118] Specifically, by using advanced data cleaning tools and techniques, such as data deduplication algorithms, missing value filling methods, and outlier detection mechanisms, etc. In addition, to meet the requirements of effective causal entropy calculation, continuous time series data is discretized through binning or symbolic recoding methods, converting continuous data into discrete categories to prepare a formatted data set for the calculation of causal entropy.

[0119] The causal entropy calculation module is used to calculate the effective causal entropy based on the preprocessed data, forming a causal relationship network of electricity consumption characteristics in the industrial chain; the causal entropy calculation module adopts the principle of transfer entropy and proposes the concept of effective causal entropy. By randomizing the original data. This module calculates the transfer entropy of the original data to evaluate the direct causal relationship between variables. Subsequently, a data randomization link is introduced. By randomly shuffling the data sequence, the time series dependence is eliminated, thereby reducing the impact of data bias on the causal entropy calculation. Calculate the causal entropy again from the randomized data and subtract it from the causal entropy of the original data to obtain the effective causal entropy, eliminating the non-random dependence between the data, thus accurately calculating the difference in causal entropy. As the sample size increases, this module uses customized algorithms and software implementations to ensure the accuracy and efficiency of the calculation. As the sample size increases, the estimation method of effective causal entropy is continuously optimized to improve its stability and reliability, gradually improving the recognition accuracy of causal relationships, broadening the breadth of analysis, and providing a powerful analysis tool for revealing potential causal connections in the industrial chain.

[0120] The centrality analysis module is used to perform centrality analysis on the causal relationship network of the electricity consumption characteristics of the industrial chain based on the effective causal entropy by using the eigenvector centrality algorithm and the PageRank algorithm respectively, identify the key nodes in the network, and obtain the causal analysis results of the upstream and downstream industries' electricity consumption characteristics. The centrality analysis module introduces the effective causal entropy to improve the centrality and PageRank algorithms, dynamically adjusts the weighted eigenvector centrality and the PageRank algorithm, and quantifies the importance of the nodes in the causal relationship network. This module can introduce the effective causal entropy considered by the causal analysis module and the content characteristics of the network nodes themselves, assign personalized weights to them, and then conduct in-depth centrality analysis to identify industry nodes with high influence. These nodes may be the main driving forces of electricity demand and play a decisive role in the stability and efficiency of the industrial chain. By cooperating with power grid enterprises to formulate resource allocation strategies and risk management measures, and giving priority to ensuring the power supply of these high-centrality industries, the synergy effect of the entire industrial chain can be promoted, ensuring its stable and efficient operation.

[0121] Embodiment 3 of the present invention provides a terminal, including a processor and a storage medium; the storage medium is used to store instructions;

[0122] The processor is used to operate according to the instructions to execute the steps of the method.

[0123] Embodiment 4 of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps of the method.

[0124] Compared with the prior art, the beneficial effects of the present invention at least include:

[0125] 1. The present invention systematically processes the electricity consumption data of enterprises in the upstream and downstream industries of the power grid industrial chain, obtains accurate information on the key electricity consumption characteristics in the industrial chain. These information provide a decision-making basis for power grid enterprises to optimize the allocation of power resources, enabling the power supply to more accurately meet the key needs of the industrial chain, and improving the overall operation efficiency and economy of the power grid.

[0126] 2. The effective causal entropy proposed by the present invention uses the concept of transfer entropy to quantify the non-linear causal influence of electricity consumption characteristics between upstream and downstream industries, and eliminates the non-random dependence between data by randomizing the original data, thereby accurately calculating the difference in causal entropy. The proposed effective causal entropy not only reveals the direct and indirect electricity dependence relationships between industries, but also can quantify the intensity of these relationships. As the sample size increases, the accuracy of identifying causal relationships can be gradually improved, and the breadth of analysis can be broadened, providing a powerful analysis method for revealing potential causal connections in the industrial chain, thereby realizing the refined management of power grid load.

[0127] 3. The present invention constructs an industrial chain causal relationship network based on effective causal entropy. The obtained network is a weighted directed network, where the direction of the link reflects the direction of the causal relationship, and the weight of the link indicates the strength of the causal relationship measured by effective causal entropy, significantly improving the prediction accuracy and response speed of the power grid to fluctuations in electricity demand.

[0128] 4. The present invention introduces effective causal entropy and combines eigenvector centrality and PageRank algorithm to deeply analyze the causal network. Through network analysis, the direct and indirect associations of electricity consumption characteristics between different industries can be revealed, and the importance of nodes in the causal relationship network can be quantified. The present invention can identify key industry nodes with high centrality, predict their potential impact on the stability of the power grid. These nodes may be the main driving factors of electricity demand, play a decisive role in the stability and efficiency of the industrial chain, and can help power grid enterprises prioritize the power supply to high-centrality industries that have a decisive impact on the stability and efficiency of the industrial chain and formulate corresponding power dispatching strategies to ensure the resilience and reliability of the power grid in the face of sudden demand changes, promote the synergy effect of the entire industrial chain, and ensure its stable and efficient operation.

[0129] The present disclosure may be a system, method, and / or computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.

[0130] The computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. The computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punch card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium used herein is not construed as an instantaneous signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0131] The computer-readable program instructions described herein can be downloaded to various computing / processing devices from a computer-readable storage medium or downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.

[0132] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present disclosure.

[0133] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific embodiments of the present invention, and any modifications or equivalent replacements that do not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.

Claims

1. The causal analysis method of electricity consumption characteristics of upstream and downstream industries based on effective causal entropy is characterized by: include: Obtain electricity consumption data of upstream and downstream industries in the industrial chain, including enterprise electrical quantity data and enterprise category data; Perform data cleaning, data integration, outlier processing and data discretization preprocessing on the acquired data; Calculate the effective causal entropy based on the preprocessed data to form a causal relationship network of electricity consumption characteristics of the industrial chain; The eigenvector centrality algorithm and PageRank algorithm were used respectively to conduct centrality analysis on the causal relationship network of electricity consumption characteristics of the industrial chain based on effective causal entropy, identify the key nodes in the network, and obtain the causal analysis results of electricity consumption characteristics of upstream and downstream industries.

2. The method for causal analysis of electricity consumption characteristics of upstream and downstream industries based on effective causal entropy according to claim 1 is characterized by: The enterprise electrical quantity data includes the enterprise's electricity consumption during peak and off-peak periods, overall electricity load and power quality indicators; The enterprise category data include the enterprise's industry affiliation, upstream and downstream industrial chain positioning and enterprise scale.

3. The method for causal analysis of electricity consumption characteristics of upstream and downstream industries based on effective causal entropy according to claim 1 is characterized by: The data cleaning, data integration, outlier processing and data discretization preprocessing of the acquired data include: Remove erroneous, duplicate or incomplete data records, merge data sets from different sources to prevent data conflicts, delete, fill or interpolate missing data, and discretize data in continuous data sets through binning functions.

4. The method for causal analysis of electricity consumption characteristics of upstream and downstream industries based on effective causal entropy according to claim 1 is characterized by: The effective causal entropy is calculated based on the preprocessed data to form a causal relationship network of electricity consumption characteristics of the industrial chain, specifically including: Take the preprocessed data as the original data and calculate the causal entropy C of the original data Z→X|Y (k,l,m); Randomly disrupt the sequence of original data to obtain randomized data; Calculating Causal Entropy of Randomized Data and the causal entropy C of the original data Z→X|Y (k, l, m) combined to obtain the effective causal entropy EC Z→X|Y (k,l,m).

5. The method for causal analysis of electricity consumption characteristics of upstream and downstream industries based on effective causal entropy according to claim 4 is characterized by: The causal entropy C Z→X|Y The calculation formula for (k,l,m) is as follows: Among them, C Z→X|Y (k,l,m) is the original data S t The causal entropy is the causal entropy from process Z to process X based on process Y; k, l, m are k, l, m time nodes before time t; is the k-order time-delay subsequence of the known process X and the l-order time-delay subsequence of process Y The conditional entropy of is the k-order time-delay subsequence of the known process X l-order time-delay subsequence of process Y and the m-order time-delay subsequence of process Z The conditional entropy of .

6. The method for causal analysis of electricity consumption characteristics of upstream and downstream industries based on effective causal entropy according to claim 5 is characterized by: where p(,,) and p(|) are the joint and conditional probability densities.

7. The method for causal analysis of electricity consumption characteristics of upstream and downstream industries based on effective causal entropy according to claim 4 is characterized by: The effective causal entropy EC Z→x|Y The calculation formula for (k,l,m) is as follows: In the formula, EC Z→X|Y (k, l, m) represents the effective causal entropy from process Z to process X based on process Y; Represents the causal entropy of the randomized data.

8. The method for causal analysis of electricity consumption characteristics of upstream and downstream industries based on effective causal entropy according to claim 1 is characterized by: The formula for the characteristic vector centrality algorithm to perform centrality analysis on the electricity consumption characteristic causal relationship network of the industrial chain based on effective causal entropy is as follows: Written in matrix form as follows: Ax=κx(13) Among them, A is the adjacency matrix of the network, and its value is composed of effective causal entropy; x is the eigenvector of A, whose element x i represents the eigenvector centrality score of the i-th node in the network; x j is the eigenvector centrality score of the jth node; A ij is the element in the i-th row and j-th column of the adjacency matrix, indicating the connection status between node i and node j. If there is a connection relationship, the value is 1, otherwise it is 0; κ is the maximum eigenvalue of the adjacency matrix; n is the number of network nodes.

9. The method for causal analysis of electricity consumption characteristics of upstream and downstream industries based on effective causal entropy according to claim 1 is characterized by: The formula for the PageRank algorithm to perform centrality analysis on the causal relationship network of electricity consumption characteristics of the industrial chain based on effective causal entropy is as follows: x=(I-αAD -1 ) -1 1 (14) In the formula, α is a positive constant; 1 is a consistent vector (1,1,1,…); D is a diagonal matrix whose elements is the out-degree of the node; A is the adjacency matrix of the network, which is composed of effective causal entropy; I is the identity matrix; x is the network PageRank centrality vector.

10. A causal analysis system for electricity consumption characteristics of upstream and downstream industries based on effective causal entropy, using the method described in any one of claims 1 to 9, characterized in that: The system comprises: The data acquisition module is used to obtain electricity consumption data of upstream and downstream industries in the industrial chain, including enterprise electrical quantity data and enterprise category data; Data processing module, used for data cleaning, data integration, outlier processing and data discretization preprocessing of the acquired data; The causal entropy calculation module is used to calculate the effective causal entropy based on the preprocessed data to form a causal relationship network of electricity consumption characteristics of the industrial chain; The centrality analysis module is used to perform centrality analysis on the causal relationship network of electricity consumption characteristics of the industrial chain based on effective causal entropy using the eigenvector centrality algorithm and PageRank algorithm respectively, identify key nodes in the network, and obtain causal analysis results of electricity consumption characteristics of upstream and downstream industries.

11. A terminal comprising a processor and a storage medium; characterized in that: The storage medium is used to store instructions; The processor is configured to operate according to the instructions to execute the steps of the method according to any one of claims 1-9.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.

Citation Information

Cited By

  • Commercial bank associated party analysis method and device, electronic equipment and storage medium

    CN120975213A