Data transmission analysis method, device and equipment based on multi-scale condition mutual information

By employing a multi-scale conditional mutual information analysis method, the problems of data fragmentation and weak correlation in power grid engineering were solved, enabling accurate characterization and quantification of cross-stage data impact paths, and providing reliable technical support for the whole-process management of power grid engineering.

CN122064937APending Publication Date: 2026-05-19RES INST OF ECONOMICS & TECH STATE GRID SHANDONG ELECTRIC POWER +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
RES INST OF ECONOMICS & TECH STATE GRID SHANDONG ELECTRIC POWER
Filing Date
2025-12-16
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing technologies are insufficient to effectively process multi-source, heterogeneous, and multi-stage distributed engineering data in power grid projects, resulting in data fragmentation, weak correlation, and poor value transmission across links, which fails to meet the management needs of precise investment and full-process value optimization.

Method used

A multi-scale conditional mutual information data transmission analysis method is adopted. Through Morlet wavelet decomposition and conditional mutual information progressive strategy, key lag components are screened, the optimal embedding vector is constructed, the conditional probability density and marginal probability density are calculated, and the data transmission intensity and direction are quantified.

Benefits of technology

It enables multi-scale, adaptive, and high-precision characterization of the dynamic influence relationships between engineering data across different stages, providing quantitative basis and reliable support for engineering investment optimization and operation and maintenance risk prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122064937A_ABST
    Figure CN122064937A_ABST
Patent Text Reader

Abstract

The invention provides a data transmission analysis method, device and equipment based on multi-scale condition mutual information, and relates to the technical field of data mining. The method comprises the following steps: performing Morlet wavelet decomposition on an original sequence to obtain wavelet coefficients under a plurality of time scales, and constructing a candidate lag variable set corresponding to the time scale based on each wavelet coefficient under each time scale; screening a key lag component in each candidate lag variable set by adopting a conditional mutual information progressive strategy to obtain an optimal embedding vector corresponding to each time scale; and for each time scale, estimating conditional probability density and marginal probability density based on the corresponding optimal embedding vector, calculating conditional mutual information, and substituting the conditional mutual information into a corresponding variable embedding data transmission intensity calculation formula to obtain data transmission intensity and direction between the first original sequence and the second original sequence. According to the method, multi-scale, self-adaptive and high-precision depiction of the dynamic influence relationship among the cross-stage engineering data can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data mining technology, and in particular to a data transfer analysis method, apparatus, and equipment based on multi-scale conditional mutual information. Background Technology

[0002] With the accelerated construction of new power systems, the scale of data generated by power grid engineering in planning, investment decision-making, construction, and subsequent operation and maintenance continues to grow, exhibiting typical characteristics such as multi-source heterogeneity, cross-stage distribution, strong nonlinearity, and multi-scale time-varying. Against the backdrop of the rapid penetration of technologies such as big data, artificial intelligence, and cloud computing, power grid engineering management is transforming from traditional static experience-based management to a modern management model driven by data across the entire process and all elements. However, current engineering data still exhibits significant fragmentation between different business stages, with problems such as inconsistent data standards, weak data correlation, and poor value transmission across stages being relatively prominent, making it difficult to meet the current management needs for precise investment and full-process value optimization. How to effectively connect engineering data across multiple stages and disciplines, and accurately depict the dynamic influence relationships between different stages, has become a key challenge in the digital management of engineering projects.

[0003] In the field of scientific research, various information theory-based data transfer analysis methods have been proposed to characterize the influence paths between variables. For example, methods such as local probability density, local dynamic mapping, and conditional cumulative distribution are used to construct index systems that characterize the direction and intensity of information flow between variables. This method can capture the nonlinear characteristics of local regions of the system and reflect the fine-grained dependencies between variables through conditional probability density, thus overcoming the shortcomings of traditional mutual information measures, which can only reflect overall correlation and cannot characterize directionality. However, this method still has many limitations: most information measures are based on a single or fixed scale, making it difficult to reflect the coupling characteristics of engineering time series at different time scales; the original methods mainly deal with relationships between single or few variables, lacking effective high-dimensional conditional embedding mechanisms for typical multi-department, multi-field, and multi-stage multi-variable systems in engineering scenarios. Summary of the Invention

[0004] This invention provides a data transmission analysis method, apparatus, and device based on multi-scale conditional mutual information to solve the problems of unclear cross-stage transmission paths, complex multi-scale relationships, and obvious multi-variable coupling characteristics of engineering data.

[0005] In a first aspect, embodiments of the present invention provide a data transfer analysis method based on multi-scale conditional mutual information, comprising: The original sequence is decomposed by Morlet wavelet to obtain wavelet coefficients at multiple time scales, and a set of candidate lag variables corresponding to each time scale is constructed based on the wavelet coefficients at each time scale. The original sequence includes a first original sequence and a second original sequence, which are time series data of different stages in the entire life cycle of the power grid project. A conditional mutual information progressive strategy is used to screen key lagged components from each set of candidate lagged variables to obtain the optimal embedding vector for each time scale. For each time scale, the conditional probability density and marginal probability density are estimated based on the corresponding optimal embedding vector, and the conditional mutual information is calculated and substituted into the corresponding variable embedding data transmission strength calculation formula to obtain the data transmission strength and direction between the first original sequence and the second original sequence.

[0006] In one possible implementation, a conditional mutual information progressive strategy is used to screen key lagged components from each set of candidate lagged variables to obtain the optimal embedding vector for each time scale, including: For each time scale, an empty vector is constructed as the embedding vector, and the process is repeated multiple times, with the following steps performed in each iteration: Select an element that satisfies the conditional maximum mutual information criterion from the current set of candidate lagged variables, and move this element from the current set of candidate lagged variables into the current embedding vector until the stopping criterion is met. Then, take the current embedding vector as the optimal embedding vector for this time scale.

[0007] In one possible implementation, the maximum criterion is:

[0008] in, For given conditions Next element and elements Mutual information between them For the first Time scale during the second loop The corresponding embedding vector, For the first Time scale during the second loop The corresponding set of candidate lagged variables, for any element in, To add Element; The stopping criteria are:

[0009] in, for and Mutual information between them for and Mutual information between them This is the significance threshold.

[0010] In one possible implementation, for each time scale, the conditional probability density and marginal probability density are estimated based on the corresponding optimal embedding vector, including: Based on the nearest neighbor evolution of the state space vector, the evolution rule of the system is locally approximated by the first-order Taylor expansion, and the probability density of the first original sequence at each time scale is calculated by combining the Sigmoid-type conditional cumulative distribution function, under the condition that the corresponding optimal embedding vector is known. The probability density of the first original sequence at future time steps is estimated using the k-th nearest neighbor density estimator.

[0011] In one possible implementation, the set of candidate lagged variables is characterized by:

[0012] in, Time scale The corresponding set of candidate lagged variables, For the first original sequence on the time scale Next time lag wavelet coefficients, For the second original sequence on the time scale Next time lag wavelet coefficients; The formula for calculating the data transfer strength of variable embedding is:

[0013] in, Time scale Second original sequence With the first original sequence Data transmission strength between them For given conditions Down and Mutual information between them for and Mutual information between them For time scale The future embedding vector, In order to be on the time scale The following is obtained from the historical lag variables of the second original sequence Y, containing The optimal embedding vector for each component In order to be on the time scale The following is obtained from the historical lag variables of the first original sequence X, containing The optimal embedding vector for each component.

[0014] In one possible implementation, the original sequence also includes a third original sequence, in which case the set of candidate lagged variables is:

[0015] in, Time scale The corresponding set of candidate lagged variables, For the first original sequence on the time scale Next time lag wavelet coefficients, For the second original sequence on the time scale Next time lag wavelet coefficients, For the third original sequence on the time scale Next time lag wavelet coefficients; Correspondingly, after substituting the calculated conditional mutual information into the corresponding variable embedding data transmission strength calculation formula, it also includes: Substitute the conditional mutual information into the corresponding extended variable embedding data transfer strength calculation formula; whereby the extended variable embedding data transfer strength calculation formula is:

[0016] in, Time scale Second original sequence With the first original sequence Data transmission strength between them For given conditions Down and Mutual information between them for and Mutual information between them For time scale The future embedding vector, In order to be on the time scale The following is obtained from the historical lag variables of the second original sequence Y, containing The optimal embedding vector for each component In order to be on the time scale The following is obtained from the historical lag variables of the first original sequence X, containing The optimal embedding vector for each component In order to be on the time scale The following is obtained from the historical lagged variables of the third original sequence, containing The optimal embedding vector for each component.

[0017] In one possible implementation, the method further includes: By randomly shuffling the time indices of the first original sequence multiple times, multiple first replacement sequences are obtained; By randomly shuffling the time indices of the second original sequence multiple times, multiple second substitution sequences are obtained; For each time scale, the data transfer intensity of each first substitution sequence and the second original sequence, and the data transfer intensity of the first original sequence and each second substitution sequence are calculated, and the average value is used as the bias correction term at that time scale. Accordingly, after obtaining the data transfer strength between the first and second original sequences, the process also includes: The bias-corrected data transmission strength is obtained by subtracting the bias correction term for that time scale from the data transmission strength corresponding to each time scale.

[0018] Secondly, embodiments of the present invention provide a data transmission and analysis apparatus based on multi-scale conditional mutual information, comprising: The decomposition construction module is used to perform Morlet wavelet decomposition on the original sequence to obtain wavelet coefficients at multiple time scales, and to construct a set of candidate lag variables corresponding to each time scale based on the wavelet coefficients at each time scale; wherein, the original sequence includes a first original sequence and a second original sequence, which are time series data of different stages in the entire life cycle of the power grid project; The component filtering module is used to filter key lagged components from each set of candidate lagged variables using a conditional mutual information progressive strategy, so as to obtain the optimal embedding vector for each time scale. The modeling and calculation module is used to estimate the conditional probability density and marginal probability density based on the corresponding optimal embedding vector for each time scale, and calculate the conditional mutual information. Substituting this information into the corresponding variable embedding data transmission strength calculation formula, the data transmission strength and direction between the first original sequence and the second original sequence are obtained.

[0019] Thirdly, embodiments of the present invention provide an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method described in the first aspect or any possible implementation thereof.

[0020] The data transmission analysis method, apparatus, and equipment based on multi-scale conditional mutual information provided in this invention perform multi-scale wavelet decomposition on the original time-series data. This separates the overlapping long-term trends, stage fluctuations, and short-term noise into independent analyses at different time scales, avoiding information distortion caused by feature mixing in single-scale analysis. Employing a progressive filtering strategy driven by conditional mutual information, it automatically identifies and retains key lag components that truly contribute incremental information to the predicted future state of the target at each time scale, while eliminating redundancy and noise to construct the optimal embedding vector. Finally, based on the optimal embedding vector, conditional mutual information is calculated by estimating the conditional probability density and marginal probability density, thereby quantifying the data transmission strength. This method fully considers the nonlinearity and local distribution characteristics of engineering data, making it more accurate and robust than estimations based on a simple global distribution. Overall, this solution achieves multi-scale, adaptive, and high-precision characterization of the dynamic influence relationships between cross-stage engineering data, providing a quantitative basis for engineering investment optimization, full-process quality management, and operation and maintenance risk prediction. Attached Figure Description

[0021] Figure 1 This is a flowchart illustrating the implementation of a data transfer analysis method based on multi-scale conditional mutual information according to an embodiment of the present invention. Figure 2 This is a schematic diagram of the structure of a data transmission and analysis device based on multi-scale conditional mutual information provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0022] To address the challenges of unclear cross-stage data transmission paths, complex multi-scale relationships, and significant multivariate coupling characteristics in engineering data, this invention proposes an engineering data transmission analysis method that integrates multi-scale wavelet representation, conditional mutual information embedding, multivariate extension (MVE), and local probability density modeling. Through this multi-module collaborative approach, the data interaction chain between planning, construction, and operation and maintenance stages can be clearly depicted, providing quantifiable decision-making support for whole-process investment optimization, cost deviation diagnosis, and post-project evaluation, thus achieving data-driven collaborative management of the entire engineering lifecycle.

[0023] This invention addresses the problems of data fragmentation, weak correlation, and unclear action chains in the entire lifecycle of power grid engineering, including planning, construction, and operation and maintenance. It aims to solve the technical bottleneck of difficulty in identifying the true transmission relationships between cross-departmental, multi-stage, and multi-variable engineering data. Most existing engineering data management systems rely on static statistics or field matching methods, which can only reflect surface-level correlations and cannot characterize the dynamic impact structure of data across time scales, business stages, and local areas. Furthermore, traditional information transmission or mutual information methods have limitations such as single-scale analysis, fixed embedding structures, and inability to handle high-dimensional multivariables and nonlinear local features, making it difficult to meet the needs of identifying the coupling characteristics of engineering data at different stages.

[0024] Therefore, the core technical problems to be solved by this invention include: how to construct a data transmission and identification mechanism capable of simultaneously handling multi-scale time-varying characteristics, nonlinear dynamic characteristics, and multivariate coupling relationships; how to automatically select key lagged variables from massive engineering data to form an optimal embedding structure that accurately characterizes the cross-stage data influence path; and how to achieve accurate calculation of conditional mutual information based on local probability density, thereby obtaining directional, quantifiable, and interpretable multi-scale data transmission strength and path. By solving these problems, digital and intelligent support will be achieved for the interconnection of engineering data, the transparency of value transmission, and cross-departmental business collaboration, providing a reliable technical foundation for engineering investment optimization and operation and maintenance risk prediction.

[0025] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0026] See Figure 1 The document illustrates a flowchart of the data transfer analysis method based on multi-scale conditional mutual information provided in an embodiment of the present invention, detailed below: Step 101: Perform Morlet wavelet decomposition on the original sequence to obtain wavelet coefficients at multiple time scales, and construct a set of candidate lag variables corresponding to each time scale based on the wavelet coefficients at each time scale; wherein, the original sequence includes a first original sequence and a second original sequence, and the first original sequence and the second original sequence are time series data of different stages in the entire life cycle of the power grid project.

[0027] In this embodiment, power grid engineering generates massive amounts of multi-source, heterogeneous time-series data throughout its entire lifecycle, including planning, design, construction, and operation and maintenance. Examples include load forecasts during the planning phase, daily investment completions during the construction phase, and equipment failure records during the operation and maintenance phase. These data exhibit complex dynamic influence relationships between different stages, but traditional methods struggle to quantify their cross-stage transmission paths and intensity.

[0028] To extract multi-scale features from cross-disciplinary time-series data, the original sequence is first scaled using Morlet wavelets. This step only completes the multi-scale expansion of the data, providing a unified basic representation for subsequent embedding construction and data transfer rate calculation.

[0029] Conditional mutual information (CMI) in hybrid embeddings is a time-domain nonlinear data transfer metric, described by its wavelet spread as follows. Assume X and Y are two arbitrary time series of length N and L. X and L Y This represents the maximum time delay between X and Y. To compute the VE from Y to X, the X and Y time series are converted into MORLET wavelet coefficients using a generating function.

[0030]

[0031]

[0032] Where ω0∈[5,6] is the normalized frequency, and the time lag τ X (1≤τ) X ≤L X ) and τ Y (1≤τ) Y ≤L Y ) is the translation parameter used to locate wavelet s, s i (1≤i≤m, where m is the total number of time scales) is the time scale that determines the wavelet width and resolution, and * represents complex conjugate. This wavelet setting is the same as WTE.

[0033] Data transmission strength (VE) measures "how much uncertainty about the future of X is reduced by embedding which aspects of Y's history, and in what way, into the system state." This information content depends entirely on the composition of the embedded variables, so VE is also called variable embedding information content. VE is calculated for each time scale. i (1≤i≤m). For each time scale s i (1≤i≤m), time horizon T () The future embedding vector is defined as

[0034]

[0035] Furthermore, MVE, as a multivariate extension of VE, also begins its expansion at this scale. Without loss of generality, assume that X, Y, and Z are three arbitrary time series L of length N and L, respectively. X L Y L Z This represents the maximum time lag of the three time series. MVE first converts the X, Y, and Z time series into MORLET wavelet coefficients:

[0036]

[0037]

[0038] Similarly, the future embedding vector at scale s i The definition is as follows:

[0039] After obtaining the multi-scale representation, a set of candidate lagged variables is constructed at each scale si, and the lagged components that can provide the greatest information contribution are selected from the candidate lagged variable sets using the progressive embedding strategy of CMI.

[0040] The set of candidate lagged variables is defined as follows:

[0041] Step 102: Use a conditional mutual information progressive strategy to screen key lagged components from each set of candidate lagged variables to obtain the optimal embedding vector for each time scale.

[0042] In this embodiment, the core idea of ​​the strategy is to automatically and iteratively select the key historical components that provide the most informational increment for predicting the future state of the target variable from a large candidate set, while eliminating redundant information. Specifically, at each time scale, the algorithm starts with an empty embedding vector. In each iteration, it evaluates each remaining variable in the current candidate set and calculates the conditional mutual information between the remaining variable and the target future state, given that the selected embedding vectors are already included. Then, the variable that provides the maximum conditional mutual information value is added to the embedding vector and removed from the candidate set. This process is repeated until a preset stopping condition is met (e.g., the information contribution of the newly added variable is no longer significant). The resulting embedding vector is the optimal embedding vector at that scale, representing the most concise and predictive combination of historical information.

[0043] The progressive scheme of Conditional Mutual Information (CMI) for each time scale s i Both are applicable. The asymptotic approach starts from an empty vector b. In the first iteration cycle, VE... To find the element that satisfies the maximum criterion :

[0044] in yes and The mutual information rate between them. Select elements that satisfy the maximum criterion. To connect form , and then from Remove from And obtained .

[0045] In the k-th iteration cycle, VE searches for elements in the remainder set B. (Originally obtained from the k-th iteration) Satisfies the maximum criterion.

[0046]

[0047] and element x′ from arrive To obtain the expanded embedding vector and .

[0048] The progressive scheme stops after k+1 iterations and uses As the final selected embedding vector, it stops if the following criteria are met:

[0049] Here, A is the significance threshold (between 0 and 1), which controls the inclusion of embedded components. This stopping criterion ensures the inclusion of contributing components while preventing the addition of useless components. The progressive scheme stops if it fails to provide significant data when a new component is included.

[0050] Time scale s i VE is calculated in the following way:

[0051] Among them, the VE data transfer between coupled time series is evaluated at the same time scale. i (i=1,2,...,64) Unlike VE, MVE's candidate component set is multivariate and is defined by all time series data in the system.

[0052] The initially selected embedding vector is still an empty vector. .

[0053] MVE follows the same progressive scheme and maximum criterion as VE, the only difference being the collective set of candidate components. MVE selects candidate components from all, rather than just two, variables in the system that facilitate inference of direct data transfer.

[0054] If the progressive scheme stops and is used in the k+1th iteration cycle... As the final selected embedding vector. At time scale s i Below, the MVE representation of Y->X is as follows:

[0055] in , and They are The X, Y, and Z components. Furthermore, the MVE assesses data transfer between wavelet coefficients of the same scale.

[0056] Step 103: For each time scale, estimate the conditional probability density and marginal probability density based on the corresponding optimal embedding vector, and calculate the conditional mutual information. Substitute this information into the corresponding variable embedding data transmission strength calculation formula to obtain the data transmission strength and direction between the first original sequence and the second original sequence.

[0057] In this embodiment, the core task is to estimate two key probability densities for each time scale: first, the marginal probability density of the target variable (such as the first original sequence X) at future moments, which describes the statistical regularity of the future value of X itself; and second, the conditional probability density of the future value of X given the optimal embedding vector (which includes the history of the first sequence itself, the history of the second sequence Y, etc.), which reflects the uncertainty of the future of X in a specific historical context. By comparing these two probability distributions, the conditional mutual information can be calculated, which quantifies the amount of "additional" information about the future of X provided solely by the historical part of Y, given all the optimal historical information (including the history of X itself). Finally, substituting this conditional mutual information value into the formula for calculating the data transfer strength of variable embedding yields a scalar result that directly characterizes the data transfer strength from the second original sequence Y to the first original sequence X at that time scale, and its non-zero value itself implies the transfer direction from Y to X.

[0058] Conditional probability density is used to describe the probability density function of a given upstream variable. In the case of downstream variables The possible values ​​and their probability distribution reflect the degree of uncertainty constraint imposed by upstream data on downstream data. Different conditional probability density distributions indicate... right Future changes in x have a significant impact; if the conditional probability density does not change significantly with x, it indicates a weak correlation between the two. Local probability density modeling can be used to accurately estimate the conditional probability density. ( | Based on this, conditional entropy and conditional mutual information are obtained, thereby quantifying the actual impact of data from one stage on data from the next stage. Marginal probability density ( The conditional probability density (CPD) is used to describe the distribution characteristics of a variable, reflecting its typical value range, frequency of occurrence, and overall uncertainty level across the entire sample. Its significance lies in providing the basic uncertainty under "unconditional conditions," enabling subsequent comparison of the difference between conditional probability density and marginal probability density to determine whether one variable reduces the uncertainty of another. Marginal probability density is an important component in calculating entropy and joint entropy; in this invention, it is used to construct the denominator benchmark for mutual information and conditional mutual information, mathematically supporting a quantitative assessment of the intensity of cross-departmental data transfer.

[0059] This invention employs multi-scale wavelet decomposition of the original time-series data to separate the overlapping long-term trends, stage fluctuations, and short-term noise for independent analysis at different time scales, avoiding information distortion caused by feature mixing in single-scale analysis. A conditional mutual information-driven progressive filtering strategy is used to automatically identify and retain key lag components that truly contribute incremental information to the predicted future state of the target at each time scale, while eliminating redundancy and noise to construct the optimal embedding vector. Finally, based on the optimal embedding vector, conditional mutual information is calculated by estimating the conditional probability density and marginal probability density, thereby quantifying the data transmission strength. This method fully considers the nonlinearity and local distribution characteristics of engineering data, making it more accurate and robust than estimations based on a simple global distribution. Overall, this scheme achieves multi-scale, adaptive, and high-precision characterization of the dynamic influence relationships between cross-stage engineering data, providing a quantitative basis for engineering investment optimization, full-process quality management, and operation and maintenance risk prediction.

[0060] In one possible implementation, a conditional mutual information progressive strategy is used to screen key lagged components from each set of candidate lagged variables to obtain the optimal embedding vector for each time scale, including: For each time scale, an empty vector is constructed as the embedding vector, and the process is repeated multiple times, with the following steps performed in each iteration: Select an element that satisfies the conditional maximum mutual information criterion from the current set of candidate lagged variables, and move this element from the current set of candidate lagged variables into the current embedding vector until the stopping criterion is met. Then, take the current embedding vector as the optimal embedding vector for this time scale.

[0061] In this embodiment, for each timescale to be analyzed, the algorithm initialization phase first constructs an empty vector as the initial embedding vector, and simultaneously uses the constructed set of candidate lag variables for that scale as the initial candidate pool. Subsequently, the algorithm enters an iterative process consisting of multiple loops. In each loop, the algorithm performs a core step: based on the conditional mutual information maximization criterion, it evaluates and compares all elements in the current candidate pool, selecting the element that can maximize the predictive ability for the target's future state. Once this element is identified, it is removed from the current candidate lag variable set and added to the end of the currently constructed embedding vector, thereby updating and expanding the embedding vector. This "evaluation-selection-transfer" loop continues until a predefined stopping criterion is triggered. The stopping criterion is used to determine whether to continue adding new variables; its purpose is usually to prevent overfitting and ensure that the components included in the embedding vector are those with substantial informational contributions. When the loop terminates, the embedding vector obtained from the last update is determined as the final optimal embedding vector for that timescale, which, in a data-driven manner, condenses the most influential historical lag patterns.

[0062] It should be noted that in this embodiment, conditional mutual information is used as an evaluation criterion for variable screening. It is used to gradually select the embedding components that can maximize the future predictability of the target variable from the set of candidate lagged variables. This stage focuses on relative information gain rather than final numerical accuracy.

[0063] In one possible implementation, the maximum criterion is:

[0064] in, For given conditions Next element and elements Mutual information between them For the first Time scale during the second loop The corresponding embedding vector, For the first Time scale during the second loop The corresponding set of candidate lagged variables, for any element in, To add Element; The stopping criteria are:

[0065] in, for and Mutual information between them for and Mutual information between them This is the significance threshold.

[0066] In this embodiment, the maximum criterion aims to find the variable that provides the maximum information gain in each iteration. This criterion ensures that each newly added variable is the most information-potential under the current conditions.

[0067] Regarding stopping criteria, an effective design is based on the principle of diminishing marginal returns of information contribution. For example, one can compare the change in the overall explanatory power of the embedding vector for the target's future state before and after the inclusion of the k-th element. This is the estimated mutual information value after assuming the inclusion of the next optimal element. This criterion means that when the information gain from the newly added variable falls below a certain threshold (1-A), its contribution is considered insignificant, thus terminating the screening process. The threshold A can be set by domain experts based on the need for a balance between model simplification and predictive power.

[0068] In one possible implementation, for each time scale, the conditional probability density and marginal probability density are estimated based on the corresponding optimal embedding vector, including: Based on the nearest neighbor evolution of the state space vector, the evolution rule of the system is locally approximated by the first-order Taylor expansion, and the probability density of the first original sequence at each time scale is calculated by combining the Sigmoid-type conditional cumulative distribution function, under the condition that the corresponding optimal embedding vector is known. The probability density of the first original sequence at future time steps is estimated using the k-th nearest neighbor density estimator.

[0069] In this embodiment, based on a local dynamical system model and a nonparametric density estimation method, the key components in the calculation of data transfer volume—conditional probability density and marginal probability density—are estimated. The goal is to provide a numerical basis for the mutual information term in the VE / MVE formula in the previous section.

[0070] First, the joint probability is decomposed into marginal probability and conditional probability:

[0071] Subsequently, the conditional probability density and marginal probability density are estimated by modeling the local nonlinear dynamical system and estimating the marginal density function, respectively. The prediction quality of the deterministic model is then used to estimate the conditional probability density. , , and , , , these probabilities are related to the occurrence when given the previous state of the system and the likelihood of. Specifically, the method is based on an estimate of and the nearest neighbors of. The marginal probability density functions and can also be approximated using the same strategy. To this end, start by constructing the cumulative distribution function in order to derive the required probability density from it. This function is, by definition, a real-valued and strictly increasing function of the random variable X, and is usually denoted as:

[0072] where the right side refers to the probability that the random variable X takes a value less than or equal to x. Therefore, when a < b, the probability that x falls in the semi-closed interval (a, b] is:

[0073] The cumulative density function of a continuous random variable X can be expressed as the integral of its probability density function as follows : H<00004 >Similarly, for the conditional cumulative distribution function, we have:

[0075] Therefore, the conditional probability density (CPD) function can be written as:

[0076] If the conditional cumulative function in is chosen as a sigmoid function , whose argument is , like the following formula:

[0077] where f is a system model constructed from the data; this allows for distinguishing between two extreme cases of the distribution: one is strictly deterministic, i.e., without error, where , F is expressed as:

[0078] where θ is the Heaviside function, and the other is the probabilistic case, where the parameter . In a sense, r is a measure of the randomness of the process X.

[0079] From the formulas ​​It is evident that if there is an approximation of the dynamic rules governing system evolution, then an appropriate cumulative distribution function can be constructed to estimate the probability density required for calculating data transfer. In this case, it is recommended to choose a sigmoid function where the parameter r is proportional to 1 / σ, where σ is the standard deviation of the modeling error, because the slope of the transition segment of the sigmoid function determines the width of the associated probability distribution.

[0080] At this point, the above can be summarized into two main ideas: First, Schreiber's data transfer, along with some other estimators, can be represented using conditional probability density (CPD); second, CPD can be estimated using deterministic models of the dynamic rules governing the data generation. The remainder of this section aims to implement these models and demonstrate how they can be used to construct estimators of data transfer, with computational costs comparable to determining the nearest quantity for k=1.

[0081] Suppose we have a dynamical system state sequence ,in These data either come from timing measurements of the system's d state components or are obtained by reconstructing the state space using partial state information based on Takens' theorem. In this method, the evolution rule of the dynamical system is estimated through a local approximation of f. This method can be algorithmically described as follows: 1. Given the i-th data point , where i = 1, 2, ..., N.

[0082] 2. Determination and Status The m closest neighbors Using Euclidean distance as a metric, and according to... Sort by distance from smallest to largest. It is the index of the j-th neighbor of the i-th data point, with a value range of [1, N]. 1).

[0083] 3. By analyzing f in its nearest neighbor... A first-order Taylor expansion at that point yields an approximation:

[0084] in It is the Jacobian matrix of f.

[0085] 4. Finally, under the zeroth-order approximation, the evolution of the system is approximated by the nearest neighbor evolution, that is: If a first-order approximation is used, then calculation is required. .

[0086] The marginal probability density functions p(x) and p(y) can be estimated using the k-th nearest neighbor density estimator.

[0087] To graphically illustrate the method for estimating probability density, Figure 1 The conditional probability density function and marginal probability density function are shown when the number of data points N=500 and the skewtent map parameter a=0.65.

[0088]

[0089] Therefore, the dynamic rule f in the above formula is approximated by the zeroth order, and the formula is used. and The conditional probability density (CPD) function can be written as:

[0090] In one possible implementation, the set of candidate lagged variables is characterized by:

[0091] in, Time scale The corresponding set of candidate lagged variables, For the first original sequence on the time scale Next time lag wavelet coefficients, For the second original sequence on the time scale Next time lag wavelet coefficients; The formula for calculating the data transfer strength of variable embedding is:

[0092] in, Time scale Second original sequence With the first original sequence Data transmission strength between them For given conditions Down and Mutual information between them for and Mutual information between them For time scale The future embedding vector, In order to be on the time scale The following is obtained from the historical lag variables of the second original sequence Y, containing The optimal embedding vector for each component In order to be on the time scale The following is obtained from the historical lag variables of the first original sequence X, containing The optimal embedding vector for each component.

[0093] In this embodiment, under the bivariate analysis scenario, the candidate lag variable set consists of two parts of historical wavelet coefficients: the first part originates from the first original sequence X itself, containing coefficients from the current time... Backtracking to maximum lag The coefficients, the second part of which comes from the second original sequence Y, include its backtracking to the maximum lag. The coefficient.

[0094] The optimal embedding vector obtained based on this set It can be further divided into components that mainly come from the history of X. and the main component from the history of Y At this point, the variable embedding formula used to calculate the transmission strength from Y to X is specified as follows: Among them, molecules Calculated the history of controlling X itself Under these conditions, the history of Y For X Future The contribution (i.e., conditional mutual information). Denominator This is the total mutual information, used to normalize the transmission strength, making its value range easier to interpret. This formula clearly quantifies the net information transmission strength of Y to X after excluding the effects of X's autocorrelation.

[0095] In one possible implementation, the original sequence also includes a third original sequence, in which case the set of candidate lagged variables is:

[0096] in, Time scale The corresponding set of candidate lagged variables, For the first original sequence on the time scale Next time lag wavelet coefficients, For the second original sequence on the time scale Next time lag wavelet coefficients, For the third original sequence on the time scale Next time lag wavelet coefficients; Correspondingly, after substituting the calculated conditional mutual information into the corresponding variable embedding data transmission strength calculation formula, it also includes: Substitute the conditional mutual information into the corresponding extended variable embedding data transfer strength calculation formula; whereby the extended variable embedding data transfer strength calculation formula is:

[0097] in, Time scale Second original sequence With the first original sequence Data transmission strength between them For given conditions Down and Mutual information between them for and Mutual information between them For time scale The future embedding vector, In order to be on the time scale The following is obtained from the historical lag variables of the second original sequence Y, containing The optimal embedding vector for each component In order to be on the time scale The following is obtained from the historical lag variables of the first original sequence X, containing The optimal embedding vector for each component In order to be on the time scale The following is obtained from the historical lagged variables of the third original sequence, containing The optimal embedding vector for each component.

[0098] In this embodiment, when the analysis scenario involves more than two variables, it is necessary to control the influence of other potential confounding variables in order to more accurately capture direct data transfer relationships. In this scenario, the original sequence includes a first original sequence X and a second original sequence Y, as well as a third original sequence Z, which is any original sequence other than the first original sequence X and the second original sequence Y. At this time, regarding the time scale... Constructed candidate lagged variable set It needs to be expanded by including a third sequence Z, which backtracks from the current time to the maximum lag. The wavelet coefficients, i.e. The optimal embedding vector obtained from this extended set. Its historical information can be categorized into three parts: originating from X's own history. Originating from the history of Y and the history derived from Z (and other possible variables) Accordingly, the calculation of data transmission strength needs to be upgraded to a multivariate embedding formula: The key difference between this formula and the bivariate formula lies in the conditional term: the conditional mutual information in the numerator is based on the known history of X itself. and the history of all other variables Under the premise of calculating the history of Y The information contribution to the future of X. This is equivalent to statistically excluding possible mediating or confounding effects such as variable Z, thus assessing the strength of data transfer from Y to X that is closer to a direct causal relationship. The results are more robust and explanatory than bivariate analysis.

[0099] In one possible implementation, the method further includes: By randomly shuffling the time indices of the first original sequence multiple times, multiple first replacement sequences are obtained; By randomly shuffling the time indices of the second original sequence multiple times, multiple second substitution sequences are obtained; For each time scale, the data transfer intensity of each first substitution sequence and the second original sequence, and the data transfer intensity of the first original sequence and each second substitution sequence are calculated, and the average value is used as the bias correction term at that time scale. Accordingly, after obtaining the data transfer strength between the first and second original sequences, the process also includes: The bias-corrected data transmission strength is obtained by subtracting the bias correction term for that time scale from the data transmission strength corresponding to each time scale.

[0100] In this embodiment, to verify the statistical significance of the data transfer analysis results and avoid spurious transfer signals due to the inherent characteristics of the sequence, this method introduces a bias correction procedure based on substitute data. Specifically, a time-shift substitution index is used to verify the significance of the results. Taking VE as an example, let... and The MORLET wavelet coefficients, i.e., VE, represent any time series X and Y. (s i ) indicates that on the time scale s i The following is the VE data transfer from X to Y. Maintain... Keep it unchanged, and randomly shuffle it. The time index, thus obtaining The replacement sequence. Then, the VE method is applied to the original sequence respectively. With alternative series The result is denoted as VE. (s i , q), where q represents the number of the substitution sequence. Therefore, X→Y at scale s i The deviation correction VE data transmission is as follows:

[0101] Use in the following contexts To represent at scale s i Below is the deviation correction from X to Y. The definition method is the same for the deviation correction VE in the reverse direction, as well as the MVE and WTE after deviation correction.

[0102] The bias correction was applied to the data transfer volume at all scales, and the results are as follows:

[0103] It can also further calculate and compare the results of the data transfer:

[0104] As shown above, the engineering data transmission analysis method provided in this embodiment, based on multi-scale wavelet decomposition, conditional mutual information embedding construction, and local probability density modeling, enables the identification of dynamic interaction relationships among engineering data from multiple stages, including planning, construction, and operation and maintenance. Addressing the characteristics of engineering data such as multi-source heterogeneity, complex cross-stage coupling, and significant local nonlinearity, this method first constructs a unified multi-scale data representation system through wavelet decomposition to extract key structural features from different stages. Then, it employs a conditional mutual information-driven embedding construction strategy to automatically identify lagged variables that have a real impact on downstream stages, achieving effective feature extraction across departments and multiple variables. Finally, it obtains an accurate estimate of the conditional mutual information based on local probability density modeling, substitutes it into the VE / MVE formula, and obtains the transmission strength and direction of engineering data at each scale. This method can be used to characterize the impact path of data throughout the entire process of planning, construction, and operation and maintenance, providing a quantitative basis for engineering investment optimization, full-process quality management, and operation and maintenance risk prediction. The workflow is shown below, with the specific steps divided into the following three steps: Step 1: Perform Morlet wavelet decomposition on the engineering data to obtain feature sequences at each scale and construct a unified data representation space.

[0105] Step 2: Using a CMI / PCMI progressive strategy, key lagged components are selected from multi-scale candidate variables to form the optimal conditional embedding structure, and then extended to multivariate MVE modeling.

[0106] Step 3: Based on the optimal embedding vector obtained in Step 2, the marginal probability density and conditional probability density are constructed using a local dynamic model and a non-parametric density estimation method, respectively, and the conditional mutual information is calculated accordingly. This mutual information is then substituted into the multi-scale definition of VE / MVE to obtain the transmission strength and direction of engineering data at different scales, reflecting the actual impact of cross-stage and cross-professional data on future behavior.

[0107] Compared with existing engineering data management methods based on single-scale mutual information, simple correlation analysis, or empirical statistics, this invention, through a combined technical approach of "multi-scale wavelet representation + conditional mutual information embedding construction + multivariate MVE + local probability density modeling," has the following beneficial effects in characterizing cross-stage data transfer in power grid engineering: This invention enables multi-scale characterization of data throughout the entire engineering process, overcoming the distortion problem of single-scale analysis. First, it utilizes Morlet wavelets to decompose data from planning, construction, and operation phases into coefficient sequences at different scales, making the original non-stationary sequences approximate locally stationary processes at each scale. This separates the previously intertwined long-term trends, periodic fluctuations, and short-term disturbances onto different scales. Subsequent embedding construction and information transfer calculations are performed separately at each scale, effectively modeling the impact of "different time characteristics" independently. Theoretically, this avoids the underestimation or overestimation of mutual information caused by mixing long-term trends with short-term noise in calculations, thereby improving the accuracy and interpretability of engineering data transfer analysis.

[0108] This invention achieves automatic screening of "true influencing factors" through conditional mutual information-driven embedding. It constructs a set of candidate lagged variables at each scale and maximizes conditional mutual information. The recursive criterion is used to select new variables, and... As a stopping condition, this process is equivalent to retaining only variables that significantly increase the explanatory power of data for future stages, while excluding redundant or highly overlapping fields from the embedding vector, given the "set of considered variables." Therefore, compared to traditional methods of fixing lag orders or manually selecting fields, this invention can automatically extract a set of variables from massive engineering data that contribute "incremental information" to downstream stages, ensuring that the information transmission reflects effective impacts rather than simple correlations, and reducing estimation bias caused by the curse of dimensionality.

[0109] This invention employs local probability density modeling to improve the accuracy of mutual information estimation and adapt to the nonlinearity and locality of engineering data. In mutual information calculation, this invention no longer assumes a globally linear or simple distribution, but instead uses local dynamic mapping and a sigmoid-type conditional cumulative distribution to construct the conditional probability density. And the edge density is estimated using the k-nearest neighbor method. Compared to mutual information calculation methods based on global histograms or kernel density, this invention is more stable in engineering scenarios with limited sample size and uneven distribution.

[0110] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0111] The following are device embodiments of the present invention. For details not described in detail, please refer to the corresponding method embodiments described above.

[0112] Figure 2 A schematic diagram of the data transfer and analysis device based on multi-scale conditional mutual information provided in an embodiment of the present invention is shown. For ease of explanation, only the parts related to the embodiment of the present invention are shown, and are described in detail below: like Figure 2 As shown, the data transfer and analysis device 2 based on multi-scale conditional mutual information includes: The decomposition and construction module 21 is used to perform Morlet wavelet decomposition on the original sequence to obtain wavelet coefficients at multiple time scales, and to construct a set of candidate lag variables corresponding to each time scale based on the wavelet coefficients at each time scale; wherein, the original sequence includes a first original sequence and a second original sequence, and the first original sequence and the second original sequence are time series data of different stages in the entire life cycle of the power grid project; The component screening module 22 is used to screen key lagged components from each candidate lagged variable set using a conditional mutual information progressive strategy to obtain the optimal embedding vector for each time scale. The modeling and calculation module 23 is used to estimate the conditional probability density and marginal probability density based on the corresponding optimal embedding vector for each time scale, and calculate the conditional mutual information and substitute it into the corresponding variable embedding data transmission strength calculation formula to obtain the data transmission strength and direction between the first original sequence and the second original sequence.

[0113] In one possible implementation, the component filtering module 22 is specifically used for: For each time scale, an empty vector is constructed as the embedding vector, and the process is repeated multiple times, with the following steps performed in each iteration: Select an element that satisfies the conditional maximum mutual information criterion from the current set of candidate lagged variables, and move this element from the current set of candidate lagged variables into the current embedding vector until the stopping criterion is met. Then, take the current embedding vector as the optimal embedding vector for this time scale.

[0114] In one possible implementation, the maximum criterion is:

[0115] in, For given conditions Next element and elements Mutual information between them For the first Time scale during the second loop The corresponding embedding vector, For the first Time scale during the second loop The corresponding set of candidate lagged variables, for any element in, To add Element; The stopping criteria are:

[0116] in, for and Mutual information between them for and Mutual information between them This is the significance threshold.

[0117] In one possible implementation, the modeling calculation module 23 is specifically used for: Based on the nearest neighbor evolution of the state space vector, the evolution rule of the system is locally approximated by the first-order Taylor expansion, and the probability density of the first original sequence at each time scale is calculated by combining the Sigmoid-type conditional cumulative distribution function, under the condition that the corresponding optimal embedding vector is known. The probability density of the first original sequence at future time steps is estimated using the k-th nearest neighbor density estimator.

[0118] In one possible implementation, the set of candidate lagged variables is characterized by:

[0119] in, Time scale The corresponding set of candidate lagged variables, For the first original sequence on the time scale Next time lag wavelet coefficients, For the second original sequence on the time scale Next time lag wavelet coefficients; The formula for calculating the data transfer strength of variable embedding is:

[0120] in, Time scale Second original sequence With the first original sequence Data transmission strength between them For given conditions Down and Mutual information between them for and Mutual information between them For time scale The future embedding vector, In order to be on the time scale The following is obtained from the historical lag variables of the second original sequence Y, containing The optimal embedding vector for each component In order to be on the time scale The following is obtained from the historical lag variables of the first original sequence X, containing The optimal embedding vector for each component.

[0121] In one possible implementation, the original sequence also includes a third original sequence, in which case the set of candidate lagged variables is:

[0122] in, Time scale The corresponding set of candidate lagged variables, For the first original sequence on the time scale Next time lag wavelet coefficients, For the second original sequence on the time scale Next time lag wavelet coefficients, For the third original sequence on the time scale Next time lag wavelet coefficients; Correspondingly, the modeling and calculation module 23 is also used for: After substituting the calculated conditional mutual information into the corresponding variable embedding data transfer strength calculation formula, the conditional mutual information is then substituting into the corresponding extended variable embedding data transfer strength calculation formula; whereby the extended variable embedding data transfer strength calculation formula is:

[0123] in, Time scale Second original sequence With the first original sequence Data transmission strength between them For given conditions Down and Mutual information between them for and Mutual information between them For time scale The future embedding vector, In order to be on the time scale The following is obtained from the historical lag variables of the second original sequence Y, containing The optimal embedding vector for each component In order to be on the time scale The following is obtained from the historical lag variables of the first original sequence X, containing The optimal embedding vector for each component In order to be on the time scale The following is obtained from the historical lagged variables of the third original sequence, containing The optimal embedding vector for each component.

[0124] In one possible implementation, the component filtering module 22 is further configured to: By randomly shuffling the time indices of the first original sequence multiple times, multiple first replacement sequences are obtained; By randomly shuffling the time indices of the second original sequence multiple times, multiple second substitution sequences are obtained; For each time scale, the data transfer intensity of each first substitution sequence and the second original sequence, and the data transfer intensity of the first original sequence and each second substitution sequence are calculated, and the average value is used as the bias correction term at that time scale. Accordingly, after obtaining the data transfer strength between the first and second original sequences, the process also includes: The bias-corrected data transmission strength is obtained by subtracting the bias correction term for that time scale from the data transmission strength corresponding to each time scale.

[0125] This invention employs multi-scale wavelet decomposition of the original time-series data to separate the overlapping long-term trends, stage fluctuations, and short-term noise for independent analysis at different time scales, avoiding information distortion caused by feature mixing in single-scale analysis. A conditional mutual information-driven progressive filtering strategy is used to automatically identify and retain key lag components that truly contribute incremental information to the predicted future state of the target at each time scale, while eliminating redundancy and noise to construct the optimal embedding vector. Finally, based on the optimal embedding vector, conditional mutual information is calculated by estimating the conditional probability density and marginal probability density, thereby quantifying the data transmission strength. This method fully considers the nonlinearity and local distribution characteristics of engineering data, making it more accurate and robust than estimations based on a simple global distribution. Overall, this scheme achieves multi-scale, adaptive, and high-precision characterization of the dynamic influence relationships between cross-stage engineering data, providing a quantitative basis for engineering investment optimization, full-process quality management, and operation and maintenance risk prediction.

[0126] Figure 3This is a schematic diagram of an electronic device provided in an embodiment of the present invention. Figure 3 As shown, the electronic device 3 of this embodiment includes a processor 30 and a memory 31. The memory 31 stores a computer program 32. When the processor 30 executes the computer program 32, it implements the steps in the various method embodiments described above. Alternatively, when the processor 30 executes the computer program 32, it implements the functions of each module / unit in the various device embodiments described above.

[0127] For example, computer program 32 may be divided into one or more modules / units, which are stored in memory 31 and executed by processor 30 to complete the present invention. The one or more modules / units may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of computer program 32 in electronic device 3.

[0128] Electronic device 3 may include, but is not limited to, processor 30 and memory 31. Those skilled in the art will understand that... Figure 3 This is merely an example of electronic device 3 and does not constitute a limitation on electronic device 3. It may include more or fewer components than shown, or combine certain components, or different components. For example, electronic device 3 may also include input / output devices, network access devices, buses, etc.

[0129] For the sake of simplicity and clarity, only the above-described functional modules / units are used as examples. In practical applications, the functions described above can be assigned to different functional modules / units as needed. These modules / units can be implemented in hardware, software, or a combination of both.

[0130] In the above embodiments, the descriptions of each embodiment have their own emphasis. Parts not detailed or described in a particular embodiment can be referred to in the relevant descriptions of other embodiments. Unless otherwise specified or in conflict with logic, the terminology and / or descriptions between different embodiments are consistent and can be referenced interchangeably. Technical features in different embodiments can be combined to form new embodiments based on their inherent logical relationships.

[0131] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A data transfer analysis method based on multi-scale conditional mutual information, characterized in that, include: The original sequence is decomposed using Morlet wavelet decomposition to obtain wavelet coefficients at multiple time scales, and a set of candidate lag variables corresponding to each time scale is constructed based on the wavelet coefficients at each time scale; wherein, the original sequence includes a first original sequence and a second original sequence, and the first original sequence and the second original sequence are time series data of different stages in the entire life cycle of the power grid project; A conditional mutual information progressive strategy is used to screen key lagged components from each set of candidate lagged variables to obtain the optimal embedding vector for each time scale. For each time scale, the conditional probability density and marginal probability density are estimated based on the corresponding optimal embedding vector, and the conditional mutual information is calculated and substituted into the corresponding variable embedding data transmission strength calculation formula to obtain the data transmission strength and direction between the first original sequence and the second original sequence.

2. The data transfer analysis method based on multi-scale conditional mutual information according to claim 1, characterized in that, The stepwise strategy employing conditional mutual information to screen key lagged components from each set of candidate lagged variables, obtaining the optimal embedding vector for each time scale, includes: For each time scale, an empty vector is constructed as the embedding vector, and the process is repeated multiple times, with the following steps performed in each iteration: Select an element that satisfies the conditional maximum mutual information criterion from the current set of candidate lagged variables, and move this element from the current set of candidate lagged variables into the current embedding vector until the stopping criterion is met. Then, take the current embedding vector as the optimal embedding vector for this time scale.

3. The data transfer analysis method based on multi-scale conditional mutual information according to claim 2, characterized in that, The maximum criterion is: in, For given conditions Next element and elements Mutual information between them For the first Time scale during the next loop The corresponding embedding vector, For the first Time scale during the next loop The corresponding set of candidate lagged variables, for any element in, To add Element; The stopping criteria are as follows: in, for and Mutual information between them for and Mutual information between them This is the significance threshold.

4. The data transfer analysis method based on multi-scale conditional mutual information according to claim 1, characterized in that, The estimation of conditional probability density and marginal probability density based on the corresponding optimal embedding vector for each time scale includes: Based on the nearest neighbor evolution of the state space vector, the system evolution rule is locally approximated by the first-order Taylor expansion, and the probability density of the first original sequence at each time scale is calculated in combination with the Sigmoid-type conditional cumulative distribution function, under the condition that the corresponding optimal embedding vector is known. The probability density of the first original sequence at future time steps is estimated using the k-th nearest neighbor density estimator.

5. The data transfer analysis method based on multi-scale conditional mutual information according to any one of claims 1 to 4, characterized in that, The set of candidate lagged variables is as follows: in, Time scale The corresponding set of candidate lagged variables, For the first original sequence on the time scale Next time lag wavelet coefficients, For the second original sequence on the time scale Next time lag wavelet coefficients; The formula for calculating the data transmission strength of the embedded variables is: in, Time scale Second original sequence With the first original sequence Data transmission strength between them For given conditions Down and Mutual information between them for and Mutual information between them For time scale The future embedding vector, In order to be on a time scale The following is obtained from the historical lag variables of the second original sequence Y, containing The optimal embedding vector for each component In order to be on a time scale The following is obtained from the historical lag variables of the first original sequence X, containing The optimal embedding vector for each component.

6. The data transfer analysis method based on multi-scale conditional mutual information according to claim 5, characterized in that, The original sequence also includes a third original sequence, at which point the set of candidate lagged variables is: in, Time scale The corresponding set of candidate lagged variables, For the first original sequence on the time scale Next time lag wavelet coefficients, For the second original sequence on the time scale Next time lag wavelet coefficients, For the third original sequence on the time scale Next time lag wavelet coefficients; Accordingly, after substituting the calculated conditional mutual information into the corresponding variable embedding data transmission strength calculation formula, the method further includes: Substitute the conditional mutual information into the corresponding extended variable embedding data transmission strength calculation formula; wherein, the extended variable embedding data transmission strength calculation formula is: in, Time scale Second original sequence With the first original sequence Data transmission strength between them For given conditions Down and Mutual information between them for and Mutual information between them For time scale The future embedding vector, In order to be on a time scale The following is obtained from the historical lag variables of the second original sequence Y, containing The optimal embedding vector for each component In order to be on a time scale The following is obtained from the historical lag variables of the first original sequence X, containing The optimal embedding vector for each component In order to be on a time scale The following is obtained from the historical lagged variables of the third original sequence, containing The optimal embedding vector for each component.

7. The data transfer analysis method based on multi-scale conditional mutual information according to claim 5, characterized in that, The method further includes: By randomly shuffling the time index of the first original sequence multiple times, multiple first replacement sequences are obtained; By randomly shuffling the time index of the second original sequence multiple times, multiple second substitution sequences are obtained; For each time scale, the data transfer intensity of each first substitution sequence and the second original sequence, and the data transfer intensity of the first original sequence and each second substitution sequence are calculated, and the average value is used as the deviation correction term at that time scale. Accordingly, after obtaining the data transfer strength between the first original sequence and the second original sequence, the method further includes: The bias-corrected data transmission strength is obtained by subtracting the bias correction term for that time scale from the data transmission strength corresponding to each time scale.

8. A data transmission and analysis device based on multi-scale conditional mutual information, characterized in that, include: The decomposition construction module is used to perform Morlet wavelet decomposition on the original sequence to obtain wavelet coefficients at multiple time scales, and to construct a set of candidate lag variables corresponding to each time scale based on the wavelet coefficients at each time scale; wherein, the original sequence includes a first original sequence and a second original sequence, and the first original sequence and the second original sequence are time series data of different stages in the entire life cycle of the power grid project; The component filtering module is used to filter key lagged components from each set of candidate lagged variables using a conditional mutual information progressive strategy, so as to obtain the optimal embedding vector for each time scale. The modeling and calculation module is used to estimate the conditional probability density and marginal probability density based on the corresponding optimal embedding vector for each time scale, and calculate the conditional mutual information and substitute it into the corresponding variable embedding data transmission strength calculation formula to obtain the data transmission strength and direction between the first original sequence and the second original sequence.

9. The data transmission and analysis device based on multi-scale conditional mutual information according to claim 8, characterized in that, The component filtering module is specifically used for: For each time scale, an empty vector is constructed as the embedding vector, and the process is repeated multiple times, with the following steps performed in each iteration: Select an element that satisfies the conditional maximum mutual information criterion from the current set of candidate lagged variables, and move this element from the current set of candidate lagged variables into the current embedding vector until the stopping criterion is met. Then, take the current embedding vector as the optimal embedding vector for this time scale.

10. An electronic device, characterized in that, It includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method as described in any one of claims 1 to 7.